Voice wake-up function control method and apparatus, electronic device, medium, and product

CN122821940APending Publication Date: 2026-09-25BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510352824.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,相关技术中,电子设备设备容易被周围环境中其他电子设备发出的唤醒词误唤醒

Benefits of technology

[0018]根据本公开一些实施例,提供一种存储介质,所述存储介质存储有计算机程序或指令,当所述存储介质中的所述计算机程序或指令由电子设备的处理器执行时,使得电子设备能够执行第一方面或第一方面任一实施例所述的语音唤醒功能的控制方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821940A_ABST
    Figure CN122821940A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a voice wake-up function control method and device, electronic equipment, medium and product. The voice wake-up function control method comprises: in response to detecting that an input voice meets a voice wake-up function execution condition, performing sound source detection on the input voice; if it is detected that the sound source of the input voice is a sound emitted by a non-target object, the voice wake-up function is prevented from being executed; wherein the sound source detection is used to determine whether the input voice is a sound emitted by a target object or a sound emitted by a non-target object. Through the present disclosure, the voice wake-up function is prevented from being falsely woken up by a sound emitted by a non-target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio processing, and more particularly to a control method, apparatus, electronic device, medium, and product for a voice wake-up function. Background Technology

[0002] As the functionality of electronic devices such as smart cockpits, smartphones, tablets, and smart TVs continues to improve, voice wake-up has become a convenient human-computer interaction method. Users can wake up the device to perform corresponding operations through specific voice commands.

[0003] However, in related technologies, electronic devices are easily mistakenly awakened by wake words emitted by other electronic devices in the surrounding environment. When the wake word emitted by an electronic device is exactly the same as the pre-programmed wake-up voice, the device may be woken up unexpectedly, leading to increased battery consumption, unnecessary actions, and a severely negative impact on user experience. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this disclosure provides a control method, device, electronic device, medium, and product for voice wake-up function.

[0005] According to some embodiments of this disclosure, a control method for a voice wake-up function is provided, comprising: in response to detecting input voice that satisfies the function of performing a voice wake-up function, performing sound source detection on the input voice; if the input voice is detected to be a sound emitted by a non-target object, then preventing the execution of the voice wake-up function; wherein the sound source detection is used to determine whether the input voice is a sound emitted by a target object or a sound emitted by a non-target object.

[0006] In some embodiments, the method further includes: if the input voice is detected to be a voice emitted by the target object, then performing the voice wake-up function.

[0007] In some embodiments, the sound source detection of the input speech includes: performing sound source detection on the input speech based on a sound source detection model; wherein the sound source detection model is trained in the following manner: obtaining a first pre-trained model, wherein the input data of the first pre-trained model is audio, the output data of the first pre-trained model is the audio type of the audio, the first pre-trained model includes an output layer, and the output layer of the first pre-trained model is initially a multi-head output layer; adjusting the output layer of the first pre-trained model to a single-head output layer so that the output content of the adjusted first pre-trained model represents the sound emitted by the target object or the sound emitted by a non-target object, and / or, iteratively training the first pre-trained model based on the features of the sound training data until the loss value of the first pre-trained model meets the loss threshold, so as to adjust the model parameters of the first pre-trained model; obtaining the sound source detection model based on the adjusted first pre-trained model; wherein the sound training data includes speech and the label corresponding to the speech, the label representing that the speech is the sound emitted by the target object or the sound emitted by a non-target object.

[0008] In some embodiments, the first pre-trained model is obtained based on knowledge distillation of the second pre-trained model, the number of stacked layers of the encoder and decoder of the first pre-trained model is lower than the number of stacked layers of the encoder and decoder of the second pre-trained model, and the model parameters of the first pre-trained model are the same as the model parameters of the second pre-trained model.

[0009] In some embodiments, obtaining the sound source detection model based on the adjusted first pre-trained model includes: quantizing the model parameters in the adjusted first pre-trained model so that the accuracy of the model parameters of the quantized first pre-trained model is lower than the accuracy of the model parameters of the first pre-trained model before quantization; and using the quantized first pre-trained model as the sound source detection model.

[0010] In some embodiments, the sound emitted by the target object is the sound emitted by the user, and the sound emitted by the non-target object is an electronic sound emitted by an electronic device.

[0011] According to some embodiments of this disclosure, a control device for a voice wake-up function is provided, comprising: a processing unit, configured to perform sound source detection on the input voice in response to detecting input voice that satisfies the function of performing a voice wake-up function; and an execution unit, configured to prevent the performance of the voice wake-up function if the sound source of the input voice is detected to be a sound emitted by a non-target object; wherein the sound source detection is used to determine whether the input voice is a sound emitted by a target object or a sound emitted by a non-target object.

[0012] In some embodiments, the execution unit is further configured to: if it is detected that the input voice is a voice emitted by the target object, then execute the voice wake-up function.

[0013] In some embodiments, the processing unit performs sound source detection on the input speech in the following manner: Based on a sound source detection model, the input speech is detected; wherein the sound source detection model is trained as follows: a first pre-trained model is obtained, the input data of the first pre-trained model is audio, the output data of the first pre-trained model is the audio type of the audio, the first pre-trained model includes an output layer, and the output layer of the first pre-trained model is initially a multi-head output layer; the output layer of the first pre-trained model is adjusted to a single-head output layer so that the output content of the adjusted first pre-trained model represents the sound emitted by the target object or the sound emitted by a non-target object, and / or, based on the features of the sound training data, the first pre-trained model is iteratively trained until the loss value of the first pre-trained model meets a loss threshold, so as to adjust the model parameters of the first pre-trained model; based on the adjusted first pre-trained model, the sound source detection model is obtained; wherein the sound training data includes speech and the corresponding label of the speech, the label representing that the speech is the sound emitted by the target object or the sound emitted by a non-target object.

[0014] In some embodiments, the first pre-trained model is obtained based on knowledge distillation of the second pre-trained model, the number of stacked layers of the encoder and decoder of the first pre-trained model is lower than the number of stacked layers of the encoder and decoder of the second pre-trained model, and the model parameters of the first pre-trained model are the same as the model parameters of the second pre-trained model.

[0015] In some embodiments, obtaining the sound source detection model based on the adjusted first pre-trained model includes: quantizing the model parameters in the adjusted first pre-trained model so that the accuracy of the model parameters of the quantized first pre-trained model is lower than the accuracy of the model parameters of the first pre-trained model before quantization; and using the quantized first pre-trained model as the sound source detection model.

[0016] In some embodiments, the sound emitted by the target object is the sound emitted by the user, and the sound emitted by the non-target object is an electronic sound emitted by an electronic device.

[0017] According to some embodiments of this disclosure, an electronic device is provided, including: a processor; a memory for storing processor-executable computer programs or instructions; wherein the processor is configured to: execute the computer programs or instructions to implement a control method for the voice wake-up function as described in the first aspect or any embodiment of the first aspect.

[0018] According to some embodiments of this disclosure, a storage medium is provided that stores a computer program or instructions, which, when executed by a processor of an electronic device, enable the electronic device to perform the control method for the voice wake-up function described in the first aspect or any embodiment of the first aspect.

[0019] According to some embodiments of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a control method for the voice wake-up function as described in the first aspect or any embodiment of the first aspect.

[0020] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: by performing sound source detection on the input voice that satisfies the voice wake-up function, it is determined whether the voice segment is a sound emitted by the target object, so as to ensure that the input voice that can perform the voice wake-up function is emitted by the user, and avoid the voice wake-up function being mistakenly awakened by a sound emitted from a non-target object.

[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0023] Figure 1 This is a schematic diagram illustrating an audio detection process according to an exemplary embodiment.

[0024] Figure 2 This is a flowchart illustrating a control method for a voice wake-up function according to some embodiments of the present disclosure.

[0025] Figure 3 This is a flowchart illustrating a control method for a voice wake-up function according to some embodiments of the present disclosure.

[0026] Figure 4 This is a schematic diagram illustrating the control process of a voice wake-up function according to an exemplary embodiment.

[0027] Figure 5 This is a flowchart illustrating a control method for a voice wake-up function according to some embodiments of the present disclosure.

[0028] Figure 6 This is a flowchart illustrating a sound source detection model training method according to some embodiments of the present disclosure.

[0029] Figure 7This is a flowchart illustrating a sound source detection model training method according to some embodiments of the present disclosure.

[0030] Figure 8 This is a schematic diagram illustrating the training process of a sound source detection model according to an exemplary embodiment.

[0031] Figure 9 This is a block diagram illustrating a control device for a voice wake-up function according to some embodiments of the present disclosure.

[0032] Figure 10 This is a block diagram illustrating an electronic device for controlling a voice wake-up function according to some embodiments of the present disclosure. Detailed Implementation

[0033] Some embodiments of this disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.

[0034] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0035] This disclosure provides a control method for voice wake-up function in some embodiments, applicable to scenarios where it is necessary to verify whether audio is an electronic sound, such as in a smart cockpit or smart home scenario. After a smart device in a smart vehicle or home detects that the audio played by an electronic device includes wake-up information that meets the wake-up conditions, it further detects whether the audio segment containing the wake-up information is an electronic sound. Here, the electronic sound is the audio emitted by the electronic device through a speaker.

[0036] In related technologies, the main focus is on detecting whether the audio contains wake-up keywords, and analyzing the frequency, energy, and waveform patterns of audio segments containing wake-up keywords to determine whether the audio segment containing the wake-up keyword matches the audio emitted by the target object, thereby determining whether to wake up the target device and / or target application. Taking a smart cockpit's in-vehicle infotainment system as an example, when passengers use electronic devices to play videos, they may play wake-up words that activate the system, making it easy to wake up. Similarly, when passengers use electronic devices to make voice calls, the other party may also utter words containing the wake-up word, easily waking up the system. Once awakened, the system may execute unnecessary commands, impacting the user experience. Another example is smart devices in a smart home. When users use mobile phones, tablets, or other electronic devices at home, or watch television, these devices may emit sounds containing wake-up words, thereby waking up smart devices in the home and even executing unnecessary commands.

[0037] In one exemplary embodiment, the audio detection process in the related art is as follows: Figure 1 As shown. Figure 1 This is a schematic diagram illustrating an audio detection process according to an exemplary embodiment. Audio detection in related technologies can include three steps: sound acquisition, front-end signal processing, and wake-up event detection. The sound acquisition step primarily uses a microphone array built into the electronic device to continuously acquire ambient sound at a sampling frequency of 48000Hz and a resolution of 24 bits to accurately capture sound details. The front-end signal processing step mainly performs signal processing operations on the sound signals acquired by the microphone array. These operations include at least echo cancellation, noise reduction, dereverberation, and beamforming to eliminate echoes, noise, and interference. Wake-up event detection mainly performs real-time feature analysis on the front-end processed audio, including but not limited to frequency characteristics and energy distribution. Algorithms such as the Fast Fourier Transform (FFT) are used to convert the time-domain signal into a frequency-domain signal to clearly analyze the frequency components of the sound. The analyzed audio features are compared with a preset voice wake-up feature template. The voice wake-up feature template includes key features such as the standard frequency range, energy threshold, and unique waveform pattern of the wake-up voice set by the user.

[0038] However, the relevant technologies only involve determining whether the audio contains a wake-up keyword and whether the audio matches the audio entered by the user to determine whether the audio meets the wake-up conditions. When electronic devices use speakers to play audio containing a wake-up keyword uttered by the user, the relevant technologies may also wake up the target device and / or target application, causing the target device and / or target application to be falsely woken up, affecting the user experience.

[0039] In view of this, some embodiments of this disclosure provide a control method for a voice wake-up function. By detecting the input speech, when it is determined that the input speech meets the conditions for generating a voice wake-up event, the source of the input speech that meets the conditions for executing the voice wake-up function is detected to determine whether the speech segment is a sound emitted by the target object or a sound emitted by a non-target object, thereby realizing the judgment of speech that meets the conditions for generating a voice wake-up event and reducing the probability of false wake-up.

[0040] Figure 2 This is a flowchart illustrating a control method for a voice wake-up function according to some embodiments of the present disclosure, such as... Figure 2 As shown, it includes the following steps.

[0041] In step S11, in response to detecting input voice that satisfies the function of voice wake-up, the sound source of the input voice is detected.

[0042] In this embodiment of the disclosure, a voice wake-up event is used to execute the voice wake-up function. The conditions for generating the voice wake-up event are determined based on voice wake-up feature templates. Voice wake-up feature templates typically refer to feature templates used to identify specific wake-up words in a voice wake-up system; these feature templates are constructed by extracting feature parameters from the speech signal.

[0043] In step S12, if the sound source of the input voice is detected to be a sound not emitted by the target object, the voice wake-up function is prevented from being executed.

[0044] In this embodiment, sound source detection is performed on the input voice that satisfies the voice wake-up function to determine whether the input voice is a sound emitted by the target object or a sound emitted by a non-target object. If it is a sound emitted by a non-target object, such as a sound emitted by an electronic device, the voice wake-up function is prevented from being executed. This ensures that the input voice that can execute the voice wake-up function is emitted by the user, and avoids the voice wake-up function being mistakenly activated by a sound emitted from an electronic device.

[0045] In this embodiment of the present disclosure, after detecting that the voice segment is a sound emitted by a non-target object, the voice wake-up function corresponding to the voice wake-up event can be prevented from being executed, thereby reducing the occurrence of the voice wake-up function being falsely triggered.

[0046] Figure 3 This is a flowchart illustrating a control method for a voice wake-up function according to some embodiments of the present disclosure, such as... Figure 3 As shown, it includes the following steps.

[0047] In step S21, in response to detecting input voice that satisfies the function of voice wake-up, the sound source of the input voice is detected.

[0048] In step S22, if the input voice is detected to be the voice of the target object, the voice wake-up function is executed.

[0049] In this embodiment of the disclosure, if the detected voice segment is the voice emitted by the target object, such as the voice emitted by the user, then the instruction related to the voice wake-up event is executed to perform the voice wake-up function and realize the control of the wake-up voice in the electronic device.

[0050] In an exemplary embodiment, the audio detection process provided in this disclosure is as follows: Figure 4 As shown. Figure 4 This is a schematic diagram illustrating a control process for a voice wake-up function according to an exemplary embodiment. The audio detection process provided in this embodiment may include four steps: sound acquisition, front-end signal processing, wake-up event detection, and sound source detection and suppression. It should be understood that... Figure 4 The three steps of sound acquisition, front-end signal processing, and wake-up event detection are: Figure 1 The three steps of sound acquisition, front-end signal processing, and wake-up event detection in the related technologies are the same, and will not be repeated here. After a wake-up event is detected, the sound source detection and suppression step is performed. In the sound source detection and suppression step, the sound source of the speech segment including the wake word is detected. If it is determined that the speech segment is not a sound emitted by the target object, the wake-up operation is suppressed, and the instructions related to the wake word included in the speech segment are not executed.

[0051] It should be understood that sound source detection can detect whether all audio information in a speech segment is a sound produced by a non-target object, or it can detect whether a portion of the audio information in a speech segment is a sound produced by a non-target object. For example, if a speech segment includes both speech and music audio information, the sound source detection operation can detect whether both the speech and music audio information are sounds produced by a non-target object or by the target object, or it can detect whether the speech audio information is a sound produced by a non-target object or by the target object.

[0052] In this embodiment of the disclosure, a pre-trained sound source detection model can be used to detect the sound source of the input speech that satisfies the function of voice wake-up.

[0053] Figure 5 This is a flowchart illustrating a control method for a voice wake-up function according to some embodiments of the present disclosure, such as... Figure 5 As shown, it includes the following steps.

[0054] In step S31, in response to detecting input voice that satisfies the function of voice wake-up, the sound source of the input voice is detected.

[0055] In step S32, based on the sound source detection model and the speech segment, a detection result is generated, which indicates whether the speech segment is a sound emitted by the target object or a sound emitted by a non-target object.

[0056] In this embodiment of the disclosure, the input to the sound source detection model can be speech, and the output of the sound source detection model can be the detection result corresponding to the speech. Based on the detection result generated by the sound source detection model, it is determined whether the input audio is a sound emitted by a non-target object.

[0057] In this embodiment of the disclosure, the sound source detection model used to determine whether the audio is a sound emitted by a non-target object can be obtained by adjusting and training a pre-trained model.

[0058] Figure 6 This is a flowchart illustrating a sound source detection model training method according to some embodiments of this disclosure, such as... Figure 6 As shown, it includes the following steps.

[0059] In step S41, a first pre-trained model is obtained. The input data of the first pre-trained model is audio, and the output data of the first pre-trained model is the audio type corresponding to the audio.

[0060] In this embodiment of the disclosure, the first pre-trained model can be a model trained using training data with weak labels. Weak labeling primarily identifies the audio type corresponding to the audio, such as alarm sounds, conversations, crying, wind sounds, and thunder sounds, while not labeling the detailed information corresponding to each sound category, such as the start and end times of each sound category. It should be understood that an audio file can correspond to multiple audio types.

[0061] In step S42, the output layer of the first pre-trained model is adjusted to a single-head output layer so that the output content of the adjusted first pre-trained model represents the sound emitted by the target object or the sound emitted by a non-target object, and / or, based on the features of the sound training data, the first pre-trained model is iteratively trained until the loss value of the first pre-trained model meets the loss threshold, so as to adjust the model parameters of the first pre-trained model.

[0062] In this embodiment of the disclosure, by setting the output layer of the first pre-trained model to a single-head output layer, the output content of the adjusted first pre-trained model represents the wake-up information as either a sound emitted by the target object or a sound emitted by a non-target object. Furthermore, by repeatedly training, calculating loss, and performing direction propagation on the first pre-trained model using features from the sound training data, the model parameters of the first pre-trained model can be adjusted based on the features of the sound training data.

[0063] In this embodiment of the disclosure, by adjusting the model structure and / or model parameters of the pre-trained first pre-trained model, the training speed of the sound source detection model is improved. Furthermore, by fine-tuning the already trained model, the model can retain the content learned during the pre-training process, thereby obtaining a high-performance sound source detection model.

[0064] In this embodiment of the disclosure, the sound training data includes audio and corresponding labels, where the labels indicate whether the audio is a sound emitted by the target object or a sound emitted by a non-target object.

[0065] In step S43, a sound source detection model is obtained based on the adjusted first pre-trained model.

[0066] In this embodiment of the disclosure, after adjusting the first pre-trained model to obtain the sound source detection model, information preventing the execution of the wake-up information matching function operation can be recorded during the process of using the sound source detection model to detect audio segments including wake-up information, and the sound source detection model can be adjusted based on the recorded information. The recorded information may include, for example, the features of the audio detected by the sound source detection model, the generated detection results, and the time when the detection results were generated.

[0067] In this embodiment of the disclosure, the first pre-trained model can be obtained by training based on a pre-trained dataset or by knowledge distillation.

[0068] In one exemplary embodiment, taking a second pre-trained model as the teacher model and using a teacher-student model training mode to perform knowledge distillation on the second pre-trained model to obtain a first pre-trained model, the acquisition of the first pre-trained model is explained. By compressing the encoder and decoder layers in the second pre-trained model, selecting a portion of the encoding layers in the encoder, and a portion of the decoding layers in the decoder, the model architecture of the first pre-trained model is constructed. This results in the first pre-trained model having a lower stacking number of encoder and decoder layers than the second pre-trained model, reducing the deployment difficulty and computational complexity of the first pre-trained model. Furthermore, by reusing the model parameters in the second pre-trained model as the initial parameters of the first pre-trained model, the model parameters of the first pre-trained model are identical to those of the second pre-trained model, reducing the training cost of the first pre-trained model.

[0069] In another exemplary embodiment, the first pre-trained model can be quantized after the model parameters of the first pre-trained model are adjusted, in order to reduce the model size and computational power consumption.

[0070] Figure 7This is a flowchart illustrating a sound source detection model training method according to some embodiments of this disclosure, such as... Figure 7 As shown, it includes the following steps.

[0071] In step S51, the model parameters in the adjusted first pre-trained model are quantized so that the accuracy of the model parameters of the quantized first pre-trained model is lower than that of the model parameters of the first pre-trained model before quantization.

[0072] In this embodiment of the disclosure, the quantization process mainly converts high-precision parameter data in the model into low-precision parameter data.

[0073] In step S52, the first pre-trained model after quantization is used as the sound source detection model.

[0074] In this embodiment of the disclosure, by quantizing the first pre-trained model and using the quantized first pre-trained model as the sound source detection model, the storage requirements of the model are reduced, and the difficulty of deploying the model on electronic devices is lowered.

[0075] In an exemplary embodiment, the training method of the sound source detection model is as follows: Figure 8 As shown. Figure 8 This is a schematic diagram illustrating the training process of a sound source detection model according to an exemplary embodiment. Figure 8 In this model, the second pre-trained model can be called the teacher (large model), and the first pre-trained model can be called the student (small model). The training process for the teacher (large model) is as follows: A neural network (transformer) is trained using thousands of hours of weakly labeled training data. This training data includes thousands of hours of video data from various scenarios, all recorded by real users. Audio data (corresponding to speech) is extracted from each video and weakly labeled. Weak labeling primarily identifies the sound categories within the audio data, such as alarm sounds, human voices, crying sounds, wind sounds, and thunder sounds. Detailed start and end times for each category in the audio data are not labeled during weak labeling. This weakly labeled audio data is then used as the training data for the teacher (large model), also known as pre-training data. For the pre-training data, a frame-by-frame windowing process is used to perform Fourier transform and extract spectral features. For example, for 16kHz audio, a frame can be set to 32ms, a rectangular window can be added, and after Fourier transform, the 64-dimensional spectral energy (fbank) characteristics can be obtained through 64 frequency domain filter banks.

[0076] After training the large teacher model, the number of transformer structure layers (i.e., the number of stacked layers of the encoder and decoder) of the large teacher model is compressed, and the parameters of the first 3 layers of the large teacher model are used as the initialization parameters of the small student model. The small student model is trained through the teacher-student training mode to obtain the trained small student model.

[0077] However, the training objectives of the trained student mini-model and the sound source detection model differ. The trained student mini-model cannot be applied to the task of discriminating sounds emitted by non-target objects and requires further fine-tuning. First, the student mini-model's structure is modified by converting the final multi-head output layer into a single-head output layer, so that the student mini-model's output represents whether the sound is emitted by a non-target object. Recorded audio training data, after feature extraction, is used as the fine-tuning training set. This training set is then used to fine-tune the trained mini-model. The scenarios in the fine-tuning training set should be highly diverse, comprehensively covering various real-world scenarios, and clearly labeled to indicate whether the sound is emitted by a non-target object. Recording scenarios include, but are not limited to: playing audio containing wake words using different models of mobile phones, tablets, smart TVs, etc.; audio playback content including but not limited to short videos, long videos, terminal calls, application voice, and walkie-talkie voice; and audio played at different distances, locations, and volume levels.

[0078] After fine-tuning the student small model after training, 8-bit quantization can be performed on the model. This allows the model size to be reduced to one-quarter of its original size while ensuring that the model performance remains basically stable. At the same time, it reduces the model's computational power consumption, making the final trained model (corresponding to the sound source detection model) suitable for the deployment requirements of mobile terminals.

[0079] In this embodiment of the disclosure, the sound emitted by the target object can be the sound emitted by the user, i.e., the human voice, while the sound emitted by the non-target object can be the electronic sound emitted by the electronic device or the sound emitted by the animal.

[0080] In this embodiment, by performing sound source detection on the input voice that meets the requirements for performing the voice wake-up function, and determining whether the voice segment is an electronic sound, the problem of smart devices being mistakenly awakened by wake words emitted by other electronic devices is effectively solved, thus improving the accuracy and reliability of the voice wake-up function. Furthermore, through precise sound feature analysis and a reasonable suppression decision mechanism, unnecessary device wake-up times are reduced, power consumption is lowered, and the user experience of using the voice wake-up function is improved, avoiding operational errors and privacy leaks caused by accidental wake-ups.

[0081] Figure 9This is a block diagram illustrating a control device 100 for a voice wake-up function according to some embodiments of the present disclosure. (Refer to...) Figure 9 The device includes a processing unit 101 and an execution unit 102.

[0082] The processing unit 101 is configured to perform sound source detection on the input voice in response to the detection of input voice that satisfies the function of performing voice wake-up.

[0083] The execution unit 102 is configured to prevent the execution of the voice wake-up function if it is detected that the sound source of the input voice is a sound emitted by a non-target object; wherein, the sound source detection is used to determine whether the input voice is a sound emitted by the target object or a sound emitted by a non-target object.

[0084] In one embodiment, the execution unit 102 is further configured to: if the input voice is detected to be the voice emitted by the target object, then execute the voice wake-up function.

[0085] In one embodiment, the processing unit 101 performs sound source detection on the input speech in the following manner: Based on a sound source detection model, the input speech is detected; wherein the sound source detection model is trained as follows: a first pre-trained model is obtained, the input data of the first pre-trained model is audio, the output data of the first pre-trained model is the audio type of the audio, the first pre-trained model includes an output layer, and the output layer of the first pre-trained model is initially a multi-head output layer; the output layer of the first pre-trained model is adjusted to a single-head output layer so that the output content of the adjusted first pre-trained model represents the sound emitted by the target object or the sound emitted by a non-target object, and / or, based on the features of the sound training data, the first pre-trained model is iteratively trained until the loss value of the first pre-trained model meets the loss threshold, so as to adjust the model parameters of the first pre-trained model; based on the adjusted first pre-trained model, a sound source detection model is obtained; wherein the sound training data includes speech and the corresponding label of the speech, the label representing whether the speech is emitted by the target object or the sound emitted by a non-target object.

[0086] In one embodiment, the first pre-trained model is obtained by knowledge distillation of the second pre-trained model, the number of stacked layers of the encoder and decoder of the first pre-trained model is lower than the number of stacked layers of the encoder and decoder of the second pre-trained model, and the model parameters of the first pre-trained model are the same as the model parameters of the second pre-trained model.

[0087] In one embodiment, obtaining a sound source detection model based on an adjusted first pre-trained model includes: quantizing the model parameters in the adjusted first pre-trained model so that the accuracy of the model parameters of the quantized first pre-trained model is lower than the accuracy of the model parameters of the first pre-trained model before quantization; and using the quantized first pre-trained model as the sound source detection model.

[0088] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0089] Figure 10 This is a block diagram illustrating an electronic device 200 for controlling a voice wake-up function according to some embodiments of the present disclosure. For example, the electronic device 200 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0090] Reference Figure 10 The electronic device 200 may include one or more of the following components: processing component 202, memory 204, power component 206, multimedia component 208, audio component 210, input / output (I / O) interface 212, sensor component 214, and communication component 216.

[0091] Processing component 202 typically controls the overall operation of electronic device 200, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 202 may include one or more modules to facilitate interaction between processing component 202 and other components. For example, processing component 202 may include a multimedia module to facilitate interaction between multimedia component 208 and processing component 202.

[0092] Memory 204 is configured to store various types of data to support the operation of electronic device 200. Examples of such data include instructions for any application or method operating on electronic device 200, contact data, phonebook data, messages, pictures, videos, etc. Memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0093] Power component 206 provides power to various components of electronic device 200. Power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 200.

[0094] Multimedia component 208 includes a screen that provides an output interface between the electronic device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 208 includes a front-facing camera and / or a rear-facing camera. When the electronic device 200 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0095] Audio component 210 is configured to output and / or input audio signals. For example, audio component 210 includes a microphone (MIC) configured to receive external audio signals when electronic device 200 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 204 or transmitted via communication component 216. In some embodiments, audio component 210 also includes a speaker for outputting audio signals.

[0096] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0097] Sensor assembly 214 includes one or more sensors for providing state assessments of various aspects of electronic device 200. For example, sensor assembly 214 can detect the on / off state of electronic device 200, the relative positioning of components such as the display and keypad of electronic device 200, changes in position of electronic device 200 or a component of electronic device 200, the presence or absence of user contact with electronic device 200, orientation or acceleration / deceleration of electronic device 200, and temperature changes of electronic device 200. Sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 214 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 214 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0098] Communication component 216 is configured to facilitate wired or wireless communication between electronic device 200 and other devices. Electronic device 200 can access wireless networks based on communication standards, such as WiFi, 3G, 4G, 5G, other communication standards, or combinations thereof. In some embodiments of this disclosure, communication component 216 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of this disclosure, communication component 216 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0099] In some embodiments of this disclosure, the electronic device 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0100] In some embodiments of this disclosure, a storage medium including instructions is also provided, such as a memory 204 including instructions, which can be executed by a processor 220 of an electronic device 200 to perform the above-described method. For example, the storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0101] In some embodiments of this disclosure, a storage medium is provided, which may be a non-transitory computer-readable storage medium.

[0102] In some embodiments of this disclosure, when instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the control method for the voice wake-up function described above.

[0103] In some embodiments of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the control method for the voice wake-up function involved in any of the above embodiments.

[0104] In one exemplary embodiment, the processor executing the computer program may be deployed in an electronic device, for example.

[0105] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the described functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0106] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. Rather, these terms are used only to distinguish one component, part, region, layer, or section from another. Therefore, without departing from the teachings of the examples described herein, a first component, part, region, layer, or section mentioned in the examples may also be referred to as a second component, part, region, layer, or section. Furthermore, the terms “first” and “second” are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as “first” or “second” may explicitly or implicitly include at least one of that feature.

[0107] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.

[0108] In this description, "multiple" means at least two, referring to two or more, such as two, three, etc., unless otherwise explicitly specified. Other quantifiers are similar. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, unless otherwise specified or clearly indicated from the context, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more."

[0109] It should be understood that, unless otherwise specifically indicated, features of various embodiments of this disclosure described herein can be combined with each other. As used herein, the term "and / or" includes any one of the related listed items and any combination of two or more; "and / or" describes the association relationship between related objects, indicating that three relationships may exist, for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Similarly, "at least one of..." includes any one of the related listed items and any combination of two or more.

[0110] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.

[0111] Furthermore, the term "exemplary" is used herein to indicate that it serves as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous compared to other aspects or designs. Rather, the use of the term "exemplary" is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to indicate an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies A or B" is intended to indicate any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then applying A or B satisfies the condition under any of the foregoing instances.

[0112] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if it is not structurally equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous to any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in this disclosure, such terms are intended to be inclusive in a manner similar to the term “including.”

[0113] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0114] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for controlling a voice wake-up function, characterized in that, include: In response to detecting input voice that satisfies the function of voice wake-up, sound source detection is performed on the input voice; If the input voice is detected to be a voice not emitted by the target object, the voice wake-up function is prevented from being executed; The sound source detection is used to determine whether the input speech is a sound emitted by the target object or a sound emitted by a non-target object.

2. The method according to claim 1, characterized in that, The method further includes: If the input voice is detected to be the voice of the target object, the voice wake-up function is executed.

3. The method according to claim 2, characterized in that, The step of detecting the sound source of the input speech includes: Based on the sound source detection model, the input speech is subjected to sound source detection; The sound source detection model is trained in the following manner: Obtain a first pre-trained model, the input data of the first pre-trained model is audio, the output data of the first pre-trained model is the audio type of the audio, the first pre-trained model includes an output layer, and the output layer of the first pre-trained model is initially a multi-head output layer; The output layer of the first pre-trained model is adjusted to a single-head output layer so that the output content of the adjusted first pre-trained model represents the sound emitted by the target object or the sound emitted by a non-target object, and / or, based on the features of the sound training data, the first pre-trained model is iteratively trained until the loss value of the first pre-trained model meets the loss threshold, so as to adjust the model parameters of the first pre-trained model. The sound source detection model is obtained based on the adjusted first pre-trained model; The sound training data includes speech and corresponding labels, wherein the labels indicate whether the speech is a sound produced by a target object or a sound produced by a non-target object.

4. The method according to claim 3, characterized in that, The first pre-trained model is obtained by knowledge distillation of the second pre-trained model. The number of stacked layers of the encoder and decoder of the first pre-trained model is lower than that of the second pre-trained model. The model parameters of the first pre-trained model are the same as those of the second pre-trained model.

5. The method according to claim 4, characterized in that, The sound source detection model, obtained based on the adjusted first pre-trained model, includes: The model parameters in the adjusted first pre-trained model are quantized so that the accuracy of the model parameters in the quantized first pre-trained model is lower than that of the model parameters in the first pre-trained model before quantization. The first pre-trained model after quantization is used as the sound source detection model.

6. The method according to any one of claims 1-5, characterized in that, The sound emitted by the target object is the sound emitted by the user, and the sound emitted by the non-target object is the electronic sound emitted by the electronic device.

7. A control device for voice wake-up function, characterized in that, include: The processing unit is configured to perform sound source detection on the input voice in response to detecting input voice that satisfies the function of performing voice wake-up; An execution unit is configured to prevent the execution of the voice wake-up function if it is detected that the sound source of the input voice is a sound emitted by a non-target object. The sound source detection is used to determine whether the input speech is a sound emitted by the target object or a sound emitted by a non-target object.

8. The apparatus according to claim 7, characterized in that, The execution unit is also used for: If the input voice is detected to be the voice of the target object, the voice wake-up function is executed.

9. The apparatus according to claim 8, characterized in that, The processing unit performs sound source detection on the input speech in the following manner: Based on the sound source detection model, the input speech is subjected to sound source detection; The sound source detection model is trained in the following manner: Obtain a first pre-trained model, the input data of the first pre-trained model is audio, the output data of the first pre-trained model is the audio type of the audio, the first pre-trained model includes an output layer, and the output layer of the first pre-trained model is initially a multi-head output layer; The output layer of the first pre-trained model is adjusted to a single-head output layer so that the output content of the adjusted first pre-trained model represents the sound emitted by the target object or the sound emitted by a non-target object, and / or, based on the features of the sound training data, the first pre-trained model is iteratively trained until the loss value of the first pre-trained model meets the loss threshold, so as to adjust the model parameters of the first pre-trained model. The sound source detection model is obtained based on the adjusted first pre-trained model; The sound training data includes speech and corresponding labels, wherein the labels indicate whether the speech is a sound produced by a target object or a sound produced by a non-target object.

10. The apparatus according to claim 9, characterized in that, The first pre-trained model is obtained by knowledge distillation of the second pre-trained model. The number of stacked layers of the encoder and decoder of the first pre-trained model is lower than that of the second pre-trained model. The model parameters of the first pre-trained model are the same as those of the second pre-trained model.

11. The apparatus according to claim 10, characterized in that, The sound source detection model, obtained based on the adjusted first pre-trained model, includes: The model parameters in the adjusted first pre-trained model are quantized so that the accuracy of the model parameters in the quantized first pre-trained model is lower than that of the model parameters in the first pre-trained model before quantization. The first pre-trained model after quantization is used as the sound source detection model.

12. The apparatus according to any one of claims 7-11, characterized in that, The sound emitted by the target object is the sound emitted by the user, and the sound emitted by the non-target object is the electronic sound emitted by the electronic device.

13. An electronic device, characterized in that, include: processor; Memory used to store computer programs or instructions that can be executed by a processor; The processor is configured to execute the computer program or instructions to implement the steps of the control method for the voice wake-up function according to any one of claims 1 to 6.

14. A storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by the processor of the electronic device, enable the electronic device to perform the control method for the voice wake-up function according to any one of claims 1 to 6.

15. A computer program product, characterized in that, The system includes a computer program that, when executed by a processor, implements a control method for the voice wake-up function as described in any one of claims 1 to 6.