Voice wake-up method and apparatus, electronic device, and storage medium
By adjusting the amplitude of the voice signal and using a pre-trained voice wake-up model to identify wake-up keywords, the problem of low wake-up rate in quiet environments is solved, and efficient device wake-up is achieved under different vocal conditions.
Patent Information
- Application Number
- PCT/CN2025/103729
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-06-26
- Publication Date
- 2026-01-08
AI Technical Summary
In quiet environments, voice wake-up solutions struggle to effectively detect wake-up keywords, resulting in low wake-up rates.
By determining the voice signal acquired by the target device, the voice amplitude is identified and adjusted to the reference voice amplitude range. A pre-trained voice wake-up model is then used to identify preset wake-up keywords to wake up the device.
It improves the wake-up rate in quiet environments, enhances the adaptability and robustness of the voice wake-up function, ensures accurate wake-up of the device under various sound conditions, and improves the user experience.
Smart Images

Figure CN2025103729_08012026_PF_FP_ABST
Abstract
Description
Voice wake-up method and device, electronic device, and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202410882914.9, filed on July 2, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to a voice wake-up method, device, electronic device, and storage medium. BACKGROUND
[0003] Generally, a device needs to be woken up from a sleep state to a working state to process instructions normally, such as wake-up methods including touch wake-up (such as a lock screen key), timing wake-up (such as an alarm clock), passive wake-up (such as a phone), and the like. In order to make device wake-up more convenient, voice wake-up technology is gradually used to wake up a device by voice to switch the device from a sleep state to a working state. However, the related voice wake-up scheme has difficulty in effectively solving the wake-up problem in a soft voice environment, specifically, when a soft voice is used to say a wake-up word, it is difficult to detect the wake-up keyword, resulting in a low wake-up rate. SUMMARY
[0004] The present disclosure provides a voice wake-up method, device, electronic device, and storage medium to solve the problem of low wake-up rate caused by difficulty in detecting a wake-up keyword in a soft voice environment.
[0005] In a first aspect, embodiments of the present disclosure provide a voice wake-up method, which includes:
[0006] determining a first voice signal obtained by a target device, the first voice signal supporting carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device;
[0007] performing voice amplitude adjustment on the first voice signal to obtain a second voice signal, a voice amplitude of the second voice signal being within a reference voice amplitude interval, the reference voice amplitude interval being used to indicate a voice amplitude range that should be met by a voice signal used when supporting switching the target device or the target application from a sleep state to a working state by voice;
[0008] controlling the target device or the target application to wake up based on the second voice signal.
[0009] In a second aspect, embodiments of the present disclosure also provide a voice wake-up device, which includes:
[0010] determining a first voice signal obtained by the target device, the first voice signal carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device;
[0011] adjusting the first voice signal to obtain a second voice signal, a voice amplitude of the second voice signal being within a reference voice amplitude interval, the reference voice amplitude interval indicating a voice amplitude range that should be met by a voice signal used to switch the target device or the target application from a sleep state to a working state through voice;
[0012] controlling the target device or the target application based on the second voice signal.
[0013] In a third aspect, the present disclosure provides an electronic device, comprising:
[0014] one or more processors;
[0015] a storage device configured to store one or more programs,
[0016] when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the voice wake-up method provided by any of the embodiments of the present disclosure.
[0017] In a fourth aspect, the present disclosure provides a storage medium containing computer executable instructions for performing the voice wake-up method provided by any of the embodiments of the present disclosure when executed by a computer processor.
[0018] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0020] FIG. 1 is a flow diagram of a voice wake-up method according to an embodiment of the present disclosure;
[0021] FIG. 2 is a flow diagram of another voice wake-up method according to an embodiment of the present disclosure;
[0022] FIG. 3 is a structural schematic diagram of a voice wake-up device according to an embodiment of the present disclosure;
[0023] FIG. 4 is a structural schematic diagram of an electronic device for implementing a voice wake-up method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are merely for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0025] It should be understood that each of the steps described in the method embodiments of the present disclosure can be performed in different orders, and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0026] The term "comprising" and variations thereof as used in the present disclosure are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.
[0027] It should be noted that the "first", "second", and the like concepts mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0028] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0029] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0030] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.
[0031] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware, such as an electronic device, an application, a server or a storage medium, performing the operation of the technical solution of the present disclosure according to the prompt information.
[0032] As an optional but non-limiting implementation, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0033] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure. Other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0034] It can be understood that the data (including but not limited to the data itself, the obtaining or use of the data) involved in the technical solution should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0035] FIG. 1 is a flow diagram of a voice wake-up method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of waking up a device from a sleep state to a working state by a voice manner. The voice wake-up method can be executed by a voice wake-up apparatus, which can be implemented in the form of software and / or hardware and is generally integrated on any electronic device with network communication function, such as a mobile terminal, a PC terminal or a server.
[0036] As shown in FIG. 1, the voice wake-up method of the embodiment of the present disclosure can include the following processes:
[0037] S110, determining a first voice signal obtained by a target device, the first voice signal supporting carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device.
[0038] The target device can refer to a device that is switched from a sleep state to a working state through a specific voice instruction including a preset wake-up keyword. The target device needs to receive a specific voice instruction including a preset wake-up keyword to start or activate its main function when it is in a standby, sleep or inactivated state. For example, a smart speaker is in a low-power standby state when it is not woken up. Only when the user speaks the preset wake-up word, the smart speaker will start to respond and execute the user's subsequent voice instructions, such as playing music, querying the weather, etc. For another example, some smart home appliances, such as smart TVs or smart air conditioners, can also be set to a voice wake-up mode. Only after receiving the correct voice wake-up instruction including the preset wake-up keyword, the smart home appliances will be switched from a sleep state to a working state and be ready to receive further operation instructions.
[0039] For another example, the target device can be a terminal device such as a mobile phone, which can be installed with a target application program (such as a voice assistant) that can be started and controlled through voice instructions. The target application program can be woken up by a wake-up word received by the terminal device. For another example, the target device can be a headset connected to a smart device such as a mobile phone, which can be installed with a target application program (such as a voice assistant) that can be started through voice instructions. The user can input a wake-up word to the voice assistant through the headset to wake up the voice assistant, input various voice instructions to the voice assistant through the headset, and obtain voice responses provided by the voice assistant through the headset.
[0040] The first voice signal can include sound information obtained within a preset distance range around the location of the target device. The sound information corresponding to the first voice signal contains various voice contents, which can be a speech emitted by a sound source close to the target device, a mixture of noise and voice in the environment of the location where the target device is located, or a preset keyword related to waking up the target device or the target application. Therefore, the first voice signal supports carrying a preset wake-up keyword for waking up the target device or the target application. The preset wake-up keyword can be a specific voice instruction or phrase, which is used to trigger the device to switch from a sleep or standby state to a working state and be ready to receive and process subsequent voice commands or perform related operations. When the target device or the target application receives the preset wake-up keyword through voice, it can be switched from a sleep state to a working state.
[0041] Optionally, the first voice signal can include sound information acquired within a preset distance range of the surrounding of the position where the target device is located in a reference low voice environment. The reference low voice environment can refer to a sound environment with relatively low sound intensity, small volume and low noise level. The sound intensity of light voice is relatively weak, usually lower than the volume of standard pronunciation. Sound level meter and other devices can be used to measure the intensity of the sound to determine the sound intensity range of light voice. In such a low voice environment, the speaking voice is usually soft and the background noise is relatively weak, and the overall acoustic environment is quiet, for example, a library, a quiet conference room, a private study room, etc. can be regarded as a low voice environment, which will make it more difficult to detect keywords in the voice signal in a low voice environment, thereby affecting the voice wake-up rate.
[0042] The above scheme determines the first voice signal acquired by the target device and supports the carrying of the preset wake-up keyword, thereby providing a basis for voice wake-up, and then determining whether to wake up the device based on the detailed analysis and processing of the first voice signal.
[0043] As an optional but non-limiting implementation, determining the first voice signal acquired by the target device includes the following steps A1-A2:
[0044] Step A1, using a microphone configured on the target device to acquire voice signals in a preset distance range of the target device in real time.
[0045] Step A2, obtaining the first voice signal by preprocessing the real-time acquired voice signal, the preprocessing including at least one of noise removal and filtering processing.
[0046] The microphone configured on the target device can be a device capable of converting the sound in the surrounding environment of the position where the target device is located into an electrical signal. The microphone configured on the target device has a specific sensitivity and receiving range, and the sound information within the preset distance range of the surrounding of the position where the target device is located can be acquired by using the microphone installed on the target device. Furthermore, at least one preprocessing such as noise removal and filtering processing can be performed on the real-time acquired voice signal to obtain the first voice signal.
[0047] The preset distance range can be a pre-set distance area. For example, if the preset distance range is 3 meters, the microphone will collect the sound within a radius of 3 meters centered on the target device. The microphone configured on the target device will continuously and immediately convert the sound information generated within the preset distance range of the position where the target device is located into an electrical signal and transmit it to the target device or the target application for subsequent analysis, processing or execution of corresponding operations, such as waking up the device, to realize timely acquisition of voice information within a specific distance around the target device and provide data support for various voice-related functions of the device.
[0048] S120, the first voice signal is adjusted in voice amplitude to obtain a second voice signal, and the voice amplitude of the second voice signal is within a reference voice amplitude interval, the reference voice amplitude interval being used to indicate a voice amplitude range that should be met by a voice signal used to support switching of the target device or the target application from a dormant state to an active state by voice.
[0049] In actual applications, voice signals often have large fluctuations, which can be too strong or too weak, affecting the quality and intelligibility of voice signals, especially in a low voice environment, the sound of the first voice signal corresponding to the target device is relatively low, and it is difficult to detect the preset wake-up keyword from the first voice signal, resulting in voice wake-up failure. Therefore, after receiving the first voice signal collected by the target device, the first voice signal can be adjusted in voice amplitude, so that the voice amplitude of the adjusted second voice signal is within the reference voice amplitude interval.
[0050] Voice amplitude can refer to the strength or energy size of a voice signal, and voice amplitude reflects the loudness or volume of the corresponding sound. When the voice amplitude of the first voice signal is weak, the voice amplitude of the first voice signal is automatically increased, the voice signal is amplified, and the adjusted second voice signal is more easily perceived and processed; when the voice amplitude of the first voice signal is strong, the voice amplitude of the first voice signal is reduced to avoid voice signal overload and distortion. Through such dynamic adjustment, it is helpful to keep the voice amplitude of the second voice signal used for final voice wake-up within a relatively appropriate reference voice amplitude interval, improving voice recognition performance and stability.
[0051] The reference voice amplitude interval is an index for indicating a voice signal amplitude range, and is a voice amplitude range required to be met by a voice signal when supporting switching of the target device or the target application from a dormant state to an active state by voice. Specifically, when the target device or the target application is in a dormant state, it needs to be woken up and switched to an active state by receiving a specific voice signal, and the reference voice amplitude interval specifies the range that the voice signal should meet in amplitude.
[0052] In actual applications, the microphone of the target device will collect the surrounding voice signal and convert it into an electrical signal for processing. If the amplitude of the collected voice signal is within the reference voice amplitude interval and carries the preset wake-up keyword, the target device or the target application is likely to be woken up and enter an active state. If the amplitude of the collected voice signal is not within the reference voice amplitude interval, even if it carries the preset wake-up keyword, the target device or the target application is likely to be unable to be woken up.
[0053] It can be understood that different target devices or target applications can have different reference voice amplitude intervals and preset wake-up keywords, and specific settings will be determined according to the design and application scene of the target device or target application. In addition, factors such as environmental noise, voice volume and tone may also affect the effect of voice wake-up.
[0054] By adjusting the first voice signal to the second voice signal with an amplitude in the reference interval, even if the amplitude of the original voice signal is not suitable, the requirement of switching the device from the sleep state to the working state can be met, thereby improving the probability of successfully waking up the target device or target application. Whether the user speaks loudly or softly, the voice signal can be adjusted to an appropriate amplitude range, making the voice wake-up function of the device more adaptable to various speaking situations, solving the problem that the device is difficult to wake up due to the voice amplitude not meeting the requirements, making the user more smooth and convenient when using the voice wake-up function, without deliberately increasing or decreasing the volume, so that the target device or target application can be accurately woken up in more complex sound environments.
[0055] As an optional but non-limiting implementation, the voice amplitude adjustment of the first voice signal to obtain the second voice signal includes the following steps B1-B2:
[0056] Step B1, continuously detecting the voice amplitude of the first voice signal, and determining the gain value corresponding to the first voice signal according to the voice amplitude of the first voice signal.
[0057] Step B2, amplifying or attenuating the voice amplitude of the first voice signal according to the gain value corresponding to the first voice signal to obtain the second voice signal.
[0058] The first voice signal within a preset distance range of the position where the target device is located is obtained through a microphone or other audio input device configured on the target device, the input first voice signal is monitored in real time for amplitude, and the change of its strength is analyzed. According to the detected voice amplitude of the first voice signal and the preset gain control strategy, the appropriate gain value corresponding to the first voice signal is calculated. For example, the gain value corresponding to the first voice signal can be determined according to the average amplitude, peak amplitude, dynamic range and other factors of the first voice signal.
[0059] Further, the calculated appropriate gain value corresponding to the first voice signal can be applied to the first voice signal to amplify or attenuate the voice amplitude of the first voice signal, and the first voice signal after the voice amplitude gain adjustment is processed subsequently to obtain a second voice signal, such as filtering, encoding, etc., to meet specific application requirements. In the entire processing process, the change of the first voice signal is continuously monitored, and the gain value is dynamically recalculated and adjusted according to the amplitude of the first voice signal, so as to ensure that the amplitude of the output second voice signal always remains in the expected range.
[0060] By using the above scheme, through such a processing process, the automatic gain of the voice signal can effectively improve the quality and intelligibility of the voice signal, and adapt to different input conditions and application scenarios.
[0061] In S130, the target device or target application is controlled to wake up based on the second voice signal.
[0062] By performing voice amplitude gain adjustment on the first voice signal, the voice amplitude of the output second voice signal can be adjusted to a suitable range, which can reduce the loss of key information caused by unsuitable voice amplitude. The acoustic features corresponding to the second voice signal not only include acoustic features such as prosody and timbre, but also include acoustic features for describing a preset wake-up keyword. By identifying the acoustic features corresponding to the second voice signal to determine whether the preset wake-up keyword exists in the second voice signal, the target device or target application can be controlled to wake up.
[0063] By adaptively adjusting the amplitudes of voice signals with different voice amplitudes, the voice habits of different users and the changes of the sound environment can be adapted, so that the wake-up keyword can be more accurately recognized, and the accuracy of wake-up can be improved. Whether speaking softly or in a noisy environment, the voice signal after amplitude adjustment can be better processed and recognized, enhancing the robustness of the system. Users do not need to deliberately control the volume of the voice, and the device can be successfully woken up in the case of natural voice, making the voice wake-up function more convenient and humanized, and improving the user experience.
[0064] Furthermore, since the second voice signal is a voice signal after voice amplitude adjustment, even if the first voice signal is a voice signal collected in a low voice environment, it can still be ensured to be in a suitable voice amplitude interval after amplitude adjustment, which can effectively highlight the key voice features and help to accurately control the target device or target application for voice wake-up in a complex low voice environment. Especially for users with weak voice or voice disorders, voice amplitude adjustment can increase the possibility of successfully waking up the device by voice.
[0065] The technical scheme of the embodiment of the present disclosure provides a basis for realizing voice wake-up by determining the first voice signal obtained by the target device and supporting the preset wake-up keyword carried in the first voice signal. Then, the first voice signal is adjusted in voice amplitude to obtain the second voice signal located in the reference voice amplitude interval. This adjustment enables the voice amplitude to meet the voice amplitude range required for switching the target device or the target application from the sleep state to the working state, effectively solving the wake-up problem caused by inappropriate voice amplitude. The technical scheme can improve the detection capability of the soft wake-up keyword by adjusting the voice amplitude when the wake-up keyword is spoken softly, thereby significantly improving the wake-up rate under soft voice and enabling the voice wake-up function to reliably and effectively control the device in a wider voice amplitude range.
[0066] FIG. 2 is a flowchart of another voice wake-up method provided by the embodiment of the present disclosure. The technical scheme of the present embodiment further optimizes the process of waking up the target device or the target application based on the second voice signal in the foregoing embodiment based on the technical scheme of the foregoing embodiment. The present embodiment can be combined with each optional scheme in one or more of the foregoing embodiments.
[0067] As shown in FIG. 2, the voice wake-up method of the embodiment of the present disclosure can include the following processes:
[0068] S210, determining a first voice signal obtained by a target device, the first voice signal supporting carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device.
[0069] S220, adjusting the first voice signal in voice amplitude to obtain a second voice signal, the voice amplitude of the second voice signal being located in a reference voice amplitude interval, the reference voice amplitude interval being used to indicate a voice amplitude range that should be met by a voice signal used for supporting switching the target device or the target application from a sleep state to a working state by voice.
[0070] S230, inputting the second voice signal to a pre-trained voice wake-up model; wherein the pre-trained voice wake-up model can support identifying acoustic features corresponding to the voice signal and detecting the possibility of the presence of the preset wake-up keyword in the voice signal based on the acoustic features corresponding to the voice signal, and the pre-trained voice wake-up model only supports inputting the voice signal located in the reference voice amplitude interval.
[0071] Optionally, the pre-trained voice wake-up model can identify the acoustic features of the voice signal corresponding to the reference low voice environment in which the voice distortion interference exists in the voice signal corresponding to the voice amplitude of the voice emitted is less than the preset voice amplitude, and detect the possibility of the presence of the preset wake-up keyword in the voice signal based on the acoustic features of the voice signal corresponding to the voice signal. For example, the reference low voice environment can refer to a voice environment in which the voice intensity is relatively low, the volume is small, and the noise level is not high.
[0072] The second voice signal is input into the pre-trained voice wake-up model, which analyzes and processes the voice signal and outputs a probability value of voice wake-up. The voice wake-up model is usually trained based on deep learning, such as recurrent neural network (RNN) or long short-term memory network (LSTM), etc. These models can learn the features in the voice signal and determine whether there is a wake-up word or voice instruction. For example, if the probability value output by the wake-up model exceeds a preset threshold, it is considered that the voice wake-up is successful, and the target device or target application will switch from the sleep state to the working state.
[0073] The voice wake-up model can extract acoustic features from the input second voice signal, which can represent the spectral characteristics of the voice signal. Further, the extracted features are analyzed using the voice wake-up model, which can determine whether there is a wake-up word and the possibility of the presence of the preset wake-up keyword based on the learned mapping relationship between the voice signal and the possibility of the presence of the wake-up keyword in the voice signal. If it is determined that there is a wake-up word, the corresponding wake-up operation is triggered, such as waking up the device or executing a specific instruction.
[0074] With the above scheme, the specific processed voice signal can be input into the specially trained voice wake-up model, which can identify the acoustic features of the voice and detect the possibility of the presence of the preset wake-up keyword, but has specific requirements for the amplitude of the input voice signal. When identifying, it is necessary to ensure that the voice amplitude of the voice signal input into the voice wake-up model is within the reference voice amplitude interval.
[0075] As an optional but non-limiting implementation, the determination process of the pre-trained voice wake-up model provided in the embodiments of the present disclosure includes the following steps C1-C2:
[0076] Step C1: Determine the reference training data to be used when training the voice wake-up model. The reference training data is generated by superimposing the speech distortion characteristics of the third speech signal to simulate the first speech amplitude. The third speech signal includes a speech signal generated by attenuating the speech amplitude of the fourth speech signal with the second speech amplitude. The first speech amplitude is less than the preset speech amplitude, and the second speech amplitude is greater than the preset speech amplitude. The preset speech amplitude is the speech amplitude that the speech signal needs to be distinguished and understood.
[0077] Step C2: Based on the reference training data, control the voice wake-up model to be trained to perform a voice wake-up training task, so as to adjust the parameters of the voice wake-up model to be trained and obtain a pre-trained voice wake-up model.
[0078] During the training of the voice wake-up model, considering the need for the model to accurately recognize and detect speech signals in low-noise environments, reference training data was generated based on simulated speech distortions that may exist in low-noise environments. For example, speech distortion can refer to the distortion or deformation that occurs during the transmission, processing, or storage of speech signals. Possible speech distortions mean that under specific conditions, speech signals may exhibit various types of distortion phenomena.
[0079] Optionally, when speaking in a neutral tone environment, the speech signal may undergo at least one of the following speech distortions: pitch change, duration change, timbre change, and tone change. For example, pitch change manifests as the pitch of a neutral tone being lower than that of a normal pronunciation, and there may be instances of pitch drop or instability; duration change manifests as the duration of a neutral tone being shorter than that of a normal pronunciation, and there may be instances of shortened or incomplete syllables; timbre change manifests as the timbre of a neutral tone becoming blurred or unclear, consonants may be weakened or omitted, and vowels may become less full; tone change manifests as in some languages the neutral tone potentially causing changes or disappearance of tone.
[0080] To this end, in the training process of the voice wake-up model, a large number of fourth voice signals of the second voice amplitude can be obtained, which can be understood as voices with normal sound size and voice distortion that may exist in the low-voice environment. To this end, the fourth voice signal of the second voice amplitude can be first attenuated in voice amplitude to generate a third voice signal, and then the voice distortion is superimposed according to the voice feature of the first voice amplitude simulated by the third voice signal to generate reference training data. Then, the voice wake-up model to be trained can be controlled to perform a voice wake-up training task with reference to the reference training data, so as to ensure that the voice wake-up model can fully consider the influence of voice distortion in the low-voice environment on the voice signal when performing voice recognition, and avoid reducing the voice detection accuracy due to voice distortion in the low-voice environment.
[0081] It can be understood that voice distortion is a characteristic of soft voice, but the specific performance will vary depending on language, dialect and individual pronunciation habits. In voice processing and voice recognition, the characteristics of soft voice need to be considered to improve the understanding and recognition accuracy of soft voice. At the same time, for occasions that require accurate expression of semantics, such as speeches, broadcasts, etc., soft voice should be avoided as much as possible to ensure the clarity and intelligibility of the voice.
[0082] By using the above scheme, the model is allowed to contact data containing various voice distortion conditions, so that the model can learn the voice feature rules under different distortion modes, thereby accurately recognizing and processing new and unseen voice distortions, enhancing the adaptability and generalization ability of the model to different scenarios; and the model can stably perform the voice wake-up task when facing voice distortions caused by environmental, device or voice emission methods, etc., without being easily disturbed or making mistakes, thereby improving the reliability and stability of the model, reducing misjudgment and omission caused by distortion, and significantly improving the accuracy of voice wake-up, ensuring that the device can accurately respond to the user's wake-up instruction.
[0083] As an optional but non-limiting implementation manner, the reference training data to be used by the voice wake-up model to be trained in the training process includes the following steps D1-D3:
[0084] Step D1, simulating the voice distortion of the voice feature of the first voice amplitude on the third voice signal to obtain a reference voice signal, the third voice signal being a voice signal with the first voice amplitude obtained by attenuating the voice amplitude of the fourth voice signal of the second voice amplitude.
[0085] Step D2, adjusting the voice amplitude of the reference voice signal to obtain a target voice signal, the voice amplitude of the target voice signal being within the reference voice amplitude interval.
[0086] Step D3, generating the speech wake-up model to be trained according to the target speech signal, wherein the reference training data to be used when training the speech wake-up model.
[0087] In an optional example, the speech distortion superposition on the third speech signal simulating the voice production feature of the first speech amplitude comprises: using an audio special effect processor to simulate the voice distortion superposition on the voice production feature of the first speech amplitude. For example, the audio special effect processor can perform various processing on the audio signal, such as distortion, compression, filtering, etc. By adjusting the parameters of the special effect processor, the voice distortion effect that may occur when speaking softly can be simulated.
[0088] In an optional example, the speech distortion superposition on the third speech signal simulating the voice production feature of the first speech amplitude comprises: adjusting the audio parameters to simulate the voice distortion superposition on the voice production feature of the first speech amplitude. For example, adjusting the audio parameters means using audio editing software, which can adjust the volume, pitch, timbre, etc. of the audio. Reducing the volume can simulate soft voice production, and appropriately adjusting other parameters can also produce some distortion effects.
[0089] In an optional example, the speech distortion superposition on the third speech signal simulating the voice production feature of the first speech amplitude comprises: applying an audio filter to simulate the voice distortion superposition on the voice production feature of the first speech amplitude. For example, applying an audio filter can change the frequency response of the audio, thereby producing different sound effects. Some audio filters that simulate distortion, saturation or other distortion effects are tried to simulate the voice distortion under soft voice production.
[0090] In an optional example, the speech distortion superposition on the third speech signal simulating the voice production feature of the first speech amplitude comprises: superimposing noise or environmental sound effects to simulate the voice distortion superposition on the voice production feature of the first speech amplitude. For example, superimposing noise or environmental sound effects means superimposing appropriate noise or environmental sound effects, such as light wind or background noise, on the audio of soft voice production, which can increase the sense of reality and distortion.
[0091] In an optional example, the speech distortion superposition on the third speech signal simulating the voice production feature of the first speech amplitude comprises: using audio synthesis technology to simulate the voice distortion superposition on the voice production feature of the first speech amplitude. For example, using audio synthesis technology, a sound with specific distortion characteristics is created through an audio synthesis tool or software, and it is mixed with the original audio to simulate the voice distortion under soft voice production.
[0092] In an optional example, the voice distortion superposition of the voice signal simulating the vocalization feature of the first voice amplitude comprises: the voice distortion superposition of the voice signal simulating the vocalization feature of the first voice amplitude is simulated with reference to the actual recording. For example, the actual recording is simulated by recording some light vocalization samples and analyzing their spectrum and characteristics, and the above method is used to simulate similar distortion effects to make them closer to the real situation.
[0093] By using the above scheme, by simulating different vocalization features and superimposing voice distortion, the model can learn more diverse voice patterns, thereby better adapting to various real vocalization situations, and adjusting the amplitude of the voice signal to the reference interval, so that the training data has consistency in amplitude, which is beneficial to model learning and optimization. And using the processed and generated reference training data for training can make the model more accurately identify and wake up, while enhancing its stability in different environments.
[0094] As an optional but non-limiting implementation, through dynamic data augmentation, the existing voice signal can be adjusted to generate or adjust the training data for voice wake-up model training in the data enhancement process. Dynamic data augmentation is more flexible and adaptive, and can be adjusted in real time according to the characteristics of the data and the needs of the model. During model training, by adding a dynamic data augmentation of voice amplitude, the proportion of small volume light voice data in the model training data is increased to improve the light voice wake-up rate of the model.
[0095] Specifically, for conventional voice wake-up data, the operation of multiplying a relatively low signal gain is used to effectively simulate the audio samples when speaking softly. This method can more realistically restore the changes in voice amplitude when people communicate softly. In order to more realistically simulate the voice distortion that may occur when speaking softly, the above generated soft voice samples will be further processed by an audio migrator. This audio migrator will make fine adjustments and optimizations to the samples based on acoustic principles and actual voice characteristics to make them as close as possible to the voice data characteristics of normal soft voice in real scenarios, thereby increasing the diversity and authenticity of the data. Using the above reference training data, a high-performance voice wake-up model is trained. During training, the parameters of the model are continuously adjusted to accurately capture the characteristics and rules of soft voice, thereby improving the recognition ability of soft voice wake-up.
[0096] S240, according to the possibility that the second voice signal output by the pre-trained voice wake-up model exists a preset wake-up keyword, the target device or target application is controlled to wake up.
[0097] As an optional but non-limiting implementation, the possibility of the second voice signal output by the pre-trained voice wake-up model containing the preset wake-up keyword is determined, and the wake-up control of the target device or the target application comprises the following steps E1-E2:
[0098] Step E1, if the possibility of the second voice signal containing the preset wake-up keyword is greater than the preset probability threshold, the target device or the target application is switched from the sleep state to the working state.
[0099] Step E2, if the possibility of the second voice signal containing the preset wake-up keyword is not greater than the preset probability threshold, the target device or the target application is not switched from the sleep state to the working state.
[0100] By setting the preset probability threshold to determine whether to start the state switching of the device, the truly effective wake-up instruction can be accurately recognized, the false wake-up is avoided, and the accuracy of the wake-up is improved. When the possibility is greater than the threshold, the switching is started, so that only the second voice signal highly possibly containing the preset wake-up keyword can wake up the device, and unnecessary state switching caused by misjudgment is reduced. When the possibility is not greater than the threshold, the switching is not started, and the target device or the target application continues to wait for the effective wake-up instruction input, so that the device is prevented from being incorrectly woken up from the sleep state due to uncertain or incorrect judgment, and device resources and energy are saved. According to the possibility of the preset wake-up keyword in the voice signal, whether to switch the state of the device is determined, so that the resources and energy consumption of the device can be reasonably allocated. Through the above scheme, the effective wake-up instruction of the user can be responded in time, and the disturbance caused by false wake-up is avoided, so that the user feels convenient and comfortable when using the voice wake-up function.
[0101] The technical scheme of the embodiment of the present disclosure provides a basis for realizing voice wake-up by determining the first voice signal obtained by the target device and supporting the preset wake-up keyword carried therein. Then, the first voice signal is adjusted in voice amplitude to obtain the second voice signal located in the reference voice amplitude interval. This adjustment makes the amplitude of the voice signal meet the required voice amplitude range for switching the target device or the target application from the sleep state to the working state, effectively solves the wake-up problem caused by inappropriate voice amplitude, and improves the detection capability of the soft or light wake-up word through the adjustment of the voice amplitude, thereby significantly improving the wake-up rate under soft or light voice, so that the voice wake-up function can reliably and effectively control the device in a wider voice amplitude range.
[0102] FIG. 3 is a structural schematic diagram of a voice wake-up device provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to a case where a device is woken up by voice to switch from a sleep state to a working state. The voice wake-up device can be implemented in the form of software and / or hardware, and is generally integrated in any electronic device with network communication function, such as a mobile terminal, a PC terminal, or a server.
[0103] As shown in FIG. 3, the voice wake-up device of the embodiment of the present disclosure can include the following:
[0104] The determining module 310 is configured to determine a first voice signal acquired by a target device, the first voice signal supporting carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device;
[0105] The adjusting module 320 is configured to perform voice amplitude adjustment on the first voice signal to obtain a second voice signal, the voice amplitude of the second voice signal being within a reference voice amplitude interval, the reference voice amplitude interval being used to indicate a voice amplitude range that should be met by a voice signal used when supporting switching the target device or the target application from a sleep state to a working state by voice;
[0106] The control module 330 is configured to perform wake-up control on the target device or the target application based on the second voice signal.
[0107] On the basis of the above embodiment, optionally, the first voice signal acquired by the target device includes:
[0108] Real-time acquisition of voice signals within a preset distance range of the target device is performed by using a sound pickup device configured on the target device;
[0109] The first voice signal is obtained after pre-processing of the real-time acquired voice signals, the pre-processing including at least one of noise removal and filtering processing.
[0110] On the basis of the above embodiment, optionally, the voice amplitude adjustment on the first voice signal to obtain the second voice signal includes:
[0111] The voice amplitude of the first voice signal is continuously detected, and a gain value corresponding to the first voice signal is determined according to the voice amplitude of the first voice signal;
[0112] The first voice signal is amplified or attenuated in voice amplitude according to the gain value corresponding to the first voice signal to obtain the second voice signal.
[0113] On the basis of the above embodiment, optionally, the wake-up control on the target device or the target application based on the second voice signal includes:
[0114] input the second voice signal to a pre-trained voice wake-up model; wherein the pre-trained voice wake-up model can support identification of acoustic features corresponding to the voice signal, and can support detection of a possibility that a preset wake-up keyword exists in the voice signal based on the acoustic features corresponding to the voice signal, and the pre-trained voice wake-up model only supports input of a voice signal located within a reference voice amplitude interval;
[0115] wake up control is performed on the target device or the target application according to the possibility that the preset wake-up keyword exists in the second voice signal output by the pre-trained voice wake-up model.
[0116] On the basis of the above-mentioned embodiments, optionally, the determination process of the pre-trained voice wake-up model comprises:
[0117] determining reference training data to be used by the voice wake-up model to be trained when training, the reference training data being generated after voice distortion superposition is performed on a voice feature simulating a first voice amplitude of a third voice signal, the third voice signal comprising a voice signal generated by performing voice amplitude attenuation on a fourth voice signal of a second voice amplitude, the first voice amplitude being less than a preset voice amplitude, the second voice amplitude being greater than the preset voice amplitude, and the preset voice amplitude being a voice amplitude required to be satisfied for voice signals to be distinguished and understood;
[0118] controlling the voice wake-up model to be trained to perform a voice wake-up training task based on the reference training data, so as to adjust parameters of the voice wake-up model to be trained to obtain the pre-trained voice wake-up model.
[0119] On the basis of the above-mentioned embodiments, optionally, the reference training data to be used by the voice wake-up model to be trained when training comprises:
[0120] obtaining a reference voice signal after voice distortion superposition is performed on a voice feature simulating a first voice amplitude of a third voice signal, the third voice signal being a voice signal with the first voice amplitude obtained by performing voice amplitude attenuation on a fourth voice signal of a second voice amplitude;
[0121] performing voice amplitude adjustment on the reference voice signal to obtain a target voice signal, the voice amplitude of the target voice signal being within a reference voice amplitude interval;
[0122] generating the reference training data to be used by the voice wake-up model to be trained when training according to the target voice signal.
[0123] On the basis of the above-mentioned embodiments, optionally, the wake-up control performed on the target device or the target application according to the possibility that the preset wake-up keyword exists in the second voice signal output by the pre-trained voice wake-up model comprises:
[0124] If the possibility of the preset wake-up keyword existing in the second voice signal is greater than the preset probability threshold, the target device or the target application is switched from the sleep state to the working state;
[0125] If the possibility of the preset wake-up keyword existing in the second voice signal is not greater than the preset probability threshold, the target device or the target application is not switched from the sleep state to the working state.
[0126] The technical scheme of the embodiment of the present disclosure provides a basis for realizing voice wake-up by determining the first voice signal obtained by the target device and supporting the preset wake-up keyword carried therein. Then, the first voice signal is adjusted in voice amplitude to obtain the second voice signal located in the reference voice amplitude interval. This adjustment enables the voice amplitude to meet the voice amplitude range required for switching the target device or the target application from the sleep state to the working state, effectively solves the wake-up problem caused by inappropriate voice amplitude, and improves the detection capability of the soft wake-up keyword through the adjustment of the voice amplitude when the wake-up keyword is spoken softly, thereby significantly improving the wake-up rate under soft voice and enabling the voice wake-up function to reliably and effectively work in a wider voice amplitude range.
[0127] The voice wake-up apparatus provided by the embodiment of the present disclosure can execute the voice wake-up method provided by any embodiment of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.
[0128] It should be noted that each unit and module included in the above apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for convenient mutual distinction, and does not limit the protection scope of the embodiments of the present disclosure.
[0129] FIG. 4 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Referring to FIG. 4, a structural schematic diagram of an electronic device (for example, a terminal device or a server in FIG. 4) 400 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include but is not limited to a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle-mounted terminal (for example, a vehicle-mounted navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 4 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0130] As shown in FIG. 4, the electronic device 400 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 402 or loaded into a random access memory (RAM) 403 from a storage device 408. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0131] Generally, the following devices can be connected to the I / O interface 405: input devices 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 409. The communication devices 409 can allow the electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG. 4 shows the electronic device 400 with various devices, it should be understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0132] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 409, or installed from the storage devices 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0133] The names of messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0134] The electronic device provided by the embodiments of the present disclosure and the voice wake-up method provided by the above-described embodiments belong to the same inventive concept, and technical details not described in detail in the present embodiments can be referred to the above-described embodiments, and the present embodiments have the same beneficial effects as the above-described embodiments.
[0135] The embodiments of the present disclosure provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the voice wake-up method provided by the above-described embodiments.
[0136] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer readable program code is contained. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF, etc., or any suitable combination thereof.
[0137] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0138] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and can be accessed via the electronic device.
[0139] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: determine a first voice signal acquired by a target device, the first voice signal supporting carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device; perform voice amplitude adjustment on the first voice signal to obtain a second voice signal, a voice amplitude of the second voice signal being within a reference voice amplitude interval, the reference voice amplitude interval being used to indicate a voice amplitude range that should be met by a voice signal used to support switching the target device or the target application from a dormant state to an active state through voice; and perform wake-up control on the target device or the target application based on the second voice signal.
[0140] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages or combinations of languages including object or visual programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0141] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified functions. It should also be noted that, in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0142] The units described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the names of the units do not constitute a limitation on the units themselves. For example, the first obtaining unit can also be described as a unit that obtains at least two Internet protocol addresses.
[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include: Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0145] The above description is merely the preferred embodiments of the present disclosure and the explanation of the principles of the applied technology. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above technical features can be replaced with the technical features disclosed in the present disclosure (but not limited to) that have similar functions to form technical solutions.
[0146] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0147] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A voice wake-up method, comprising: determining a first voice signal obtained by a target device, the first voice signal supporting carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device; performing voice amplitude adjustment on the first voice signal to obtain a second voice signal, the voice amplitude of the second voice signal being within a reference voice amplitude interval, the reference voice amplitude interval being used to indicate a voice amplitude range that a voice signal should satisfy when supporting switching the target device or the target application from a sleep state to a working state by voice; controlling the target device or the target application to wake up based on the second voice signal.
2. The method of claim 1, wherein, The determination of the first voice signal obtained by the target device comprises: obtaining a voice signal in a preset distance range of the target device in real time by using a microphone configured on the target device; obtaining the first voice signal by preprocessing the voice signal obtained in real time, the preprocessing comprising at least one of noise removal and filtering processing.
3. The method of claim 1 or 2, wherein, The voice amplitude adjustment on the first voice signal to obtain the second voice signal comprises: continuously detecting the voice amplitude of the first voice signal, and determining a gain value corresponding to the first voice signal according to the voice amplitude of the first voice signal; amplifying or attenuating the voice amplitude of the first voice signal according to the gain value corresponding to the first voice signal to obtain the second voice signal.
4. The method according to any one of claims 1 to 3, wherein, The control of the target device or the target application to wake up based on the second voice signal comprises: inputting the second voice signal into a pre-trained voice wake-up model; wherein the pre-trained voice wake-up model can support identifying acoustic features corresponding to a voice signal, and detecting a possibility that a preset wake-up keyword exists in the voice signal based on the acoustic features corresponding to the voice signal, the pre-trained voice wake-up model only supporting inputting a voice signal within the reference voice amplitude interval; controlling the target device or the target application to wake up according to the possibility that the preset wake-up keyword exists in the second voice signal output by the pre-trained voice wake-up model.
5. The method of claim 4, wherein, The determination process of the pre-trained voice wake-up model comprises: determining reference training data to be used by a voice wake-up model to be trained when training, the reference training data being generated by performing voice distortion superposition on voice features simulated according to a third voice signal at a first voice amplitude, the third voice signal comprising a voice signal generated by performing voice amplitude attenuation on a fourth voice signal at a second voice amplitude, the first voice amplitude being less than a preset voice amplitude, the second voice amplitude being greater than the preset voice amplitude, the preset voice amplitude being a voice amplitude required to be satisfied for voice signals to be distinguished and understood; controlling the voice wake-up model to be trained to perform a voice wake-up training task based on the reference training data, so as to adjust parameters of the voice wake-up model to be trained to obtain the pre-trained voice wake-up model.
6. The method of claim 5, wherein, The determination of the reference training data to be used by the to-be-trained voice wake-up model during training comprises: The third voice signal is a voice signal with the first voice amplitude obtained by performing voice amplitude attenuation on the fourth voice signal with the second voice amplitude; The reference voice signal is obtained by performing voice amplitude adjustment on the third voice signal; The target voice signal is generated according to the target voice signal.
7. The method of claim 4, wherein, The possibility that the second voice signal output by the pre-trained voice wake-up model contains a preset wake-up keyword is used to control the wake-up of the target device or the target application, comprising: If the possibility that the second voice signal contains a preset wake-up keyword is greater than a preset probability threshold, the target device or the target application is switched from a sleep state to a working state; If the possibility that the second voice signal contains a preset wake-up keyword is not greater than the preset probability threshold, the target device or the target application is not switched from a sleep state to a working state.
8. A voice wake-up device, comprising: A determination module configured to determine a first voice signal obtained by a target device, wherein the first voice signal supports carrying a preset wake-up keyword used to wake up the target device or a target application associated with the target device; An adjustment module configured to perform voice amplitude adjustment on the first voice signal to obtain a second voice signal, wherein the voice amplitude of the second voice signal is within a reference voice amplitude interval, and the reference voice amplitude interval is used to indicate a voice amplitude range that should be met by a voice signal used to support switching the target device or the target application from a sleep state to a working state through voice; A control module configured to control the wake-up of the target device or the target application based on the second voice signal.
9. An electronic device, comprising: One or more processors; A storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the voice wake-up method of any one of claims 1-7.
10. A storage medium containing computer-executable instructions, wherein, The computer executable instructions, when executed by a computer processor, are used to perform the voice wake-up method of any one of claims 1-7.
Citation Information
Patent Citations
Method and system for setting terminal wakeup, and wakeup method and system
CN106131292A
Voice wake-up method and device, apparatus and storage medium
CN111192590A
Voice increasing method, system and device and storage medium
CN111599371A
Voice wake-up method and device, electronic equipment, medium and program product
CN114283793A
A computer implemented method and an apparatus for silence detection in speech recognition
US20230326481A1