Self-adaptive multistage noise reduction voice wake-up method and system and electronic equipment

By adaptively selecting the appropriate noise reduction model for noise reduction processing, a single noise reduction model is solved to solve the problem of difficult performance equalization in high and low signal-to-noise ratio scenarios, achieving lower voice signal distortion and higher wake-up accuracy.

CN120220683APending Publication Date: 2025-06-27AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510498042.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, a single noise reduction model is difficult to meet the performance balance in high signal-to-noise ratio and low signal-to-noise ratio scenarios at the same time, resulting in large distortion of voice signals, increasing the possibility of false judgment of wake-up models.

Method used

Adaptive multi-stage noise reduction method is adopted, and the received audio signal is estimated by configuring a generalized noise reduction model and multiple different volumes of signal-to-noise ratio, and adaptively selecting an appropriate noise reduction model for noise reduction processing, including using a single noise reduction model at high signal-to-noise ratio and a composite noise reduction model at low signal-to-noise ratio.

Benefits of technology

It effectively reduces the distortion of voice signal during noise reduction, improves the robustness and accuracy of low signal-to-noise ratio voice wake-up, and reduces the possibility of false wake-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220683A_ABST
    Figure CN120220683A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice wake-up method and system for adaptive multistage noise reduction and electronic equipment. The method comprises the following steps: configuring noise reduction models for local equipment, wherein the noise reduction models comprise a generalization noise reduction model and a plurality of noise reduction models of different volumes aiming at a signal-to-noise ratio; performing signal-to-noise ratio estimation on the received audio signal, and based on the estimated signal-to-noise ratio score, performing adaptive noise reduction model decision from a plurality of noise reduction models configured in local equipment to obtain noise reduction voice with reserved human voice; and performing voice wake-up judgment on the noise reduction voice, and if the noise reduction voice accords with a preset wake-up threshold, waking up the local device. According to the embodiment of the invention, a mode of mixing clean voice and noisy voice pair data and high-signal-to-noise-ratio voice and low-signal-to-noise-ratio voice pair data is used, the robustness of a low-resource noise reduction model can be improved, and noise reduction distortion is reduced. A self-adaptive noise reduction model is adopted, so that obvious voice signal distortion cannot be caused on the basis of ensuring noise elimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent voice, and in particular to a voice wake-up method, system and electronic device with adaptive multi-level noise reduction. Background Art

[0002] The low signal-to-noise ratio voice wake-up scheme usually consists of a noise reduction model and a wake-up model. The noise reduction model is used to obtain the audio with reduced noise, and the wake-up model is used to detect whether the score of the input audio is greater than or equal to a preset threshold. Specifically, the noisy audio is input, and the noise-reduced audio is obtained by the noise reduction model; the wake-up model calculates the score of the noise-reduced audio and compares it with the preset wake-up threshold. If the score is greater than or equal to the preset threshold, a wake-up instruction is output, otherwise no instruction is output.

[0003] In the process of implementing the present invention, the inventor found that there are at least the following problems in the related art: For the low signal-to-noise ratio scenarios in the real environment, a single noise reduction model cannot be well compatible, and there will be a serious distortion phenomenon. Specifically, because a single noise reduction model is used to fit the data with high signal-to-noise ratio and low signal-to-noise ratio at the same time, the noise reduction model will produce over-attenuation, and it cannot meet the performance balance in both high signal-to-noise ratio and low signal-to-noise ratio scenarios. The voice signal is distorted greatly, resulting in a relatively high possibility of misjudgment of the wake-up model. This defect has long existed in the field of low signal-to-noise ratio voice wake-up. Summary of the Invention

[0004] In order to at least solve the problem that the noise reduction model in the prior art cannot meet the performance balance in both high signal-to-noise ratio and low signal-to-noise ratio scenarios, resulting in large distortion of the voice signal and misjudgment of the wake-up model.

[0005] In a first aspect, an embodiment of the present invention provides a voice wake-up method with adaptive multi-level noise reduction, including: Configuring at least one noise reduction model trained with clean and noisy voice pairs and high signal-to-noise ratio and low signal-to-noise ratio voice pairs that are aligned in the time dimension and have the same voice content for the local device, including: a generalization noise reduction model, and multiple noise reduction models with different volumes for different signal-to-noise ratios; Performing signal-to-noise ratio estimation on the received audio signal, and based on the estimated signal-to-noise ratio value, making an adaptive noise reduction model decision from the multiple noise reduction models configured in the local device to obtain the noise-reduced voice with the human voice retained, including: - When the signal-to-noise ratio value meets the first threshold interval, it is decided to use the generalization noise reduction model for independent noise reduction, - When the signal-to-noise ratio value meets the second threshold interval, it is decided to repeatedly select two noise reduction models from the generalization noise reduction model and the noise reduction models with different volumes for different signal-to-noise ratios for composite noise reduction; Perform a voice wake-up judgment on the noise-reduced speech, and if it meets the preset wake-up threshold, wake up the local device.

[0006] In a second aspect, an embodiment of the present invention provides a voice wake-up system with adaptive multi-level noise reduction, including: A model configuration module, configured to configure at least one noise reduction model for the local device, which is trained with clean and noisy speech pairs that are aligned in the time dimension and have the same speech content, as well as speech pairs with high and low signal-to-noise ratios, including: a generalization noise reduction model, and multiple noise reduction models with different volumes for the signal-to-noise ratio; A signal-to-noise ratio estimation module, configured to estimate the signal-to-noise ratio of the received audio signal, based on the estimated signal-to-noise ratio value; A decision-making module, configured to adaptively make a decision on the noise reduction model from multiple noise reduction models already configured in the local device; A noise reduction module, configured to, when the signal-to-noise ratio value meets the first threshold range, decide to use the generalization noise reduction model for independent noise reduction, and when the signal-to-noise ratio value meets the second threshold range, decide to repeatedly select two noise reduction models from the generalization noise reduction model and noise reduction models with different volumes for the signal-to-noise ratio for composite noise reduction, to obtain noise-reduced speech with human voices retained; A wake-up module, configured to perform a voice wake-up judgment on the noise-reduced speech, and if it meets the preset wake-up threshold, wake up the local device.

[0007] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the steps of the adaptive multi-level noise reduction voice wake-up method according to any embodiment of the present invention.

[0008] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the steps of the adaptive multi-level noise reduction voice wake-up method according to any embodiment of the present invention.

[0009] In a fifth aspect, an embodiment of the present invention provides a computer program product, including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, they implement the steps of the adaptive multi-level noise reduction voice wake-up method according to any embodiment of the present invention.

[0010] The beneficial effects of the embodiments of the present invention are as follows: By using pairs of clean speech and noisy speech, as well as pairs of high signal-to-noise ratio speech and low signal-to-noise ratio speech data, the robustness of the low-resource noise reduction model can be improved, and the distortion of noise reduction can be reduced. Moreover, a noise reduction decision is introduced. Different noise reduction modules are configured according to the performance of the device. When the signal-to-noise ratio of the input audio is relatively high, only a single noise reduction model is used, which will not cause obvious distortion of the speech signal. When the signal-to-noise ratio of the input audio is relatively low, an adaptive two-stage noise reduction model is adopted. On the basis of ensuring noise cancellation, it will not cause obvious distortion of the speech signal either. It can well solve the problem of false wake-up caused by distortion brought by the noise reduction module. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a flowchart of a voice wake-up method with adaptive multi-stage noise reduction provided by an embodiment of the present invention; Figure 2 It is a specific schematic flowchart of a voice wake-up method with adaptive multi-stage noise reduction provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a voice wake-up system with adaptive multi-stage noise reduction provided by an embodiment of the present invention; Figure 4 It is a schematic structural diagram of an embodiment of an electronic device with adaptive multi-stage noise reduction for voice wake-up provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0014] As Figure 1 shown, it is a flowchart of a voice wake-up method with adaptive multi-stage noise reduction provided by an embodiment of the present invention, including the following steps: S11: Configure at least one noise reduction model for the local device, which is trained with clean and noisy speech pairs and speech pairs with high and low signal-to-noise ratios that are aligned in the time dimension and have the same speech content. The noise reduction models include: a generalized noise reduction model and multiple noise reduction models with different scales for the signal-to-noise ratio. S12: Estimate the signal-to-noise ratio of the received audio signal. Based on the estimated signal-to-noise ratio value, make an adaptive decision on the noise reduction model from the multiple noise reduction models configured on the local device to obtain noise-reduced speech with the human voice retained, including: - When the signal-to-noise ratio value meets the first threshold range, decide to use the generalized noise reduction model for independent noise reduction. - When the signal-to-noise ratio value meets the second threshold range, decide to repeatedly select two noise reduction models from the generalized noise reduction model and the noise reduction models with different scales for the signal-to-noise ratio for composite noise reduction. S13: Perform a voice wake-up judgment on the noise-reduced speech. If it meets the preset wake-up threshold, wake up the local device.

[0015] This application takes into account that in the prior art, in order to improve the wake-up rate, the number of model parameters and the amount of computation are usually increased to enhance the performance and robustness of the system in low signal-to-noise ratio scenarios. However, in local network-free wake-up devices, such as smart curtains and smart lights, these devices are usually implemented by an electric light / curtain + voice control central control module. They do not have a high-performance processing chip and have limited computing power, so the number of parameters and the amount of computation cannot be increased. In real usage scenarios, there are inevitably environmental noise, human voice interference, and possible device self-noise. For example, the rotation of a washing machine, the operation of a floor cleaning robot, and the sound played by a TV in a family, or in a vehicle environment, the noise of an engine will cause a low signal-to-noise ratio and affect the user's wake-up speech.

[0016] Due to limited computing power, there are problems of difficult wake-up and false wake-up during wake-up. However, to solve the problem of difficult wake-up, the wake-up threshold is usually reduced. Although this solves the problem of difficult wake-up, due to the reduction of the wake-up threshold, false wake-up will be caused. On the contrary, to solve false wake-up, the wake-up threshold is increased, or multiple verifications for wake-up are added, which will cause difficult wake-up. Therefore, it is difficult to optimize both difficult wake-up and false wake-up problems in the prior art.

[0017] This method proposes a low signal-to-noise ratio voice wake-up system with adaptive multi-level noise reduction. For the case where both memory and computing power are limited, without increasing the computing power and memory of the local device, the wake-up performance and robustness of the system are improved. And for local devices with relatively abundant computing power, it can adaptively provide a suitable wake-up model to enhance the wake-up performance and robustness of the system. Simply put, this method consists of adaptive noise reduction and wake-up. Specifically, it includes: a signal-to-noise ratio estimation module, a noise reduction decision module, a noise reduction model, and a wake-up model.

[0018] For step S11, considering that local voice devices in real scenarios are different from each other and the performance of each local voice device is completely different, the present application configures multiple different noise reduction models for different local devices, including: a generalization noise reduction model and multiple noise reduction models with different scales for signal-to-noise ratio.

[0019] As an implementation manner, the multiple noise reduction models configured for the local device are determined by the memory and computing power of the local device.

[0020] Specifically, for example, in a common device scoring method, based on the local voice memory and computing power, the computing power score of the local device is obtained. There are different evaluation methods for calculating the computing power score of the device, which are not limited here and will not be elaborated in detail.

[0021] For example, when the computing power score of the local device is low, only one generalization noise reduction model is configured for it; while when the computing power score of the local device is high, two generalization noise reduction models and multiple noise reduction models with different scales can be configured for it. The specific scores and model selections will not be elaborated here and can be adaptively adjusted according to actual needs.

[0022] As an implementation manner, training multiple noise reduction models with clean and noisy speech pairs and high and low signal-to-noise ratio speech pairs that are aligned in the time dimension and have the same speech content includes: Inputting the clean and noisy speech pairs and high and low signal-to-noise ratio speech pairs into the noise reduction model, respectively determining the clean, noisy, high signal-to-noise ratio, and low signal-to-noise ratio audio features, and constructing a high-dimensional feature mapping network; Using the feature mapping relationship network to estimate the predicted speech of the noisy speech, and training the noise reduction model with the error between the clean speech and the predicted speech until the error is less than a preset target.

[0023] In this implementation manner, regarding the training of the noise reduction model, the noise reduction model can adopt DCCRN (Deep Complex Convolution Recurrent Network). This method uses clean and noisy speech pairs and high and low signal-to-noise ratio speech pairs that are aligned in the time dimension and have the same speech content to train the noise reduction model. These data can be obtained by adding noise with different amplitudes to the clean speech and processing, so as to obtain clean and noisy speech pairs and high and low signal-to-noise ratio speech pairs that are aligned in the time dimension and have the same speech content. The pre-trained models are trained separately, rather than training multiple models at one time.

[0024] After being input into the model, the clean, noisy, high signal-to-noise ratio, and low signal-to-noise ratio audio features of the same content are respectively determined, and a high-dimensional feature mapping network between the features is constructed. During noise reduction, the noisy speech is received and the predicted clean speech after noise reduction is output. The noise reduction model is trained through the error between the predicted clean speech and the true clean speech in the training data until the error is less than the preset target. In this way, a generalized independent noise reduction model is obtained.

[0025] Further, the training further includes: Inputting a first quantity of clean and noisy speech pairs and a second quantity of high signal-to-noise ratio and low signal-to-noise ratio speech pairs into the noise reduction model to train the noise reduction model for the signal-to-noise ratio, where the second quantity is greater than the first quantity.

[0026] In this embodiment, considering adapting to more different local devices, clean, noisy, and a larger quantity of high signal-to-noise ratio and low signal-to-noise ratio audio pair features can be selected to construct a mapping relationship network between the features. That is to say, in this training, the data volume of the high signal-to-noise ratio and low signal-to-noise ratio audio pairs is greater than that of the clean and noisy speech pairs, so that a noise reduction model for the signal-to-noise ratio can be trained. The data volume can also be further adjusted according to requirements to train multiple noise reduction models of different sizes.

[0027] For step S12, after configuring an adaptive noise reduction model for different local devices, voice wake-up is performed. The received audio signal is sent to the signal-to-noise ratio estimation module, and the signal-to-noise ratio score of the audio signal can be obtained. The signal-to-noise ratio score of the audio signal is sent to the noise reduction decision module for adaptive noise reduction model decision.

[0028] The decision is divided into two parts, relatively high signal-to-noise ratio and relatively low signal-to-noise ratio. For example, a signal-to-noise ratio greater than or equal to 70 is set as the first threshold interval, and a signal-to-noise ratio less than 70 is set as the second threshold interval.

[0029] If the signal-to-noise ratio of the received audio signal conforms to the first threshold interval, it means that there is noise but relatively little, so it is decided to use the generalized noise reduction model for independent noise reduction.

[0030] If the signal-to-noise ratio of the received audio signal conforms to the second threshold interval, it means that there is more noise and the sound quality is poor, and further noise reduction is required. However, considering the performance of the local device and the marginal effect, two noise reduction models can be repeatedly selected from the configured noise reduction models for composite noise reduction.

[0031] Specifically, the adaptive noise reduction model decision based on the estimated signal-to-noise ratio score includes: - When the signal-to-noise ratio score conforms to the second threshold interval and is in the first sub-interval, it is decided to repeatedly use the generalized noise reduction model for two-time composite noise reduction; - When the signal-to-noise ratio value meets the second threshold range and is in the second sub-range, it is decided to perform composite noise reduction using the generalized noise reduction model and the noise reduction models of the first volume or the second volume; - When the signal-to-noise ratio value meets the second threshold range and is in the third sub-range, it is decided to perform composite noise reduction using the noise reduction model of the first volume and the noise reduction model of the second volume.

[0032] In this embodiment, within the second threshold range, the method further divides into multiple sub-ranges, which takes into account both the models contained in the device itself and the noise reduction effect. For example, the signal-to-noise ratio is divided into a first sub-range of 60 - 70, a second sub-range of 40 - 60, and a third sub-range of 0 - 40.

[0033] If the signal-to-noise ratio of the audio signal is in the first sub-range, according to the performance of the local device, the same generalized noise reduction model can be selected to perform noise reduction twice repeatedly, or two identical generalized noise reduction models can be connected in series to perform noise reduction twice to achieve composite noise reduction. After the audio signal is input, the first generalized noise reduction model is used to reduce the noise of the speech, and the denoised speech is then input into the second generalized noise reduction model. Although the two noise reduction models are the same, just like AI removing enlarged images, the same algorithm and the same input may output different contents. Therefore, to a certain extent, two generalized noise reduction models connected in series can further improve the noise reduction effect. However, after all, they are the same generalized noise reduction models. Even if more are connected in series, the additional computing power consumed will not bring greater improvement to the noise reduction. Considering further compressing the space occupied by the noise reduction model, one generalized noise reduction model can also be used to perform noise reduction twice repeatedly. In this case, even if the memory of the local device is small enough to only configure one generalized noise reduction model, multiple decision-making methods are also realized, achieving a better noise reduction effect.

[0034] If the signal-to-noise ratio of the audio signal is in the second sub-range, according to the performance of the local device, one generalized noise reduction model and a noise reduction model for the signal-to-noise ratio of the first volume (relatively small) or the second volume (relatively large) can be selected. In this way, after the audio signal is input into the generalized noise reduction model for generalized noise reduction, it is then input into another different noise reduction model for low signal-to-noise ratio. The noise reduction effect at low signal-to-noise ratio is further improved.

[0035] If the signal-to-noise ratio of the audio signal is in the second sub-range, it indicates that the noise of the audio signal is very large at this time and enhanced noise reduction is required. At this time, it is decided to perform composite noise reduction using the noise reduction model of the first volume and the noise reduction model of the second volume to more comprehensively improve the noise reduction effect at low signal-to-noise ratio.

[0036] Finally, the noise-reduced speech with the human voice retained is obtained.

[0037] For step S13, the denoised speech after denoising processing is input into the wake-up module to calculate the wake-up score. The wake-up module consists of a wake-up model, generally using FSMN (Feedforward Sequential Memory Networks), with the input being audio and the output being the wake-up score. The obtained wake-up score is compared with a preset wake-up threshold. If the wake-up score is greater than or equal to the preset wake-up threshold, the system of the local device is activated for voice interaction; otherwise, it is not activated.

[0038] As Figure 2 shown in the simplified flowchart of the above method, the steps can be simplified as follows: Step 1: Obtain paired training data. The paired training data refers to a pair of clean speech and noisy speech or a pair of high signal-to-noise ratio speech and low signal-to-noise ratio speech, whose speech content is exactly the same and is completely aligned in the time dimension; Step 2: Use the paired training data to train a denoising model. The denoising model generally uses DCCRN (DeepComplex Convolution Recurrent Network), with the input being noisy speech and the output being enhanced speech with reduced noise; Step 3: Obtain the received audio signal; Step 4: Send the audio signal into the signal-to-noise ratio estimation module to obtain the signal-to-noise ratio score of the audio signal; Step 5: Send the signal-to-noise ratio score of the audio signal in Step 4 into the denoising decision module to determine whether the signal-to-noise ratio score is greater than or equal to a preset signal-to-noise ratio threshold. If the signal-to-noise ratio score is greater than or equal to the preset signal-to-noise ratio threshold, use a single denoising model to eliminate noise; otherwise, use a two-stage denoising model to eliminate noise and retain the human voice part; Step 6: According to the decision of the denoising decision module in Step 5, use a single denoising model or a two-stage denoising model to eliminate noise and retain the human voice part to obtain the denoised audio. The single denoising model is the trained denoising model obtained in Step 2, and the two-stage denoising model is composed of two single denoising models connected in series; Step 7: Input the denoised audio obtained in Step 6 into the wake-up module to calculate the wake-up score. The wake-up module consists of a wake-up model, generally using FSMN (Feedforward Sequential Memory Networks), with the input being audio and the output being the wake-up score; Step 8: Compare the wake-up score obtained in Step 7 with a preset wake-up threshold. If the wake-up score is greater than or equal to the preset wake-up threshold, activate the system for voice interaction; otherwise, do not activate the system.

[0039] It can be seen from this embodiment that when the method is training a noise reduction model, by using a method of mixing clean speech and noisy speech pairs, as well as high signal-to-noise ratio speech and low signal-to-noise ratio speech pairs of data, the robustness of the low-resource noise reduction model can be improved, and the distortion of noise reduction can be reduced. And a noise reduction decision is introduced. Different noise reduction modules are configured according to the performance of the device. When the signal-to-noise ratio of the input audio is relatively high, only a single noise reduction model is used, which will not cause obvious distortion of the speech signal. When the signal-to-noise ratio of the input audio is relatively low, an adaptive two-stage noise reduction model is adopted. On the basis of ensuring noise cancellation, it will not cause obvious distortion of the speech signal either. It can well solve the problem of false wake-up caused by the distortion brought by the noise reduction module.

[0040] As Figure 3 shown in the structural schematic diagram of a voice wake-up system with adaptive multi-stage noise reduction provided by an embodiment of the present invention. This system can execute the voice wake-up method with adaptive multi-stage noise reduction described in any of the above embodiments and is configured in a terminal.

[0041] A voice wake-up system 10 with adaptive multi-stage noise reduction provided in this embodiment includes: a model configuration module 11, a signal-to-noise ratio estimation module 12, a decision module 13, a noise reduction module 14, and a wake-up module 15.

[0042] Among them, the model configuration module 11 is used to configure at least one noise reduction model trained with clean and noisy speech pairs and high signal-to-noise ratio and low signal-to-noise ratio speech pairs that are aligned in the time dimension and have the same speech content for the local device, including: a generalized noise reduction model, and multiple noise reduction models with different scales for the signal-to-noise ratio; the signal-to-noise ratio estimation module 12 is used to estimate the signal-to-noise ratio of the received audio signal, and based on the estimated signal-to-noise ratio value, estimate the signal-to-noise ratio of the received audio signal, based on the estimated signal-to-noise ratio value; the decision module 13 is used to make an adaptive noise reduction model decision from the multiple noise reduction models already configured in the local device; the noise reduction module 14 is used to decide to use the generalized noise reduction model for independent noise reduction when the signal-to-noise ratio value meets the first threshold range, and when the signal-to-noise ratio value meets the second threshold range, decide to repeatedly select two noise reduction models from the generalized noise reduction model and noise reduction models with different scales for the signal-to-noise ratio for composite noise reduction to obtain noise-reduced speech with human voice retained; the wake-up module 15 is used to perform a voice wake-up judgment on the noise-reduced speech, and if it meets the preset wake-up threshold, wake up the local device.

[0043] The embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and these computer-executable instructions can execute the voice wake-up method with adaptive multi-stage noise reduction in any of the above method embodiments; As an implementation, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are configured as follows: Configure at least one noise reduction model trained with clean and noisy speech pairs and speech pairs with high and low signal-to-noise ratios that are aligned in the time dimension and have the same speech content for the local device, including: a generalization noise reduction model, and multiple noise reduction models with different sizes for the signal-to-noise ratio; Perform signal-to-noise ratio estimation on the received audio signal, and based on the estimated signal-to-noise ratio value, make an adaptive noise reduction model decision from the multiple noise reduction models configured on the local device to obtain noise-reduced speech with human voices retained, including: - When the signal-to-noise ratio value meets the first threshold interval, decide to use the generalization noise reduction model for independent noise reduction, - When the signal-to-noise ratio value meets the second threshold interval, decide to repeatedly select two noise reduction models from the generalization noise reduction model and the noise reduction models with different sizes for the signal-to-noise ratio for composite noise reduction; Perform voice wake-up judgment on the noise-reduced speech, and if it meets the preset wake-up threshold, wake up the local device.

[0044] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by a processor, perform the voice wake-up method of adaptive multi-level noise reduction in any of the above method embodiments.

[0045] Figure 4 It is a schematic diagram of the hardware structure of an electronic device for the voice wake-up method of adaptive multi-level noise reduction provided in another embodiment of the present application, as Figure 4 shown. The device includes: One or more processors 410 and a memory 420, Figure 4 Taking one processor 410 as an example. The device for the voice wake-up method of adaptive multi-level noise reduction may further include: an input device 430 and an output device 440.

[0046] The processor 410, the memory 420, the input device 430, and the output device 440 can be connected through a bus or other means, Figure 4 Taking connection through a bus as an example.

[0047] The memory 420 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice wake-up method with adaptive multi-level noise reduction in the embodiments of the present application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 420, that is, implements the voice wake-up method with adaptive multi-level noise reduction in the above method embodiments.

[0048] The memory 420 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data, etc. In addition, the memory 420 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 420 may optionally include a memory remotely disposed relative to the processor 410, and these remote memories can be connected to the mobile device through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0049] The input device 430 can receive input digital or character information. The output device 440 may include a display device such as a display screen.

[0050] The one or more modules are stored in the memory 420 and, when executed by the one or more processors 410, execute the voice wake-up method with adaptive multi-level noise reduction in any of the above method embodiments.

[0051] The above product can execute the method provided in the embodiments of the present application and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.

[0052] The non-volatile computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the device through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0053] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the voice wake-up method for adaptive multi-level noise reduction according to any embodiment of the present invention.

[0054] The electronic devices in the embodiments of the present application exist in multiple forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0055] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as tablet computers.

[0056] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.

[0057] (4) Other electronic devices with data processing functions.

[0058] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also include other elements not explicitly listed, or also include elements inherent to such a process, method, article, or device. Without more limitations, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article, or device including the said elements.

[0059] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0060] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voice wake-up method with adaptive multi-level noise reduction, comprising: Configure at least one noise reduction model trained by clean and noisy speech pairs and high signal-to-noise ratio and low signal-to-noise ratio speech pairs aligned in time dimension and with consistent speech content for the local device, including: a generalized noise reduction model, and noise reduction models of multiple different sizes for signal-to-noise ratio; The received audio signal is estimated to have a signal-to-noise ratio, and based on the estimated signal-to-noise ratio value, an adaptive noise reduction model decision is made from a plurality of noise reduction models configured on the local device to obtain noise-reduced speech that retains human voice, including: -When the signal-to-noise ratio value meets the first threshold interval, a decision is made to use the generalized denoising model for independent denoising, -When the signal-to-noise ratio value meets the second threshold interval, a decision is made to repeatedly select two noise reduction models from the generalized noise reduction model and the noise reduction models for different volumes of the signal-to-noise ratio to perform composite noise reduction; Perform voice wake-up judgment on the noise reduction voice, and if it meets a preset wake-up threshold, wake up the local device.

2. The method according to claim 1, wherein: If the multiple noise reduction models configured on the local device include: a generalized noise reduction model, a noise reduction model of a first volume for signal-to-noise ratio, and a noise reduction model of a second volume for signal-to-noise ratio, the adaptive noise reduction model decision based on the estimated signal-to-noise ratio value includes: -When the signal-to-noise ratio value meets the second threshold interval and is in the first sub-interval, a decision is made to repeatedly use the generalized denoising model to perform composite denoising twice; -When the signal-to-noise ratio value meets the second threshold interval and is in the second sub-interval, a decision is made to use the generalized denoising model and the denoising model of the first volume or the second volume to perform composite denoising; - When the signal-to-noise ratio value meets the second threshold interval and is in the third sub-interval, a decision is made to perform composite noise reduction using the noise reduction model of the first volume and the noise reduction model of the second volume.

3. The method according to claim 1, wherein: The multiple noise reduction models configured for the local device are determined by the memory and computing power of the local device.

4. The method according to claim 1, wherein: The training of multiple noise reduction models using clean and noisy speech pairs and high signal-to-noise ratio and low signal-to-noise ratio speech pairs aligned in time dimension and with consistent speech content includes: Inputting the clean and noisy speech pairs and the high signal-to-noise ratio and low signal-to-noise ratio speech pairs into a noise reduction model, determining the clean, noisy, high signal-to-noise ratio, and low signal-to-noise ratio audio features respectively, and constructing a high-dimensional feature mapping network; The predicted speech of the noisy speech is estimated using the feature mapping relationship network, and the noise reduction model is trained using the error between the clean speech and the predicted speech until the error is less than a preset target.

5. The method according to claim 4, wherein: The training also includes: A first number of clean and noisy speech pairs and a second number of high signal-to-noise ratio and low signal-to-noise ratio speech pairs are input into a denoising model to train the denoising model for signal-to-noise ratio, wherein the second number is greater than the first number.

6. The method according to claim 1, wherein: The performing speech wake-up judgment on the noise reduction speech includes: performing wake-up judgment using a feedforward sequential memory network.

7. A voice wake-up method with adaptive multi-level noise reduction, comprising: A model configuration module, configured to configure at least one noise reduction model trained by clean and noisy speech pairs and high signal-to-noise ratio and low signal-to-noise ratio speech pairs aligned in time dimension and with consistent speech content for a local device, including: a generalized noise reduction model, and noise reduction models of multiple different sizes for signal-to-noise ratios; a signal-to-noise ratio estimation module, for estimating a signal-to-noise ratio of a received audio signal based on an estimated signal-to-noise ratio value; A decision module, used for making an adaptive noise reduction model decision from a plurality of noise reduction models configured in the local device; A noise reduction module, configured to, when the signal-to-noise ratio value meets a first threshold interval, decide to use the generalized noise reduction model for independent noise reduction, and when the signal-to-noise ratio value meets a second threshold interval, decide to repeatedly select two noise reduction models from the generalized noise reduction model and noise reduction models with different volumes for signal-to-noise ratios for composite noise reduction to obtain noise-reduced speech that retains human voice; The wake-up module is used to perform voice wake-up judgment on the noise reduction voice, and wake up the local device if it meets a preset wake-up threshold.

8. A storage medium having a computer program product stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

9. A computer program product having instructions embedded on a storage medium, wherein the instructions implement the steps of the method according to any one of claims 1 to 6.

10. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 6.