Voice wake-up method and system based on noise reduction separation and electronic equipment

By performing noise reduction and audio feature judgment on local devices, combined with cloud separation and verification, the voice wake-up problem affected by external noise is solved, and the system wake-up accuracy and robustness are improved.

CN120340504APending Publication Date: 2025-07-18AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510482778.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the case of large external noise in daily life, the performance of the target wake-up model will decline, the noise affects the accuracy of the system, and the judgment of a single module is easy to misjudgment.

Method used

Noise reduction is carried out in local devices with limited computing power, and the noise reduction voice of the target speaker and the interfering speaker is retained. The wake-up words are judged through the audio characteristics, and the voice is sent to the cloud server for separation and verification, and a secondary verification is carried out in combination with local and cloud models.

Benefits of technology

It improves the wake-up performance in noise situations, enhances the system's anti-external noise and anti-interference ability, reduces the misjudgment of wake-up events, and improves the robustness and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340504A_ABST
    Figure CN120340504A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice wake-up method and system based on noise reduction separation and electronic equipment. The method comprises the following steps: performing noise reduction on voice in local equipment with limited computing power to obtain noise-reduced voice in which environmental noise is eliminated and the sound of a target speaker and the sound of an interference speaker are reserved; extracting audio features of the noise reduction voice on the local equipment, judging whether a wake-up word exists in the noise reduction voice or not through the audio features, and sending the noise reduction voice to a cloud server if the wake-up word exists; separating the noise-reduced voice according to the number of the speakers at the cloud server, and outputting multi-channel voice data of each channel representing one speaker; and judging whether the voice data of each channel contains a wake-up word from the multi-channel voice data, and if so, activating and waking up the local equipment. According to the embodiment of the invention, the local noise reduction module suppresses the environmental noise, improves the wake-up performance under the condition of noise, and improves the external noise and interference resistance of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent voice, and in particular to a voice wake-up method, system and electronic device based on noise reduction and separation. Background Art

[0002] In a voice interaction with a corresponding target speaker, the device is usually in a standby state and, when awakened by the target speaker, conducts a voice interaction with the target speaker. To achieve the purpose of identifying and responding to the target speaker, the prior art needs to extract the voice of the target speaker, so it is necessary to build a target speaker registration library, usually using the target person to register in advance, or using the wake-up word audio to register.

[0003] Specifically, the target user (target speaker) is determined according to the wake-up voice, and the voice feature value of the target user is obtained; then, a voice separation operation is performed on the received voice command to obtain at least one human voice command; then, according to the voice feature value corresponding to each human voice command and the voice feature value corresponding to the wake-up voice command, the similarity corresponding to each human voice command is determined; finally, if it is detected that the highest similarity is greater than or equal to a preset similarity, semantic recognition is performed on the human voice command with the highest similarity to identify the target command.

[0004] In the process of implementing the present invention, the inventor found that there are at least the following problems in the related art: The prior art tends to ideal recognition. In the case of large external noise in daily life, this causes the performance of the target wake-up model to decline. At the same time, the wake-up voice segment is not processed, which will introduce noise and affect the accuracy of the system, and using a single module to judge similarity has a relatively high possibility of misjudgment. Summary of the Invention

[0005] To at least solve the problem of voice wake-up recognition in the prior art. In a first aspect, an embodiment of the present invention provides a voice wake-up method based on noise reduction and separation, including: Noise reduction is performed on the voice in a local device with limited computing power to obtain noise-reduced voice that eliminates environmental noise and retains the voice of the target speaker and the voice of the interfering speaker; Audio features of the noise-reduced voice are extracted on the local device with limited computing power, and whether there is a wake-up word in the noise-reduced voice is judged through the audio features. If the wake-up word exists, the noise-reduced voice is sent to the cloud server; The noise-reduced voice is separated according to the number of speakers on the cloud server, and multi-channel voice data representing each speaker corresponding to each channel is output; It is judged whether each channel of voice data in the multi-channel voice data contains a wake-up word. If the wake-up word is included, the local device is activated and awakened.

[0006] In a second aspect, an embodiment of the present invention provides a voice wake-up system based on noise reduction and separation, including: A noise reduction module, configured to perform noise reduction on the voice in a local device with limited computing power to obtain noise-reduced voice that eliminates environmental noise and retains the voice of the target speaker and the interfering speaker; A wake-up module, configured to extract audio features of the noise-reduced voice on the local device with limited computing power, and determine whether there is a wake-up word in the noise-reduced voice through the audio features. If there is the wake-up word, the noise-reduced voice is sent to the cloud server; A cloud separation module, configured to separate the noise-reduced voice according to the number of speakers on the cloud server, and output multi-channel voice data in which each channel represents a corresponding speaker; A verification module, configured to determine whether each channel of voice data contains a wake-up word from the multi-channel voice data. If the wake-up word is included, the local device is activated and woken up.

[0007] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the voice wake-up method based on noise reduction and separation according to any embodiment of the present invention.

[0008] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and is characterized in that when the program is executed by a processor, the steps of the voice wake-up method based on noise reduction and separation according to any embodiment of the present invention are implemented.

[0009] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and is characterized in that when the computer program / instructions are executed by a processor, the steps of the voice wake-up method based on noise reduction and separation according to any embodiment of the present invention are implemented.

[0010] The beneficial effects of the embodiments of the present invention are as follows: The combination of the local module and the cloud model improves the overall system performance without increasing the local computing power. The local noise reduction module suppresses environmental noise, improving the wake-up performance in the presence of noise. The further separation of human voices by the cloud separation model improves the anti-noise and anti-interference capabilities of the system. In addition, the primary judgment of the local wake-up model combined with the secondary verification of the cloud separation + verification module improves the accuracy of wake-up judgment, reduces the misjudgment of wake-up events, and improves the robustness of the system. Further, the wake-up of the target person is also considered, improving the functionality of voice wake-up. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 is a flowchart of a voice wake-up method based on noise reduction and separation provided by an embodiment of the present invention; Figure 2 is a specific example flowchart of a voice wake-up method based on noise reduction and separation provided by an embodiment of the present invention; Figure 3 is a schematic structural diagram of a voice wake-up system based on noise reduction and separation provided by an embodiment of the present invention; Figure 4 is a schematic structural diagram of an embodiment of an electronic device for voice wake-up based on noise reduction and separation provided by an embodiment of the present invention. Detailed implementation manners

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0014] As Figure 1 shown is a flowchart of a voice wake-up method based on noise reduction and separation provided by an embodiment of the present invention, including the following steps: S11: Perform noise reduction on the voice in a local device with limited computing power to obtain noise-reduced voice that eliminates environmental noise and retains the voice of the target speaker and the voice of the interfering speaker; S12: Extract audio features of the noise-reduced voice on the local device with limited computing power, and determine whether there is a wake-up word in the noise-reduced voice based on the audio features. If the wake-up word exists, send the noise-reduced voice to the cloud server; S13: Separate the noise-reduced voice on the cloud server according to the number of speakers, and output multi-channel voice data where each channel represents a corresponding speaker; S14: Determine whether each channel of voice data in the multi-channel voice data contains a wake-up word. If the wake-up word is included, activate and wake up the local device.

[0015] This method is for completing the voice wake-up function and performing voice interaction in a low signal-to-noise ratio scenario. Voice recognition usually may be processed in the cloud, but wake-up mostly occurs locally on the device. However, the computing power and resources of local devices are often limited. Therefore, the ability to handle external noise and human voice interference is relatively weak. In real usage scenarios, it is inevitable that there will be external noise and human voice interference. In high-noise situations, existing voice interaction systems often have problems such as being unable to wake up and false wake-up, which will affect the user experience. Therefore, this method considers determining how to improve the system accuracy under the condition of limited computing power to ensure normal operation on the device.

[0016] Briefly speaking, limited by the real scenarios in reality and also considering the computing power of local devices, this method has a local device for primary noise reduction wake-up, and the cloud system for secondary separation, and then a verification system. Thus, it can improve the accuracy of system wake-up without increasing the computing power of local devices.

[0017] For step S11, the voice is sent to the noise reduction module of the local device. This module consists of a noise reduction model with low computing power, which performs noise reduction processing on the input audio to remove background noise in the real environment, such as the sounds of household appliances, walking, music, etc. in a home environment. For example, the sounds of wind, honking, music, etc. in a vehicle environment. These environmental noises are eliminated through the noise reduction module. At the same time, the noise-reduced voice with human voices is retained, including the target human voice and interfering human voices.

[0018] It should be noted that noise reduction must be performed locally, even though the cloud has better performance in the real scenario. For those skilled in the art, the cloud model can use greater computing power to obtain a better noise reduction effect, and at the same time simplify the local module, which can adapt to more local device situations and is more convenient for transplantation. However, since no noise reduction is done, this places higher requirements on the ability of the local wake-up model to process noisy audio. Limited by the limited computing power of local devices, misjudgment occurs at the wake-up module stage and it cannot enter the cloud for processing.

[0019] For step S12, after being processed by the noise reduction module, the noise-reduced audio is sent to the wake-up module. This module detects the input audio to determine whether there is a wake-up word. Specifically, first, audio features are extracted from the output audio, and then the audio features are sent to the wake-up model. The wake-up model adopts a CNN (Convolutional Neural Network) structure, and the model outputs the probability of wake-up. Finally, the final wake-up probability is obtained through the post-processing module to determine whether there is a wake-up word.

[0020] As an implementation, after determining whether there is a wake-up word in the noise-reduced voice by using the audio features, if there is a wake-up word, the method further includes: According to the wake-up point time information, expand the preset time forward and backward respectively, and intercept the noise-reduced voice segment and send it to the cloud server.

[0021] Specifically, if there is a wake-up word, according to the wake-up point time information, expand forward by F seconds and backward by B seconds respectively, intercept the wake-up audio time segment, and then send the audio to the cloud. (Generally, F is in [1 - 2], and B is in [0.5 - 1]) For step S13, after determining the local wake-up audio, send it to the separation module of the cloud server. The cloud separation module is composed of a separation model. The separation model can adopt structures such as TCN (Temporal Convolutional Network), DPRNN (Dual-Path Recurrent Neural Network), TFPSNet (Time-Frequency Domain Path Scanning Network), etc., and separate the input audio according to the preset number of speakers, output multi-channel data, where each channel represents a speaker. If there is a channel without a speaker, it means that the input audio has fewer speakers than the preset number. Then input the separated multi-channel audio into the verification module.

[0022] It should be noted that those skilled in the art usually think of local verification + cloud secondary verification to obtain a more accurate wake-up result for verification. Integration is indeed a wake-up solution with straightforward logic, simple implementation, and lower cost. However, this application finds that not only environmental noise but also other human voice interferences affect the wake-up accuracy. Therefore, separation is more important, and the detection of wake-up is relatively simple. Whether multi-channel voice data can be separated with each channel representing a corresponding speaker is particularly important. Therefore, this method is local verification + cloud separation.

[0023] For step S14, after obtaining multi-channel voice data with each channel representing a corresponding speaker, perform verification to further confirm whether the wake-up word is included. If the wake-up word is included, activate and wake up the local device. This module can be composed of a local wake-up model or a cloud recognition model.

[0024] As an implementation manner, judging whether each channel of the multi-channel voice data contains a wake-up word is judged in the local device, including: The local device receives the multi-channel voice data sent from the cloud server; Perform active voice detection on the multi-channel voice data, and only perform wake-up word judgment verification on the channel voice data detected to contain human voices. If the wake-up word is included, activate and wake up the local device.

[0025] In this embodiment, since the wake-up verification computing power is relatively simple, it can be verified again by the wake-up module of the local device. That is to say, the same module verifies twice. Further, in order to save computing resources, before inputting the multi-channel audio, a VAD (Voice Activity Detection) model will be used to detect whether each channel contains human voices. Channels that do not contain human voices will not be sent to the wake-up model for verification. Through the secondary verification of the separated audio, it is finally confirmed whether the wake-up word is contained. If the wake-up word is contained, the local device will be activated and woken up. Additionally, a wake-up model with greater performance can also be used for verification.

[0026] As an implementation manner, it is judged in the cloud server whether each channel voice data in the multi-channel voice data contains the wake-up word, including: The cloud server performs wake-up word judgment verification on the multi-channel voice data. If the wake-up word is included, the single-channel voice data containing the wake-up word and the verification result are sent to the local device.

[0027] In this embodiment, this method also takes into account that the cloud server has stronger capabilities and can perform wake-up word verification, specifically as Figure 2 shown. However, after the verification passes, the single-channel voice data of the wake-up word and the verification result need to be sent to the local device. The purpose of doing this is that this method also takes into account the wake-up of the target person.

[0028] Nowadays, users pay more attention to data security. The personal information of users is stored locally instead of in the cloud. If the wake-up judgment is made locally, the local voice library can be used for judgment. Since the cloud does not store the voice information of users, it can only judge which voice channel has been woken up. Therefore, if the wake-up judgment is made in the cloud, after confirming that there is a wake-up word, the local device still needs to judge whether the single-channel voice data is the voice of the target speaker. If it is the voice of the target speaker, the local device will be activated and woken up.

[0029] Therefore, the voice wake-up of this method is divided into two types. One is that hitting the wake-up word can wake up. Additionally, it will only wake up when the target speaker hits the wake-up word. As an implementation, it is judged from the multi-channel voice data whether the voice of the target speaker contains the wake-up word. If the wake-up word is included, the local device is activated and woken up. If the voice of the interfering speaker contains the wake-up word, the local device is not activated and woken up. Furthermore, it is judged from the multi-channel voice data whether the voice of the target speaker contains the wake-up word. If the wake-up word is included, the local device is activated and woken up. If the voice of the interfering speaker contains the wake-up word, the local device is not activated and woken up.

[0030] The wake-up of the target speaker in this method is mainly divided into two parts: local noise reduction wake-up and cloud separation, and then local / cloud verification is performed. The combination of the local module and the cloud model improves the overall system performance without increasing the local computing power. The local noise reduction module suppresses environmental noise and improves the wake-up performance in the presence of noise. The further separation of human voices by the cloud separation model improves the system's ability to resist external noise and interference. In addition, the primary judgment of the local wake-up model combined with the secondary verification of the cloud separation + verification module improves the accuracy of wake-up judgment, reduces the misjudgment of wake-up events, and improves the system's robustness. Furthermore, the wake-up of the target person is also considered to enhance the functionality of voice wake-up.

[0031] As Figure 3 shown is a schematic structural diagram of a voice wake-up system based on noise reduction and separation provided by an embodiment of the present invention. This system can execute the voice wake-up method based on noise reduction and separation described in any of the above embodiments and is configured in a terminal.

[0032] A voice wake-up system 10 based on noise reduction and separation provided in this embodiment includes: a noise reduction module 11, a wake-up module 12, a cloud separation module 13, and a verification module 14.

[0033] Among them, the noise reduction module 11 is used to perform noise reduction on the voice in a local device with limited computing power to obtain noise-reduced voice that eliminates environmental noise and retains the voice of the target speaker and the voice of the interfering speaker; the wake-up module 12 is used to extract the audio features of the noise-reduced voice on the local device with limited computing power, and judge whether there is a wake-up word in the noise-reduced voice through the audio features. If there is the wake-up word, the noise-reduced voice is sent to the cloud server; the cloud separation module 13 is used to separate the noise-reduced voice according to the number of speakers in the cloud server and output multi-channel voice data in which each channel represents a corresponding speaker; the verification module 14 is used to judge whether each channel of voice data contains the wake-up word from the multi-channel voice data. If the wake-up word is included, the local device is activated and woken up.

[0034] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by a processor, the voice wake-up method based on noise reduction and separation in any of the above method embodiments is executed.

[0035] Figure 4 FIG. is a schematic hardware structure diagram of an electronic device for the voice wake-up method based on noise reduction and separation provided in another embodiment of the present application. As Figure 4 shown, the device includes: One or more processors 410 and a memory 420. Figure 4 Taking one processor 410 as an example. The device for the voice wake-up method based on noise reduction and separation may further include: an input device 430 and an output device 840.

[0036] The processor 410, the memory 420, the input device 430, and the output device 840 may be connected through a bus or other means. Figure 4 Taking the connection through a bus as an example.

[0037] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice wake-up method based on noise reduction and separation in the embodiments of the present application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 420, that is, implementing the voice wake-up method based on noise reduction and separation in the above method embodiments.

[0038] The memory 420 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data, etc. In addition, the memory 420 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 420 may optionally include a memory remotely set relative to the processor 410, and these remote memories may be connected to the mobile device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0039] The input device 430 can receive input digital or character information. The output device 840 may include a display device such as a display screen.

[0040] The one or more modules are stored in the memory 420 and, when executed by the one or more processors 410, perform the voice wake-up method based on noise reduction separation in any of the above method embodiments.

[0041] The above product can execute the method provided by the embodiments of the present application, and has functional modules and beneficial effects corresponding to the execution of the method. For technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiments of the present application.

[0042] The non-volatile computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0043] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the voice wake-up method based on noise reduction separation according to any embodiment of the present invention.

[0044] The electronic device of the embodiments of the present application exists in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0045] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as tablet computers.

[0046] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.

[0047] (4) Other electronic devices with data processing functions.

[0048] In this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" not only include those elements, but also include other elements not expressly listed, or also include elements inherent to such a process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the said element.

[0049] The device embodiments described above are merely illustrative, where the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0050] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voice wake-up method based on noise reduction and separation, comprising: Performing noise reduction on the voice in a local device with limited computing power to obtain noise-reduced voice that eliminates environmental noise and retains the voice of the target speaker and the voice of the interfering speaker; Extracting the audio features of the noise-reduced voice on the local device with limited computing power, and determining whether there is a wake-up word in the noise-reduced voice through the audio features. If there is the wake-up word, sending the noise-reduced voice to the cloud server; Separating the noise-reduced voice according to the number of speakers in the cloud server, and outputting multi-channel voice data with each channel representing a corresponding speaker; Determining whether each channel of voice data in the multi-channel voice data contains a wake-up word. If the wake-up word is included, activating and waking up the local device.

2. The method according to claim 1, wherein, After outputting the multi-channel voice data with each channel representing a corresponding speaker, the method further includes: Determining whether the voice of the target speaker in the multi-channel voice data contains a wake-up word. If the wake-up word is included, activating and waking up the local device. If the voice of the interfering speaker contains the wake-up word, not activating and waking up the local device.

3. The method according to claim 1, wherein, The determination of whether each channel of voice data in the multi-channel voice data contains a wake-up word is performed in the local device, and includes: The local device receives the multi-channel voice data sent from the cloud server; Performing active voice detection on the multi-channel voice data, and only performing wake-up word judgment and verification on the channel voice data detected to contain human voice. If the wake-up word is included, activating and waking up the local device.

4. The method according to claim 1, wherein, The determination of whether each channel of voice data in the multi-channel voice data contains a wake-up word is performed in the cloud server, and includes: The cloud server performs wake-up word judgment and verification on the multi-channel voice data. If the wake-up word is included, sending the single-channel voice data containing the wake-up word and the verification result to the local device.

5. The method according to claim 4, wherein After the local device receives the single-channel voice data and the verification result, the method further includes: Determining whether the single-channel voice data is the voice of the target speaker. If it is the voice of the target speaker, activating and waking up the local device.

6. The method according to claim 1, wherein After determining whether there is a wake-up word in the noise-reduced voice through the audio features, if there is the wake-up word, the method further includes: According to the wake-up point time information, expanding a preset time forward and backward respectively, and intercepting a noise-reduced voice segment to send to the cloud server.

7. A voice wake-up system based on noise reduction and separation, comprising: A noise reduction module for performing noise reduction on the voice in a local device with limited computing power to obtain noise-reduced voice that eliminates environmental noise and retains the voice of the target speaker and the voice of the interfering speaker; A wake-up module for extracting the audio features of the noise-reduced voice on the local device with limited computing power, and determining whether there is a wake-up word in the noise-reduced voice through the audio features. If there is the wake-up word, sending the noise-reduced voice to the cloud server; A cloud separation module, configured to separate the noise-reduced speech according to the number of speakers in the cloud server, and output multi-channel speech data where each channel represents a corresponding speaker; A verification module, configured to determine whether the speech data of each channel contains a wake word from the multi-channel speech data, and if the wake word is included, activate and wake up the local device.

8. A storage medium having a computer program product stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method described in any one of claims 1-6.

9. A computer program product, which embeds instructions on a storage medium, and the instructions implement the steps of the method described in any one of claims 1-6.

10. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method described in any one of claims 1-6.