Multi-channel speech noise reduction method and system, electronic device and storage medium

By splitting the traditional multi-channel speech noise reduction model and using speech activity detection and sound source localization information to generate vectors, noise reduction of specified angles or speech segments of multi-channel speech is achieved, solving the problem that the existing model cannot adapt to environmental changes and improving the flexibility and adaptability of the model.

CN118865995BActive Publication Date: 2025-08-12MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411237494.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-08-12
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

Existing multi-channel speech noise reduction models cannot achieve accurate noise reduction for specified angles or specified speech segments, especially when the environment changes and the model needs to be retrained.

Method used

By acquiring multi-channel speech signals and converting them into frequency domain signals, a vector is generated using voice activity detection and sound source localization information, which is then input into a multi-channel speech noise reduction model to achieve target noise reduction. The VAD and DOA modules in the traditional model are split to meet user needs.

Benefits of technology

There is no need to retrain the model in any scenario. Only the voice activity detection and sound source localization information need to be adjusted to achieve precise noise reduction at specified angles or speech segments, which improves the flexibility and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118865995B_ABST
    Figure CN118865995B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-channel speech noise reduction method and system, electronic device and storage medium. The multi-channel speech noise reduction method includes: obtaining a multi-channel speech signal; processing the multi-channel speech signal to obtain a multi-channel frequency domain signal; determining first speech activity detection information; determining first sound source localization information; obtaining a first vector based on the first speech activity detection information; obtaining a second vector based on the first sound source localization information; inputting the first vector and the second vector into a multi-channel speech noise reduction model to obtain a target noise reduction model; inputting the multi-channel frequency domain signal into the target noise reduction model for noise reduction. In the present invention, the speech activity detection information and the sound source localization information are first determined, and then the target noise reduction model is obtained based on the speech activity detection information and the sound source localization information. Finally, the target noise reduction model is used to perform noise reduction, thereby achieving noise reduction of multi-channel speech under specified angle information or specified speech segments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multi-channel speech noise reduction method and system, an electronic device, and a storage medium. Background Art

[0002] With the development of deep neural networks, their use in multi-channel speech noise reduction is increasing. Existing multi-channel speech noise reduction models rely on enhancing speech signals. When multiple speech targets are present, the louder signal is enhanced. However, when specific angles or speech segments require targeted enhancement, traditional multi-channel speech noise reduction models are unable to achieve noise reduction for multi-channel speech. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art or related art.

[0004] To this end, a first aspect of the present invention provides a multi-channel speech noise reduction method.

[0005] A second aspect of the present invention provides a multi-channel speech noise reduction system.

[0006] A third aspect of the present invention provides an electronic device.

[0007] A fourth aspect of the present invention provides a storage medium.

[0008] A fifth aspect of the present invention provides a computer program product.

[0009] In view of this, according to a first aspect of the present invention, a method for denoising multi-channel speech is proposed, comprising: obtaining a multi-channel speech signal; processing the multi-channel speech signal to obtain a multi-channel frequency domain signal; determining first voice activity detection information, wherein the first voice activity detection information corresponds to the multi-channel speech signal; determining first sound source localization information, wherein the first sound source localization information corresponds to the multi-channel speech signal; obtaining a first vector based on the first voice activity detection information; obtaining a second vector based on the first sound source localization information, wherein the first vector and the second vector are vectors of the same dimension; inputting the first vector and the second vector into a multi-channel speech noise reduction model to obtain a target noise reduction model; and inputting the multi-channel frequency domain signal into the target noise reduction model for noise reduction.

[0010] The multi-channel speech noise reduction method provided by the present invention mainly includes: first obtaining a multi-channel speech signal to be denoised, wherein the multi-channel speech signal can be input by the user or collected and obtained by a data model. After obtaining the multi-channel speech signal, the multi-channel speech signal is processed to obtain a multi-channel frequency domain signal of the multi-channel speech signal. The multi-channel frequency domain signal includes a combination of different frequency components of the multi-channel speech signal in the frequency domain. By processing the multi-channel speech signal to obtain the multi-channel frequency domain signal, subsequent noise reduction of the multi-channel audio signal is facilitated. Then, first voice activity detection information and first sound source localization information are determined, wherein the first voice activity detection information and the first sound source localization information both correspond to the multi-channel speech signal, wherein the first voice activity detection information and the first sound source localization information can be determined based on the multi-channel speech signal or input by the user. The first voice activity detection information is obtained using voice activity detection (VAD) technology. VAD technology can automatically detect active (sound) and inactive (silent) portions of a speech signal. Therefore, the first voice activity detection information can also be referred to as VAD information. By determining the first voice activity detection information, it is possible to determine which segment of the multi-channel speech signal the user wishes to enhance and reduce noise for. The first sound source localization information is obtained using direction of arrival (DOA) technology. DOA technology can determine the direction of a sound source by analyzing the arrival time differences of sound signals in different microphone arrays. Therefore, the first sound source localization information can also be referred to as DOA information or angle information. By determining the first sound source localization information, it is possible to determine the direction of the speech signal in the multi-channel speech signal the user wishes to enhance and reduce noise for. After determining the first voice activity detection information and the first sound source localization information, the first voice activity detection information is converted into a first vector, and the first sound source localization information is converted into a second vector, where the first and second vectors are vectors of the same dimension. The first vector and the second vector are numerical vectors, and the first voice activity detection information and the first sound source localization information are converted into numerical vectors to adapt to different outputs. The conversion method may use embedding (a method of mapping high-dimensional data into a continuous vector representation in a low-dimensional space) technology.After obtaining the first vector and the second vector, the first vector and the second vector are input into the multi-channel speech noise reduction model, converting the multi-channel speech noise reduction model into a target noise reduction model. It can be understood that the first vector represents the segment of the multi-channel speech signal that the user wishes to enhance and reduce noise, and the second vector represents the direction of the speech signal in the multi-channel speech signal that the user wishes to enhance and reduce noise. Therefore, after inputting the first vector and the second vector into the multi-channel speech noise reduction model, a target noise reduction model that meets the user's noise reduction requirements is obtained. This means that when the environment changes, the multi-channel speech noise reduction model does not need to be retrained; only the first sound source localization information and the first voice activity detection information need to be changed. Finally, the multi-channel frequency domain signal is input into the target noise reduction model for noise reduction. The present invention decomposes the traditional multi-channel speech noise reduction model, moving the process of processing the multi-channel speech signal to obtain voice activity detection information and sound source localization information before the multi-channel speech signal is input into the multi-channel speech noise reduction model. This separates the VAD module, DOA module, and multi-channel speech noise reduction module in the traditional multi-channel speech noise reduction model. Furthermore, the VAD module and DOA module can not only determine the corresponding information from the multi-channel speech signal, but also determine user-specified VAD information and DOA information according to user requirements, thereby achieving the technical effect of noise reduction for multi-channel speech under specified angle information or specified speech segments. Furthermore, the present invention converts the first voice activity detection information into a first vector and the first sound source localization information into a second vector, and then inputs the first and second vectors into the multi-channel speech noise reduction model, thereby converting the multi-channel speech noise reduction model into a target noise reduction model that meets the user's requirements. This eliminates the need for retraining the multi-channel speech noise reduction model in any scenario, and only requires changing the first voice activity detection information and first sound source localization information according to the scenario.

[0011] The multi-channel speech noise reduction method according to the present invention may also have the following technical features:

[0012] In some technical solutions, optionally, the step of determining the first voice activity detection information includes: determining the second voice activity detection information based on the multi-channel frequency domain signal; determining whether third voice activity detection information is obtained, wherein the third voice activity detection information corresponds to the multi-channel voice signal; based on not obtaining the third voice activity detection information, using the second voice activity detection information as the first voice activity detection information; based on obtaining the third voice activity detection information, using the third voice activity detection information as the first voice activity detection information.

[0013] In this technical solution, the step of determining first voice activity detection information includes: first, determining second voice activity detection information based on a multi-channel frequency domain signal. The method may be to input the multi-channel frequency domain signal into a neural network model to obtain the second voice activity detection information, or to perform signal processing on the multi-channel frequency domain signal to obtain the second voice activity detection information. Then, determining whether third voice activity detection information has been acquired, wherein the third voice activity detection information corresponds to the multi-channel voice signal, i.e., the third voice activity detection information is user-specified voice activity detection information. By determining whether the third voice activity detection information has been acquired, it is possible to determine whether the user has specified voice activity detection information. If the third voice activity detection information has not been acquired, the second voice activity detection information is used as the first voice activity detection information. That is, if the user has not specified voice activity detection information, the multi-channel audio signal is denoised using the voice activity detection information determined in the multi-channel frequency domain signal. If the third voice activity detection information is acquired, the third voice activity detection information is used as the first voice activity detection information. That is, if the user has specified voice activity detection information, the multi-channel voice signal is denoised using the user-specified voice activity detection information. In the present invention, when the user has specified voice activity detection information, the user-specified voice activity detection information is used to reduce the noise of the multi-channel voice signal. When the user has not specified voice activity detection information, the voice activity detection information in the multi-channel voice signal is used to reduce the noise of the multi-channel voice signal. This achieves the technical effect of being able to reduce the noise of the multi-channel voice signal in any scenario.

[0014] In some technical solutions, optionally, the step of determining the second voice activity detection information based on the multi-channel frequency domain signal includes: determining a voice activity detection model, wherein the voice activity detection model is a neural network model; and inputting the multi-channel frequency domain signal into the voice activity detection model to obtain the second voice activity detection information.

[0015] In this technical solution, the step of determining second voice activity detection information based on a multi-channel frequency domain signal includes: first determining a voice activity detection model, wherein the voice activity detection model may be a neural network model, and the voice activity detection model is capable of detecting voice activity detection information, namely, automatically detecting active (sound) and inactive (silent) portions of the multi-channel voice signal. The multi-channel frequency domain signal is then input into the determined voice activity detection model to obtain the second voice activity detection information. By using the voice activity detection model to detect the second voice activity detection information, the accuracy of the second voice activity detection information is improved.

[0016] In some technical solutions, optionally, the step of determining a voice activity detection model includes: determining an application scenario; determining signal-to-noise ratio information and / or computing power information based on the application scenario; and determining a voice activity detection model based on the signal-to-noise ratio information and / or computing power information.

[0017] In this technical solution, the step of determining the voice activity detection model includes: first determining an application scenario, such as an application scenario such as household appliances, then determining the signal-to-noise ratio information and / or computing power information based on the application scenario, and then selecting the corresponding voice activity detection model based on the signal-to-noise ratio information and / or computing power information. For example, when there is sufficiently large computing power information in the current application scenario, it is necessary to select a voice activity detection model with large computing power. When the signal-to-noise ratio information is high in the current application scenario, a voice activity detection model with small computing power is selected.

[0018] In some technical solutions, optionally, the step of determining the first sound source localization information includes: determining the second sound source localization information based on the multi-channel frequency domain signal; determining whether third sound source localization information is acquired, wherein the third sound source localization information corresponds to the multi-channel speech signal; based on the acquisition of the third sound source localization information, using the third sound source localization information as the first sound source localization information; and based on the failure to acquire the third sound source localization information, using the second sound source localization information as the first sound source localization information.

[0019] In this technical solution, the step of determining first sound source localization information includes: first, determining second sound source localization information based on the multi-channel frequency domain signal. This method may be to input the multi-channel frequency domain signal into a neural network model to obtain the second sound source localization information, or to perform signal processing on the multi-channel frequency domain signal to obtain the second sound source localization information. Then, determining whether third sound source localization information has been obtained, wherein the third sound source localization information corresponds to the multi-channel speech signal, i.e., the third sound source localization information is user-specified sound source localization information. By determining whether the third sound source localization information has been obtained, it is possible to determine whether the user has specified sound source localization information. If the third sound source localization information has not been obtained, the second sound source localization information is used as the first sound source localization information. That is, if the user has not specified sound source localization information, the multi-channel audio signal is denoised using the sound source localization information determined in the multi-channel frequency domain signal. If the third sound source localization information has been obtained, the third sound source localization information is used as the first sound source localization information. That is, if the user has specified sound source localization information, the multi-channel speech signal is denoised using the user-specified sound source localization information. In the present invention, when the user has specified sound source localization information, the noise of the multi-channel speech signal is reduced by using the sound source localization information specified by the user. When the user does not specify the sound source localization information, the noise of the multi-channel speech signal is reduced by using the sound source localization information in the multi-channel speech signal, thereby achieving the technical effect of being able to reduce the noise of the multi-channel speech signal in any scenario.

[0020] In some technical solutions, optionally, the step of determining the second sound source localization information based on the multi-channel frequency domain signal includes: determining a sound source localization model, wherein the sound source localization model is a neural network model; inputting the multi-channel frequency domain signal into the sound source localization model to obtain the second sound source localization information.

[0021] In this technical solution, the step of determining second sound source localization information based on the multi-channel frequency domain signal includes: first, determining a sound source localization model, which can be a neural network model. The sound source localization model is capable of detecting sound source localization information by analyzing the arrival time differences of speech signals at different microphone arrays to determine the direction of the sound source. The multi-channel frequency domain signal is then input into the determined sound source localization model to obtain the second sound source localization information. By using the sound source localization model to detect the second sound source localization information, the accuracy of the second sound source localization information is improved.

[0022] In some technical solutions, optionally, the step of determining the sound source localization model includes: determining an application scenario; determining signal-to-noise ratio information and / or computing power information based on the application scenario; and determining the sound source localization model based on the signal-to-noise ratio information and / or computing power information.

[0023] In this technical solution, the step of determining the sound source localization model includes: first determining an application scenario, such as an application scenario such as household appliances, then determining the signal-to-noise ratio information and / or computing power information based on the application scenario, and then selecting the corresponding sound source localization model based on the signal-to-noise ratio information and / or computing power information. For example, when there is sufficiently large computing power information in the current application scenario, it is necessary to select a sound source localization model with large computing power. When the signal-to-noise ratio information is high in the current application scenario, a sound source localization model with small computing power is selected.

[0024] In some technical solutions, optionally, the step of determining the first sound source localization information further includes: acquiring image information, wherein the image information is related to the multi-channel speech signal; and determining the first sound source localization information based on the image information.

[0025] In this technical solution, the step of determining the first sound source localization information further includes: first obtaining image information, where the image information is related to the multi-channel voice signal. For example, the image information may be an image of the user transmitting the multi-channel voice signal. The first sound source localization information is determined by identifying and analyzing the image information and analyzing the angle of the user's mouth in the image information. Furthermore, the first voice activity detection information may be determined by analyzing the user's lip movements in the image information.

[0026] In some technical solutions, optionally, the step of processing the multi-channel speech signal to obtain a multi-channel frequency domain signal includes: determining a multi-channel time domain signal based on the multi-channel speech signal, wherein the multi-channel time domain signal corresponds to the multi-channel speech signal; and determining a multi-channel frequency domain signal based on the multi-channel time domain signal.

[0027] In this technical solution, the step of processing a multi-channel voice signal to obtain a multi-channel frequency domain signal includes: first, performing signal processing on the multi-channel voice signal to obtain a multi-channel time domain signal, wherein the multi-channel time domain signal can be obtained by sampling and quantizing the multi-channel voice signal, and the multi-channel time domain signal is a digital representation of the multi-channel voice signal. By converting the multi-channel voice signal into a multi-channel time domain signal, a basis is provided for subsequent processing and analysis of the multi-channel voice signal. After obtaining the multi-channel time domain signal, a multi-channel frequency domain signal is determined based on the multi-channel time domain signal, wherein the multi-channel time domain signal can be converted into a multi-channel frequency domain signal by Fourier transform. By determining the multi-channel frequency domain signal, it is convenient to perform noise reduction on the multi-channel voice signal.

[0028] According to a second aspect of the present invention, a multi-channel speech noise reduction system is proposed, wherein the multi-channel speech noise reduction system includes: a first acquisition module, the first acquisition module is used to acquire a multi-channel speech signal; a first processing module, the first processing module is used to process the multi-channel speech signal to obtain a multi-channel frequency domain signal; a first determination module, the first determination module is used to determine first voice activity detection information, wherein the first voice activity detection information corresponds to the multi-channel speech signal; a second determination module, the second determination module is used to determine first sound source localization information, wherein the first sound source localization information corresponds to the multi-channel speech signal; a second processing module, the second processing module is used to obtain a first vector based on the first voice activity detection information; a third processing module, the third processing module is used to obtain a second vector based on the first sound source localization information, wherein the first vector and the second vector are vectors of the same dimension; a first input module, the first input module is used to input the first vector and the second vector into a multi-channel speech noise reduction model to obtain a target noise reduction model; and a second input module, the second input module is used to input the multi-channel frequency domain signal into the target noise reduction model for noise reduction.

[0029] The multi-channel speech noise reduction system provided by the present invention mainly includes: a first acquisition module, a first processing module, a first determination module, a second determination module, a second processing module, a third processing module, a first input module, and a second input module. The first acquisition module can acquire a multi-channel speech signal to be denoised, wherein the multi-channel speech signal can be input by a user or acquired by a data model. After obtaining the multi-channel speech signal, the first processing module processes the multi-channel speech signal to obtain a multi-channel frequency domain signal of the multi-channel speech signal. The multi-channel frequency domain signal includes a combination of different frequency components of the multi-channel speech signal in the frequency domain. By processing the multi-channel speech signal to obtain the multi-channel frequency domain signal, subsequent noise reduction of the multi-channel audio signal is facilitated. The first determination module then determines first voice activity detection information, and the second determination module determines first sound source localization information. The first voice activity detection information and the first sound source localization information both correspond to the multi-channel speech signal. The first voice activity detection information and the first sound source localization information can be determined based on the multi-channel speech signal or input by the user. The first voice activity detection information is obtained using voice activity detection (VAD) technology. VAD technology can automatically detect active (sound) and inactive (silent) portions of a speech signal. Therefore, the first voice activity detection information can also be referred to as VAD information. By determining the first voice activity detection information, it is possible to determine which segment of the multi-channel speech signal the user wishes to enhance and reduce noise. The first sound source localization information is obtained using direction of arrival (DOA) technology. DOA technology can determine the direction of a sound source by analyzing the arrival time differences of sound signals in different microphone arrays. Therefore, the first sound source localization information can also be referred to as DOA information or angle information. By determining the first sound source localization information, it is possible to determine the direction of the speech signal in the multi-channel speech signal the user wishes to enhance and reduce noise. After determining the first voice activity detection information and the first sound source localization information, the second processing module converts the first voice activity detection information into a first vector, and the third processing module converts the first sound source localization information into a second vector, where the first and second vectors are vectors of the same dimension. The first vector and the second vector are numerical vectors, and the first voice activity detection information and the first sound source localization information are converted into numerical vectors to adapt to different outputs. The conversion method may use embedding (a method of mapping high-dimensional data into a continuous vector representation in a low-dimensional space) technology.After obtaining the first vector and the second vector, the first input module inputs the first vector and the second vector into the multi-channel speech noise reduction model, converting the multi-channel speech noise reduction model into a target noise reduction model. It can be understood that the first vector represents which segment of the multi-channel speech signal the user wishes to enhance and reduce noise, and the second vector represents which direction of the speech signal in the multi-channel speech signal the user wishes to enhance and reduce noise. Therefore, after inputting the first vector and the second vector into the multi-channel speech noise reduction model, a target noise reduction model that meets the user's noise reduction requirements will be obtained. Therefore, when the environment changes, the multi-channel speech noise reduction model does not need to be retrained, and only the first sound source localization information and the first voice activity detection information need to be changed. Finally, the second input module inputs the multi-channel frequency domain signal into the target noise reduction model for noise reduction. The present invention splits the traditional multi-channel speech noise reduction model, moving the process of processing the multi-channel speech signal to obtain voice activity detection information and sound source localization information before the multi-channel speech signal is input into the multi-channel speech noise reduction model. This separates the VAD module, DOA module, and multi-channel speech noise reduction module in the traditional multi-channel speech noise reduction model. Furthermore, the VAD module and DOA module can not only determine the corresponding information from the multi-channel speech signal, but also determine user-specified VAD information and DOA information according to user requirements, thereby achieving the technical effect of noise reduction for multi-channel speech under specified angle information or specified speech segments. Furthermore, the present invention converts the first voice activity detection information into a first vector and the first sound source localization information into a second vector, and then inputs the first and second vectors into the multi-channel speech noise reduction model, thereby converting the multi-channel speech noise reduction model into a target noise reduction model that meets the user's requirements. This eliminates the need for retraining the multi-channel speech noise reduction model in any scenario, requiring only the first voice activity detection information and the first sound source localization information to be modified according to the scenario.

[0030] According to a third aspect of the present invention, an electronic device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the multi-channel speech noise reduction method as described above are implemented.

[0031] The electronic device provided by the present invention, in which the processor implements the steps of the above-mentioned multi-channel speech noise reduction method when executing the computer program, can achieve the technical effects of any of the above-mentioned technical solutions, and will not be repeated here.

[0032] According to a fourth aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned multi-channel speech noise reduction methods are implemented.

[0033] The storage medium provided by the present invention and the computer program, when executed by a processor, implement the steps of the above-mentioned multi-channel speech noise reduction method, which can achieve the technical effects of any of the above-mentioned technical solutions and will not be repeated here.

[0034] According to a fifth aspect of the present invention, a computer program product is proposed, comprising a computer program, which, when executed by a processor, implements the steps of the multi-channel speech noise reduction method in any of the above technical solutions.

[0035] The computer program product provided by this technical solution implements the steps of the multi-channel speech noise reduction method as in any technical solution of the present invention, and thus has all the beneficial effects of the multi-channel speech noise reduction method as in any technical solution of the present invention, which will not be repeated here.

[0036] Additional aspects and advantages of the invention will become apparent from the description which follows, or may be learned by practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0038] Figure 1 A schematic flow chart showing a method for reducing noise of multi-channel speech according to an embodiment of the present invention is shown;

[0039] Figure 2 A schematic flow chart showing a step of determining first voice activity detection information in a multi-channel speech noise reduction method according to an embodiment of the present invention;

[0040] Figure 3 A schematic flow chart illustrating a step of determining second voice activity detection information based on multi-channel frequency domain signals in a multi-channel speech noise reduction method according to an embodiment of the present invention;

[0041] Figure 4 A schematic flow chart showing the steps of determining a voice activity detection model in a multi-channel speech noise reduction method according to an embodiment of the present invention is shown;

[0042] Figure 5 A flowchart illustrating a step of determining first sound source localization information in a multi-channel speech noise reduction method according to an embodiment of the present invention is shown;

[0043] Figure 6 A schematic flow chart illustrating the step of determining second sound source localization information based on multi-channel frequency domain signals in a multi-channel speech noise reduction method according to an embodiment of the present invention is shown;

[0044] Figure 7A schematic flow chart showing the steps of determining a sound source localization model in a multi-channel speech noise reduction method according to an embodiment of the present invention is shown;

[0045] Figure 8 A second flow chart illustrating the step of determining first sound source localization information in the multi-channel speech noise reduction method according to an embodiment of the present invention is shown;

[0046] Figure 9 A schematic flow chart showing the steps of processing a multi-channel speech signal to obtain a multi-channel frequency domain signal in a multi-channel speech noise reduction method according to an embodiment of the present invention;

[0047] Figure 10 A schematic diagram showing the principle of a multi-channel speech noise reduction method according to an embodiment of the present invention is shown;

[0048] Figure 11 A schematic block diagram of a multi-channel speech noise reduction system according to an embodiment of the present invention is shown;

[0049] Figure 12 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0050] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features therein can be combined with each other without conflict.

[0051] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0052] Figure 1 A flow chart of a multi-channel speech noise reduction method according to an embodiment of the present invention is shown. The method includes:

[0053] Step 102: Acquire a multi-channel speech signal;

[0054] Step 104: Process the multi-channel speech signal to obtain a multi-channel frequency domain signal;

[0055] Step 106: Determine first voice activity detection information, wherein the first voice activity detection information corresponds to the multi-channel voice signal;

[0056] Step 108: determining first sound source localization information, wherein the first sound source localization information corresponds to the multi-channel speech signal;

[0057] Step 110: Obtain a first vector according to the first voice activity detection information;

[0058] Step 112: Obtain a second vector according to the first sound source localization information, wherein the first vector and the second vector are vectors of the same dimension;

[0059] Step 114: Input the first vector and the second vector into a multi-channel speech noise reduction model to obtain a target noise reduction model;

[0060] Step 116: Input the multi-channel frequency domain signal into the target noise reduction model for noise reduction.

[0061] The multi-channel speech noise reduction method provided by the present invention mainly includes: first obtaining a multi-channel speech signal to be denoised, wherein the multi-channel speech signal can be input by the user or collected and obtained by a data model. After obtaining the multi-channel speech signal, the multi-channel speech signal is processed to obtain a multi-channel frequency domain signal of the multi-channel speech signal. The multi-channel frequency domain signal includes a combination of different frequency components of the multi-channel speech signal in the frequency domain. By processing the multi-channel speech signal to obtain the multi-channel frequency domain signal, subsequent noise reduction of the multi-channel audio signal is facilitated. Then, first voice activity detection information and first sound source localization information are determined, wherein the first voice activity detection information and the first sound source localization information both correspond to the multi-channel speech signal, wherein the first voice activity detection information and the first sound source localization information can be determined based on the multi-channel speech signal or input by the user. The first voice activity detection information is obtained using voice activity detection (VAD) technology. VAD technology can automatically detect active (sound) and inactive (silent) portions of a speech signal. Therefore, the first voice activity detection information can also be referred to as VAD information. By determining the first voice activity detection information, it is possible to determine which segment of the multi-channel speech signal the user wishes to enhance and reduce noise for. The first sound source localization information is obtained using direction of arrival (DOA) technology. DOA technology can determine the direction of a sound source by analyzing the arrival time differences of sound signals in different microphone arrays. Therefore, the first sound source localization information can also be referred to as DOA information or angle information. By determining the first sound source localization information, it is possible to determine the direction of the speech signal in the multi-channel speech signal the user wishes to enhance and reduce noise for. After determining the first voice activity detection information and the first sound source localization information, the first voice activity detection information is converted into a first vector, and the first sound source localization information is converted into a second vector, where the first and second vectors are vectors of the same dimension. The first vector and the second vector are numerical vectors, and the first voice activity detection information and the first sound source localization information are converted into numerical vectors to adapt to different outputs. The conversion method may use embedding (a method of mapping high-dimensional data into a continuous vector representation in a low-dimensional space) technology.After obtaining the first vector and the second vector, the first vector and the second vector are input into the multi-channel speech noise reduction model, converting the multi-channel speech noise reduction model into a target noise reduction model. It can be understood that the first vector represents the segment of the multi-channel speech signal that the user wishes to enhance and reduce noise, and the second vector represents the direction of the speech signal in the multi-channel speech signal that the user wishes to enhance and reduce noise. Therefore, after inputting the first vector and the second vector into the multi-channel speech noise reduction model, a target noise reduction model that meets the user's noise reduction requirements is obtained. This means that when the environment changes, the multi-channel speech noise reduction model does not need to be retrained; only the first sound source localization information and the first voice activity detection information need to be changed. Finally, the multi-channel frequency domain signal is input into the target noise reduction model for noise reduction. The present invention decomposes the traditional multi-channel speech noise reduction model, moving the process of processing the multi-channel speech signal to obtain voice activity detection information and sound source localization information before the multi-channel speech signal is input into the multi-channel speech noise reduction model. This separates the VAD module, DOA module, and multi-channel speech noise reduction module in the traditional multi-channel speech noise reduction model. The VAD module and DOA module can not only determine the corresponding information from the multi-channel speech signal but also determine user-specified VAD information and DOA information according to user requirements, thereby achieving the technical effect of noise reduction for multi-channel speech under specified angle information or specified speech segments. Furthermore, the present invention converts the first voice activity detection information into a first vector and the first sound source localization information into a second vector, and then inputs the first and second vectors into the multi-channel speech noise reduction model. This converts the multi-channel speech noise reduction model into a target noise reduction model that meets the user's requirements. This eliminates the need for retraining the multi-channel speech noise reduction model in any scenario; instead, the first voice activity detection information and the first sound source localization information need to be modified according to the scenario.

[0062] Figure 2 A schematic flow chart illustrating a step of determining first voice activity detection information in a multi-channel speech noise reduction method according to an embodiment of the present invention is provided. The step of determining the first voice activity detection information includes:

[0063] Step 202: Determine second voice activity detection information based on the multi-channel frequency domain signal;

[0064] Step 204: Determine whether third voice activity detection information is acquired, wherein the third voice activity detection information corresponds to the multi-channel voice signal;

[0065] Step 206: Based on the failure to obtain the third voice activity detection information, the second voice activity detection information is used as the first voice activity detection information;

[0066] Step 208: Based on the acquired third voice activity detection information, the third voice activity detection information is used as the first voice activity detection information.

[0067] In this embodiment, the step of determining first voice activity detection information includes: first, determining second voice activity detection information based on the multi-channel frequency domain signal. This method may be to input the multi-channel frequency domain signal into a neural network model to obtain the second voice activity detection information, or to perform signal processing on the multi-channel frequency domain signal to obtain the second voice activity detection information. Then, determining whether third voice activity detection information has been acquired, wherein the third voice activity detection information corresponds to the multi-channel voice signal, i.e., the third voice activity detection information is user-specified voice activity detection information. By determining whether the third voice activity detection information has been acquired, it is possible to determine whether the user has specified voice activity detection information. If the third voice activity detection information has not been acquired, the second voice activity detection information is used as the first voice activity detection information. That is, if the user has not specified voice activity detection information, the multi-channel audio signal is denoised using the voice activity detection information determined from the multi-channel frequency domain signal. If the third voice activity detection information is acquired, the third voice activity detection information is used as the first voice activity detection information. That is, if the user has specified voice activity detection information, the multi-channel voice signal is denoised using the user-specified voice activity detection information. In the present invention, when the user has specified voice activity detection information, the user-specified voice activity detection information is used to reduce the noise of the multi-channel voice signal. When the user has not specified voice activity detection information, the voice activity detection information in the multi-channel voice signal is used to reduce the noise of the multi-channel voice signal. This achieves the technical effect of being able to reduce the noise of the multi-channel voice signal in any scenario.

[0068] Figure 3 A schematic flow chart illustrating the step of determining second voice activity detection information based on multi-channel frequency domain signals in a multi-channel speech noise reduction method according to an embodiment of the present invention is provided. The step of determining the second voice activity detection information based on the multi-channel frequency domain signals includes:

[0069] Step 302: Determine a voice activity detection model, wherein the voice activity detection model is a neural network model;

[0070] Step 304: Input the multi-channel frequency domain signal into a voice activity detection model to obtain second voice activity detection information.

[0071] In this embodiment, the step of determining second voice activity detection information based on the multi-channel frequency domain signal includes: first determining a voice activity detection model, wherein the voice activity detection model may be a neural network model, and the voice activity detection model is capable of detecting voice activity detection information, namely, automatically detecting active (sound) and inactive (silent) portions in the multi-channel voice signal. The multi-channel frequency domain signal is then input into the determined voice activity detection model to obtain the second voice activity detection information. By using the voice activity detection model to detect the second voice activity detection information, the accuracy of the second voice activity detection information is improved.

[0072] Figure 4 A flowchart illustrating the step of determining a voice activity detection model in a multi-channel speech noise reduction method according to an embodiment of the present invention is shown. The step of determining the voice activity detection model includes:

[0073] Step 402: Determine the application scenario;

[0074] Step 404: Determine signal-to-noise ratio information and / or computing power information according to the application scenario;

[0075] Step 406: Determine a voice activity detection model based on the signal-to-noise ratio information and / or the computing power information.

[0076] In this embodiment, the step of determining a voice activity detection model includes: first determining an application scenario, such as an application scenario for household appliances, then determining signal-to-noise ratio information and / or computing power information based on the application scenario, and then selecting a corresponding voice activity detection model based on the signal-to-noise ratio information and / or computing power information. For example, when the computing power information is sufficiently large in the current application scenario, a voice activity detection model with high computing power needs to be selected; when the signal-to-noise ratio information is high in the current application scenario, a voice activity detection model with low computing power needs to be selected. For example, in a scenario with high noise but sufficient computing power (such as a vacuum cleaner or range hood), more accurate VAD and DOA information is required. In this case, a sound source localization model and voice activity detection model with higher computing power can be switched to a more effective and efficient sound source localization model and voice activity detection model, but the core multi-channel voice noise reduction model does not need to be retrained. In the case of a high signal-to-noise ratio and limited computing power, a sound source localization model and voice activity detection model with lower computing power can be switched to a less efficient sound source localization model and voice activity detection model, but the core multi-channel voice noise reduction model does not need to be retrained.

[0077] Figure 5 A flowchart of a step of determining first sound source localization information in a multi-channel speech noise reduction method according to an embodiment of the present invention is shown; wherein the step of determining the first sound source localization information includes:

[0078] Step 502: Determine second sound source localization information based on the multi-channel frequency domain signal;

[0079] Step 504: Determine whether third sound source localization information is obtained, wherein the third sound source localization information corresponds to the multi-channel speech signal;

[0080] Step 506: Based on the acquired third sound source localization information, the third sound source localization information is used as the first sound source localization information;

[0081] Step 508: Based on the failure to obtain the third sound source localization information, the second sound source localization information is used as the first sound source localization information.

[0082] In this embodiment, the step of determining first sound source localization information includes: first, determining second sound source localization information based on the multi-channel frequency domain signal. This method may be to input the multi-channel frequency domain signal into a neural network model to obtain the second sound source localization information, or to perform signal processing on the multi-channel frequency domain signal to obtain the second sound source localization information. Then, determining whether third sound source localization information has been obtained, wherein the third sound source localization information corresponds to the multi-channel speech signal, i.e., the third sound source localization information is user-specified sound source localization information. By determining whether the third sound source localization information has been obtained, it is possible to determine whether the user has specified sound source localization information. If the third sound source localization information has not been obtained, the second sound source localization information is used as the first sound source localization information. That is, if the user has not specified sound source localization information, the multi-channel audio signal is denoised using the sound source localization information determined in the multi-channel frequency domain signal. If the third sound source localization information has been obtained, the third sound source localization information is used as the first sound source localization information. That is, if the user has specified sound source localization information, the multi-channel speech signal is denoised using the user-specified sound source localization information. In the present invention, when the user has specified sound source localization information, the noise of the multi-channel speech signal is reduced by using the sound source localization information specified by the user. When the user does not specify the sound source localization information, the noise of the multi-channel speech signal is reduced by using the sound source localization information in the multi-channel speech signal, thereby achieving the technical effect of being able to reduce the noise of the multi-channel speech signal in any scenario.

[0083] Figure 6 A schematic flow chart illustrating the step of determining second sound source localization information based on multi-channel frequency domain signals in a multi-channel speech noise reduction method according to an embodiment of the present invention is provided. The step of determining second sound source localization information based on multi-channel frequency domain signals includes:

[0084] Step 602: Determine a sound source localization model, wherein the sound source localization model is a neural network model;

[0085] Step 604: Input the multi-channel frequency domain signal into the sound source localization model to obtain second sound source localization information.

[0086] In this embodiment, the step of determining second sound source localization information based on the multi-channel frequency domain signal includes: first determining a sound source localization model, wherein the sound source localization model may be a neural network model. The sound source localization model is capable of detecting sound source localization information by analyzing the arrival time differences of speech signals at different microphone arrays to determine the direction of the sound source. The multi-channel frequency domain signal is then input into the determined sound source localization model to obtain the second sound source localization information. By using the sound source localization model to detect the second sound source localization information, the accuracy of the second sound source localization information is improved.

[0087] Figure 7 A schematic flow chart illustrating the steps of determining a sound source localization model in a multi-channel speech noise reduction method according to an embodiment of the present invention is provided. The steps of determining the sound source localization model include:

[0088] Step 702: Determine the application scenario;

[0089] Step 704: Determine signal-to-noise ratio information and / or computing power information according to the application scenario;

[0090] Step 706: Determine a sound source localization model based on the signal-to-noise ratio information and / or the computing power information.

[0091] In this embodiment, the step of determining the sound source localization model includes: first determining an application scenario, such as an application scenario such as household appliances, then determining the signal-to-noise ratio information and / or computing power information based on the application scenario, and then selecting the corresponding sound source localization model based on the signal-to-noise ratio information and / or computing power information. For example, when there is sufficient computing power information in the current application scenario, it is necessary to select a sound source localization model with high computing power. When the signal-to-noise ratio information is high in the current application scenario, a sound source localization model with low computing power is selected. For example, in a scene with high noise but sufficient computing power (sweeping machine, range hood), more accurate VAD information and DOA information are needed. At this time, a sound source localization model and voice activity detection model with better effect and higher computing power can be switched, but the core multi-channel speech noise reduction model does not need to be retrained. In the case of a relatively high signal-to-noise ratio and tight computing power resources, we can switch to a sound source localization model and voice activity detection model with lower computing power, but the core multi-channel speech noise reduction model can also be retrained.

[0092] Figure 8 A second flow chart illustrating the step of determining first sound source localization information in the multi-channel speech noise reduction method according to an embodiment of the present invention is provided. The step of determining the first sound source localization information further includes:

[0093] Step 802: Acquire image information, wherein the image information is related to the multi-channel speech signal;

[0094] Step 804: Determine first sound source localization information according to the image information.

[0095] In this embodiment, the step of determining the first sound source localization information further includes: first obtaining image information, where the image information is related to the multi-channel voice signal. For example, the image information may be an image of a user transmitting the multi-channel voice signal. The first sound source localization information is determined by identifying and analyzing the image information and analyzing the angle of the user's mouth in the image information. Furthermore, the first voice activity detection information may be determined by analyzing the user's lip movements in the image information. In other words, the first voice activity detection information is determined based on the image information.

[0096] Figure 9 A flow chart illustrating the steps of processing a multi-channel speech signal to obtain a multi-channel frequency domain signal in a multi-channel speech noise reduction method according to an embodiment of the present invention is shown; wherein the steps of processing the multi-channel speech signal to obtain a multi-channel frequency domain signal include:

[0097] Step 902: Determine a multi-channel time domain signal according to the multi-channel speech signal, wherein the multi-channel time domain signal corresponds to the multi-channel speech signal;

[0098] Step 904: Determine a multi-channel frequency domain signal according to the multi-channel time domain signal.

[0099] In this embodiment, the step of processing a multi-channel voice signal to obtain a multi-channel frequency domain signal includes: first, performing signal processing on the multi-channel voice signal to obtain a multi-channel time domain signal, wherein the multi-channel time domain signal can be obtained by sampling and quantizing the multi-channel voice signal, and the multi-channel time domain signal is a digital representation of the multi-channel voice signal. By converting the multi-channel voice signal into a multi-channel time domain signal, a basis is provided for subsequent processing and analysis of the multi-channel voice signal. After obtaining the multi-channel time domain signal, a multi-channel frequency domain signal is determined based on the multi-channel time domain signal, wherein the multi-channel time domain signal can be converted into a multi-channel frequency domain signal by Fourier transform. By determining the multi-channel frequency domain signal, noise reduction of the multi-channel voice signal is facilitated.

[0100] Figure 10 The schematic diagram of the principle of the multi-channel speech noise reduction method of an embodiment of the present invention is shown. In the present invention, the traditional integrated multi-channel speech noise reduction model is divided into multiple modules for training. And using the embedding method, the VAD information and angle information, i.e., DOA information, are converted into embeddings (vectors of uniform dimension) and input into the multi-channel speech noise reduction model. Figure 10As shown, the multi-channel speech signal is first converted into a multi-channel time domain signal, and then converted into a multi-channel frequency domain signal based on the multi-channel time domain signal, and the multi-channel frequency domain signal is respectively input into the VAD module and the VAD module. By inputting the multi-channel frequency domain signal into the VAD module, it is possible to detect which segment of the multi-channel speech signal is the audio of human speech, that is, to obtain VAD information. The VAD module can be a neural network module or a signal processing module. After obtaining the VAD information through the VAD module, it is necessary to determine whether there is specified VAD information, wherein the specified VAD information refers to the VAD information input by the user, rather than the VAD information determined by the multi-channel audio signal. If the specified VAD information exists, the specified VAD information is converted into a numerical vector and input into the multi-channel speech noise reduction model, thereby completing VAD embedding. If the specified VAD information does not exist, the VAD information obtained by the VAD module is converted into a numerical vector and input into the multi-channel speech noise reduction model, thereby completing VAD embedding. In the DOA module, by inputting a multi-channel frequency domain signal, the strongest angle information of the current speech, namely the DOA information, can be calculated, where the angle range is 0 degrees to 360 degrees. The module can be a neural network module or a signal processing module. After obtaining the DOA information through the DOA module, it is also necessary to determine whether there is specified angle information, where the specified angle information refers to the angle information input by the user, rather than the angle information determined by the multi-channel audio signal. If the specified angle information exists, the specified angle information is converted into a numerical vector and input into the multi-channel speech noise reduction model, thereby completing the DOA embedding. If the specified angle information does not exist, the DOA information obtained through the DOA module is converted into a numerical vector and input into the multi-channel speech noise reduction model, thereby completing the DOA embedding. After completing the DOA embedding and VAD embedding, a multi-channel speech noise reduction model with VAD embedding and DOA embedding is obtained, and then the multi-channel frequency domain signal is input into it to obtain the enhanced and denoised single-channel speech signal.

[0101] For example, in the process of multi-channel voice signal noise reduction, if the application scenario is a scene with a camera, the user-specified angle information and the changes in the lips in the image can be used to calculate the user-specified VAD information, and then the specified angle information and the specified VAD information are sent to the multi-channel voice noise reduction model for noise reduction. When there is voiceprint information, the VAD model of the voiceprint information can be used to obtain the signal of the specified voice segment, that is, the specified VAD information, and then the specified VAD information is sent to the multi-channel voice noise reduction model for noise reduction. At the same time, there are two scenarios in voice interaction: wake-up and recognition. It is generally believed that the angle of a person will not change after wake-up. Then in the subsequent noise reduction, the angle information after wake-up can be fixed, that is, the angle information after wake-up can be used as the specified angle information, and the specified angle information is sent to the multi-channel voice noise reduction model for noise reduction.

[0102] Figure 11 A schematic block diagram of a multi-channel speech noise reduction system 1100 according to an embodiment of the present invention is shown; wherein the multi-channel speech noise reduction system 1100 includes:

[0103] A first acquisition module 1102 is configured to acquire a multi-channel speech signal;

[0104] A first processing module 1104 is configured to process the multi-channel speech signal to obtain a multi-channel frequency domain signal;

[0105] a first determining module 1106, configured to determine first voice activity detection information, wherein the first voice activity detection information corresponds to the multi-channel voice signal;

[0106] a second determining module 1108, configured to determine first sound source localization information, wherein the first sound source localization information corresponds to the multi-channel speech signal;

[0107] A second processing module 1110 is configured to obtain a first vector according to the first voice activity detection information;

[0108] A third processing module 1112 is configured to obtain a second vector according to the first sound source localization information, wherein the first vector and the second vector are vectors of the same dimension;

[0109] A first input module 1114 is configured to input the first vector and the second vector into a multi-channel speech noise reduction model to obtain a target noise reduction model;

[0110] The second input module 1116 is configured to input the multi-channel frequency domain signal into the target noise reduction model for noise reduction.

[0111] The multi-channel speech noise reduction system 1100 provided by the present invention mainly includes: a first acquisition module 1102, a first processing module 1104, a first determination module 1106, a second determination module 1108, a second processing module 1110, a third processing module 1112, a first input module 1114 and a second input module 1116. Among them, the first acquisition module 1102 can obtain a multi-channel speech signal that needs to be denoised, wherein the multi-channel speech signal can be input by a user or collected and obtained by a data model. After obtaining the multi-channel speech signal, the first processing module 1104 processes the multi-channel speech signal to obtain a multi-channel frequency domain signal of the multi-channel speech signal. The multi-channel frequency domain signal includes a combination of different frequency components of the multi-channel speech signal in the frequency domain. By processing the multi-channel speech signal to obtain the multi-channel frequency domain signal, it is convenient to subsequently perform noise reduction on the multi-channel audio signal. Then, the first determination module 1106 determines first voice activity detection information, and the second determination module 1108 determines first sound source localization information. Both the first voice activity detection information and the first sound source localization information correspond to the multi-channel speech signal. The first voice activity detection information and the first sound source localization information can be determined based on the multi-channel speech signal or input by the user. The first voice activity detection information is obtained using voice activity detection (VAD) technology, which automatically detects active (sound) and inactive (silent) portions of a speech signal. Therefore, the first voice activity detection information can also be referred to as VAD information. By determining the first voice activity detection information, it can be determined which segment of the multi-channel speech signal the user desires to enhance and reduce noise. The first sound source localization information is obtained using DOA (Direction of Arrival) technology. DOA technology can determine the direction of the sound source by analyzing the arrival time differences of sound signals in different microphone arrays. Therefore, the first sound source localization information can also be called DOA information or angle information. By determining the first sound source localization information, it can be determined which direction of the multi-channel speech signal the user wishes to enhance and reduce noise. After determining the first voice activity detection information and the first sound source localization information, the second processing module 1110 converts the first voice activity detection information into a first vector, and the third processing module 1112 converts the first sound source localization information into a second vector, where the first vector and the second vector are vectors of the same dimension. The first vector and the second vector are numerical vectors. By converting the first voice activity detection information and the first sound source localization information into numerical vectors, different outputs can be adapted. The conversion method can use embedding (a continuous vector representation that maps high-dimensional data into a low-dimensional space) technology.After obtaining the first vector and the second vector, the first input module 1114 inputs the first vector and the second vector into the multi-channel speech noise reduction model, converting the multi-channel speech noise reduction model into a target noise reduction model. It is understood that the first vector represents which segment of the multi-channel speech signal the user wishes to enhance and reduce noise, and the second vector represents which direction of the multi-channel speech signal the user wishes to enhance and reduce noise. Therefore, after inputting the first vector and the second vector into the multi-channel speech noise reduction model, a target noise reduction model that meets the user's noise reduction requirements is obtained. This means that when the environment changes, the multi-channel speech noise reduction model does not need to be retrained; only the first sound source localization information and the first voice activity detection information need to be changed. Finally, the second input module 1116 inputs the multi-channel frequency domain signal into the target noise reduction model for noise reduction. The present invention decomposes the traditional multi-channel speech noise reduction model, moving the process of processing the multi-channel speech signal to obtain voice activity detection information and sound source localization information before the multi-channel speech signal is input into the multi-channel speech noise reduction model. This separates the VAD module, DOA module, and multi-channel speech noise reduction module in the traditional multi-channel speech noise reduction model. The VAD module and DOA module can not only determine the corresponding information from the multi-channel speech signal but also determine user-specified VAD information and DOA information according to user requirements, thereby achieving the technical effect of noise reduction for multi-channel speech under specified angle information or specified speech segments. Furthermore, the present invention converts the first voice activity detection information into a first vector and the first sound source localization information into a second vector, and then inputs the first and second vectors into the multi-channel speech noise reduction model. This converts the multi-channel speech noise reduction model into a target noise reduction model that meets the user's requirements. This eliminates the need for retraining the multi-channel speech noise reduction model in any scenario; instead, the first voice activity detection information and the first sound source localization information need to be modified according to the scenario.

[0112] In some embodiments, optionally, the first determination module 1106 is used to determine second voice activity detection information based on the multi-channel frequency domain signal; determine whether third voice activity detection information is obtained, wherein the third voice activity detection information corresponds to the multi-channel voice signal; based on not obtaining the third voice activity detection information, use the second voice activity detection information as the first voice activity detection information; based on obtaining the third voice activity detection information, use the third voice activity detection information as the first voice activity detection information.

[0113] In this embodiment, the first determination module 1106 is configured to determine second voice activity detection information based on the multi-channel frequency domain signal. This can be achieved by inputting the multi-channel frequency domain signal into a neural network model to obtain the second voice activity detection information, or by performing signal processing on the multi-channel frequency domain signal to obtain the second voice activity detection information. The module then determines whether third voice activity detection information has been acquired. The third voice activity detection information corresponds to the multi-channel voice signal, i.e., the third voice activity detection information is user-specified voice activity detection information. By determining whether the third voice activity detection information has been acquired, it is possible to determine whether the user has specified voice activity detection information. If the third voice activity detection information has not been acquired, the second voice activity detection information is used as the first voice activity detection information. That is, if the user has not specified voice activity detection information, the multi-channel audio signal is denoised using the voice activity detection information determined in the multi-channel frequency domain signal. If the third voice activity detection information is acquired, the third voice activity detection information is used as the first voice activity detection information. That is, if the user has specified voice activity detection information, the multi-channel voice signal is denoised using the user-specified voice activity detection information. In the present invention, when the user has specified voice activity detection information, the user-specified voice activity detection information is used to reduce the noise of the multi-channel voice signal. When the user has not specified voice activity detection information, the voice activity detection information in the multi-channel voice signal is used to reduce the noise of the multi-channel voice signal. This achieves the technical effect of being able to reduce the noise of the multi-channel voice signal in any scenario.

[0114] In some embodiments, optionally, the first determining module 1106 is further specifically configured to determine a voice activity detection model, wherein the voice activity detection model is a neural network model; and input the multi-channel frequency domain signal into the voice activity detection model to obtain second voice activity detection information.

[0115] In this embodiment, the first determination module 1106 is further specifically configured to determine a voice activity detection model, which may be a neural network model. The voice activity detection model is capable of detecting voice activity detection information, namely, automatically detecting active (sound) and inactive (silent) portions in a multi-channel speech signal. The multi-channel frequency domain signal is then input into the determined voice activity detection model to obtain second voice activity detection information. By using the voice activity detection model to detect the second voice activity detection information, the accuracy of the second voice activity detection information is improved.

[0116] In some embodiments, optionally, the first determination module 1106 is further specifically used to determine an application scenario; determine signal-to-noise ratio information and / or computing power information according to the application scenario; and determine a voice activity detection model according to the signal-to-noise ratio information and / or computing power information.

[0117] In this embodiment, the first determination module 1106 needs to first determine the application scenario, such as a household appliance application scenario, and then determine the signal-to-noise ratio information and / or computing power information based on the application scenario. Then, the corresponding voice activity detection model is selected based on the signal-to-noise ratio information and / or computing power information. For example, when the computing power information is sufficiently large in the current application scenario, a voice activity detection model with high computing power needs to be selected. When the signal-to-noise ratio information is high in the current application scenario, a voice activity detection model with low computing power is selected. In some embodiments, the second determination module 1108 is optionally configured to determine second sound source localization information based on the multi-channel frequency domain signal; determine whether third sound source localization information is obtained, wherein the third sound source localization information corresponds to the multi-channel speech signal; based on the third sound source localization information being obtained, use the third sound source localization information as the first sound source localization information; based on the failure to obtain the third sound source localization information, use the second sound source localization information as the first sound source localization information.

[0118] In this embodiment, the second determination module 1108 is configured to determine second sound source localization information based on the multi-channel frequency domain signal. This can be achieved by inputting the multi-channel frequency domain signal into a neural network model or by performing signal processing on the multi-channel frequency domain signal. A determination is then made as to whether third sound source localization information has been acquired. The third sound source localization information corresponds to the multi-channel speech signal, i.e., the third sound source localization information is user-specified sound source localization information. By determining whether the third sound source localization information has been acquired, it is possible to determine whether the user has specified sound source localization information. If the third sound source localization information has not been acquired, the second sound source localization information is used as the first sound source localization information. That is, if the user has not specified sound source localization information, the multi-channel audio signal is denoised using the sound source localization information determined in the multi-channel frequency domain signal. If the third sound source localization information has been acquired, the third sound source localization information is used as the first sound source localization information. That is, if the user has specified sound source localization information, the multi-channel speech signal is denoised using the user-specified sound source localization information. In the present invention, when the user has specified sound source localization information, the noise of the multi-channel speech signal is reduced by using the sound source localization information specified by the user. When the user does not specify the sound source localization information, the noise of the multi-channel speech signal is reduced by using the sound source localization information in the multi-channel speech signal, thereby achieving the technical effect of being able to reduce the noise of the multi-channel speech signal in any scenario.

[0119] In some embodiments, optionally, the second determination module 1108 is further specifically configured to determine a sound source localization model, wherein the sound source localization model is a neural network model; and input the multi-channel frequency domain signal into the sound source localization model to obtain second sound source localization information.

[0120] In this embodiment, the second determination module 1108 is further specifically configured to determine a sound source localization model. The sound source localization model may be a neural network model that can detect sound source localization information. Specifically, the sound source localization model determines the direction of the sound source by analyzing the arrival time differences of speech signals at different microphone arrays. The multi-channel frequency domain signal is then input into the determined sound source localization model to obtain second sound source localization information. By using the sound source localization model to detect the second sound source localization information, the accuracy of the second sound source localization information is improved.

[0121] In some embodiments, optionally, the second determination module 1108 is further specifically used to determine an application scenario; determine signal-to-noise ratio information and / or computing power information according to the application scenario; and determine a sound source localization model according to the signal-to-noise ratio information and / or computing power information.

[0122] In this embodiment, the second determination module 1108 needs to first determine an application scenario, such as an application scenario such as household appliances, and then determine the signal-to-noise ratio information and / or computing power information based on the application scenario, and then select the corresponding sound source localization model based on the signal-to-noise ratio information and / or computing power information. For example, when there is sufficiently large computing power information in the current application scenario, it is necessary to select a sound source localization model with large computing power. When the signal-to-noise ratio information is high in the current application scenario, a sound source localization model with small computing power is selected.

[0123] In some embodiments, optionally, the second determining module 1108 is further configured to obtain image information, where the image information is related to the multi-channel speech signal; and determine the first sound source localization information according to the image information.

[0124] In this embodiment, the second determination module 1108 also needs to first obtain image information, where the image information is related to the multi-channel voice signal. For example, the image information may be an image of the user sending the multi-channel voice signal. By identifying and analyzing the image information and analyzing the angle of the user's mouth in the image information, the first sound source localization information is determined. Furthermore, the first voice activity detection information can be determined by analyzing the user's lip movements in the image information.

[0125] In some embodiments, optionally, the first processing module 1104 is configured to determine a multi-channel time domain signal based on the multi-channel speech signal, wherein the multi-channel time domain signal corresponds to the multi-channel speech signal; and determine a multi-channel frequency domain signal based on the multi-channel time domain signal.

[0126] In this embodiment, the first processing module 1104 is used to perform signal processing on the multi-channel voice signal to obtain a multi-channel time domain signal, wherein the multi-channel time domain signal can be obtained by sampling and quantizing the multi-channel voice signal. The multi-channel time domain signal is a digital representation of the multi-channel voice signal. By converting the multi-channel voice signal into the multi-channel time domain signal, a basis is provided for subsequent processing and analysis of the multi-channel voice signal. After obtaining the multi-channel time domain signal, a multi-channel frequency domain signal is determined based on the multi-channel time domain signal, wherein the multi-channel time domain signal can be converted into a multi-channel frequency domain signal by Fourier transform. By determining the multi-channel frequency domain signal, noise reduction of the multi-channel voice signal is facilitated.

[0127] Figure 12 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown; wherein, the electronic device 120 includes a memory 1202, a processor 1204, and a computer program stored in the memory 1202 and executable on the processor 1204. When the processor 1204 executes the computer program, the steps of the multi-channel speech noise reduction method as described above are implemented.

[0128] The electronic device 120 provided by the present invention, in which the processor 1204 implements the steps of the above-mentioned multi-channel speech noise reduction method when executing the computer program, can achieve the technical effects of any of the above-mentioned embodiments and will not be repeated here.

[0129] One embodiment of the present invention provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the steps of any of the above-mentioned multi-channel speech noise reduction methods.

[0130] The storage medium provided by the present invention and the computer program, when executed by a processor, implement the steps of the above-mentioned multi-channel speech noise reduction method, which can achieve the technical effects of any of the above-mentioned embodiments and will not be repeated here.

[0131] One embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the computer program implements the steps of the multi-channel speech noise reduction method in any of the above embodiments.

[0132] The computer program product provided in this embodiment implements the steps of the multi-channel speech noise reduction method as in any embodiment of the present invention, and thus has all the beneficial effects of the multi-channel speech noise reduction method as in any embodiment of the present invention, which will not be described in detail here.

[0133] In this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance, unless otherwise expressly specified or limited. Terms such as "connect," "install," and "fix" should be interpreted broadly. For example, "connect" can refer to a fixed connection, a detachable connection, or an integral connection; and can be directly connected or indirectly connected through an intermediary. Those skilled in the art will understand the specific meanings of these terms in the present invention based on specific circumstances.

[0134] Throughout this specification, terms such as "one embodiment," "some embodiments," and "specific embodiments" mean that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0135] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for noise reduction of multi-channel speech, characterized in that: include: Acquire multi-channel speech signals; Processing the multi-channel speech signal to obtain a multi-channel frequency domain signal; determining first voice activity detection information, wherein the first voice activity detection information corresponds to the multi-channel voice signal; determining first sound source localization information, wherein the first sound source localization information corresponds to the multi-channel speech signal; obtaining a first vector according to the first voice activity detection information; Obtaining a second vector according to the first sound source localization information, wherein the first vector and the second vector are vectors of the same dimension; Inputting the first vector and the second vector into a multi-channel speech noise reduction model to obtain a target noise reduction model; Inputting the multi-channel frequency domain signal into the target noise reduction model for noise reduction; The step of determining first voice activity detection information includes: determining whether third voice activity detection information is acquired, wherein the third voice activity detection information corresponds to the multi-channel voice signal and is user-specified voice activity detection information; Based on the acquired third voice activity detection information, using the third voice activity detection information as the first voice activity detection information; The step of determining the first sound source localization information includes: determining whether third sound source localization information is obtained, wherein the third sound source localization information corresponds to the multi-channel speech signal and is user-specified sound source localization information; Based on the acquired third sound source localization information, the third sound source localization information is used as the first sound source localization information; The step of processing the multi-channel speech signal to obtain a multi-channel frequency domain signal comprises: Performing sampling and quantization operations on the multi-channel speech signal to obtain a multi-channel time domain signal, wherein the multi-channel time domain signal corresponds to the multi-channel speech signal; The multi-channel time domain signal is converted into the multi-channel frequency domain signal by Fourier transform.

2. The multi-channel speech noise reduction method according to claim 1, characterized in that: The step of determining the first voice activity detection information further includes: determining second voice activity detection information according to the multi-channel frequency domain signal; Based on the failure to obtain the third voice activity detection information, the second voice activity detection information is used as the first voice activity detection information.

3. The multi-channel speech noise reduction method according to claim 2, characterized in that: The step of determining second voice activity detection information based on the multi-channel frequency domain signal includes: determining a voice activity detection model, wherein the voice activity detection model is a neural network model; The multi-channel frequency domain signal is input into the voice activity detection model to obtain the second voice activity detection information.

4. The multi-channel speech noise reduction method according to claim 3, characterized in that: The step of determining the voice activity detection model includes: Determine the application scenario; Determining signal-to-noise ratio information and / or computing power information according to the application scenario; The voice activity detection model is determined according to the signal-to-noise ratio information and / or the computing power information.

5. The multi-channel speech noise reduction method according to claim 1, characterized in that: The step of determining the first sound source localization information further includes: Determine second sound source localization information according to the multi-channel frequency domain signal; Based on the fact that the third sound source localization information is not obtained, the second sound source localization information is used as the first sound source localization information.

6. The multi-channel speech noise reduction method according to claim 5, characterized in that: The step of determining second sound source localization information according to the multi-channel frequency domain signal comprises: Determining a sound source localization model, wherein the sound source localization model is a neural network model; The multi-channel frequency domain signal is input into the sound source localization model to obtain the second sound source localization information.

7. The multi-channel speech noise reduction method according to claim 6, characterized in that: The step of determining the sound source localization model includes: Determine the application scenario; Determining signal-to-noise ratio information and / or computing power information according to the application scenario; The sound source localization model is determined according to the signal-to-noise ratio information and / or the computing power information.

8. The multi-channel speech noise reduction method according to claim 1, characterized in that: The step of determining the first sound source localization information further includes: Acquiring image information, wherein the image information is related to the multi-channel voice signal; First sound source localization information is determined according to the image information.

9. A multi-channel speech noise reduction system, characterized in that: include: A first acquisition module, configured to acquire a multi-channel speech signal; A first processing module, configured to process the multi-channel speech signal to obtain a multi-channel frequency domain signal; a first determining module, configured to determine first voice activity detection information, wherein the first voice activity detection information corresponds to the multi-channel voice signal; a second determining module, configured to determine first sound source localization information, wherein the first sound source localization information corresponds to the multi-channel speech signal; a second processing module, configured to obtain a first vector according to the first voice activity detection information; a third processing module, configured to obtain a second vector according to the first sound source localization information, wherein the first vector and the second vector are vectors of the same dimension; a first input module, configured to input the first vector and the second vector into the multi-channel speech noise reduction model to obtain a target noise reduction model; A second input module, the second input module is used to input the multi-channel frequency domain signal into the target noise reduction model for noise reduction; The first determining module is used for: determining whether third voice activity detection information is acquired, wherein the third voice activity detection information corresponds to the multi-channel voice signal and is user-specified voice activity detection information; Based on the acquired third voice activity detection information, using the third voice activity detection information as the first voice activity detection information; The second determining module is used for: determining whether third sound source localization information is obtained, wherein the third sound source localization information corresponds to the multi-channel speech signal and is user-specified sound source localization information; Based on the acquired third sound source localization information, the third sound source localization information is used as the first sound source localization information; The first processing module is further configured to: Performing sampling and quantization operations on the multi-channel speech signal to obtain a multi-channel time domain signal, wherein the multi-channel time domain signal corresponds to the multi-channel speech signal; The multi-channel time domain signal is converted into the multi-channel frequency domain signal by Fourier transform.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the multi-channel speech noise reduction method according to any one of claims 1 to 8 are implemented.

11. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-channel speech noise reduction method according to any one of claims 1 to 8 are implemented.

12. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the multi-channel speech noise reduction method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Audio source enhancement facilitated by using video data

    CN112151063A