Directional speech enhancement methods, electronic devices, and storage media

By performing neural network processing on the enhanced speech signal, the probability of the speaker's speech in the target area is determined and separated, thus solving the problem of abrupt noise in directional sound pickup technology and improving the clarity of the target speaker's speech.

CN115620739BActive Publication Date: 2026-03-06AISPEECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing directional sound pickup technology cannot effectively handle abrupt noise changes, resulting in a reduced speech signal-to-noise ratio and affecting the accuracy of target sound source direction calculation.

Method used

By acquiring the enhanced speech signal and inputting it into a pre-trained neural network model, the probability of each speaker's speech in the target area is determined, and speech separation is performed using speech masking values ​​to achieve speech enhancement for the target speaker.

Benefits of technology

It improves the clarity and intelligibility of the target speaker's voice and effectively suppresses interfering human voices and environmental noise in non-target areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620739B_ABST
    Figure CN115620739B_ABST
Patent Text Reader

Abstract

This invention discloses a method, electronic device, and storage medium for directional speech enhancement. The method includes: acquiring enhancement results for each region of a speech signal; inputting the enhancement results for each region into a pre-trained neural network model to obtain speech masking values ​​for each region; determining the speech presence probability of each speaker within a target region based on the enhancement results and the speech masking values, wherein each speaker includes target speakers and non-target speakers; and performing speech separation on the enhancement results for each region based on the speech presence probabilities to obtain the enhancement result of the target speaker in the target region. This invention, through judging the enhancement results and speech masking values ​​for each region of a speech signal, determines the speech presence probability of each speaker within a target region, and then performs speech separation on the enhancement results for each region based on the speech presence probabilities, thereby achieving speech enhancement for the target speaker in the target region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech recognition technology, and particularly relates to a method for enhancing speech in a specified direction, as well as an electronic device and storage medium. Background Technology

[0002] In existing directional sound pickup technologies, multiple target sounds and user-input directional sound pickup commands are acquired, and time delay compensation is applied to these multiple target sounds to ensure their timing consistency. The target sound corresponding to the directional sound pickup command among the multiple target sounds is taken as the directional sound, and noise reduction is applied to the directional sound. The MCRA (minima-controlled recursive averaging algorithms) noise estimation method is a traditional signal processing method that can only estimate relatively stable noise and cannot track and estimate abrupt noise changes, such as mouse and keyboard clicks, music ringtones, and door closing sounds.

[0003] In existing technologies, a method for directional sound pickup of target sound sources is proposed. This method involves: 1) acquiring sound source signals within a preset sound intensity level range to obtain a sound source signal observation matrix; 2) processing the sound source observation matrix through filtering and framing, and calculating the short-time spectrum; 3) using the TDOA (time different of arrival) method to determine the approximate location of the target sound source corresponding to the peak position with the minimum delay on the cross-correlation curve; 4) within the approximate location range of the target sound source, using the MVDR (minimum variance distortionless response) method to determine the accurate location of the target sound source; 5) directionally acquiring the target sound source signal based on the accurate location of the target sound source; and 6) when there are two or more target sound sources, repeating steps 3 to 5 at the remaining peak positions of the original cross-correlation curve until directional sound pickup of all target sound sources is completed. The accuracy of the DOA (direction of arrival) algorithm decreases in the presence of environmental noise, with higher noise levels resulting in lower accuracy, and errors in direction calculation significantly impair the pickup effect. This existing solution cannot handle noise with the same direction as the target sound source.

[0004] The inventors discovered that in directional sound pickup technology, traditional noise estimation algorithms have a slower rate of noise change than speech, making it impossible to accurately and timely estimate sudden, non-steady-state noise. For directional sound pickup methods targeting sound sources, the presence of environmental noise reduces the speech signal-to-noise ratio, affecting the results of the correlation matrix. Consequently, the obtained signal and noise subspaces deviate from the true values, leading to a deviation between the calculated signal direction and the true sound source direction. Summary of the Invention

[0005] The embodiments of the present invention are intended to solve at least one of the above-mentioned technical problems.

[0006] In a first aspect, embodiments of the present invention provide a method for enhancing speech in a specified direction, comprising: acquiring enhancement results for each region in a speech signal; inputting the enhancement results for each region into a pre-trained neural network model to obtain a speech masking value for each region, wherein each region includes a target region and / or a non-target region, and the target region is a region within a given angle range; determining the speech presence probability of each speaker within the target region based on the enhancement results and the speech masking value, wherein each speaker includes a target speaker and a non-target speaker; and performing speech separation on the enhancement results of each region based on the speech presence probability to obtain the enhancement result of the target speaker in the target region.

[0007] Secondly, embodiments of the present invention provide a speech enhancement device for a specified direction, comprising: an acquisition module configured to acquire enhancement results of each region in a speech signal, input the enhancement results of each region into a pre-trained neural network model to obtain a speech masking value for each region, wherein each region includes a target region and / or a non-target region, and the target region is a region within a given angle range; a judgment module configured to determine the speech presence probability of each speaker within the target region based on the enhancement results and the speech masking value, wherein each speaker includes a target speaker and a non-target speaker; and a separation module configured to perform speech separation on the enhancement results of each region based on the speech presence probability to obtain the enhancement result of the target speaker in the target region.

[0008] Thirdly, embodiments of the present invention provide an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the speech enhancement method of any of the above-described directions of the present invention.

[0009] Fourthly, embodiments of the present invention provide a storage medium storing one or more programs including execution instructions, the execution instructions being readable and executable by an electronic device (including but not limited to a computer, server, or network device, etc.) to perform the speech enhancement method of any of the above-described directions of the present invention.

[0010] Fifthly, embodiments of the present invention also provide a computer program product, the computer program product including a computer program stored on a storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform any of the above-mentioned speech enhancement methods in the specified direction.

[0011] This invention determines the probability of speech presence of each speaker in the target area by judging the enhancement results and speech masking values ​​of each region in the speech signal, and then performs speech separation on the enhancement results of each region based on the speech presence probability, thereby achieving speech enhancement of the target speaker in the target area. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating an embodiment of the speech enhancement method for a specified direction according to the present invention;

[0014] Figure 2 A flowchart illustrating another embodiment of the speech enhancement method for a specified direction according to the present invention;

[0015] Figure 3 A flowchart illustrating yet another embodiment of the speech enhancement method for a specified direction according to the present invention;

[0016] Figure 4 A flowchart illustrating yet another embodiment of the speech enhancement method for a specified direction according to the present invention;

[0017] Figure 5 A schematic diagram of a speech enhancement device for a specified direction provided by the present invention;

[0018] Figure 6 This is a schematic diagram of the structure of an embodiment of the electronic device of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0021] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, elements, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0022] In this invention, terms such as "module," "device," and "system" refer to relevant entities applied to a computer, such as hardware, combinations of hardware and software, software, or software in execution. More specifically, for example, an element can be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program, and / or a computer. Furthermore, an application program or script running on a server, and the server itself, can also be an element. One or more elements may be in an execution process and / or thread, and elements may be localized on a single computer and / or distributed across two or more computers, and may be run on various computer-readable media. Elements can also communicate via local and / or remote processes based on signals having one or more data packets, for example, signals from data interacting with another element in a local system, a distributed system, and / or interacting with other systems via signals over a network of the Internet.

[0023] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0024] This invention provides a method for enhancing speech in a specified direction, which can be applied to electronic devices. The electronic device can be a computer, server, or other electronic product, etc., and this invention does not limit this to any particular device.

[0025] Please refer to Figure 1 This illustrates a method for enhancing speech in a specified direction according to an embodiment of the present invention.

[0026] like Figure 1As shown, in step 101, the enhancement results of each region in the speech signal are obtained, and the enhancement results of each region are input into a pre-trained neural network model to obtain the speech masking value of each region. The regions include target regions and / or non-target regions, and the target regions are regions within a given angle range.

[0027] In step 102, the probability of the presence of each speaker's voice in the target area is determined based on the enhancement result and the voice masking value, wherein each speaker includes the target speaker and non-target speakers;

[0028] In step 103, speech separation is performed on the enhancement results of each region based on the speech presence probability to obtain the enhancement result of the target speaker in the target region.

[0029] In this embodiment, for step 101, the speech signals of each region are acquired, and the enhancement results of each region in the speech signal are calculated using a microphone array beamforming algorithm. Each region includes a target region and / or a non-target region. The enhancement results of the target region and the non-target region calculated in the speech signal are respectively input into a pre-trained neural network model, and the results of the neural network model are output to obtain the speech masking value of each region, which is equivalent to obtaining the speech masking values ​​of the target region and the non-target region. The speech signal of the target region is set by the angle range. For example, by giving the angle range of the target region, the microphone array is weighted to obtain the weight vector 0. The weight vector 0 is applied to the original signal of the microphone array to obtain the enhancement result of the target region. For non-target regions, after dividing them into N equal angular ranges, weight constraints are applied to obtain weight vectors 1...N. These weight vectors are then applied to the microphone array signal to obtain the enhancement results for the non-target regions. To further increase the distinguishability between the enhancement results for target and non-target regions, a combination of one or more algorithms can be used, such as LCMV (linear constraint minimal variance), GSC (generalized sidelobe cancellation), and TBRR (transient beam-to-reference ratio). Features for each frame are calculated separately for the enhancement results of both target and non-target regions. These features are kept completely consistent with those used during model training, including frequency range, dimension, pre-emphasis, frame stitching, and CMVN (cepstral mean and variance normalization). These features are then input into a pre-trained neural network model, which calculates the speech masking value for each frequency point in each frame.

[0030] Next, in step 102, the probability of each speaker's speech presence in the target area is determined based on the enhancement results of the target area and the enhancement results of the non-target area, as well as the speech masking values ​​of the target area and the non-target area. This probability is determined using methods such as energy / signal-to-interference ratio (SINR). First, for noise in the target area, the neural network estimates the speaker's speech masking value accurately enough, and then simple algorithms such as OMLSA (optimally modified log-spectral amplitude estimator) can achieve good noise suppression. Second, for interfering human voices / noise in the non-target area, a method similar to the (TBRR transient beam-to-reference ratio) ratio of the current beam to the reference beam is used, combined with thresholds such as energy / SINR. For example, the energy / SINR of the target area and the non-target area at the frequency level is compared with thresholds 1 and 2, where threshold 1 is smaller than threshold 2. Frequency points below threshold 1 are considered to be non-target area interference with human voice / noise, and the probability of the speaker's voice in the target area is 0; frequency points above threshold 2 are considered to be only the target speaker's frequency points in the target area, without interference with human voice, and the probability of the target speaker's voice in the target area is 1; frequency points between thresholds 1 and 2 are approximated by smooth interpolation to finally obtain the probability of the target speaker's voice in the target area.

[0031] Finally, for step 103, speech separation is performed on the enhancement results of the target region and the non-target region based on the final probability of the speaker's speech presence within the target region. After speech separation, the enhanced result of the target speaker in the target region is obtained. Speech separation algorithms typically employ blind source separation, a widely used method in signal processing for accurately extracting multiple source signals from mixed signals. This method assumes the independence of the target speech signal and the interfering speech / noise signal, maximizing this independence as the objective function and criterion for signal separation performance. Using the probability of the speaker's speech presence within the target region, an iterative method is used to estimate the noise covariance (including interfering voices from the non-target region, noise, and noise from the target region). The noise covariance is then substituted into the blind source separation framework to solve for the optimal separation matrix. This optimal separation matrix is ​​used to separate the enhancement results of the target and non-target regions, resulting in the enhanced speech result of the target region. Speech separation further suppresses interfering voices from the non-target region and environmental noise from various directions, improving the intelligibility and clarity of the target speaker's speech in the target region.

[0032] The method in this application embodiment determines the probability of speech presence of each speaker in the target area by judging the enhancement results and speech masking values ​​of each region in the speech signal, and then performs speech separation on the enhancement results of each region according to the speech presence probability, so as to realize speech enhancement of the target speaker in the target area.

[0033] It should be noted that, to further enhance the distinction between the enhancement results of the target area and the non-target area, a directional microphone (the microphone described above is an omnidirectional microphone) can be used. The advantage of a directional microphone is that each microphone unit has a beamforming-like effect; this advantage becomes more pronounced when the number of microphones in the array is smaller. Disadvantages: a) Limited directional types, such as cardioid and figure-eight patterns, cannot meet the needs of arbitrary angle ranges; b) Certain requirements are placed on the microphone pickup channel structure (the microphone's back cavity needs to be hollowed out for pickup), which is not supported by all device form factors.

[0034] Please refer to Figure 2 This illustrates another method for directional speech enhancement provided by an embodiment of the present invention. The flowchart is primarily a summary of the flowchart. Figure 1 The flowchart further defines the step 101 of "inputting the enhancement results of each region into a pre-trained neural network model to obtain the speech masking value of each region".

[0035] like Figure 2 As shown, in step 201, the features of each frame in the enhancement results of each region are calculated respectively, and the features of each frame are input into the pre-trained neural network model.

[0036] In step 202, the speech masking value at each frequency point in each frame is obtained by forward calculation via the neural network model.

[0037] In this embodiment, for step 201, the features of each frame in the speech enhancement results in the target region and the non-target region are calculated respectively, and the features of each frame in the speech enhancement results in the target region and the non-target region are input into the trained neural network model, wherein the features of each frame in the speech enhancement results in the target region and the non-target region are completely consistent with those during the training of the neural network model.

[0038] For step 202, the speech masking value at each frequency point in each frame is obtained by forward computation using a neural network model. The neural network model first collects near-field clean speech (e.g., professional studio recordings) and pure noise (excluding speech) data from various noise scenarios. Then, the near-field clean speech is modulated using a large number of different room impulse responses and spatial frequency responses, and then superimposed with various noises within a set signal-to-noise ratio range to obtain noisy speech. Speech features of the noisy speech are extracted, such as amplitude / complex spectrum of narrowband / subband FFT (fast fourier transform), Mel-domain / Bark-domain (fbank filter bank) filter banks, MFCC, or combinations of other features, as input for model training. Because speech has temporal continuity, the model performance is improved after frame concatenation. However, backward frame concatenation introduces latency issues; therefore, more frames can be concatenated forward and fewer backward. For example, within a range imperceptible to the human ear, the latency should be maximized to ensure performance, typically within tens of milliseconds. For example, a 10ms frame shift is used to stitch two frames forward, one frame forward, and one frame backward, resulting in a total of four frames of feature input. The delay at this point is the sum of the one frame from the signal processing overlap and the one frame from the backward stitch, totaling two frames and 20ms. This delay is small and will not cause significant differences in the listening experience or affect near-far dual-talk scenarios.

[0039] The method in this application embodiment calculates the features of each frame in the enhancement results of the target region and the non-target region, and inputs the features of each frame into a neural network model for calculation, so as to obtain the speech masking value at each frequency point of each frame.

[0040] It should be noted that the provided neural network model can optionally perform CMVN (cepstral mean and variance normalization) on the input features. Enabling this function makes the model less sensitive to the absolute amplitude of the input data, which is beneficial for model convergence and also for far-field speech with small amplitude. A variety of model types are available, such as DNN (deep neural networks), CNN (convolutional neural networks), LSTM (long short-term memory), FSMN (feedforward sequential memory networks), RNN (recurrent neural networks), GRU (gate recurrent unit), DCNN (deconvolutional neural networks), and combinations thereof. Since some devices, such as portable devices (e.g., headphones, watches), have relatively limited computing power and storage space, the model type and parameter count need to be determined based on the specific circumstances.

[0041] It should be noted that the energy ratio, amplitude spectrum, complex spectrum, and masking value of speech and / or noise are typically chosen as labels for model training. Then, L1 / smooth L1 / L2 norms are calculated on the labels and the energy ratio, amplitude spectrum, complex spectrum, and masking value of the speech and / or noise output by the model in the loss function, or end-to-end metrics such as speech signal-to-noise ratio, objective speech quality assessment, and short-term objective intelligibility are selected. Finally, the model is trained on a large amount of data (usually over 1000 hours) using various deep learning tools and optimizers, and converges after multiple rounds. The converged model has an accurate estimation ability for speaker speech / environmental noise under various reverberation and different signal-to-noise ratio environments, and the masking value of clean speech can be obtained through simple conversion.

[0042] In some optional embodiments, the probability of speech presence of each speaker in the target region is determined based on the speech masking values ​​of the target and non-target regions, as well as the enhancement results of the target and non-target regions. Each speaker in the target region includes the target speaker and interfering voices. For noise in the target region, the speech masking values ​​of the speakers in the target region are estimated using a neural network. After determining the speech masking values ​​of the speakers in the target region, different algorithms are used for calculation to suppress the noise in the target region. For dry noise in the non-target region, a method similar to TBRR is used, combined with thresholds such as energy / information-to-interference ratio (SINR). For example, the energy / SINR of the target and non-target regions at the frequency level is compared with thresholds 1 and 2, where threshold 1 is smaller than threshold 2. Frequency points below threshold 1 are considered to be interfering voices / noise frequencies in the non-target region, and thus the probability of speech presence of the target speaker in the target region is 0. Frequency points above threshold 2 are considered to be frequencies only present in the target region, and thus the probability of speech presence of the target speaker in the target region is 1. Frequency points between thresholds 1 and 2 are approximated by smooth interpolation to finally obtain the probability of speech presence of the target speaker in the target region.

[0043] Please refer to this again. Figure 3 This illustrates another method for directional speech enhancement provided by an embodiment of the present invention. The flowchart is primarily a summary of the flowchart. Figure 1 A flowchart further defining the steps.

[0044] like Figure 3 As shown, in step 301, the preset parameters of the frequency point of the current region are compared with the first preset threshold and the second preset threshold respectively, wherein the first preset threshold is less than the second preset threshold, and the preset parameter is energy or signal-to-interference ratio;

[0045] In step 302, if the preset parameter is lower than the first preset threshold, then the frequency point of the current area is the frequency point of noise / interference voice in the non-target area, and the probability of the speech of the target speaker in the current area is 0.

[0046] In step 303, if the preset parameter is higher than the second preset threshold, then the frequency point of the current region is the frequency point of the target speaker in the target region, and the probability of the speech of the target speaker in the current region is equal to 1.

[0047] In this embodiment, for step 301, the preset parameters of the frequency points of the current region are compared with the first preset threshold and the second preset threshold respectively. The current region includes the target region and the non-target region. The preset first preset threshold is less than the preset second preset threshold. The preset parameters of the frequency points of the target region and the non-target region are the energy or signal-to-interference ratio of the frequency points of the target region and the non-target region. For example, the energy / signal-to-interference ratio of the target region and the non-target region at the frequency point level is compared with threshold 1 and 2. Threshold 1 is smaller than threshold 2.

[0048] Then, for step 302, when the energy / information-to-interference ratio at the frequency level of the target region or the non-target region is lower than the preset first preset threshold, the frequency of the current region is the frequency of noise in the non-target region or the frequency of interfering human voice in the non-target region, and the probability of the speech of the target speaker in the current region is 0.

[0049] Finally, for step 303, if the energy / information-to-interference ratio of the target region or non-target region at the frequency level is higher than the preset second preset threshold, then the frequency of the current region is the frequency of the target speaker in the target region, and the probability of the speech of the target speaker in the current region is equal to 1.

[0050] The method in this application embodiment compares the preset parameters of the frequency point of the current region with the first preset threshold and the second preset threshold respectively to determine whether the frequency point of the current region is the frequency point of the target speaker in the target region or the frequency point of noise or interference voice in the non-target region. If it is the frequency point of the speaker in the target region, the probability of the speech of the target speaker in the target region is equal to 1. If it is the frequency point of noise or interference voice in the non-target region, the probability of the speech of the target speaker in the non-target region is 0.

[0051] In some optional embodiments, if the preset parameters of the frequency points in the current region are greater than or equal to the first preset threshold and less than or equal to the second preset threshold, the frequency points in the current region are smoothly interpolated and estimated to obtain the probability of the presence of the speech of the target speaker in the current region. For example, if the preset parameters of the frequency points in the current region are greater than or equal to the first preset threshold and less than or equal to the second preset threshold, the frequency points between the first preset threshold 1 and the second preset threshold 2 are smoothly interpolated and approximated to obtain the probability of the presence of the speech of the target speaker in the target region.

[0052] Please refer to this again. Figure 4 This illustrates another method for directional speech enhancement provided by an embodiment of the present invention. The flowchart is primarily a summary of the flowchart. Figure 1 The flowchart further defines the step in step 103 of the process of “performing speech separation on the enhancement results of each region based on the speech presence probability to obtain the enhancement results of the target speaker in the target region”.

[0053] like Figure 4 As shown, in step 401, the noise covariance in the current region is estimated using an iterative method based on the probability of the target speaker's speech presence in the current region.

[0054] Then, in step 402, the noise and covariance of non-target human voices in each region are substituted into the speech separation algorithm to obtain the separation matrix;

[0055] Finally, in step 403, the enhancement results of each region are separated using the separation matrix to obtain the speech enhancement results of the target speaker in the target region.

[0056] For step 401, based on the probability of the target speaker's voice presence in the current region, the noise covariance in the current region is estimated using an iterative method. For example, using the probability of the target speaker's voice presence in the target region, the noise covariance (including interfering human voices in non-target regions, noise, and noise in the target region) is estimated using an iterative method. For step 402, the noise and non-target human voice covariance of each region are substituted into the speech separation algorithm. The speech separation algorithm usually selects blind source separation. The noise covariance is substituted into the blind source separation framework to obtain the separation matrix. For step 403, the enhancement results of the target and non-target regions are separated using the separation matrix to obtain the speech enhancement result of the target region. The speech enhancement result of the target region is the speech enhancement result of the target speaker in the target region.

[0057] The method in this application embodiment, by using speech separation, can further suppress interfering human voices in non-target areas and environmental noise from all directions, thereby improving the clarity of the speech of the target speaker in the target area.

[0058] In some optional embodiments, by setting a given angle range, the enhancement results for the target region and non-target regions are calculated using a microphone array beamforming algorithm. By giving the angle range of the target region, weight constraints are applied to the microphone array to obtain weight vector 0. Weight vector 0 is applied to the original signal of the microphone array to obtain the enhancement result for the target region. For the non-target region, it is divided into N equal angle ranges, and weight constraints are similarly applied to obtain weight vectors 1...N. Weight vectors 1...N are applied to the microphone array signal to obtain the enhancement result for the non-target region.

[0059] It should be noted that the speech separation algorithm in this application typically uses blind source separation. This algorithm models the distribution of multiple sounds and then solves for the separation matrix using an iterative formula. However, blind source separation has a problem with being "blind"—it doesn't know which of the separated sounds is the target human voice. Therefore, it's necessary to use angular information (i.e., the target region enhancement result) to determine the target human voice.

[0060] It should be noted that when generating training data, neural networks consider a sufficient number of scenarios, such as various room impulse responses, speech direct mixing ratios at various distances, various types of environmental noise, and various speech signal-to-noise ratios / signal-to-interference ratios. The more generalizable the trained model is, the more accurate the estimation of speech masking values ​​in various real-world scenarios will be. Combined with the stability and universality advantages of traditional signal processing (speech presence probability estimation, speech separation, etc.), a very good speech enhancement effect can be obtained.

[0061] It should be noted that this application also provides an alternative scheme, utilizing a multi-channel neural network for directional speech enhancement. Data simulation of the multi-channel microphone array involves: convolving the target speaker's speech with the multi-channel room impulse response 1, the interfering speaker's speech with the multi-channel room impulse response 2, and the environmental noise with the multi-channel room impulse response 3 to N, then averaging the results. Finally, these three are superimposed according to a set signal-to-interference-ratio (SIR) and signal-to-noise ratio (SNR) to obtain the noisy signal from the microphone array. The neural network model uses the real / imaginary parts, cosine / sine IPD, and amplitude spectrum of all channel microphone data after FFT transformation as model input, and the target speaker's speech as the model label for model training. The model exhibits significant suppression effects for noise sources / interfering human voices incident from a single direction.

[0062] Please refer to Figure 5 The illustration shows a speech enhancement device 500 for a specified direction provided by an embodiment of the present invention, including an acquisition module 510, a judgment module 520 and a separation module 530.

[0063] The acquisition module 510 is configured to acquire the enhancement results of each region in the speech signal, input the enhancement results of each region into a pre-trained neural network model, and obtain the speech masking value of each region. Each region includes a target region and / or a non-target region, where the target region is a region within a given angle range. The judgment module 520 is configured to determine the speech presence probability of each speaker within the target region based on the enhancement results and the speech masking value. Each speaker includes a target speaker and a non-target speaker. The separation module 530 is configured to perform speech separation on the enhancement results of each region based on the speech presence probability to obtain the enhancement result of the target speaker in the target region.

[0064] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0065] In some embodiments, the present invention provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions, which can be read and executed by an electronic device (including but not limited to a computer, server, or network device, etc.) to perform the speech enhancement method of any of the above-described directions of the present invention.

[0066] In some embodiments, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform any of the above-described speech enhancement methods in the specified direction.

[0067] In some embodiments, the present invention also provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a speech enhancement method in a specified direction.

[0068] Figure 6 This is a schematic diagram of the hardware structure of an electronic device for performing a speech enhancement method in a specified direction, as provided in another embodiment of this application. Figure 6 As shown, the device includes:

[0069] One or more processors 610 and memory 620, Figure 6 Take the 610 processor as an example.

[0070] The device for performing the speech enhancement method in a specified direction may further include an input device 630 and an output device 640.

[0071] The processor 610, memory 620, input device 630, and output device 640 can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0072] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the speech enhancement method in the specified direction in the embodiments of this application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, thereby implementing the speech enhancement method in the specified direction described in the above method embodiments.

[0073] The memory 620 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the voice enhancement device in the specified direction. Furthermore, the memory 620 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 620 may optionally include memory remotely located relative to the processor 610, and this remote memory can be connected to the voice enhancement device in the specified direction via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0074] Input device 630 can receive input digital or character information, and generate signals related to user settings and function control of the voice enhancement device in a specified direction. Output device 640 may include a display device such as a display screen.

[0075] The one or more modules are stored in the memory 620, and when executed by the one or more processors 610, they perform the speech enhancement method in the specified direction in any of the above method embodiments.

[0076] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0077] The electronic devices in this application embodiments exist in various forms, including but not limited to:

[0078] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.

[0079] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include: PDAs, MIDs, and UMPCs, etc.

[0080] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.

[0081] (4) Other airborne electronic devices with data interaction capabilities, such as vehicle-mounted systems installed on vehicles.

[0082] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A directional speech enhancement method, comprising: obtaining enhancement results of each region in a speech signal, wherein the each region comprises a target region and at least one non-target region, and the target region is a region of a given angle range; inputting the enhancement results of the each region into a pre-trained neural network model to obtain speech mask values of the each region; judging speech existence probabilities of a target speaker and a non-target speaker in the target region based on the enhancement results of the each region and the speech mask values; performing speech separation on the enhancement results of the each region based on the speech existence probabilities to obtain an enhancement result of the target speaker in the target region, wherein the speech separation comprises iteratively estimating noise covariance and substituting into a blind source separation algorithm to obtain a separation matrix, and performing speech separation by using the separation matrix.

2. The method of claim 1, wherein, The inputting the enhancement results of the each region into a pre-trained neural network model to obtain speech mask values of the each region comprises: calculating features of each frame in the enhancement results of the each region respectively, and inputting the features of each frame into a pre-trained neural network model; obtaining speech mask values of each frequency point of each frame through forward calculation of the neural network model.

3. The method of claim 1, wherein, The judging speech existence probabilities of a target speaker and a non-target speaker in the target region based on the enhancement results of the each region and the speech mask values comprises: determining the speech existence probabilities of the target speaker and the non-target speaker in the target region according to the speech mask values of the each region and the enhancement results of the each region.

4. The method of claim 3, wherein, The determining the speech existence probabilities of the target speaker and the non-target speaker in the target region comprises: comparing a preset parameter of a frequency point of a current region with a first preset threshold and a second preset threshold respectively, wherein the first preset threshold is smaller than the second preset threshold, and the preset parameter is energy or signal-to-interference ratio; if the preset parameter is lower than the first preset threshold, the frequency point of the current region is a frequency point of noise / interference voice in a non-target region, and the speech existence probability of the target speaker in the current region is 0; if the preset parameter is higher than the second preset threshold, the frequency point of the current region is a frequency point of the target speaker in a target region, and the speech existence probability of the target speaker in the current region is equal to 1.

5. The method of claim 4, wherein, The method further comprises: if the preset parameter of the frequency point of the current region is greater than or equal to the first preset threshold and less than or equal to the second preset threshold, performing smooth interpolation estimation on the frequency point of the current region to obtain the speech existence probability of the target speaker in the current region.

6. The method of claim 1, wherein, The performing speech separation on the enhancement results of the each region based on the speech existence probabilities to obtain an enhancement result of the target speaker in the target region comprises: estimating noise covariance in a current region by using an iterative method based on the speech existence probability of the target speaker in the current region; inputting covariance of noise and non-target voice of each region into a speech separation algorithm to obtain a separation matrix; performing speech separation on the enhancement results of the each region by using the separation matrix to obtain the speech enhancement result of the target speaker in the target region.

7. The method of claim 1, wherein, The enhancement result of each region in the speech signal comprises: A given angle range is set, and the enhancement result of the target region and the non-target region is obtained by a microphone array beamforming algorithm.

8. A speech enhancement device of a specified direction, comprising: An acquisition module configured to acquire an enhancement result of each region in a speech signal, wherein the each region comprises a target region and at least one non-target region, and the target region is a region of a given angle range; and input the enhancement result of the each region into a pre-trained neural network model to obtain a speech masking value of the each region; A judgment module configured to judge a speech existence probability of a target speaker and a non-target speaker in the target region based on the enhancement result of the each region and the speech masking value; A separation module configured to perform speech separation on the enhancement result of the each region based on the speech existence probability to obtain an enhancement result of a target speaker of the target region, wherein the speech separation comprises iterative estimation of noise covariance and substitution into a blind source separation algorithm to obtain a separation matrix, and speech separation is performed by using the separation matrix.

9. An electronic device comprising: At least one processor, and a memory connected with the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method of any one of claims 1 to 7.

10. A storage medium having stored thereon a computer program, characterized in that The program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice enhancing algorithm based on second-order differential microphone array

    CN110310650A

  • Voice call processing method and related device

    CN111179957A

  • Speech enhancement method and device, equipment, storage medium and program

    CN113223552A