Voice processing method and apparatus
By echo cancellation and beamforming of voice signals, interfering candidate beams are obtained and filtered, the problem of low success rates of speech recognition and voice wake-up in harsh scenarios is solved, and signal quality and recognition accuracy are improved.
Patent Information
- Application Number
- CN202011065103.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-09-30
AI Technical Summary
The success rate of speech recognition and voice wake-up in harsh scenarios in the prior art is low, and the existing speech enhancement methods cannot effectively improve signal quality.
By echo cancellation and beamforming of multiple voice signals, interference candidate beams are obtained and filtered to improve signal quality, including technical means such as dereverberation and blind source separation.
Improve the signal quality of the target voice signal and improve the accuracy of speech recognition and voice wake-up.
Smart Images

Figure CN114333870B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing, and in particular, to a speech processing method and apparatus. Background Art
[0002] In the existing technical solutions for speech recognition and speech wake-up, a method of performing speech enhancement on the received speech signal is usually adopted to improve the success rate of recognition or speech wake-up. However, in harsh scenarios (such as strong external noise scenarios), the existing speech enhancement methods usually cannot perform good enhancement processing on the signal, resulting in a low success rate in speech recognition or speech wake-up. Summary of the Invention
[0003] Embodiments of the present invention provide a speech processing method and apparatus, which can perform echo cancellation on multiple paths of speech signals, and process the speech signals after echo cancellation according to the beam formed by beamforming to obtain a target speech signal, thereby improving the signal quality of the target speech signal.
[0004] In a first aspect, an embodiment of the present invention provides a speech processing method, the method comprising:
[0005] Performing echo cancellation on n paths of first speech signals to obtain n paths of second speech signals, where the first speech signals are speech signals collected;
[0006] Performing beamforming on the n paths of second speech signals to obtain m paths of first beams;
[0007] Obtaining interference candidate beams from the m paths of first beams;
[0008] Processing the n paths of second speech signals according to the interference candidate beams to obtain a target speech signal.
[0009] In this example, echo cancellation is performed on the first speech signals to obtain second speech signals, and the second speech signals are processed according to the interference candidate beams after beamforming of the second speech signals to obtain a target speech signal. Therefore, the second speech signals can be processed through the interference candidate beams to obtain a target speech signal, and the signal quality of the target speech signal can be improved.
[0010] In combination with the first aspect, in a possible implementation manner, the obtaining interference candidate beams from the m paths of first beams includes:
[0011] Obtaining the frame-level energy corresponding to the m paths of first beams to obtain m first frame-level energy values;
[0012] Determining a beam average energy value according to the m first frame-level energy values;
[0013] Determine the count values corresponding to the m first beams according to the m first frame-level energy values and the beam average energy value;
[0014] Determine the beam corresponding to the maximum count value among the count values corresponding to the m first beams as the interference candidate beam.
[0015] In this example, the first beams are counted by the frame-level energy corresponding to each first beam and the average energy of the beam, and the beam corresponding to the maximum count value is determined as the interference candidate beam. The interference candidate beam is obtained from the perspectives of frame-level energy and average energy, improving the accuracy when obtaining the interference candidate.
[0016] Combined with the first aspect, in a possible implementation manner, the processing of the n second voice signals according to the interference candidate beam to obtain a target voice signal includes:
[0017] Obtain the smoothed energy value of the interference candidate beam;
[0018] Determine the filtering intensity value according to the smoothed energy value;
[0019] Perform filtering processing on the n second voice signals according to the filtering intensity value and the interference candidate beam to obtain n third voice signals;
[0020] Process the n third voice signals to obtain a target voice signal.
[0021] In this example, the filtering intensity value determined by the smoothed energy and the interference beam are used to perform filtering processing on the second voice signal to obtain a third filtered signal, and the third filtered signal is processed to obtain a target voice signal. Since the second voice is filtered by the filtering intensity value and the interference beam, the quality of the third signal after filtering can be improved, and thus the signal quality of the target voice signal is improved.
[0022] Combined with the first aspect, in a possible implementation manner, the processing of the n third voice signals to obtain a target voice signal includes:
[0023] Obtain the i-th third voice signal and the j-th third voice signal from the n third voice signals;
[0024] Perform dereverberation and blind source separation on the i-th third voice signal and the j-th third voice signal to obtain a first candidate voice signal and a second candidate voice signal;
[0025] Perform beamforming on the n third voice signals to obtain m second beams;
[0026] Obtain a first target beam from the m second beams;
[0027] Perform at least noise reduction processing on the first target beam to obtain a processed first target beam;
[0028] Determine the target speech signal according to the first candidate speech signal, the second candidate speech signal, and the processed first target beam.
[0029] In this example, perform dereverberation and blind source separation on the i-th third speech signal and the j-th third speech signal to obtain a first candidate speech signal and a second candidate speech signal, and determine the target speech signal according to the first candidate speech signal, the second candidate speech signal, and the first target beam after noise reduction, which can improve the signal quality of the target speech signal.
[0030] Combined with the first aspect, in a possible implementation manner, the processing the n second speech signals according to the interference candidate beam to obtain a target speech signal includes:
[0031] Determine whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, perform filtering processing on the n second speech signals according to the interference candidate beam and a preset filtering intensity to obtain n fourth speech signals;
[0032] Perform beamforming on the n fourth speech signals to obtain m third beams;
[0033] Obtain a second target beam from the m third beams;
[0034] Perform at least noise reduction processing on the second target beam to obtain a processed second target beam;
[0035] Obtain the h-th second speech signal and the k-th second speech signal from the n second speech signals;
[0036] Perform filtering and blind source separation on the h-th second speech signal and the k-th second speech signal to obtain a third candidate speech signal and a fourth candidate speech signal;
[0037] Determine the target speech signal according to the third candidate speech signal, the fourth candidate speech signal, and the processed second target beam.
[0038] In this example, after determining that the interference candidate beam is a preset beam, the second speech signal is filtered according to the interference candidate beam and the preset filtering intensity, and at least noise reduction is performed on the second target beam, and dereverberation and blind source separation are performed on the h-th second speech signal and the k-th second speech signal to obtain a third candidate speech signal and a fourth candidate speech signal, so as to determine the target signal, which can improve the signal quality of the target speech signal.
[0039] Combined with the first aspect, in a possible implementation manner, the method further includes:
[0040] If the interference candidate beam is not a preset beam, the target speech signal is determined according to the interference candidate beam, the third candidate speech signal, and the fourth candidate speech signal.
[0041] If the interference candidate beam is not a preset beam, the interference candidate beam, the third candidate speech signal, and the fourth candidate speech signal are used to determine the target speech signal, which can reduce the resource consumption when determining the target speech signal.
[0042] In a second aspect, an embodiment of the present application provides a speech processing apparatus, where the apparatus includes:
[0043] An elimination unit, configured to perform echo elimination on n first speech signals to obtain n second speech signals, where the first speech signals are the collected speech signals;
[0044] A beamforming unit, configured to perform beamforming on the n second speech signals to obtain m first beams;
[0045] An acquisition unit, configured to acquire an interference candidate beam from the m first beams;
[0046] A processing unit, configured to process the n second speech signals according to the interference candidate beam to obtain a target speech signal.
[0047] Combined with the second aspect, in a possible implementation manner, the acquisition unit is specifically configured to:
[0048] Acquire the frame-level energy corresponding to the m first beams to obtain m first frame-level energy values;
[0049] Determine the beam average energy value according to the m first frame-level energy values;
[0050] Determine the count value corresponding to the m first beams according to the m first frame-level energy values and the beam average energy value;
[0051] Determine the beam corresponding to the maximum count value among the count values corresponding to the m first beams as the interference candidate beam.
[0052] In combination with the second aspect, in a possible implementation manner, the processing unit is specifically configured to:
[0053] Obtain the smoothed energy value of the interference candidate beam;
[0054] Determine a filtering intensity value according to the smoothed energy value;
[0055] Perform filtering processing on the n-channel second voice signals according to the filtering intensity value and the interference candidate beam to obtain n-channel third voice signals;
[0056] Process the n-channel third voice signals to obtain a target voice signal.
[0057] In combination with the second aspect, in a possible implementation manner, in the aspect of processing the n-channel third voice signals to obtain a target voice signal, the processing unit is specifically configured to:
[0058] Obtain the i-th third voice signal and the j-th third voice signal from the n-channel third voice signals;
[0059] Perform dereverberation and blind source separation on the i-th third voice signal and the j-th third voice signal to obtain a first candidate voice signal and a second candidate voice signal;
[0060] Perform beamforming on the n-channel third voice signals to obtain m-channel second beams;
[0061] Obtain a first target beam from the m-channel second beams;
[0062] Perform at least noise reduction processing on the first target beam to obtain a processed first target beam;
[0063] Determine the target voice signal according to the first candidate voice signal, the second candidate voice signal, and the processed first target beam.
[0064] In combination with the second aspect, in a possible implementation manner, the processing unit is configured to:
[0065] Determine whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, perform filtering processing on the n-channel second voice signals according to the interference candidate beam and a preset filtering intensity to obtain n-channel fourth voice signals;
[0066] Perform beamforming on the n-channel fourth voice signals to obtain m-channel third beams;
[0067] Obtain a second target beam from the m-channel third beams;
[0068] Perform at least noise reduction processing on the second target beam to obtain the processed second target beam;
[0069] Obtain the h-th second voice signal and the k-th second voice signal from the n-way second voice signals;
[0070] Perform filtering and blind source separation on the h-th second voice signal and the k-th second voice signal to obtain a third candidate voice signal and a fourth candidate voice signal;
[0071] Determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam.
[0072] Combined with the second aspect, in a possible implementation manner, the processing unit is further configured to:
[0073] If the interference candidate beam is not a preset beam, determine the target voice signal according to the interference candidate beam, the third candidate voice signal, and the fourth candidate voice signal.
[0074] In a third aspect, an embodiment of the present invention provides a voice processing device, including:
[0075] A memory for storing instructions; and
[0076] At least one processor coupled to the memory;
[0077] Wherein, when the at least one processor executes the instructions, the instructions cause the processor to execute all or part of the methods as shown in the first aspect.
[0078] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute all or part of the methods as shown in the first aspect.
[0079] These aspects or other aspects of the present invention will be more clearly understood in the following description of the embodiments. Description of the Drawings
[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0081] Figure 1This application provides a schematic diagram of speech enhancement in an embodiment;
[0082] Figure 2 This application provides a schematic flowchart of a speech processing method in an embodiment;
[0083] Figure 3 This invention provides a schematic structural diagram of a speech processing apparatus in an embodiment;
[0084] Figure 4 This invention provides a schematic structural diagram of another speech processing apparatus in an embodiment. Detailed implementation manners
[0085] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0086] First, the main current method of introducing speech enhancement for speech wake-up and speech enhancement processing will be introduced. As Figure 1 shown, speech enhancement is performed on the input signal, and speech wake-up or speech recognition is performed through the enhanced speech signal. The application scenarios of speech wake-up and speech recognition can be human-machine speech interaction scenarios. In human-machine speech interaction scenarios, more and more people use speech interaction to wake up electronic devices through speech or send instructions to electronic devices through speech, etc.
[0087] The electronic device of the present application can be a terminal device such as a mobile phone, a tablet, a speaker, a TV, etc. For example, as a mobile phone, a tablet, a speaker, a TV, etc., the speech processing method can process the signal with interference picked up by the microphone and output a clean speech signal as the input to the wake-up and recognition engines.
[0088] The process of the speech processing method will be specifically introduced below.
[0089] Refer to Figure 2 , Figure 2 This application provides a schematic flowchart of a speech processing method in an embodiment. As Figure 2 shown, the speech processing method includes:
[0090] S201. Perform echo cancellation on n first speech signals to obtain n second speech signals, where the first speech signals are the speech signals collected.
[0091] The n first speech signals can be n speech signals obtained through a microphone array / pickup array, and the microphone array may include n microphones, etc. n is a positive integer greater than or equal to 2.
[0092] The method for echo cancellation of n first voice signals can be to perform echo cancellation on the n first voice signals according to a reference signal. Specifically, it can be to perform echo cancellation through the ACE algorithm, and the reference signal can be a preset signal.
[0093] S202. Perform beamforming on the n second voice signals to obtain m first beams.
[0094] Beamforming can be performed on the n second voice signals through a beamforming algorithm or the like to obtain m first beams. m can be a value independent of n. For example, m can be the same as n, or m can be different from n.
[0095] S203. Obtain interference candidate beams from the m first beams.
[0096] The interference candidate beams can be determined through the frame-level energy values corresponding to the m first beams. Specifically, the interference candidate beams can be determined according to the frame-level energy values corresponding to the m first beams and the beam average energy value. The beam average energy value can be understood as the beam average energy of the m beams.
[0097] S204. Process the n second voice signals according to the interference candidate beams to obtain the target voice signal.
[0098] The n second voice signals can be filtered to obtain the filtered n third voice signals, and the target voice signal can be obtained according to the processed n third voice signals. Alternatively, it is determined whether the interference candidate beam is a preset beam, and the second voice signal is processed according to the determination result to obtain the target voice signal. The preset beam can be a continuous interference beam, and the continuous interference beam can be understood as: there is continuous noise interference in the beam.
[0099] In this example, echo cancellation is performed on the first voice signal to obtain the second voice signal, and the second voice signal is processed according to the interference candidate beams after beamforming of the second voice signal to obtain the target voice signal. Therefore, the second voice signal can be processed through the interference candidate beams to obtain the target voice signal, and the signal quality of the target voice signal can be improved.
[0100] In a possible implementation manner, a possible method for obtaining interference candidate beams from the m first beams includes:
[0101] A1. Obtain the frame-level energy corresponding to the m first beams to obtain m first frame-level energy values;
[0102] A2. Determine the beam average energy value according to the m first frame-level energy values;
[0103] A3. Determine the count values corresponding to the m first beams according to the m first frame-level energy values and the beam average energy value;
[0104] A4. Determine the beam corresponding to the maximum count value among the count values corresponding to the m first beams as the interference candidate beam.
[0105] The frame-level energy value corresponding to the first beam can be obtained by according to the number of time-domain sampling points of each frame of data corresponding to the first beam. Specifically, the frame-level energy can be obtained by the method shown in the following formula:
[0106]
[0107] where i is the i-th first beam, with a value range of [1, m], K represents the number of time-domain sampling points of each frame of data, k has a value range of [1, K], and P i is the frame-level energy of the i-th first beam.
[0108] The beam average energy value can be directly obtained by performing a mean operation on the m first frame-level energy values. Specifically, the beam average energy value can be obtained by the method shown in the following formula:
[0109]
[0110] where P avg is the beam average energy value, m is the number of total frame energies, and P i is the i-th frame-level energy value.
[0111] The method for determining the count values corresponding to the m first beams according to the m first frame-level energy values and the beam average energy value can specifically be:
[0112] First, the VAD module determines whether the current frame of each beam is in a quiet state according to the data of the beam data B1, B2,..., Bm. VAD = 0 indicates that the current frame is in a quiet state, and VAD = 1 indicates that the current frame is in a non-quiet state;
[0113] When the current frame of the i-th beam is in a non-quiet state, the beam corresponding counter C i is updated as follows:
[0114]
[0115] where P avg is the beam average energy value, and P i is the i-th first frame-level energy value.
[0116] In this example, the first beam is counted based on the frame-level energy corresponding to each first beam and the average energy of the beam, and the beam corresponding to the maximum count value is determined as the interference candidate beam. By obtaining the interference candidate beam from the perspectives of frame-level energy and average energy, the accuracy of obtaining the interference candidate is improved.
[0117] In a possible implementation, a possible method for processing the n second voice signals according to the interference candidate beam to obtain a target voice signal includes:
[0118] B1. Obtain the smoothed energy value of the interference candidate beam;
[0119] B2. Determine the filtering intensity value according to the smoothed energy value;
[0120] B3. Filter the n second voice signals according to the filtering intensity value and the interference candidate beam to obtain n third voice signals;
[0121] B4. Process the n third voice signals to obtain a target voice signal.
[0122] The smoothed energy value can be determined by the smoothed energy value of the previous frame of the data frame and the maximum energy value of the beam in the current data frame. Specifically, the smoothed energy value can be determined by the method shown in the following formula:
[0123]
[0124] where is the smoothed energy value of the interference candidate beam, is the smoothed energy value of the previous frame, and P max is the maximum energy value of the beam in the current data frame.
[0125] The filtering intensity value can be determined by the maximum beam energy. Specifically, it can be determined by the method shown in the following formula:
[0126]
[0127] where β represents the filtering intensity, H represents the high threshold, and L represents the low threshold, and H > L. When exceeds the high threshold H, β = 1; when is lower than the low threshold L, β = 0. H and L are preset values, which can be set by empirical values or historical data.
[0128] The interference candidate beam can be determined as the reference beam, and the filtering gain can be adjusted through the filtering intensity value to perform filtering processing on the n second voice signals to obtain n third voice signals. When filtering the second voice signals, linear adaptive filtering can be used, for example, Kalman filtering. Specifically, with the interference candidate beam as the reference beam, the gain of the filter can be adjusted through the filtering intensity value to obtain the third voice signals.
[0129] Beamforming can be performed on the n third voice signals to obtain m second beams, the processed first target beam can be determined according to the second beams, and two voice signals can be obtained from the n third voice signals, processed to obtain candidate voice signals, and the target voice signal can be determined according to the candidate voice signals and the processed target beam.
[0130] In this example, the filtering intensity value determined by smoothing energy and the interference beam are used to filter the second voice signals to obtain the third filtered signal, and the third filtered signal is processed to obtain the target voice signal. Since the second voice is filtered by the filtering intensity value and the interference beam, the quality of the third signal after filtering can be improved, and thus the signal quality of the target voice signal is improved.
[0131] In a possible implementation, a method for processing the n third voice signals to obtain the target voice signal may include:
[0132] C1. Obtain the i-th third voice signal and the j-th third voice signal from the n third voice signals;
[0133] C2. Perform dereverberation and blind source separation on the i-th third voice signal and the j-th third voice signal to obtain a first candidate voice signal and a second candidate voice signal;
[0134] C3. Perform beamforming on the n third voice signals to obtain m second beams;
[0135] C4. Obtain the first target beam from the m second beams;
[0136] C5. Perform at least noise reduction processing on the first target beam to obtain the processed first target beam;
[0137] C6. Determine the target voice signal according to the first candidate voice signal, the second candidate voice signal, and the processed first target beam.
[0138] The method for obtaining the third voice signal of the i-th path and the third voice signal of the j-th path may be to determine the third voice signals corresponding to the two microphones with the farthest distance in the microphone array as the third voice signal of the i-th path and the third voice signal of the j-th path.
[0139] Performing dereverberation and blind source separation on the third voice signal of the i-th path and the third voice signal of the j-th path to obtain a first candidate voice signal and a second candidate voice signal may specifically be to perform cross operations on the third voice signal of the i-th path and the third voice signal of the j-th path, etc., to obtain a first candidate voice signal and a second candidate voice signal. The method for performing dereverberation and blind source separation on the third voice signal may be to obtain a first candidate voice signal and a second candidate voice signal through a dereverberation module and a blind source separation module, and the blind source separation module may be implemented by the IVA algorithm.
[0140] The method for obtaining the first target beam from the m second beams may refer to the method for obtaining the interference candidate beam in the foregoing embodiments and will not be elaborated here.
[0141] The method for determining the target voice signal according to the first candidate voice signal, the second candidate voice signal, and the processed first target beam may be to determine any one of the first candidate voice signal, the second candidate voice signal, and the processed first target beam as the target voice signal, or to determine any combination of the first candidate voice signal, the second candidate voice signal, and the processed first target beam as the target voice signal, or to perform any combination of the first candidate voice signal, the second candidate voice signal, and the processed first target beam and perform processing, and determine the signal obtained after processing as the target voice signal.
[0142] The first target beam path wake-up enhancement signal effectively improves the effects of speech enhancement and interference suppression in an external noise scenario by filtering the interference beam. The i-th and j-th wake-up enhancement signals improve the quality of target signal extraction in the same azimuth scenario of the target and interference through a blind source separation algorithm, and according to the processed first beam, first voice signal, and second voice signal, thereby improving the quality of the target voice signal. Therefore, when using the target voice signal for speech recognition and speech wake-up, the accuracy can be improved.
[0143] In a possible implementation manner, another possible method for processing the n second voice signals according to the interference candidate beam to obtain the target voice signal may be:
[0144] D1. Determine whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, filter the n second voice signals according to the interference candidate beam and a preset filtering intensity to obtain n fourth voice signals;
[0145] D2. Perform beamforming on the n-channel fourth voice signals to obtain m-channel third beams;
[0146] D3. Obtain a second target beam from the m-channel third beams;
[0147] D4. At least perform noise reduction processing on the second target beam to obtain a processed second target beam;
[0148] D5. Obtain the h-th second voice signal and the k-th second voice signal from the n-channel second voice signals;
[0149] D6. Perform filtering and blind source separation on the h-th second voice signal and the k-th second voice signal to obtain a third candidate voice signal and a fourth candidate voice signal;
[0150] D7. Determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam.
[0151] For the specific implementation manners of the above steps D2 - D4, reference may be made to the method for obtaining the first target beam in the foregoing embodiments, which will not be elaborated here.
[0152] When performing filtering and blind source separation on the h-th second voice signal and the k-th second voice signal, reverberation removal may also be performed on the h-th second voice signal and the k-th second voice signal to obtain the third candidate voice signal and the fourth candidate voice signal. The method for filtering the h-th second voice signal and the k-th second voice signal may refer to the method for filtering the n-channel second voice signals in the foregoing.
[0153] The preset beam may be a continuous interference beam, and the continuous interference beam can be understood as: there is continuous noise interference in the beam.
[0154] To obtain the h-th second voice signal and the k-th second voice signal from the n-channel second voice signals, reference may be made to the method for obtaining the i-th third voice signal and the j-th voice signal in the foregoing embodiments, which will not be elaborated here.
[0155] To determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam, reference may be made to the method for determining the target voice signal according to the first candidate voice signal, the second candidate voice signal, and the processed first target beam in the foregoing embodiments, which will not be elaborated here.
[0156] In a possible implementation manner, if the interference candidate beam is not a preset beam, the target voice signal can be obtained through the following method:
[0157] Determine the target speech signal according to the interference candidate beam, the third candidate speech signal, and the fourth candidate speech signal. Specifically, reference may be made to the method of determining the target speech signal according to the first candidate speech signal, the second candidate speech signal, and the processed first target beam in the foregoing embodiments, which will not be elaborated herein.
[0158] In this example, reverberation cancellation and blind source separation are performed on the i-th third speech signal and the j-th third speech signal to determine the target speech signal based on the obtained first candidate speech signal, second candidate speech signal, and the first target beam after noise reduction, which can improve the signal quality of the target speech signal.
[0159] See Figure 3 , Figure 3 which is a schematic structural diagram of a speech processing device provided by an embodiment of the present invention. As Figure 3 shown, the speech processing device 30 includes:
[0160] An echo cancellation unit 301, configured to perform echo cancellation on n first speech signals to obtain n second speech signals, where the first speech signals are speech signals collected.
[0161] A beamforming unit 302, configured to perform beamforming on the n second speech signals to obtain m first beams.
[0162] An acquisition unit 303, configured to acquire interference candidate beams from the m first beams.
[0163] A processing unit 304, configured to process the n second speech signals according to the interference candidate beams to obtain a target speech signal.
[0164] In a possible implementation, the acquisition unit 303 is specifically configured to:
[0165] Acquire the frame-level energy corresponding to the m first beams to obtain m first frame-level energy values;
[0166] Determine the beam average energy value according to the m first frame-level energy values;
[0167] Determine the count value corresponding to the m first beams according to the m first frame-level energy values and the beam average energy value;
[0168] Determine the beam corresponding to the maximum count value among the count values corresponding to the m first beams as the interference candidate beam.
[0169] In a possible implementation, the processing unit 304 is specifically configured to:
[0170] Obtain the smoothed energy value of the interference candidate beam;
[0171] Determine the filtering intensity value according to the smoothed energy value;
[0172] Filter the n-channel second voice signals according to the filtering intensity value and the interference candidate beam to obtain n-channel third voice signals;
[0173] Process the n-channel third voice signals to obtain the target voice signal.
[0174] In a possible implementation manner, in terms of processing the n-channel third voice signals to obtain the target voice signal, the processing unit 304 is specifically configured to:
[0175] Obtain the i-th third voice signal and the j-th third voice signal from the n-channel third voice signals;
[0176] Perform dereverberation and blind source separation on the i-th third voice signal and the j-th third voice signal to obtain a first candidate voice signal and a second candidate voice signal;
[0177] Perform beamforming on the n-channel third voice signals to obtain m-channel second beams;
[0178] Obtain a first target beam from the m-channel second beams;
[0179] Perform at least noise reduction processing on the first target beam to obtain the processed first target beam;
[0180] Determine the target voice signal according to the first candidate voice signal, the second candidate voice signal, and the processed first target beam.
[0181] In a possible implementation manner, the processing unit 304 is configured to:
[0182] Judge whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, filter the n-channel second voice signals according to the interference candidate beam and the preset filtering intensity to obtain n-channel fourth voice signals;
[0183] Perform beamforming on the n-channel fourth voice signals to obtain m-channel third beams;
[0184] Obtain a second target beam from the m-channel third beams;
[0185] Perform at least noise reduction processing on the second target beam to obtain the processed second target beam;
[0186] Obtain the h-th second voice signal and the k-th second voice signal from the n-way second voice signals;
[0187] Perform filtering and blind source separation on the h-th second voice signal and the k-th second voice signal to obtain a third candidate voice signal and a fourth candidate voice signal;
[0188] Determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam.
[0189] In a possible implementation manner, the processing unit 304 is further configured to:
[0190] If the interference candidate beam is not a preset beam, determine the target voice signal according to the interference candidate beam, the third candidate voice signal, and the fourth candidate voice signal.
[0191] In this embodiment, the voice processing device 300 is presented in the form of a unit. Here, the "unit" may refer to an application-specific integrated circuit (ASIC), a processor and a memory that execute one or more software or firmware programs, an integrated logic circuit, and / or other devices that can provide the above functions. In addition, the above elimination unit 301, beamforming unit 302, acquisition unit 303, and processing unit 304 can be implemented by Figure 4 the processor 401 of the shown voice processing device.
[0192] Such as Figure 4 the shown voice processing device 40 can be implemented in the structure of Figure 4 in which the voice processing device 400 includes at least one processor 401, at least one memory 402, and at least one communication interface 403. The processor 401, the memory 402, and the communication interface 403 are connected through the communication bus and complete communication with each other.
[0193] The processor 401 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the above solutions.
[0194] The communication interface 403 is used to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (Wireless Local Area Networks, WLAN), etc.
[0195] The memory 402 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.
[0196] Among them, the memory 402 is used to store the application program code for executing the above solution, and is controlled by the processor 401 to execute. The processor 401 is used to execute the application program code stored in the memory 402.
[0197] The code stored in the memory 402 can execute the above-provided voice processing method to perform echo cancellation on n first voice signals to obtain n second voice signals, where the first voice signals are the collected voice signals; perform beamforming on the n second voice signals to obtain m first beams; obtain interference candidate beams from the m first beams; and process the n second voice signals according to the interference candidate beams to obtain target voice signals.
[0198] The embodiment of the present invention also provides a computer-readable storage medium, where the computer-readable storage medium can store a program, and when the program is executed, it includes some or all of the steps of any one of the voice processing methods described in the above method embodiments.
[0199] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0200] In the above embodiments, the descriptions of the various embodiments each have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0201] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.
[0202] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0204] If the above-mentioned integrated units are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. And the aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0205] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (abbreviation: ROM), a random access memory (abbreviation: RAM), a magnetic disk or an optical disc, etc.
[0206] The above has introduced the embodiments of the present invention in detail. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A speech processing method, characterized in that, The method includes: Performing echo cancellation on n first voice signals to obtain n second voice signals, where the first voice signals are voice signals collected; Performing beamforming on the n second voice signals to obtain m first beams; Obtaining interference candidate beams from the m first beams; Processing the n second voice signals according to the interference candidate beams to obtain target voice signals; Wherein, the obtaining interference candidate beams from the m first beams includes: Obtaining the frame-level energy corresponding to the m first beams to obtain m first frame-level energy values; Determining a beam average energy value according to the m first frame-level energy values; Determining the count values corresponding to the m first beams according to the m first frame-level energy values and the beam average energy value; Determining the beam corresponding to the maximum count value among the count values corresponding to the m first beams as the interference candidate beam.
2. The method according to claim 1, wherein The processing the n second voice signals according to the interference candidate beams to obtain target voice signals includes: Obtaining the smoothed energy value of the interference candidate beam; Determining a filtering intensity value according to the smoothed energy value; Performing filtering processing on the n second voice signals according to the filtering intensity value and the interference candidate beam to obtain n third voice signals; Processing the n third voice signals to obtain target voice signals.
3. The method according to claim 2, wherein The processing the n third voice signals to obtain target voice signals includes: Obtaining the i-th third voice signal and the j-th third voice signal from the n third voice signals; Performing dereverberation and blind source separation on the i-th third voice signal and the j-th third voice signal to obtain a first candidate voice signal and a second candidate voice signal; Performing beamforming on the n third voice signals to obtain m second beams; Obtaining a first target beam from the m second beams; Performing at least noise reduction processing on the first target beam to obtain a processed first target beam; Determining the target voice signal according to the first candidate voice signal, the second candidate voice signal and the processed first target beam.
4. The method according to claim 1, characterized in that, The processing the n second voice signals according to the interference candidate beams to obtain target voice signals includes: Judging whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, then performing filtering processing on the n second voice signals according to the interference candidate beam and a preset filtering intensity to obtain n fourth voice signals; Performing beamforming on the n fourth voice signals to obtain m third beams; Obtaining a second target beam from the m third beams; Performing at least noise reduction processing on the second target beam to obtain a processed second target beam; Obtaining the h-th second voice signal and the k-th second voice signal from the n second voice signals; Performing filtering and blind source separation on the h-th second voice signal and the k-th second voice signal to obtain a third candidate voice signal and a fourth candidate voice signal; Determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam.
5. The method according to claim 4, wherein The method further includes: If the interference candidate beam is not a preset beam, determine the target voice signal according to the interference candidate beam, the third candidate voice signal, and the fourth candidate voice signal.
6. A voice processing method, characterized in that, The method includes: Perform echo cancellation on n first voice signals to obtain n second voice signals, where the first voice signals are voice signals collected. Perform beamforming on the n second voice signals to obtain m first beams. Obtain an interference candidate beam from the m first beams. Process the n second voice signals according to the interference candidate beam to obtain a target voice signal. Wherein, The processing the n second voice signals according to the interference candidate beam to obtain a target voice signal includes: Obtain the smoothed energy value of the interference candidate beam. Determine a filtering intensity value according to the smoothed energy value. Perform filtering processing on the n second voice signals according to the filtering intensity value and the interference candidate beam to obtain n third voice signals. Process the n third voice signals to obtain a target voice signal.
7. A voice processing method, characterized in that, The method includes: Perform echo cancellation on n first voice signals to obtain n second voice signals, where the first voice signals are voice signals collected. Perform beamforming on the n second voice signals to obtain m first beams. Obtain an interference candidate beam from the m first beams. Process the n second voice signals according to the interference candidate beam to obtain a target voice signal. Wherein, the processing the n second voice signals according to the interference candidate beam to obtain a target voice signal includes: Determine whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, perform filtering processing on the n second voice signals according to the interference candidate beam and a preset filtering intensity to obtain n fourth voice signals. Perform beamforming on the n fourth voice signals to obtain m third beams. Obtain a second target beam from the m third beams. Perform at least noise reduction processing on the second target beam to obtain a processed second target beam. Obtain the h-th second voice signal and the k-th second voice signal from the n second voice signals. Perform filtering and blind source separation on the h-th second voice signal and the k-th second voice signal to obtain a third candidate voice signal and a fourth candidate voice signal. Determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam.
8. A voice processing device, characterized in that, The apparatus includes: An elimination unit, configured to perform echo cancellation on n first voice signals to obtain n second voice signals, where the first voice signals are voice signals collected. A beamforming unit, configured to perform beamforming on the n second voice signals to obtain m first beams. An obtaining unit, configured to obtain an interference candidate beam from the m first beams. A processing unit, configured to process the n second voice signals according to the interference candidate beam to obtain a target voice signal; Wherein, the obtaining unit is specifically configured to: Obtain the frame-level energy corresponding to the m first beams to obtain m first frame-level energy values; Determine the beam average energy value according to the m first frame-level energy values; Determine the count value corresponding to the m first beams according to the m first frame-level energy values and the beam average energy value; Determine the beam corresponding to the maximum count value among the count values corresponding to the m first beams as the interference candidate beam.
9. The device according to claim 8, characterized in that, The processing unit is specifically configured to: Obtain the smoothed energy value of the interference candidate beam; Determine the filtering intensity value according to the smoothed energy value; Perform filtering processing on the n second voice signals according to the filtering intensity value and the interference candidate beam to obtain n third voice signals; Process the n third voice signals to obtain a target voice signal.
10. The device according to claim 9, characterized in that, In terms of processing the n third voice signals to obtain a target voice signal, the processing unit is specifically configured to: Obtain the i-th third voice signal and the j-th third voice signal from the n third voice signals; Perform dereverberation and blind source separation on the i-th third voice signal and the j-th third voice signal to obtain a first candidate voice signal and a second candidate voice signal; Perform beamforming on the n third voice signals to obtain m second beams; Obtain a first target beam from the m second beams; Perform at least noise reduction processing on the first target beam to obtain a processed first target beam; Determine the target voice signal according to the first candidate voice signal, the second candidate voice signal, and the processed first target beam.
11. The device according to claim 8, characterized in that, The processing unit is configured to: Judge whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, perform filtering processing on the n second voice signals according to the interference candidate beam and the preset filtering intensity to obtain n fourth voice signals; Perform beamforming on the n fourth voice signals to obtain m third beams; Obtain a second target beam from the m third beams; Perform at least noise reduction processing on the second target beam to obtain a processed second target beam; Obtain the h-th second voice signal and the k-th second voice signal from the n second voice signals; Perform filtering and blind source separation on the h-th second voice signal and the k-th second voice signal to obtain a third candidate voice signal and a fourth candidate voice signal; Determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam.
12. The device according to claim 11, wherein The processing unit is further configured to: If the interference candidate beam is not a preset beam, determine the target voice signal according to the interference candidate beam, the third candidate voice signal, and the fourth candidate voice signal.
13. A voice processing device, characterized in that, The device includes: An elimination unit, configured to perform echo cancellation on n first voice signals to obtain n second voice signals, where the first voice signals are collected voice signals; A beamforming unit, configured to perform beamforming on the n second voice signals to obtain m first beams; An obtaining unit, configured to obtain interference candidate beams from the m first beams; A processing unit, configured to process the n second voice signals according to the interference candidate beams to obtain target voice signals; Wherein, The processing unit is specifically configured to: Obtain a smoothed energy value of the interference candidate beam; Determine a filtering intensity value according to the smoothed energy value; Perform filtering processing on the n second voice signals according to the filtering intensity value and the interference candidate beam to obtain n third voice signals; Process the n third voice signals to obtain target voice signals.
14. A voice processing device, characterized in that, The apparatus includes: An echo cancellation unit, configured to perform echo cancellation on n first voice signals to obtain n second voice signals, where the first voice signals are collected voice signals; A beamforming unit, configured to perform beamforming on the n second voice signals to obtain m first beams; An obtaining unit, configured to obtain interference candidate beams from the m first beams; A processing unit, configured to process the n second voice signals according to the interference candidate beams to obtain target voice signals; Wherein, The processing unit is configured to: Determine whether the interference candidate beam is a preset beam. If the interference candidate beam is a preset beam, perform filtering processing on the n second voice signals according to the interference candidate beam and a preset filtering intensity to obtain n fourth voice signals; Perform beamforming on the n fourth voice signals to obtain m third beams; Obtain a second target beam from the m third beams; Perform at least noise reduction processing on the second target beam to obtain a processed second target beam; Obtain an h-th second voice signal and a k-th second voice signal from the n second voice signals; Perform filtering and blind source separation on the h-th second voice signal and the k-th second voice signal to obtain a third candidate voice signal and a fourth candidate voice signal; Determine the target voice signal according to the third candidate voice signal, the fourth candidate voice signal, and the processed second target beam.
15. A voice processing device, characterized in that, Includes: A memory, configured to store instructions; And At least one processor, coupled to the memory; Wherein, when the at least one processor executes the instructions, the instructions cause the processor to execute the method according to any one of claims 1-7.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Signal processor
CN109087663A