A noise reduction method, electronic device and storage medium
By performing intelligent noise reduction on the audio signal, estimating the noise energy, and using the frequency suppression ratio for signal enhancement, the problem of insufficient robustness of point noise to speaker localization algorithms is solved, and the speaker's location can be accurately identified in noisy environments.
Patent Information
- Application Number
- CN202210993912.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-22
- Filing Date
- 2022-08-18
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2042-08-18
AI Technical Summary
In audio interaction scenarios, especially in complex environments with noise, existing speaker localization algorithms are not effective at suppressing point noise, resulting in insufficient robustness.
Intelligent noise reduction technology is used to process multiple audio signals, estimate noise energy and determine the audio suppression ratio corresponding to the frequency point, and suppress noise and highlight the speaker's voice through signal enhancement processing.
Effective noise suppression improves the robustness of the speaker localization algorithm, ensuring accurate speaker location identification in noisy environments.
Smart Images

Figure CN115331692B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of audio technology, and in particular to a noise reduction method, an electronic device and a storage medium. BACKGROUND
[0002] In audio interaction scenarios such as audio-video conferencing and voice calls, speaker positioning is required, which refers to determining the position of a sound source based on audio signals received by an audio device such as a microphone array, so as to determine the position of the current speaker.
[0003] However, the audio signals received by the audio device may contain audio of the speaker and noise, and therefore, when performing speaker positioning, how to effectively suppress the noise so as to improve the robustness of the speaker positioning algorithm has become a technical problem that needs to be solved by those skilled in the art. SUMMARY
[0004] Therefore, embodiments of the present application provide a noise reduction method, an electronic device and a storage medium to effectively suppress noise and improve the robustness of the speaker positioning algorithm.
[0005] To achieve the above object, embodiments of the present application provide the following technical solutions.
[0006] In a first aspect, a noise reduction method is provided, comprising:
[0007] obtaining multiple audio signals;
[0008] determining a target audio signal for estimating noise energy based on the multiple audio signals;
[0009] performing intelligent noise reduction processing on the target audio signal to obtain an audio suppression ratio corresponding to a frequency point of the target audio signal, the audio suppression ratio corresponding to the frequency point representing the noise energy of the target audio signal;
[0010] performing signal enhancement processing on audio signals containing noise and audio signals containing speakers based on the audio suppression ratio corresponding to the frequency point;
[0011] determining a speaker positioning result based on the signal enhancement processing result.
[0012] In a second aspect, an electronic device is provided, comprising at least one memory and at least one processor, the memory storing one or more computer executable instructions, and the processor invoking the one or more computer executable instructions to perform the noise reduction method as described in the first aspect.
[0013] In a third aspect, an embodiment of the present application provides a storage medium, which stores one or more computer executable instructions. When the one or more computer executable instructions are executed, the noise reduction method according to the first aspect is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer program, which, when executed, implements the noise reduction method according to the first aspect.
[0015] The embodiment of the present application can determine a target audio signal for intelligent noise reduction processing according to the collected multi-channel audio signals after the multi-channel audio signals are collected, and perform intelligent noise reduction processing on the target audio signal to obtain an audio suppression ratio corresponding to a frequency point of the target audio signal. The audio suppression ratio corresponding to the frequency point is used to perform signal enhancement processing on the audio signal containing noise and the audio signal containing a speaker, so that the audio signal containing noise is suppressed and the audio signal containing the speaker is highlighted using the audio suppression ratio of the frequency point. Then, the speaker positioning result is determined according to the signal enhancement processing result. In the case of effectively suppressing the noise signal and effectively highlighting the speaker signal, the purpose of highlighting the speaker voice and suppressing the interference noise is achieved, and the effect of effectively suppressing the noise and improving the robustness of the speaker positioning algorithm is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.
[0017] Figure 1 An example diagram of the speaker direction and the point noise direction.
[0018] Figure 2 A flowchart of the noise reduction method provided by the embodiment of the present application.
[0019] Figure 3 Another flowchart of the noise reduction method provided by the embodiment of the present application.
[0020] Figure 4 Still another flowchart of the noise reduction method provided by the embodiment of the present application.
[0021] Figure 5 Yet another flowchart of the noise reduction method provided by the embodiment of the present application.
[0022] Figure 6A An example diagram of linear array beamforming.
[0023] Figure 6B An example diagram of beamforming for a circular array.
[0024] Figure 7A Yet another flowchart of a noise reduction method provided by an embodiment of the application.
[0025] Figure 7B An example diagram of sound source positioning implemented by an embodiment of the application.
[0026] Figure 8 A block diagram of a noise reduction device provided by an embodiment of the application.
[0027] Figure 9 A block diagram of an electronic device. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the application.
[0029] In an audio interaction scene such as an audio and video conference, accurate speaker positioning can better support audio pickup algorithms and video directing functions. Currently, speaker positioning is usually based on the time / phase difference of audio arriving at different microphones of a microphone array, and therefore speaker positioning schemes are usually established in a good acoustic environment (such as a quiet scene). However, actual audio interaction scenes are more inclined to be complex scenes containing noise, and therefore the robustness of speaker positioning schemes in complex scenes containing noise needs to be improved.
[0030] Generally, noise can be mainly divided into diffuse noise and point noise; diffuse noise refers to noise that is not generated by a sound source at a certain position and is diffuse, such as environmental noise; point noise refers to noise generated by a fixed sound source. For diffuse noise, the influence of diffuse noise on a speaker positioning algorithm is limited, because the time difference of diffuse noise arriving at different microphones of a microphone array is relatively ambiguous, and in a good audio signal-to-noise ratio, the energy of diffuse noise can be estimated to remove noise in a speaker positioning algorithm. For point noise, because point noise has a clear sound source position, point noise is easily misidentified as audio of a speaker, thereby causing a speaker positioning algorithm to provide incorrect speaker direction information; for ease of understanding, Figure 1 An example diagram of a speaker direction and a point noise direction is exemplarily shown, and it can be seen that point noise has a similar sound source position as a speaker, and has a possibility of being misidentified as a speaker.
[0031] In summary, for audio interaction scenarios such as audio and video conferencing, how to effectively suppress the point noise when performing speaker positioning is of great significance to the robust convergence of the speaker positioning algorithm. That is, robust convergence refers to how to accurately perform speaker positioning under different noise types, especially in the presence of point noise.
[0032] Currently, the speaker positioning algorithm mainly relies on traditional noise estimation algorithms to track the sound source of the point noise. For example, by estimating the energy of the point noise, the point noise is suppressed and removed in the speaker positioning algorithm. However, due to the non-stationary nature of the point noise, and the possibility of the audio energy of the point noise being greater than that of the speaker, the traditional noise estimation algorithm based on the energy of the point noise cannot accurately track and estimate the energy of the point noise, resulting in ineffective suppression of the point noise and thus providing incorrect results for the speaker positioning algorithm.
[0033] Based on this, the embodiments of the present application provide an improved noise reduction scheme to effectively suppress the point noise and improve the robustness of the speaker positioning algorithm. The embodiments of the present application can integrate intelligent noise reduction technology in the speaker positioning algorithm, so that after determining the target audio signal based on the multi-channel audio signals collected by the audio device, the embodiments of the present application can perform intelligent noise reduction processing on the target audio signal, estimate the noise energy corresponding to the target audio signal, and obtain the audio suppression ratio corresponding to the frequency point of the target audio signal; and then use the above audio suppression ratio to perform signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker, so as to suppress the noise in the audio signal containing noise and the audio signal containing the speaker and highlight the speaker's voice, thereby improving the robustness of the speaker positioning algorithm to noise.
[0034] Based on the above idea, as an optional implementation, Figure 2 An optional flowchart of the noise reduction method provided by the embodiments of the present application is shown. The method flow can be implemented by an audio device, such as a microphone array or other device with audio collection and processing capabilities. Referring to Figure 2 The method flow can include the following steps.
[0035] In step S210, a plurality of audio signals are obtained, and the plurality of audio signals are preprocessed respectively.
[0036] Optionally, the audio device can obtain a plurality of audio signals through a plurality of audio acquisition channels. For example, a microphone array with multiple microphones can obtain a plurality of audio signals through a plurality of audio acquisition channels, and one microphone of the microphone array can collect one audio signal.
[0037] After the multi-channel audio signals are collected, the embodiments of the present application can perform preprocessing on the multi-channel audio signals respectively. Optionally, the preprocessing of the audio signals includes but is not limited to: converting each channel of the audio signals from a time domain signal to a frequency domain signal (i.e., converting each channel of the audio signals from a time domain to a frequency domain), performing amplitude normalization processing on each channel of the audio signals converted to the frequency domain, and the like.
[0038] In step S211, a target audio signal is determined according to the preprocessed multi-channel audio signals.
[0039] In the case of fusing intelligent noise reduction technology in the speaker positioning algorithm, the target audio signal can be regarded as an audio signal that needs to be processed by the intelligent noise reduction of the embodiments of the present application. By performing intelligent noise reduction processing on the audio signal, the embodiments of the present application can estimate the noise energy of the audio signal, and the estimated noise energy is reflected through the audio suppression ratio of the frequency point in the audio signal, so the target audio signal can be regarded as an audio signal for which the embodiments of the present application estimate the noise energy.
[0040] The audio suppression ratio can also be referred to as Mask (masking), that is, there is a value representing the audio suppression ratio for each time-frequency (time-frequency domain) point, 0 represents all noise and needs to be suppressed in speaker positioning, 1 represents all speech and needs to be reserved in speaker positioning, and the value range of Mask is between 0.0 and 1.0.
[0041] In some embodiments, the audio signals can be divided according to the channel or the azimuth angle, and the embodiments of the present application can determine the audio signal of the channel or the azimuth angle that needs to be estimated for noise energy as the target audio signal. As an optional implementation, in the case of dividing the audio signals according to the azimuth angle, the multi-channel audio signals can be divided into multiple azimuth angles, and the target audio signal can be the audio signal of each azimuth angle or the audio signal of part of the azimuth angles in the multiple azimuth angle audio signals. As an optional implementation, in the case of dividing the audio signals according to the channel, the target audio signal can be one channel of the audio signals converted to the frequency domain in the multi-channel audio signals.
[0042] As an optional implementation, the embodiments of the present application provide multiple ways to determine the target audio signal and the corresponding noise reduction processing scheme.
[0043] In some embodiments, in the case of dividing the audio signals according to the azimuth angle, the embodiments of the present application can estimate the noise energy of the audio signal of each azimuth angle; correspondingly, the embodiments of the present application can divide the multi-channel audio signals into multiple azimuth angle audio signals, so that the audio signal of each azimuth angle is regarded as the target audio signal.
[0044] In this case, after the target audio signal is processed by the intelligent noise reduction, the noise energy of the audio signal of each azimuth angle can be estimated, which is embodied as the audio suppression ratio corresponding to each frequency point in each azimuth angle. Then, the audio signal of each azimuth angle is processed by signal enhancement using the audio suppression ratio corresponding to each frequency point in each azimuth angle, so as to suppress the noise in the audio signal of each azimuth angle and highlight the speaker voice.
[0045] In some embodiments, in the case that the audio signal is divided according to the azimuth angle, the azimuth angle where the speaker is located is generally the azimuth angle with the maximum signal peak value. Considering the interference of the noise on the speaker voice, the part of the azimuth angle with the maximum signal peak value (for example, at least two azimuth angles with the maximum signal peak value) can be determined from the divided multiple azimuth angles, so as to estimate the noise energy of the audio signal of the part of the azimuth angle. Correspondingly, the multiple audio signals can be divided into audio signals of multiple azimuth angles, and the part of the azimuth angle with the maximum signal peak value can be determined, so as to take the audio signal of the part of the azimuth angle as the target audio signal.
[0046] In this case, after the target audio signal is processed by the intelligent noise reduction, the noise energy of the audio signal of each azimuth angle can be estimated, which is embodied as the audio suppression ratio corresponding to each frequency point in each azimuth angle. Then, the audio signal of each azimuth angle is processed by signal enhancement using the audio suppression ratio corresponding to each frequency point in each azimuth angle, so as to suppress the noise in the audio signal of each azimuth angle and highlight the speaker voice.
[0047] In some embodiments, in the case that the audio signal is divided according to the azimuth angle, the azimuth angle where the speaker is located is generally the azimuth angle with the maximum signal peak value. Considering the interference of the noise on the speaker voice, the part of the azimuth angle with the maximum signal peak value (for example, at least two azimuth angles with the maximum signal peak value) can be determined from the divided multiple azimuth angles, so as to estimate the noise energy of the audio signal of the part of the azimuth angle. Correspondingly, the multiple audio signals can be divided into audio signals of multiple azimuth angles, and the part of the azimuth angle with the maximum signal peak value can be determined, so as to take the audio signal of the part of the azimuth angle as the target audio signal.
[0048] In step S212, the target audio signal is processed by the intelligent noise reduction to obtain the audio suppression ratio corresponding to the frequency point of the target audio signal.
[0049] The embodiments of the present application can use an intelligent noise reduction algorithm to perform intelligent noise reduction processing on the target audio signal. As an optional implementation, the embodiments of the present application can use an intelligent noise reduction algorithm suitable for a conference scenario (for example, a conference room scenario) to perform intelligent noise reduction processing on the target audio signal. Of course, the embodiments of the present application can also not limit the used intelligent noise reduction algorithm.
[0050] In the case where the audio signal has been converted into a frequency domain form, the embodiments of the present application can estimate the noise energy of the target audio signal after performing intelligent noise reduction processing on the target audio signal, and reflect the audio suppression ratio corresponding to each frequency point in the target audio signal. That is, the embodiments of the present application can obtain the audio suppression ratio corresponding to each frequency point in the target audio after performing intelligent noise reduction processing on the target audio signal.
[0051] In step S213, the audio signal containing noise and the audio signal containing the speaker are subjected to signal enhancement processing according to the audio suppression ratio corresponding to the frequency point.
[0052] After obtaining the audio suppression ratio corresponding to the frequency point of the target audio signal, the embodiments of the present application can use the audio suppression ratio corresponding to the frequency point to perform signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker, so as to suppress noise and highlight the speaker's voice in the audio signal containing noise and the audio signal containing the speaker.
[0053] As an optional implementation of signal enhancement processing, the embodiments of the present application can use the audio suppression ratio of each frequency point in the target audio signal to reweight the audio signal containing noise and the audio signal containing the speaker, so as to obtain an enhanced audio signal. For example, in the audio signal containing noise and the audio signal containing the speaker, the embodiments of the present application can use the audio suppression ratio corresponding to the frequency point of the target audio signal to reweight the audio signal corresponding to the frequency point, so as to obtain an enhanced audio signal.
[0054] It should be noted that in the sound source positioning algorithm such as speaker positioning, a formula for calculating the correlation of the audio signal in the frequency domain can be set, so that for the audio signal containing noise and the audio signal containing the speaker, the embodiments of the present application can add the audio suppression ratio corresponding to the frequency point of the target audio signal in the formula for calculating the correlation of the audio signal, so as to reweight the audio signal containing noise and the audio signal containing the speaker, thereby obtaining an enhanced audio signal.
[0055] Based on different forms of the target audio signal, the audio signal containing noise and the audio signal containing the speaker can also have different forms.
[0056] In some embodiments, the multiple azimuth angle audio signals described above can include audio signals containing noise and audio signals containing speakers. In some embodiments, the signal peak maximum partial azimuth angle audio signal described above can include audio signals containing noise and audio signals containing speakers. In some embodiments, the pre-processed multi-channel audio signal described above can include audio signals containing noise and audio signals containing speakers.
[0057] In an optional implementation, if the target audio signal is the individual azimuth angle audio signal, the audio signal containing noise and the audio signal containing the speaker can be the individual azimuth angle audio signal; that is, the embodiments of the present application can use the audio suppression ratio corresponding to the frequency point of the individual azimuth angle audio signal to perform signal enhancement processing on the individual azimuth angle audio signal.
[0058] In an optional implementation, if the target audio signal is the signal peak maximum partial azimuth angle audio signal, the audio signal containing noise and the audio signal containing the speaker can be the partial azimuth angle audio signal described above; that is, the embodiments of the present application can use the audio suppression ratio corresponding to the frequency point of the partial azimuth angle audio signal to perform signal enhancement processing on the partial azimuth angle audio signal.
[0059] In an optional implementation, if the target audio signal is the one-way audio signal converted into the frequency domain, the audio signal containing noise and the audio signal containing the speaker can be the multi-channel audio signal; that is, the embodiments of the present application can use the audio suppression ratio corresponding to the frequency point of the one-way audio signal to perform signal enhancement processing on the multi-channel audio signal.
[0060] In step S214, the speaker positioning result is determined according to the signal enhancement processing result.
[0061] In some embodiments, since the signal enhancement processing result suppresses noise and highlights speaker voice, in the case of dividing the audio signal into azimuth angles, the embodiments of the present application can take the azimuth angle with the signal peak maximum after signal enhancement processing as the azimuth angle where the speaker is located, thereby obtaining the speaker positioning result.
[0062] The embodiments of the present application can determine a target audio signal for intelligent noise reduction processing according to the collected multi-channel audio signals after collecting the multi-channel audio signals, thereby performing intelligent noise reduction processing on the target audio signal to obtain an audio suppression ratio corresponding to a frequency point of the target audio signal; and performing signal enhancement processing on the audio signal containing noise and the audio signal containing a speaker based on the audio suppression ratio corresponding to the frequency point, thereby realizing suppression of the audio signal containing noise and highlighting of the audio signal containing the speaker by using the audio suppression ratio of the frequency point. Furthermore, the speaker positioning result is determined according to the signal enhancement processing result, which can achieve the purpose of highlighting the speaker voice and suppressing the interference noise in the case of effectively suppressing the noise signal and effectively highlighting the speaker signal, thereby achieving the effect of effectively suppressing the noise and improving the robustness of the speaker positioning algorithm.
[0063] As Figure 2 As an optional implementation of the flow shown, in the case that the multi-channel audio signals can be divided into multiple azimuth angles, the target audio signal for intelligent noise reduction processing can be the audio signal of each azimuth angle. Figure 3 An exemplary another optional flowchart of the noise reduction method provided by the embodiments of the present application is shown. Referring to Figure 3 The method flow can include the following steps.
[0064] In step S310, multi-channel audio signals are acquired, and the multi-channel audio signals are respectively converted from time domain to frequency domain.
[0065] In step S311, the multi-channel audio signals converted to frequency domain are respectively subjected to amplitude normalization processing.
[0066] In step S312, the multi-channel audio signals are divided into multiple azimuth angle audio signals according to a preset azimuth angle accuracy.
[0067] After the collected multi-channel audio signals are respectively pre-processed (such as time domain to frequency domain conversion and amplitude normalization processing), the embodiments of the present application can divide the multi-channel audio signals according to the preset azimuth angle accuracy, thereby obtaining audio signals of multiple azimuth angles. Optionally, the azimuth angle accuracy can be regarded as the accuracy of interval division of the multi-channel audio signals. According to the azimuth angle accuracy (azimuth interval accuracy), the embodiments of the present application can divide the multi-channel audio signals into audio signals of multiple intervals (audio signals of multiple azimuth angles), thereby achieving fine division of the multi-channel audio signals in the form of directional vectors. In an example, the embodiments of the present application can divide the multi-channel audio signals into directional vector audio signals of multiple azimuth angles with 5 degrees as the azimuth interval accuracy, such as 5 degrees as an interval angle, thereby dividing the multi-channel audio signals into multiple audio signals with 5 degrees as the interval unit.
[0068] In the present embodiment, the audio signals of multiple azimuth angles can be used as target audio signals for intelligent noise reduction processing.
[0069] In step S313, the audio signals of multiple azimuth angles are respectively subjected to intelligent noise reduction processing, thereby obtaining audio suppression ratios of each frequency point.
[0070] After the multi-channel audio signals are divided into audio signals of multiple azimuth angles, the embodiments of the present application can use an intelligent noise reduction algorithm to respectively process the audio signals of each azimuth angle. As an optional implementation, the embodiments of the present application can use an intelligent noise reduction algorithm suitable for a conference scenario (such as a conference room scenario) to respectively process the audio signals of multiple azimuth angles. Of course, the embodiments of the present application do not limit the used intelligent noise reduction algorithm.
[0071] Since the audio signals have been converted into frequency domain form in the pre-processing process, the embodiments of the present application can obtain audio suppression ratios corresponding to each frequency point of each azimuth angle after the intelligent noise reduction processing of the audio signals of each azimuth angle (in frequency domain form). The audio suppression ratio can also be referred to as Mask (masking), that is, each time-frequency point has a value representing the audio suppression ratio, 0 represents all noise and needs to be suppressed in speaker positioning, 1 represents all speech and needs to be reserved in speaker positioning, and the value of Mask ranges from 0.0 to 1.0.
[0072] In step S314, according to the audio suppression ratios of each frequency point, the audio signals of each azimuth angle are subjected to signal enhancement processing, thereby obtaining enhanced audio signals of each azimuth angle.
[0073] After obtaining the audio suppression ratios of the frequency points, the embodiment of the present application can utilize the audio suppression ratios of the frequency points to perform signal enhancement processing on the audio signals of the divided azimuth angles, so as to obtain the enhanced audio signals of the azimuth angles after signal enhancement. As an optional implementation of the signal enhancement processing, the embodiment of the present application can utilize the audio suppression ratios of the frequency points to reweight the audio signals of the azimuth angles, so as to obtain the enhanced audio signals of the azimuth angles. For example, the embodiment of the present application can use the audio suppression ratios of the frequency points to reweight the audio signals corresponding to the frequency points at the azimuth angles, so as to obtain the enhanced audio signals of the azimuth angles. For example, an audio signal of an azimuth angle is reweighted using the audio suppression ratio of the corresponding frequency point.
[0074] It should be noted that a formula for calculating the correlation of the audio signals in the frequency domain can be set in the sound source positioning algorithm such as the speaker positioning, and the embodiment of the present application can add the audio suppression ratio of the frequency point corresponding to the audio signal in the formula for calculating the correlation of the audio signals, so as to reweight the audio signals, thereby obtaining the enhanced audio signals of the azimuth angles.
[0075] In the embodiment, the audio signals of the azimuth angles can include the audio signals containing the noise and the audio signals containing the speaker.
[0076] In step S315, the speaker positioning result is determined according to the enhanced audio signals of the azimuth angles.
[0077] After the enhanced audio signals of the azimuth angles are determined, the embodiment of the present application can determine the speaker positioning result based on the enhanced audio signals of the azimuth angles, that is, determine the positioning position of the speaker. As an optional implementation, the embodiment of the present application can determine the azimuth angle in which the signal peak meets the speaker condition according to the enhanced audio signals of the azimuth angles, so as to determine the azimuth angle as the azimuth angle of the speaker, so as to obtain the speaker positioning result. In one example, the embodiment of the present application can determine the azimuth angle with the maximum signal peak according to the enhanced audio signals of the azimuth angles, so as to determine the azimuth angle with the maximum signal peak as the azimuth angle of the speaker.
[0078] In the speaker positioning algorithm, the embodiments of the present application can perform intelligent noise reduction processing on the audio signals of each azimuth angle, so that after obtaining the audio suppression ratio of each frequency point in each azimuth angle, the audio signals of each frequency point in each azimuth angle are enhanced (for example, reweighted) using the audio suppression ratio of each frequency point in each azimuth angle, so as to obtain the enhanced audio signals of each azimuth angle. Since intelligent noise reduction has a good suppression effect on steady-state and non-steady-state point noise (if the point noise is in a certain azimuth angle that is subjected to intelligent noise reduction), the audio signals of each frequency point are enhanced using the audio suppression ratio of each frequency point after intelligent noise reduction, which can effectively suppress the peak value of the azimuth angle where the point noise is located when calculating the signal peak value of each azimuth angle, thereby effectively highlighting the signal peak value of the speaker azimuth angle, achieving the purpose of highlighting the speaker voice and suppressing the interfering point noise, and further achieving the effect of effectively suppressing the point noise and improving the robustness of the speaker positioning algorithm.
[0079] It needs to be further explained that, Figure 3 The method flow shown can achieve the effects of point noise suppression and highlighting the speaker voice, but Figure 3 The method flow shown needs to perform intelligent noise reduction processing on the audio signals of multiple azimuth angles, which greatly increases the system calculation amount. Based on this, the embodiments of the present application further provide an optional noise reduction scheme to reduce the system calculation amount while effectively suppressing the point noise.
[0080] As Figure 2 As an optional implementation manner of the flow shown, in the case where the multiple audio signals can be divided into multiple azimuth angles, the target audio signal subjected to intelligent noise reduction processing can be the audio signal of the partial azimuth angle with the maximum signal peak value. Figure 4 An exemplary flow chart of the noise reduction method provided by the embodiments of the present application is shown. Referring to Figure 4 The method flow can include the following steps.
[0081] In step S410, multiple audio signals are obtained, and the multiple audio signals are subjected to time-domain to frequency-domain conversion processing respectively.
[0082] In step S411, the multiple audio signals converted into the frequency domain are subjected to amplitude normalization processing respectively.
[0083] In step S412, the multiple audio signals are divided into multiple azimuth angle audio signals according to a preset azimuth angle accuracy.
[0084] Optionally, the introduction of steps S410 to S412 can refer to the corresponding part of the foregoing, which will not be expanded here.
[0085] In step S413, the peak values of the audio signals of the respective azimuth angles are determined, and at least two target azimuth angles with peak values meeting a preset condition are determined according to the peak values of the audio signals of the respective azimuth angles, the number of the at least two target azimuth angles being less than the number of the plurality of azimuth angles.
[0086] After the multi-channel audio signal is divided into the audio signals of the plurality of azimuth angles, the embodiments of the present application do not directly perform intelligent noise reduction processing on the audio signals of the respective azimuth angles, but calculate the peak values of the audio signals of the respective azimuth angles, so as to select, based on the peak values of the audio signals of the respective azimuth angles, a part of the azimuth angles with peak values meeting a preset condition from the plurality of azimuth angles, the number of the part of the azimuth angles being at least two and not greater than the number of the divided plurality of azimuth angles. For the convenience of description, the part of the azimuth angles selected from the plurality of azimuth angles based on the peak values of the audio signals can be referred to as target azimuth angles, and the number of the target azimuth angles is at least two.
[0087] As an optional implementation, the embodiments of the present application can determine at least two target azimuth angles with the maximum signal peak values according to the peak values of the audio signals of the respective azimuth angles. For example, when two target azimuth angles are selected, the embodiments of the present application can select the two target azimuth angles with the maximum signal peak values from the divided plurality of azimuth angles.
[0088] It should be noted that the number of the target azimuth angles is selected as two only as an optional implementation, and the embodiments of the present application can select the number of the target azimuth angles according to actual conditions in the case that the number of the target azimuth angles is set to be not less than two and less than the number of the divided plurality of azimuth angles.
[0089] In the embodiments, the audio signals of the at least two target azimuth angles can be used as target audio signals requiring intelligent noise reduction processing. Since the audio signals of the at least two target azimuth angles have the characteristic of the maximum signal peak values in the plurality of azimuth angles, the speaker voice is contained in the audio signals of the at least two target azimuth angles, and considering the interference of the noise, the noise can also exist in the audio signals of the at least two target azimuth angles; therefore, the noise needs to be suppressed and the speaker voice needs to be highlighted in the audio signals of the at least two target azimuth angles.
[0090] In step S414, the audio signals of the respective target azimuth angles are respectively subjected to intelligent noise reduction processing to obtain the audio suppression ratios of the respective frequency points in the respective target azimuth angles.
[0091] In step S415, the audio signals of the respective target azimuth angles are subjected to signal enhancement processing according to the audio suppression ratios of the respective frequency points in the respective target azimuth angles, to obtain the enhanced audio signals of the respective target azimuth angles.
[0092] In step S416, the speaker positioning result is determined according to the enhanced audio signals of the target azimuth angles.
[0093] After the at least two target azimuth angles are determined, the audio signals of the target azimuth angles can be respectively subjected to intelligent noise reduction processing, so as to obtain the audio suppression ratios of the frequency points in the target azimuth angles. Optionally, the intelligent noise reduction processing and the audio suppression ratio can refer to the corresponding part in the foregoing description, and will not be described herein. In the embodiment, the audio signals of the at least two target azimuth angles can include the audio signals containing noise and the audio signals containing the speaker.
[0094] Based on the audio suppression ratios of the frequency points in the target azimuth angles, the audio signals of the target azimuth angles can be subjected to signal enhancement processing; for example, the audio signal of one target azimuth angle can be subjected to reweighting processing by using the audio suppression ratio of the corresponding frequency point. Based on the intelligent noise reduction processing, the enhanced audio signals of the target azimuth angles can be obtained, so that the target azimuth angle with the maximum peak value of the audio signal can be selected from the at least two target azimuth angles as the azimuth angle of the speaker according to the enhanced audio signals of the target azimuth angles, and the speaker positioning result can be obtained.
[0095] As an optional implementation, assuming that a point noise source is currently interfering with the speaker, the two target azimuth angles with the maximum peak values of the audio signals can be selected as the target azimuth angles when the target azimuth angles are selected; the audio signals of the target azimuth angles can be respectively subjected to intelligent noise reduction processing, so as to obtain the audio suppression ratios of the frequency points in the target azimuth angles, and the audio signals of the target azimuth angles can be subjected to audio signal enhancement (reweighting) by using the audio suppression ratios of the frequency points, so that the peak value of the target azimuth angle in which the point noise source is located can be effectively suppressed, and the signal peak value of the target azimuth angle in which the speaker is located can be effectively highlighted, so that the point noise can be effectively suppressed, and the robustness of the speaker positioning algorithm can be improved. It should be noted that when the number of point noise sources is greater than one, the number of selected target azimuth angles can be adjusted by the embodiment, which is not strictly limited herein.
[0096] It can be seen that, after the multi-channel audio signal is divided into multiple azimuth angle audio signals, the target azimuth angle containing the point noise source and the speaker can be selected from the multiple azimuth angles according to the peak value of the audio signal, so that the audio signal of the target azimuth angle is intelligently denoised, and the azimuth angle where the speaker is located can be selected from the target azimuth angle based on the peak value of the signal after the intelligent denoising, so that the number of azimuth angles subjected to the intelligent denoising is reduced to reduce the system calculation amount, thereby reducing the system calculation amount while effectively suppressing the point noise.
[0097] To further reduce the system calculation amount, as Figure 2 As an optional implementation of the flow shown in FIG. 6, in this implementation, the target audio signal subjected to the intelligent denoising can be a single-channel audio signal converted into the frequency domain. Figure 5 An exemplary flowchart of another optional denoising method provided by the embodiments of the present application is shown. Referring to FIG. 7, Figure 5 The method flow can include the following steps.
[0098] In step S510, a multi-channel audio signal is acquired, and the multi-channel audio signal is respectively converted from the time domain to the frequency domain.
[0099] In step S511, a single-channel audio signal is selected from the multi-channel audio signal, and the single-channel audio signal is subjected to intelligent denoising to obtain an audio suppression ratio of each frequency point in the single-channel audio signal.
[0100] After the multi-channel audio signal is converted from the time domain to the frequency domain, the embodiments of the present application can select a single-channel audio signal from the multi-channel audio signal, and subject the selected single-channel audio signal to intelligent denoising to obtain an audio suppression ratio of each frequency point in the single-channel audio signal. As an optional implementation, the embodiments of the present application can select a single-channel audio signal from the multi-channel audio signal at random, or select a single-channel audio signal in the middle from the multi-channel audio signal.
[0101] In an example, the embodiments of the present application can select a single-channel audio signal collected by a single microphone from a plurality of microphones of a microphone array, and subject the single-channel audio signal collected by the single microphone to intelligent denoising. In a more specific example, assuming that the difference between the signal-to-noise ratios of the speech and noise picked up by each omnidirectional microphone is within a preset range (i.e., the difference between the signal-to-noise ratios of the speech and noise picked up by each omnidirectional microphone is not large), the embodiments of the present application can select a single-channel audio signal collected by a single microphone in the middle, and subject the single-channel audio signal collected by the single microphone in the middle to intelligent denoising.
[0102] In this embodiment, the selected single-channel audio signal can be a target audio signal subjected to signal enhancement processing.
[0103] In step S512, after the multi-channel audio signals converted into the frequency domain are respectively subjected to amplitude normalization processing, the multi-channel audio signals are subjected to signal enhancement processing according to the audio suppression ratios of the respective frequency points in the one audio signal, to obtain multi-channel enhanced audio signals.
[0104] After the selected one audio signal is subjected to intelligent noise reduction processing and the audio suppression ratios of the respective frequency points in the one audio signal are obtained, the multi-channel audio signals after the amplitude normalization processing can be subjected to signal enhancement processing (such as re-weighting) by the audio suppression ratios of the respective frequency points in the one audio signal, to obtain multi-channel enhanced audio signals. As an optional implementation, the audio suppression ratios of the respective frequency points in the one audio signal can be applied to the audio signals of the relative frequency points in the one audio signal and other audio signals, so as to implement signal enhancement processing of all audio signals by using the audio suppression ratios of the frequency points in the one audio signal. For example, the Mask of the audio signal collected by one microphone at the frequency points can be applied to the audio signals collected by other microphone channels, to implement signal enhancement processing of the multi-channel microphone collected audio signals.
[0105] In this embodiment, the preprocessed multi-channel audio signals can include audio signals containing noise and audio signals containing speakers.
[0106] In step S513, the multi-channel enhanced audio signals are divided into multiple azimuth angle enhanced audio signals according to a preset azimuth angle precision.
[0107] After the multi-channel enhanced audio signals are obtained, the embodiment of the present application can divide the multi-channel enhanced audio signals according to the preset azimuth angle precision, to obtain multiple azimuth angle enhanced audio signals. As an implementation process, the same as the description of the corresponding part in the foregoing, it will not be expanded here.
[0108] In step S514, a speaker positioning result is determined according to the respective azimuth angle enhanced audio signals.
[0109] Optionally, after the multiple azimuth angle enhanced audio signals are obtained, the embodiment of the present application can determine the azimuth angle with the maximum signal peak, and take the azimuth angle with the maximum signal peak as the azimuth angle of the speaker.
[0110] Assuming that the signal-to-noise ratio of each frequency point of each audio signal in the multi-channel audio signal collected by the audio device is similar (for example, the signal-to-noise ratio of each frequency point received by each microphone in the microphone array, especially in the case of using an omnidirectional microphone), the embodiments of the present application can select an audio signal collected (for example, an audio signal collected by a microphone in the microphone array) for intelligent noise reduction processing, thereby obtaining the audio suppression ratio of each frequency point of the audio signal, and then reweighting all audio signals (for example, all audio signals collected by the microphones) to obtain enhanced audio signals of each channel. Then, the direction angle of the speaker is selected by signal peak value. The embodiments of the present application can only perform intelligent noise reduction processing on one audio signal collected, effectively suppress point noise, and highlight the voice of the speaker, while greatly reducing the system calculation amount.
[0111] The embodiments of the present application can integrate intelligent noise reduction technology in the speaker positioning algorithm, effectively track and estimate the energy of different types of noise, especially non-stationary point noise, and control the calculation amount of the system, effectively suppress point noise, and highlight the voice of the speaker, thereby improving the robustness of the speaker positioning algorithm.
[0112] In further embodiments, in the optional implementation of obtaining the audio suppression ratio corresponding to the frequency point of the target audio signal, if the target audio signal has beam forming regions in different directions, the audio suppression ratio is a mask, and the embodiments of the present application can combine the masks corresponding to the target audio signal in different directions of the beam forming region, thereby obtaining a combined mask, which can be used as the audio suppression ratio corresponding to the frequency point of the target audio signal. Then, the embodiments of the present application can use the combined mask to perform signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker, so as to reduce the weighting of the noise direction in the speaker positioning algorithm.
[0113] It should be noted that when positioning the speaker, the embodiments of the present application can collect audio signals through an audio collection array (for example, a microphone array), and the audio collection array can be a linear array or a ring array. The linear array and the ring array can have multiple directions of beam forming, which can be M directions of beam forming, including beam forming in direction one, beam forming in direction two, and beam forming in direction M. For the linear array, as an example, Figure 6A An example of beam forming of the linear array is shown, specifically, Figure 6A An example of beam forming of the linear array in direction one, direction two, and direction M is shown, which can be referred to. For the ring array, Figure 6BAn exemplary beam forming example diagram of a circular array is shown, in particular, Figure 6B An exemplary beam forming of a circular array in direction one, direction two to direction M is shown, which can be referred to.
[0114] It should be noted that in each frequency band of each frame of the audio spectrum, the beam forming area of different directions has different masks; based on this, the embodiment of the application can combine the masks corresponding to the target audio signal in the beam forming area of different directions after determining the target audio signal to obtain a combined mask; and then use the combined mask to perform signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker, so as to further reduce the weighting of the noise direction and improve the weighting of the speaker audio direction in the speaker positioning algorithm, thereby improving the accuracy of the speaker positioning algorithm. This method can be applied to the case where the target audio signal is an audio signal of each azimuth angle, an audio signal of a part of azimuth angles with the maximum signal peak value, an audio signal converted into a frequency domain, etc.; as long as the frequency band of the target audio signal has a mask corresponding to the beam forming area of different directions.
[0115] As an optional implementation, Figure 7A An exemplary yet another optional flowchart of the noise reduction method provided by the embodiment of the application is shown. The method flow can be implemented by an audio device, such as a microphone array or other device with audio acquisition and processing capability. Referring to Figure 7A The method flow can include the following steps.
[0116] In step S710, the target audio signal is subjected to beam forming processing in multiple directions, and the mask corresponding to the beam forming area of different directions of the target audio signal is determined.
[0117] After determining the target audio signal that needs to be subjected to intelligent noise reduction processing, the embodiment of the application can subject the target audio signal to beam forming processing in multiple directions, so as to determine the mask corresponding to the beam forming area of different directions of the target audio signal.
[0118] It should be further noted that the mask can be specifically a TF-Mask, and the TF-Mask is a short name of Time-Frequency Mask (time-frequency domain mask), and the implementation of the TF-Mask can be realized based on an intelligent noise reduction algorithm. For example, the signal processing algorithm (minimum statistics, iMCRA) and other algorithms can be used for steady-state noise estimation, a deep learning based data driven method can be used to obtain steady-state noise or non-steady-state noise estimation, or both can be used, and then the fusion of noise estimation is performed to obtain the final TF-Mask.
[0119] In step S711, a combined mask is determined according to the masks corresponding to the beamforming regions in different directions.
[0120] The embodiment of the present application can perform beamforming processing on the target audio signal of each frame in different directions (for example, the audio signal of each frame is subjected to beamforming processing in direction one, direction two, direction M, and the like), so that for each frequency band of the target audio signal of each frame, the embodiment of the present application can correspond to different masks in different beamforming regions, such as for a frequency band of a target audio signal, a mask corresponding to a beamforming region in a direction. Optionally, when calculating the mask, the embodiment of the present application can use a signal processing method or a deep learning model to calculate the mask corresponding to the beamforming region in different directions for each frequency band of the audio signal of each frame.
[0121] For example, after the target audio signal is subjected to beamforming processing, the target audio signal can calculate a value of 0 to 1 on each time-frequency (time-frequency domain, that is, each frame and each frequency band), which can be regarded as a mask corresponding to a beamforming direction of the target audio signal in the frequency band. Figure 6B As shown in FIG. 7, if there are M different directions of beamforming, there are M different masks for a frequency band of a target audio signal of a frame, wherein the mask corresponding to the beamforming region in the mth direction (direction m) can be represented as Mask m (ω, n), wherein n represents the frame number of the audio signal, and ω represents the frequency band.
[0122] After obtaining the masks corresponding to the beamforming regions in different directions of the target audio signal, the embodiment of the present application can combine the masks corresponding to the beamforming regions in different directions to obtain a combined mask, which can be regarded as a time-frequency (TF) spatial (space) mask (mask). For example, for each frequency band of each frame of audio signal, the masks corresponding to the beamforming regions in different directions are combined to obtain a combined mask.
[0123] In step S712, the signal enhancement processing is performed on the audio signal containing noise and the audio signal containing the speaker according to the combined mask.
[0124] Optionally, after obtaining the combined mask, the embodiment of the present application can apply the combined mask to a sound source positioning algorithm (for example, a speaker positioning algorithm), so as to reduce the weighting of the audio signal in the noise direction and improve the weighting of the audio signal in the speaker direction based on the combined mask, to realize signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker, and thus obtain a more accurate sound source positioning result of the speaker.
[0125] For ease of understanding, Figure 7B An example diagram for realizing sound source positioning by the embodiment of the present application is shown as follows: Figure 7B As shown, after the target audio signal is subjected to beamforming in M directions, beamforming in direction 1 to direction M can be output, and each direction of beamforming determines a corresponding Mask. Then, the Mask corresponding to each direction of beamforming is combined to obtain a TF spatial Mask (time-frequency domain spatial mask, i.e., the combined mask referred to in the embodiment of the present application). The TF spatial Mask is applied to a sound source positioning algorithm with adaptive weights, so as to reduce the weighting of the audio signal in the noise direction and improve the weighting of the audio signal in the speaker direction, and thus obtain the speaker direction (i.e., the sound source positioning result of the speaker). For example, for the audio signal containing noise and the audio signal containing the speaker, the embodiment of the present application can use the combined mask to reduce the weighting of the audio signal in the noise direction and improve the weighting of the audio signal in the speaker direction, so as to realize signal enhancement processing.
[0126] As an optional implementation, Figures 2 to 5 As shown in the method flow, the audio suppression ratio of the frequency point (for example, the audio suppression ratio corresponding to the frequency point of the target audio signal) can be a separate Mask of the frequency point, or can be a combined mask based on the principle of the method flow shown in Figure 7A The combined mask is determined according to the principle of the method flow shown in Figures 2 to 5 As shown in the method flow, when the target audio signal is subjected to intelligent noise reduction processing to obtain the audio suppression ratio corresponding to the frequency point of the target audio signal, the embodiment of the present application can determine the Mask corresponding to the beamforming region in different directions of each frequency band of the target audio signal according to the principle of the method flow shown in Figure 7A The combined mask is determined according to the principle of the method flow shown in Figures 2 to 5 The method flow shown in Figure 7A The method for determining the combined mask shown in can be realized by cross-referencing the corresponding parts in the foregoing, which will not be expanded here.
[0127] The noise reduction device provided by the embodiments of the present application is introduced as follows. The noise reduction device described below can be a functional module required to be set by an electronic device (for example, an audio device such as a microphone array) to implement the noise reduction method provided by the embodiments of the present application. The device content described below can be referred to the method content described above.
[0128] As an optional implementation, Figure 8 An exemplary block diagram of the noise reduction device provided by the embodiments of the present application is shown. The device can be applied to an electronic device, such as Figure 8 As shown in the figure, the device can include:
[0129] The audio signal acquisition and preprocessing module 801 is configured to acquire multiple audio signals. Optionally, the audio signal acquisition and preprocessing module 801 can also preprocess the multiple audio signals respectively.
[0130] The target audio signal determination module 802 is configured to determine a target audio signal for estimating noise energy according to the multiple audio signals.
[0131] The intelligent noise reduction module 803 is configured to perform intelligent noise reduction processing on the target audio signal to obtain an audio suppression ratio corresponding to a frequency point of the target audio signal.
[0132] The signal enhancement module 804 is configured to perform signal enhancement processing on the audio signal containing noise and the audio signal containing a speaker according to the audio suppression ratio corresponding to the frequency point.
[0133] The result determination module 805 is configured to determine a speaker positioning result according to the signal enhancement processing result.
[0134] Optionally, the signal enhancement module 804 is configured to perform signal enhancement processing on the audio signal containing noise and the audio signal containing a speaker according to the audio suppression ratio corresponding to the frequency point, including:
[0135] In the audio signal containing noise and the audio signal containing a speaker, the audio signal corresponding to the frequency point is reweighted using the audio suppression ratio corresponding to the frequency point.
[0136] Optionally, the signal enhancement module 804 is configured to perform signal enhancement processing on the audio signal containing noise and the audio signal containing a speaker according to the audio suppression ratio corresponding to the frequency point, including:
[0137] In the audio signal containing noise and the audio signal containing a speaker, the audio signal corresponding to the frequency point is reweighted using the audio suppression ratio corresponding to the frequency point.
[0138] Optionally, the target audio signal comprises each azimuth angle audio signal in a plurality of azimuth angle audio signals, or part of the azimuth angle audio signals, or one audio signal converted into a frequency domain in the plurality of audio signals; wherein the plurality of audio signals is divided into a plurality of azimuth angles.
[0139] In one aspect, optionally, the target audio signal is each azimuth angle audio signal in a plurality of azimuth angle audio signals;
[0140] The device can also be used to divide the plurality of audio signals into a plurality of azimuth angle audio signals according to a preset azimuth angle accuracy;
[0141] The target audio signal determination module 802 is configured to determine the target audio signal for estimating the noise energy according to the plurality of audio signals, including: taking the plurality of azimuth angle audio signals as the target audio signal;
[0142] The intelligent noise reduction module 803 is configured to perform intelligent noise reduction processing on the target audio signal to obtain an audio suppression ratio corresponding to each frequency point of the target audio signal, including: performing intelligent noise reduction processing on the plurality of azimuth angle audio signals respectively to obtain an audio suppression ratio of each frequency point;
[0143] The audio suppression ratio of each frequency point is used to perform signal enhancement processing on each azimuth angle audio signal to obtain an enhanced audio signal of each azimuth angle;
[0144] The result determination module 805 is configured to determine the speaker positioning result according to the signal enhancement processing result, including: determining the speaker positioning result according to the enhanced audio signal of each azimuth angle.
[0145] In another aspect, optionally, the target audio signal is part of the azimuth angle audio signals in the plurality of azimuth angle audio signals;
[0146] The device can also be used to divide the plurality of audio signals into a plurality of azimuth angle audio signals according to a preset azimuth angle accuracy;
[0147] The target audio signal determination module 802 is configured to determine the target audio signal for estimating the noise energy according to the plurality of audio signals, including: determining a peak value of each azimuth angle audio signal, and determining at least two target azimuth angles with a peak value meeting a preset condition according to the peak value of each azimuth angle audio signal, the number of the at least two target azimuth angles being less than the plurality of azimuth angles; wherein the audio signal of the at least two target azimuth angles is taken as the target audio signal;
[0148] The intelligent noise reduction module 803 is used to perform intelligent noise reduction processing on the target audio signal to obtain the audio suppression ratio corresponding to the frequency point of the target audio signal, including: performing intelligent noise reduction processing on the audio signal of each target azimuth angle to obtain the audio suppression ratio of each frequency point in each target azimuth angle;
[0149] The audio suppression ratio of each frequency point in each target azimuth angle is used to perform signal enhancement processing on the audio signal of each target azimuth angle to obtain the enhanced audio signal of each target azimuth angle.
[0150] The result determination module 805 is used to determine the speaker location result based on the signal enhancement processing result, including: determining the speaker location result based on the enhanced audio signal of each target azimuth angle.
[0151] Alternatively, the target audio signal may be one of the multiple audio signals that has been converted to the frequency domain.
[0152] The target audio signal determination module 802 is used to determine the target audio signal for estimating noise energy based on the multiple audio signals, including: selecting one audio signal from the multiple audio signals, and the one audio signal being used as the target audio signal;
[0153] The intelligent noise reduction module 803 is used to perform intelligent noise reduction processing on the target audio signal to obtain the audio suppression ratio corresponding to the frequency point of the target audio signal, including: performing intelligent noise reduction processing on a selected audio signal to obtain the audio suppression ratio of each frequency point in the audio signal;
[0154] The audio suppression ratio at each frequency point in the audio signal is used to perform signal enhancement processing on the multiple audio signals to obtain multiple enhanced audio signals.
[0155] Furthermore, the device can also be used to: divide a multi-channel enhanced audio signal into multiple enhanced audio signals with different azimuth angles according to a preset azimuth angle accuracy;
[0156] Correspondingly, the result determination module 805 is used to determine the speaker location result based on the signal enhancement processing result, including: determining the speaker location result based on the enhanced audio signals from various directional angles.
[0157] Optionally, the intelligent noise reduction module 803 is used to perform intelligent noise reduction processing on the target audio signal to obtain the audio suppression ratio corresponding to the frequency point of the target audio signal, including:
[0158] The target audio signal is processed by beamforming in multiple directions, and the masking of the target audio signal in beamforming regions in different directions is determined based on an intelligent noise reduction algorithm.
[0159] According to the different direction beamforming area corresponding to the mask, a combined mask is determined as the audio suppression ratio corresponding to the frequency point of the target audio signal.
[0160] Optionally, the signal enhancement module 804 is configured to perform signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker according to the audio suppression ratio corresponding to the frequency point, including:
[0161] According to the combined mask, the signal enhancement processing is performed on the audio signal containing noise and the audio signal containing the speaker.
[0162] Optionally, the signal enhancement module 804 is configured to perform signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker according to the combined mask, including:
[0163] For the audio signal containing noise and the audio signal containing the speaker, the combined mask is used to reduce the weighting of the audio signal in the noise direction and to increase the weighting of the audio signal in the speaker direction.
[0164] Further, the embodiments of the present application also provide an electronic device, such as a microphone array and the like audio device, which can be provided with any one of the noise reduction devices provided by the embodiments of the present application to implement the noise reduction method provided by the embodiments of the present application. Optionally, Figure 9 An exemplary optional block diagram of an electronic device is shown, as shown in Figure 9 The electronic device can include at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4.
[0165] In the embodiments of the present application, the number of processors 1, communication interfaces 2, memories 3 and communication buses 4 is at least one, and the processor 1, the communication interface 2 and the memory 3 complete communication with each other through the communication bus 4.
[0166] Optionally, the communication interface 2 can be the interface of the communication module for network communication.
[0167] Optionally, the processor 1 can be a CPU, a GPU (Graphics Processing Unit), a NPU (Neural Processing Unit), a FPGA (Field Programmable Gate Array), a TPU (Tensor Processing Unit), an AI chip, an ASIC (Application Specific Integrated Circuit), or an integrated circuit configured to implement one or more embodiments of the present application, etc.
[0168] The memory 3 can include a high-speed RAM memory, and can also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0169] The memory 3 stores one or more computer executable instructions, and the processor 1 invokes the one or more computer executable instructions to execute the noise reduction method provided by the embodiments of the present application.
[0170] Further, the embodiments of the present application also provide a storage medium storing one or more computer executable instructions, and the one or more computer executable instructions are executed to implement the noise reduction method provided by the embodiments of the present application.
[0171] Further, the embodiments of the present application also provide a computer program, and the computer program is executed to implement the noise reduction method provided by the embodiments of the present application.
[0172] The above describes a plurality of embodiment schemes provided by the embodiments of the present application, and each optional mode introduced by each embodiment scheme can be combined, cross-referenced in the case of no conflict, thereby extending a plurality of possible embodiment schemes, which can be considered as the embodiments disclosed and disclosed by the embodiments of the present application.
[0173] Although the embodiments of the present application are disclosed as above, the present application is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and therefore the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A method of noise reduction, wherein, The method comprises: acquiring a plurality of audio signals, wherein, in the case of dividing the audio signals according to azimuth angles, the plurality of audio signals are divided into a plurality of azimuth angles; determining a target audio signal for estimating noise energy according to the plurality of audio signals, the target audio signal being an audio signal of an azimuth angle for which noise energy estimation is required, and the target audio signal comprising audio signals of part of the azimuth angles; performing intelligent noise reduction processing on the target audio signal to obtain an audio suppression ratio corresponding to a frequency point of the target audio signal, the audio suppression ratio corresponding to the frequency point reflecting noise energy of the target audio signal; performing signal enhancement processing on audio signals containing noise and audio signals containing a speaker according to the audio suppression ratio corresponding to the frequency point of the target audio signal; determining a speaker positioning result according to the signal enhancement processing result.
2. The method of claim 1, wherein, The signal enhancement processing on the audio signals containing noise and the audio signals containing the speaker according to the audio suppression ratio corresponding to the frequency point comprises: re-weighting the audio signals corresponding to the frequency point in the audio signals containing noise and the audio signals containing the speaker using the audio suppression ratio corresponding to the frequency point.
3. The method of claim 2, wherein, The re-weighting of the audio signals corresponding to the frequency point in the audio signals containing noise and the audio signals containing the speaker using the audio suppression ratio corresponding to the frequency point comprises: adding the audio suppression ratio corresponding to the frequency point to a formula for calculating the correlation of the audio signals to re-weight the audio signals containing noise and the audio signals containing the speaker.
4. The method according to any one of claims 1 to 3, wherein, The target audio signal comprises: an audio signal of each azimuth angle in a plurality of azimuth angle audio signals, or an audio signal of part of the azimuth angles, or a converted frequency domain audio signal in the plurality of audio signals; wherein the plurality of audio signals are divided into a plurality of azimuth angles.
5. The method of claim 4, wherein, The target audio signal is an audio signal of each azimuth angle in a plurality of azimuth angle audio signals. The method further comprises: dividing the plurality of audio signals into a plurality of azimuth angle audio signals according to a preset azimuth angle accuracy; determining the target audio signal for estimating noise energy according to the plurality of azimuth angle audio signals comprises taking the plurality of azimuth angle audio signals as the target audio signal; performing intelligent noise reduction processing on the plurality of azimuth angle audio signals respectively to obtain an audio suppression ratio of each frequency point; wherein the audio suppression ratio of each frequency point is used to perform signal enhancement processing on the audio signal of each azimuth angle to obtain an enhanced audio signal of each azimuth angle; determining the speaker positioning result according to the enhanced audio signal of each azimuth angle.
6. The method of claim 4, wherein, The target audio signal is an audio signal of part of the azimuth angles in a plurality of azimuth angle audio signals. The method further comprises: According to a preset azimuth angle precision, the multiple audio signals are divided into multiple azimuth angle audio signals; The target audio signal for estimating noise energy is determined according to the multiple audio signals, and the method comprises: determining the peak value of each azimuth angle audio signal, and determining at least two target azimuth angles according to the peak value of each azimuth angle audio signal, the number of the at least two target azimuth angles being less than the number of the multiple azimuth angles; wherein the audio signal of the at least two target azimuth angles is the target audio signal; the target audio signal is subjected to intelligent noise reduction processing to obtain the audio suppression ratio corresponding to the frequency point of the target audio signal; wherein the audio suppression ratio of each frequency point in each target azimuth angle is used to perform signal enhancement processing on the audio signal of each target azimuth angle to obtain an enhanced audio signal of each target azimuth angle; the speaker positioning result is determined according to the signal enhancement processing result.
7. The method of claim 4, wherein, The target audio signal is a frequency domain converted audio signal in the multiple audio signals; The target audio signal for estimating noise energy is determined according to the multiple audio signals, and the method comprises: selecting an audio signal from the multiple audio signals, the selected audio signal being the target audio signal; the selected audio signal is subjected to intelligent noise reduction processing to obtain the audio suppression ratio of each frequency point in the selected audio signal; wherein the audio suppression ratio of each frequency point in the selected audio signal is used to perform signal enhancement processing on the multiple audio signals to obtain multiple enhanced audio signals; The method further comprises: According to a preset azimuth angle precision, the multiple audio signals are divided into multiple azimuth angle audio signals; 8. The method of claim 1, wherein, The speaker positioning result is determined according to the signal enhancement processing result. The target audio signal is subjected to intelligent noise reduction processing to obtain the audio suppression ratio corresponding to the frequency point of the target audio signal, and the method comprises: The target audio signal is subjected to beamforming processing in multiple directions, and the target audio signal corresponding to the mask of the beamforming region in different directions is determined based on an intelligent noise reduction algorithm; a combined mask is determined according to the mask corresponding to the beamforming region in different directions, the combined mask being the audio suppression ratio corresponding to the frequency point of the target audio signal; The signal enhancement processing is performed on the audio signal containing noise and the audio signal containing the speaker according to the audio suppression ratio corresponding to the frequency point, and the method comprises: The signal enhancement processing is performed on the audio signal containing noise and the audio signal containing the speaker according to the combined mask.
9. The method of claim 8, wherein, The signal enhancement processing on the audio signal containing noise and the audio signal containing the speaker according to the combined mask includes: For the audio signal containing noise and the audio signal containing the speaker, the combined mask is used to reduce the weighting of the audio signal in the noise direction and to increase the weighting of the audio signal in the speaker direction.
10. An electronic device, comprising: The device comprises at least one memory and at least one processor, the memory stores one or more computer executable instructions, and the processor invokes the one or more computer executable instructions to execute the noise reduction method according to any one of claims 1-9.
11. A storage medium, wherein, The storage medium stores one or more computer executable instructions, and the one or more computer executable instructions are executed to implement the noise reduction method according to any one of claims 1-9.
Citation Information
Patent Citations
Array speech enhancement algorithm
CN109308904A
Method for sound source direction estimation based on time frequency masking and deep neural network
CN109839612A
Voice enhancing algorithm based on second-order differential microphone array
CN110310650A