Audio signal processing method and device, electronic equipment and readable storage medium

By dividing frequency bands in electronic devices and fusing transmission channel information from different microphones, the problem of poor robustness in audio signal processing is solved, achieving a more natural audio signal noise reduction effect.

CN116095565BActive Publication Date: 2025-12-05VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211095430.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-12-05
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

In the prior art, electronic devices have poor robustness in processing audio signals due to the poor reliability of single-microphone wind noise characteristics and the unevenness of dual-microphone MSC wind noise detection results.

Method used

The target frequency range is divided into different frequency bands, and the transmission channel information of different microphones in each frequency band is fused, including the superposition of information such as amplitude spectrum and wind noise gain. Noise reduction is then performed by combining the transmission channel information of different audio signals in the divided frequency bands.

Benefits of technology

It improves the robustness of electronic devices in processing audio signals, ensures that the audio signal sounds natural and continuous after noise reduction, and enhances the effectiveness of wind noise suppression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116095565B_ABST
    Figure CN116095565B_ABST
Patent Text Reader

Abstract

The application discloses an audio signal processing method and device, electronic equipment and a readable storage medium, and belongs to the technical field of audio. The method comprises the following steps: dividing a target frequency range into a first frequency band and a second frequency band according to a noise frequency band of a first audio signal and a noise frequency band of a second audio signal, the first audio signal being an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal being an audio signal obtained by collecting the target sound source by a second microphone; performing first fusion processing on transmission channel information corresponding to the first audio signal and transmission channel information corresponding to the second audio signal in the first frequency band; performing second fusion processing on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal in the second frequency band; and performing noise reduction on a target audio signal after fusion processing of the corresponding transmission channel information, the target audio signal comprising at least one of the first audio signal and the second audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of audio technology, and particularly relates to an audio signal processing method and device, electronic equipment and a readable storage medium. BACKGROUND

[0002] At present, multiple microphones are usually arranged in electronic equipment, and a user can make a call, record a sound or record a video through the multiple microphones. However, in audio processing in different scenes, environmental wind noise can greatly reduce the subjective listening of audio.

[0003] Taking an electronic device with two microphones as an example, in a traditional noise reduction method, the electronic device can use a double-microphone frequency domain magnitude-squared coherence (MSC) to realize wind noise detection, map the detected wind noise to wind noise suppression gain, and then realize wind noise suppression in combination with a single-microphone wind noise feature.

[0004] However, according to the above method, on the one hand, the single-microphone wind noise feature usually has poor reliability, and on the other hand, the wind noise detection result based on the double-microphone MSC usually includes all wind noise frequency points of the double microphones. Directly mapping to wind noise gain will damage the audio signal on the microphone with a low wind noise bandwidth, thus resulting in poor robustness of the electronic device in processing the audio signal. SUMMARY

[0005] Embodiments of the present application provide an audio signal processing method, device, electronic equipment and readable storage medium, which can solve the problem of poor robustness of the electronic device in processing the audio signal.

[0006] In a first aspect, an audio signal processing method is provided, which includes: dividing a target frequency range into a first frequency band and a second frequency band according to a noise frequency band of a first audio signal and a noise frequency band of a second audio signal, the first audio signal being an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal being an audio signal obtained by collecting the target sound source by a second microphone; performing first fusion processing on transmission channel information corresponding to the first audio signal and transmission channel information corresponding to the second audio signal in the first frequency band; performing second fusion processing on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal in the second frequency band; and performing noise reduction on a target audio signal after fusion processing of the corresponding transmission channel information, the target audio signal including at least one of the first audio signal and the second audio signal.

[0007] In a second aspect, an embodiment of the present application provides an audio signal processing apparatus, the apparatus comprising a division module, a fusion module and a noise reduction module; the division module is configured to divide a target frequency range into a first frequency band and a second frequency band according to a noise frequency band of a first audio signal and a noise frequency band of a second audio signal, the first audio signal being an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal being an audio signal obtained by collecting the target sound source by a second microphone; the fusion module is configured to perform first fusion processing on transmission channel information corresponding to the first audio signal and transmission channel information corresponding to the second audio signal in the first frequency band; the fusion module is further configured to perform second fusion processing on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal in the second frequency band; and the noise reduction module is configured to perform noise reduction on a target audio signal after fusion processing of the corresponding transmission channel information, the target audio signal comprising at least one of the first audio signal and the second audio signal.

[0008] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory, the memory storing programs or instructions executable on the processor, and the programs or instructions being executed by the processor to implement the steps of the method according to the first aspect.

[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, the readable storage medium storing programs or instructions, and the programs or instructions being executed by a processor to implement the steps of the method according to the first aspect.

[0010] In a fifth aspect, an embodiment of the present application provides a chip, the chip comprising a processor and a communication interface, the communication interface being coupled to the processor, and the processor being configured to run programs or instructions to implement the method according to the first aspect.

[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, the program product being stored in a storage medium, and the program product being executed by at least one processor to implement the method according to the first aspect.

[0012] In the embodiment of the present application, the target frequency range can be divided into a first frequency band and a second frequency band according to the noise frequency band of the first audio signal and the noise frequency band of the second audio signal, the first audio signal is an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal is an audio signal obtained by collecting the target sound source by a second microphone; and in the first frequency band, the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal are subjected to first fusion processing; and in the second frequency band, the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal are subjected to second fusion processing; and the target audio signal after the fusion processing of the corresponding transmission channel information is subjected to noise reduction, and the target audio signal includes at least one of the first audio signal and the second audio signal. Through the scheme, since the electronic device can first perform fusion processing of the transmission channel information based on the divided frequency bands and the transmission channel information corresponding to each audio signal before performing noise reduction processing on the audio signals collected by different microphones, and then performs noise reduction on the audio signal after the fusion processing of the corresponding transmission channel information, the electronic device can not need to be based on the characteristics of a single audio signal or all frequency points of multiple audio signals when processing the audio signal, but can combine the transmission channel information corresponding to different audio signals in different frequency bands, thereby improving the robustness of the electronic device in processing the audio signal. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a flowchart of an audio signal processing method provided by the embodiment of the present application;

[0014] Figure 2 is one of the schematic diagrams of the audio signal processing method provided by the embodiment of the present application;

[0015] Figure 3 is the second schematic diagram of the audio signal processing method provided by the embodiment of the present application;

[0016] Figure 4 is the third schematic diagram of the audio signal processing method provided by the embodiment of the present application;

[0017] Figure 5 is the fourth schematic diagram of the audio signal processing method provided by the embodiment of the present application;

[0018] Figure 6 is the fifth schematic diagram of the audio signal processing method provided by the embodiment of the present application;

[0019] Figure 7 is an information flow schematic diagram of the audio signal processing method provided by the embodiment of the present application applied to double-microphone stereo robust wind noise detection and suppression;

[0020] Figure 8is a schematic diagram of an audio signal processing device provided by an embodiment of the present application;

[0021] Figure 9 is a schematic diagram of an electronic device provided by an embodiment of the present application;

[0022] Figure 10 is a hardware schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0024] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a category, and are not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents a "or" relationship between the front and rear associated objects.

[0025] The audio signal processing method, device, electronic device and readable storage medium provided by the embodiments of the present application will be described in detail below with reference to the drawings, through specific embodiments and application scenarios.

[0026] In outdoor conversation or audio recording, the electronic device usually collects a large amount of environmental sound, including various stationary noises and non-stationary noises. Generally, the noise comes from various sound sources in the environment, but the wind noise in the audio collection scene is mainly caused by the turbulent airflow near the microphone membrane, which can cause the microphone to produce a relatively high signal level, and the sound source of the wind noise is near the microphone. Natural wind noise mainly occurs in the low frequency range of 1kHz, and rapidly attenuates to high frequency. Sudden gusts of wind often cause wind noise with a duration of tens to hundreds of milliseconds. And due to the suddenness of the gusts of wind, the wind noise may produce a high amplitude exceeding the expected amplitude of the collected audio, showing significant non-stationary characteristics, which can greatly reduce the subjective listening of the audio, and thus an effective wind noise suppression method is needed.

[0027] Currently, from the technical means, the wind noise suppression method includes acoustic method and signal processing method. The acoustic method is to isolate the wind noise from the physical point of view, to suppress the wind noise interference from the source of signal acquisition, such as to realize the wind noise suppression by wind shield, wind noise resistant pipe and accelerometer pickup, but the application scene of this method will be limited by physical conditions; the signal processing method is to use signal processing means to realize the suppression or separation of wind noise mixed with audio, which may also contain the reconstruction of damaged audio, which can generally cope with various wind noise scenes.

[0028] In the signal processing method, the traditional wind noise suppression strategy is usually based on a single microphone (or microphone), which realizes wind noise detection, estimation and suppression through spectral centroid method, noise template method, morphological method or deep learning method, with single-mic wind noise characteristics. However, the current smart phones or truly wireless stereo earphones usually have 2 or more microphones, based on the above wind noise formation principle, the double-mic wind noise is formed by the turbulent flow near the relatively independent microphones, and the coherence (or correlation) of the two is usually very low. The traditional double-mic wind noise suppression largely depends on this characteristic, uses the frequency domain amplitude squared coherence coefficient (Magnitude-Squared Coherence, MSC) to realize wind noise detection, and maps the detected wind noise to wind noise suppression gain. However, in the double-mic stereo, the wind noise detection result usually includes all the wind noise points of the double-mic, so the detection and estimation result may only correspond to one of the microphones, and is not suitable for the other microphone.

[0029] It can be seen that the traditional double-mic wind noise suppression signal processing method usually relies heavily on the MSC feature, and realizes wind noise suppression combined with the single-mic wind noise feature with relatively low reliability, but has the following shortcomings:

[0030] 1. The wind noise detection result based on double-mic MSC includes all the wind noise points of the double-mic, and is not suitable for both microphones. Direct mapping to wind noise gain will damage the audio on the microphone with low wind noise bandwidth.

[0031] 2. The single-mic feature usually has poor reliability, resulting in insufficient robustness of wind noise suppression.

[0032] To solve the above problems, in the audio signal processing method provided in the embodiments of the present application, the target frequency range can be divided into a first frequency band and a second frequency band according to the noise frequency band of the first audio signal and the noise frequency band of the second audio signal, the first audio signal is an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal is an audio signal obtained by collecting the target sound source by a second microphone; the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal are subjected to first fusion processing in the first frequency band; the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal are subjected to second fusion processing in the second frequency band; and the target audio signal after the fusion processing of the corresponding transmission channel information is subjected to noise reduction, the target audio signal including at least one of the first audio signal and the second audio signal. Through the scheme, since the electronic device can perform fusion processing of the transmission channel information based on the divided frequency bands and the transmission channel information corresponding to each audio signal before performing noise reduction processing on the audio signals collected by different microphones, the electronic device can perform noise reduction on the audio signal after the fusion processing of the corresponding transmission channel information, and thus the electronic device can improve the robustness of processing the audio signal without being based on the characteristics of a single audio signal or all frequency points of multiple audio signals, but can combine the transmission channel information corresponding to different audio signals in different frequency bands.

[0033] The embodiments of the present application provide an audio signal processing method, Figure 1 A flowchart of the audio signal processing method provided by the embodiments of the present application is shown. As shown in the figure, Figure 1 The audio signal processing method provided by the embodiments of the present application can include the following steps 101 to 104. The method is exemplarily described below taking that an electronic device performs the method as an example.

[0034] Step 101, the electronic device divides a target frequency range into a first frequency band and a second frequency band according to a noise frequency band of a first audio signal and a noise frequency band of a second audio signal.

[0035] In the embodiments of the present application, the first audio signal is an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal is an audio signal obtained by collecting the target sound source by a second microphone.

[0036] Optionally, in the embodiments of the present application, the first audio signal and the second audio signal are audio signals collected at the same time.

[0037] Optionally, in the embodiments of the present application, the first microphone and the second microphone can be microphones arranged in the same electronic device, or can be microphones arranged in different electronic devices.

[0038] In the embodiments of the present application, the target frequency range is a frequency range composed of the frequency of the first audio signal and the frequency of the second audio signal.

[0039] Optionally, in the embodiments of the present application, the target frequency range can further include a wind-noise-free frequency band in addition to the first frequency band and the second frequency band.

[0040] Optionally, in the embodiments of the present application, the first frequency band can be an intersection frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal.

[0041] Optionally, in the embodiments of the present application, the second frequency band can be an extended difference set frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal.

[0042] In the embodiments of the present application, the first frequency band can be the intersection frequency band, and the second frequency band can be the extended difference set frequency band, so as to improve the flexibility of the electronic device in dividing the target frequency range.

[0043] Optionally, in the embodiments of the present application, the noise frequency band of the first audio signal and the noise frequency band of the second audio signal can be obtained based on a target coherence coefficient between the first audio signal and the second audio signal.

[0044] Optionally, in the embodiments of the present application, the target coherence coefficient can include at least one of the following:

[0045] (a) a magnitude squared coherence coefficient (i.e., Magnitude-Squared Coherence);

[0046] (b) a relative deviation coefficient;

[0047] (c) a relative intensity sensitivity coefficient;

[0048] (d) a magnitude squared coherence coefficient of a magnitude spectrum;

[0049] (e) a magnitude squared coherence coefficient of a phase spectrum.

[0050] In the embodiments of the present application, the target coherence coefficient is used to indicate the coherence feature between the first audio signal and the second audio signal, and is usually generated based on dissimilarity or similarity measurement with a value between 0 and 1. The process of specifically determining the target coherence coefficient is as follows:

[0051] First, in the target frequency range, the frequency point coherence (i.e., Coherence) can be expressed as formula (1) as follows:

[0052]

[0053] wherein, P XP (ω) is the power spectral density of the first audio signal X(ω), P Y P (ω) is the power spectral density of the second audio signal Y(ω), P XY COH(ω) is the cross power spectral density between the first audio signal and the second audio signal. COH(ω) is a complex number, and |COH(ω)|≤1 holds true if and only if the first audio signal and the second audio signal are perfectly coherent. To avoid square root operations, the amplitude squared coherence coefficient (a) is often used, which can be expressed as equation (2) as follows:

[0054]

[0055] Obviously, the normalization effect of MSC(ω) is not sensitive to the relative strength of X(ω) and Y(ω), while the relative strength of the first audio signal and the second audio signal is important for the judgment of noise. Therefore, the normalized power level difference, i.e., the relative deviation coefficient (b), is defined again, which can be expressed as equation (3) as follows:

[0056]

[0057] Obviously, 0≤NPLD(ω)≤1 is a measure of the dissimilarity between the desired audio signals. In addition, COH can also be transformed into a form that is sensitive to the relative strength of the first audio signal and the second audio signal, i.e., the relative strength sensitive coefficient (c), as equation (4) as follows:

[0058]

[0059] The above equation (2) can also be transformed into a version that only considers the amplitude spectrum or the phase spectrum, respectively. The version that only considers the amplitude spectrum, i.e., the amplitude squared coherence coefficient of the amplitude spectrum (d), can be expressed as equation (5) as follows:

[0060]

[0061] Obviously, the following inequality relationship (6) can be obtained, which measures the similarity between the desired audio signals:

[0062] 0≤COH_AS(ω) 2 ≤MSC(ω)≤MSC_AMP(ω)≤1; (6)

[0063] In summary, any other similarity or dissimilarity criterion with a value between 0 and 1 is available. In this way, the target coherence coefficient between the first audio signal and the second audio signal can be determined.

[0064] In the embodiments of the present application, since the target coherence coefficient can include at least one of (a) to (e) above, the electronic device can obtain different noise frequency bands of the audio signals based on different target coherence coefficients between the first audio signal and the second audio signal, thereby further improving the flexibility of dividing the target frequency range according to the noise frequency bands.

[0065] Optionally, in the embodiments of the present application, after determining the target coherence coefficient, the electronic device can obtain the presence probability of the expected audio signal based on a linear or nonlinear combination of the target coherence coefficient which can be expressed as formula (7) as follows:

[0066]

[0067] It can be understood that, since the noise energy is concentrated in the low frequency band and rapidly attenuates when tending to the high frequency band, the electronic device can determine the noise frequency band of the first audio signal and the noise frequency band of the second audio signal according to the noise energy of the first audio signal and the noise energy of the second audio signal. The union frequency band between the noise frequency band of the first audio signal and the noise frequency band of the second audio signal can be searched and estimated from the low frequency to the high frequency.

[0068] Optionally, in the embodiments of the present application, after estimating the union frequency band, the electronic device can first correct P X (ω) and P Y (ω) based on the harmonic position of the fundamental frequency, so as to prevent overestimation of the bandwidth; then the electronic device can estimate the noise frequency band of the first audio signal and the noise frequency band of the second audio signal from the union frequency band based on the corrected P X (ω) and P Y (ω).

[0069] In the embodiments of the present application, since the noise frequency band of the first audio signal and the noise frequency band of the second audio signal can be obtained based on the target coherence coefficient between the first audio signal and the second audio signal, the accuracy of obtaining the noise frequency band of the audio signal can be improved.

[0070] The specific method of dividing the target frequency range into the first frequency band, the second frequency band and the wind-noise-free frequency band by the electronic device will be described in detail below.

[0071] Optionally, in the embodiments of the present application, after estimating the noise frequency band of the first audio signal (hereinafter referred to as noise frequency band A) and the noise frequency band of the second audio signal (hereinafter referred to as noise frequency band B) based on the target coherence coefficient, the electronic device can divide the target frequency range into:

[0072] a, the intersection of the noise frequency band A and the noise frequency band B (i.e. the first frequency band);

[0073] b, the noise frequency band A and the noise frequency band B, the extended difference set between the above intersection and the corresponding extended wind noise frequency band (i.e., the second frequency band);

[0074] c, the wind noise free frequency band.

[0075] The audio signal processing method provided by the embodiments of the present application will be exemplarily described below with reference to the accompanying drawings.

[0076] Exemplarily, as shown in the figure, Figure 2 the electronic device can first estimate the noise frequency band 25 (i.e., the above-mentioned extended wind noise frequency band) based on the noise frequency band 21 (i.e., the noise frequency band of the first audio signal) and the noise frequency band 22 (i.e., the noise frequency band of the second audio signal), and then can divide the target frequency range into the frequency band 23 (i.e., the first frequency band), the frequency band 24 (i.e., the second frequency band) and the frequency band 26 (i.e., the wind noise free frequency band); it can be seen that the frequency band 23 is the intersection of the noise frequency band 21 and the noise frequency band 22, and the frequency band 24 is the extended difference set between the noise frequency band 21 and the noise frequency band 22, and the corresponding noise frequency band 25 and the frequency band 23.

[0077] Optionally, in the embodiments of the present application, the electronic device can generate the initial gain corresponding to the first audio signal and the initial gain corresponding to the second audio signal based on the (a) amplitude square coherence coefficient and the (b) relative deviation coefficient when estimating the noise frequency band of the first audio signal and the noise frequency band of the second audio signal, to be used for noise reduction of the audio signal.

[0078] Step 102, the electronic device performs first fusion processing on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal in the first frequency band.

[0079] In the embodiments of the present application, the first audio signal and the second audio signal correspond to one transmission channel respectively.

[0080] Optionally, in the embodiments of the present application, the above-mentioned transmission channel information can include amplitude spectrum, wind noise gain and noise stabilization gain of the audio signal in the corresponding transmission channel and the like.

[0081] Optionally, in the embodiments of the present application, the above-mentioned step 102 can be implemented by the following step 102a or step 102b.

[0082] Step 102a, in the case that the noise intensity of the first sub-audio signal is less than the noise intensity of the second sub-audio signal, the electronic device superimposes the transmission channel information corresponding to the first sub-audio signal to the transmission channel information corresponding to the second sub-audio signal with a first weight.

[0083] In step 102b, the electronic device superimposes, in a case where the noise intensity of the first sub-audio signal is greater than the noise intensity of the second sub-audio signal, the transmission channel information corresponding to the second sub-audio signal onto the transmission channel information corresponding to the first sub-audio signal with a second weight.

[0084] In the embodiments of the present application, the first sub-audio signal is an audio signal of the first audio signal in the first frequency band; and the second sub-audio signal is an audio signal of the second audio signal in the first frequency band.

[0085] It can be understood that the transmission channel information corresponding to the first sub-audio signal is the transmission channel information of the transmission channel corresponding to the first audio signal in the first frequency band; and the transmission channel information corresponding to the second sub-audio signal is the transmission channel information of the transmission channel corresponding to the second audio signal in the first frequency band.

[0086] Optionally, in the embodiments of the present application, the first weight and the second weight can be the same or different.

[0087] In the embodiments of the present application, the electronic device still retains one transmission channel information after superimposing the one transmission channel information onto another transmission channel information.

[0088] In the embodiments of the present application, since the electronic device can superimpose the transmission channel information in the first frequency band in different ways according to the size relationship between the noise intensity of the first sub-audio signal and the noise intensity of the second sub-audio signal, the flexibility of the electronic device in fusing the transmission channel information can be improved.

[0089] In step 103, the electronic device performs a second fusion processing on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal in the second frequency band.

[0090] Optionally, in the embodiments of the present application, the above step 103 can be implemented by the following step 103a or step 103b.

[0091] In step 103a, the electronic device superimposes, in a case where the third sub-audio signal is a noiseless audio signal, the transmission channel information corresponding to the third sub-audio signal onto the transmission channel information corresponding to the fourth sub-audio signal with a third weight.

[0092] In step 103b, the electronic device superimposes, in a case where the fourth sub-audio signal is a noiseless audio signal, the transmission channel information corresponding to the fourth sub-audio signal onto the transmission channel information corresponding to the third sub-audio signal with a fourth weight.

[0093] In the embodiments of the present application, the third sub-audio signal is an audio signal of the first audio signal in the second frequency band; and the fourth sub-audio signal is an audio signal of the second audio signal in the second frequency band.

[0094] It can be understood that the transmission channel information corresponding to the third sub-audio signal is the transmission channel corresponding to the first audio signal, and the transmission channel information in the second frequency band; the transmission channel information corresponding to the fourth sub-audio signal is the transmission channel corresponding to the second audio signal, and the transmission channel information in the second frequency band.

[0095] Optionally, in the embodiment of the present application, the third weight and the fourth weight can be the same or can be different.

[0096] In the embodiment of the present application, since the electronic device can superimpose the transmission channel information in the second frequency band in different ways in the case that the third sub-audio signal is a noise-free audio signal or the fourth sub-audio signal is a noise-free audio signal, the flexibility of the electronic device in fusing the transmission channel information can be further improved.

[0097] Optionally, in the embodiment of the present application, the processing strength of the first fusion processing can be less than the processing strength of the second fusion processing; that is, the first weight and the second weight can both be less than the target weight, and the target weight is the minimum weight of the third weight and the fourth weight.

[0098] For example, the first weight and the second weight can both be 0.5, at this time the electronic device can complete the superposition of the transmission channel information in the first frequency band with the weight 0.5; the third weight and the fourth weight can both be 1, at this time the electronic device can complete the superposition of the transmission channel information in the second frequency band with the weight 1, that is, in the second frequency band, one transmission channel information is directly replaced by another transmission channel information.

[0099] As can be seen, the first fusion processing can realize the superposition of the transmission channel information, and the second fusion processing can realize the replacement of the transmission channel information.

[0100] In the embodiment of the present application, since the processing strength of the first fusion processing can be less than the processing strength of the second fusion processing, the fusion processing of the transmission channel information can be performed in different frequency bands with different processing strengths, so that the flexibility of the electronic device in fusing the transmission channel information can be further improved.

[0101] Step 104, the electronic device performs noise reduction on the target audio signal after the corresponding transmission channel information is fused.

[0102] In the embodiment of the present application, the target audio signal includes at least one of the first audio signal and the second audio signal.

[0103] It can be understood that the electronic device can perform noise reduction on the audio signal in the first audio signal and the second audio signal, which is fused with the corresponding transmission channel information.

[0104] Optionally, in the embodiment of the present application, the transmission channel information after the fusion processing can include the first gain and the second gain.

[0105] In the embodiment of the present application, the first gain is used for noise reduction of the first audio signal, and the second gain is used for noise reduction of the second audio signal.

[0106] Optionally, in the embodiment of the present application, at least one of the first gain and the second gain is the gain after the fusion processing of the initial gain in the transmission channel information.

[0107] Optionally, in the embodiment of the present application, if the target audio signal includes the first audio signal and the second audio signal, the electronic device can apply the first gain to the amplitude spectrum of the first audio signal and apply the second gain to the amplitude spectrum of the second audio signal to perform noise reduction on the first audio signal and the second audio signal.

[0108] Optionally, in the embodiment of the present application, the step 104 can be implemented by the following step 104a.

[0109] In step 104a, the electronic device performs noise reduction on the target audio signal in a target noise reduction mode when the signal-to-wind ratio of the target audio signal is less than or equal to a preset threshold.

[0110] In the embodiment of the present application, the target noise reduction mode is a noise reduction mode in which the target audio signal is subjected to first noise reduction processing in a third frequency band and is subjected to second noise reduction processing in a fourth frequency band.

[0111] In the embodiment of the present application, the frequency of the third frequency band is less than or equal to a first frequency threshold, and the frequency of the fourth frequency band is greater than or equal to a second frequency threshold.

[0112] Optionally, in the embodiment of the present application, the first frequency threshold and the second frequency threshold can be default of the electronic device, or can be set by the user according to actual use requirements.

[0113] In the embodiment of the present application, the processing strength of the first noise reduction processing is less than the processing strength of the second noise reduction processing.

[0114] Optionally, in the embodiment of the present application, the processing strength of the first noise reduction processing can be close to 0.

[0115] Optionally, in the embodiment of the present application, the electronic device can determine the signal-to-wind ratio of the audio signal based on the noise frequency band in the audio signal.

[0116] Optionally, in the embodiment of the present application, the preset threshold can be default of the electronic device, or can be set by the user according to actual use requirements.

[0117] It can be understood that the wind speed ratio of the audio signal is less than or equal to the preset threshold, that is, there is a noise signal with an ultra-large frequency band in the audio signal; if the audio signal is denoised, the denoising needs to be conservative, that is, the suppression of the low-frequency band noise signal is reduced, and only part of the high-frequency band noise signal is suppressed, that is, the target denoising manner is used for denoising, so as to realize the natural denoising effect of the hearing.

[0118] In the embodiments of the present application, since the electronic device can perform denoising on the target audio signal in the target denoising manner (that is, performing the first denoising processing in the low-frequency band and performing the second denoising processing with greater processing intensity in the high-frequency band) when the wind speed ratio of the target audio signal is less than or equal to the preset threshold, the hearing of the target audio signal after denoising can be ensured to be more natural.

[0119] In the audio signal processing method provided in the embodiments of the present application, since the electronic device can first perform fusion processing on the transmission channel information based on the divided frequency bands and the transmission channel information corresponding to each audio signal before performing denoising processing on the audio signals collected by different microphones, and then performs denoising on the audio signals after the fusion processing on the corresponding transmission channel information, the electronic device can process the audio signals without needing to be based on the characteristics of a single audio signal or all frequency points of multiple audio signals, but can combine the transmission channel information corresponding to different audio signals in different divided frequency bands, thereby improving the robustness of the electronic device in processing the audio signals.

[0120] Optionally, after the step 104, the audio signal processing method provided in the embodiments of the present application can further include the following step 105.

[0121] Step 105, the electronic device inserts a noise compensation audio signal in at least one target frequency band.

[0122] In the embodiments of the present application, each target frequency band is one frequency band of the audio signal in the target frequency range after denoising.

[0123] In the embodiments of the present application, the noise compensation audio signal is used to compensate the audio signal in the corresponding target frequency band.

[0124] Optionally, in the embodiments of the present application, each target frequency band can correspond to one noise compensation audio signal.

[0125] Optionally, in the embodiments of the present application, the noise compensation audio signal can be an audio signal with good continuity with the audio signal in the first target frequency band; and the first target frequency band is a frequency band adjacent to the corresponding target frequency band and not including the audio signal after denoising.

[0126] In the embodiments of the present application, since the electronic device can insert the noise compensation audio signal in at least one target frequency band, the continuity of the target audio signal after noise reduction can be improved, and thus the subjective listening of the target audio signal can be improved.

[0127] An example of the application of the audio signal processing method provided by the embodiments of the present application will be exemplarily described below with reference to the accompanying drawings.

[0128] Exemplarily, the working frequency band of the audio signal is usually within 24 kHz, Figure 3 The input spectrogram of an example audio signal is shown as Figure 3 As shown, the audio signal collected by the main microphone (hereinafter referred to as audio signal A) and the audio signal collected by the auxiliary microphone (hereinafter referred to as audio signal B) have a significantly different wind noise frequency band, and the interval 31 in the smooth power spectrum corresponding to the audio signal B is a severely contaminated interval. In order to reduce the noise of the collected audio signal, the electronic device can determine the target coherence coefficient between the two audio signals based on the audio signal A and the audio signal B.

[0129] Figure 4 The target coherence coefficient determined by the electronic device and its comprehensive effect are shown as Figure 4 As shown, the target coherence coefficient determined by the electronic device includes COH_AS 2 , MSC, MSC_AMP, NPLD (i.e. (a) to (d) in the above embodiments), and the target coherence coefficient determined by the electronic device includes COH_AS 2 , MSC, MSC_AMP, NPLD (i.e. (a) to (d) in the above embodiments), and the target coherence coefficient determined by the electronic device includes COH_AS The The corresponding smooth power spectrum is shown as the smooth power spectrum 45 in Figure 4 , and then the electronic device can generate a higher robustness expected audio presence probability according to the probability

[0130] Figure 5 The noise frequency band searched and estimated by the electronic device and the corresponding wind noise gain are shown as Figure 5As shown, the noise frequency band in audio signal A is the frequency band corresponding to curve 52, the noise frequency band in audio signal B is the frequency band corresponding to curve 53, and the frequency band corresponding to curve 51 is the union frequency band of the estimated noise frequency band in audio signal A and the noise frequency band in audio signal B. Obviously, this union frequency band is overestimated. It can be seen that each noise frequency band tightly defines the frequency band where noise exists. Among them, smooth power spectrum 54 is the smooth power spectrum of wind noise gain corresponding to the noise frequency band in audio signal A, and smooth power spectrum 55 is the smooth power spectrum of wind noise gain corresponding to the noise frequency band in audio signal B.

[0131] Figure 6 The spectrograms of audio signal A and audio signal B before and after noise reduction by an electronic device are shown, such as... Figure 6 As shown, the wind noise frequency band 61 in audio signal A is reduced to frequency band 63 after noise reduction, and the wind noise frequency band 62 in audio signal B is reduced to frequency band 64 after noise reduction. It can be seen that the strong noise in the stereo input is effectively suppressed in the stereo output. Furthermore, thanks to the fusion of transmission channel information, the audio signal at low signal-to-wind ratio is effectively protected, resulting in a continuous and natural listening experience and sound quality. This achieves robust noise reduction of audio signals, improving the noise reduction effect of electronic devices.

[0132] The information flow of the audio signal processing method provided in the embodiments of this application will be described exemplarily below with reference to the accompanying drawings.

[0133] For example, Figure 7 The diagram shows the information flow of the audio signal processing method provided in this application, applied to robust wind noise detection and suppression in dual-microphone stereo systems. Figure 7 As shown, the electronic device collects audio signals X through different microphones. i (ω) (i.e., the first audio signal) and audio signal Y i After (ω) (i.e., the second audio signal), the probability of the existence of the desired audio signal can be obtained based on the target coherence coefficient between the two audio signals. (ω), and can be based on (ω) Searches from low to high frequencies and estimates a dual-microphone union wind noise bandwidth W. union ;

[0134] The electronic device can then correct the single-microphone power spectrum based on the harmonic position of the pitch to prevent bandwidth overestimation, and adjust the power spectrum of the single microphone based on the corrected single-microphone power spectrum in W. union Internal search and estimation of single-microphone wind and noise bandwidth W X (i.e., the noise band of the first audio signal) and W Y (i.e., the noise band of the second audio signal);

[0135] Thus, electronic devices can be based on WX and W Y The frequency domain (i.e., the target frequency range) is divided into: the wind noise bandwidth intersection B meet (i.e., the first frequency band), the extended wind noise bandwidth difference set B diff (i.e., the second frequency band), and the wind noise-free frequency band B clean . For B meet , both microphones contain wind noise, but usually the wind noise intensity of one transmission channel (or microphone) is less than that of the other transmission channel, based on the wind noise intensity of the single microphone, the transmission channel information in the sub-band can be fused before wind noise suppression (i.e., the first fusion processing), that is, the weak wind noise transmission channel information (including amplitude spectrum, wind noise gain, and noise reduction gain) is superimposed on the strong wind noise transmission channel information in an arithmetic or geometric mean manner (i.e., the first weight or the second weight); for B diff , usually one transmission channel receives wind noise pollution, and the other transmission channel is not polluted by wind noise, and the transmission channel information in the sub-band is also fused before wind noise suppression (i.e., the second fusion processing), that is, the wind noise-free transmission channel information is superimposed on the wind noise transmission channel information in a larger proportion (i.e., the third weight or the fourth weight); for B clean , no wind noise suppression is performed. In addition, the electronic device can also distinguish extreme wind noise conditions based on the wind noise bandwidth of the single microphone. In the case of an occasional super-large bandwidth or gale, the original audio signal wind ratio is extremely low, and the reliability of extreme wind noise suppression is poor. At this time, wind noise suppression tends to be conservative, low-frequency wind noise suppression is reduced, and only partial high-frequency wind noise suppression is performed to achieve a more natural noise reduction effect in terms of listening.

[0136] After the electronic device performs transmission channel information fusion, the electronic device can use wind noise gain (i.e., the first gain and the second gain) to act on the transmission channel amplitude spectrum to complete wind noise suppression; however, the continuity of the audio amplitude spectrum after wind noise suppression is poor, and there may be discontinuity or fluctuation in terms of listening depending on the recording audio composition. Therefore, the electronic device can insert comfort noise (i.e., noise compensation audio signal) into the frequency band after wind noise suppression (i.e., at least one target frequency band) to compensate for a certain amount of comfort noise with better continuity with the adjacent wind noise-free audio background, thereby significantly improving the subjective listening experience. Thus, wind noise suppression can be completed to obtain the noise-reduced audio signal X o (ω) and Y o (ω).

[0137] The audio signal processing method provided in the embodiments of the present application can be executed by an audio signal processing device. In the embodiments of the present application, the audio signal processing device is taken as an example to illustrate the audio signal processing device provided in the embodiments of the present application.

[0138] In combination with Figure 8The embodiment of the present application provides an audio signal processing device 80, which can comprise a division module 81, a fusion module 82 and a noise reduction module 83. The division module 81 can be used for dividing a target frequency range into a first frequency band and a second frequency band according to a noise frequency band of a first audio signal and a noise frequency band of a second audio signal, the first audio signal being an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal being an audio signal obtained by collecting the target sound source by a second microphone. The fusion module 82 can be used for performing first fusion processing on transmission channel information corresponding to the first audio signal and transmission channel information corresponding to the second audio signal in the first frequency band. The fusion module 82 can also be used for performing second fusion processing on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal in the second frequency band. The noise reduction module 83 can be used for reducing noise of a target audio signal after fusion processing of corresponding transmission channel information, and the target audio signal comprises at least one of the first audio signal and the second audio signal.

[0139] In a possible implementation, the first frequency band can be an intersection frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal; and the second frequency band can be an extended difference set frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal.

[0140] In a possible implementation, the fusion module 82 can be specifically configured to: in a case where a noise intensity of a first sub-audio signal is less than a noise intensity of a second sub-audio signal, superimpose transmission channel information corresponding to the first sub-audio signal onto transmission channel information corresponding to the second sub-audio signal with a first weight; or in a case where the noise intensity of the first sub-audio signal is greater than the noise intensity of the second sub-audio signal, superimpose transmission channel information corresponding to the second sub-audio signal onto transmission channel information corresponding to the first sub-audio signal with a second weight. The first sub-audio signal is an audio signal of the first audio signal in the first frequency band; and the second sub-audio signal is an audio signal of the second audio signal in the first frequency band.

[0141] In a possible implementation, the fusion module 82 can be specifically configured to: in a case where a third sub-audio signal is a noise-free audio signal, superimpose transmission channel information corresponding to the third sub-audio signal onto transmission channel information corresponding to a fourth sub-audio signal with a third weight; or in a case where the fourth sub-audio signal is a noise-free audio signal, superimpose transmission channel information corresponding to the fourth sub-audio signal onto transmission channel information corresponding to the third sub-audio signal with a fourth weight. The third sub-audio signal is an audio signal of the first audio signal in the second frequency band; and the fourth sub-audio signal is an audio signal of the second audio signal in the second frequency band.

[0142] In a possible implementation, the processing intensity of the first fusion processing is less than the processing intensity of the second fusion processing.

[0143] In a possible implementation, the noise reduction module 83 can be specifically configured to perform noise reduction on the target audio signal in a target noise reduction manner in a case where the signal-to-wind ratio of the target audio signal is less than or equal to a preset threshold. The target noise reduction manner is a noise reduction manner in which the target audio signal is subjected to first noise reduction processing in a third frequency band and subjected to second noise reduction processing in a fourth frequency band, the frequency of the third frequency band is less than or equal to a first frequency threshold, the frequency of the fourth frequency band is greater than or equal to a second frequency threshold, and the processing intensity of the first noise reduction processing is less than the processing intensity of the second noise reduction processing.

[0144] In a possible implementation, the audio signal processing apparatus 80 can further include an insertion module. The insertion module can be configured to insert a noise compensation audio signal in at least one target frequency band after the noise reduction module 83 performs noise reduction on the target audio signal subjected to fusion processing of the corresponding transmission channel information. Each target frequency band is a frequency band in a target frequency range in which the audio signal subjected to noise reduction is located; and the noise compensation audio signal is used to compensate the audio signal in the corresponding target frequency band.

[0145] In a possible implementation, the noise frequency band of the first audio signal and the noise frequency band of the second audio signal are obtained based on a target coherence coefficient between the first audio signal and the second audio signal.

[0146] In a possible implementation, the target coherence coefficient can include at least one of the following: an amplitude square coherence coefficient; a relative deviation coefficient; a relative intensity sensitive coefficient; an amplitude square coherence coefficient of an amplitude spectrum; and an amplitude square coherence coefficient of a phase spectrum.

[0147] In the audio signal processing apparatus provided in the embodiments of the present application, the audio signal processing apparatus can perform fusion processing of transmission channel information based on the divided frequency bands and the transmission channel information corresponding to each audio signal before performing noise reduction processing on the audio signals collected by different microphones, and then perform noise reduction on the audio signals subjected to fusion processing of the corresponding transmission channel information. Therefore, the audio signal processing apparatus can process the audio signals without being based on the characteristics of a single audio signal or all frequency points of multiple audio signals, but can combine the transmission channel information corresponding to different audio signals in different frequency bands, thereby improving the robustness of processing the audio signals.

[0148] The audio signal processing apparatus in the embodiments of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and the like, or a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a cash register, or a self-service machine, and the like, and the embodiments of the present application are not limited thereto.

[0149] The audio signal processing apparatus in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an ios operating system, or other possible operating systems, and the embodiments of the present application are not limited thereto.

[0150] The audio signal processing apparatus provided in the embodiments of the present application can implement the method embodiments Figures 1 to 7 , and each process of the method embodiments is not repeated here to avoid repetition.

[0151] As shown in Figure 9 , the embodiments of the present application further provide an electronic device 900, which includes a processor 901 and a memory 902, and the memory 902 has a program or instructions stored thereon, which can be run on the processor 901. When the program or instructions are executed by the processor 901, each step of the above-mentioned audio signal processing method embodiments is implemented, and the same technical effects are achieved. To avoid repetition, each step is not repeated here.

[0152] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device.

[0153] Figure 10 To implement the hardware structure of an electronic device in the embodiments of the present application.

[0154] The electronic device 1000 includes, but is not limited to, a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010, etc.

[0155] Those skilled in the art can understand that the electronic device 1000 can also include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 1010 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system. Figure 10 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the figure, or combine certain components, or different component arrangements, which are not described here.

[0156] The processor 1010 can be configured to divide a target frequency range into a first frequency band and a second frequency band according to a noise frequency band of a first audio signal and a noise frequency band of a second audio signal, the first audio signal being an audio signal obtained by collecting a target sound source by a first microphone, and the second audio signal being an audio signal obtained by collecting the target sound source by a second microphone; and can be configured to perform first fusion processing on transmission channel information corresponding to the first audio signal and transmission channel information corresponding to the second audio signal in the first frequency band; and can be configured to perform second fusion processing on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal in the second frequency band; and can be configured to perform noise reduction on a target audio signal after fusion processing of the corresponding transmission channel information, the target audio signal including at least one of the first audio signal and the second audio signal.

[0157] In a possible implementation, the first frequency band can be an intersection frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal; and the second frequency band can be an extended difference set frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal.

[0158] In a possible implementation, the processor 1010 can be specifically configured to: in a case where the noise intensity of the first sub-audio signal is less than the noise intensity of the second sub-audio signal, superimpose the transmission channel information corresponding to the first sub-audio signal on the transmission channel information corresponding to the second sub-audio signal with a first weight; or in a case where the noise intensity of the first sub-audio signal is greater than the noise intensity of the second sub-audio signal, superimpose the transmission channel information corresponding to the second sub-audio signal on the transmission channel information corresponding to the first sub-audio signal with a second weight. The first sub-audio signal is an audio signal of the first audio signal in the first frequency band; and the second sub-audio signal is an audio signal of the second audio signal in the first frequency band.

[0159] In a possible implementation, the processor 1010 can be specifically configured to: in a case where the third sub-audio signal is a noise-free audio signal, superimpose the transmission channel information corresponding to the third sub-audio signal on the transmission channel information corresponding to the fourth sub-audio signal with a third weight; or in a case where the fourth sub-audio signal is a noise-free audio signal, superimpose the transmission channel information corresponding to the fourth sub-audio signal on the transmission channel information corresponding to the third sub-audio signal with a fourth weight. The third sub-audio signal is an audio signal of the first audio signal in the second frequency band; and the fourth sub-audio signal is an audio signal of the second audio signal in the second frequency band.

[0160] In a possible implementation, the processing strength of the first fusion processing is less than the processing strength of the second fusion processing.

[0161] In a possible implementation, the processor 1010 can be specifically configured to: in a case where the signal-to-wind ratio of the target audio signal is less than or equal to a preset threshold, perform noise reduction on the target audio signal in a target noise reduction mode. The target noise reduction mode is a noise reduction mode in which the target audio signal is subjected to first noise reduction processing in a third frequency band, and is subjected to second noise reduction processing in a fourth frequency band; the frequency of the third frequency band is less than or equal to a first frequency threshold, the frequency of the fourth frequency band is greater than or equal to a second frequency threshold, and the processing strength of the first noise reduction processing is less than the processing strength of the second noise reduction processing.

[0162] In a possible implementation, the processor 1010 can be further configured to: after performing noise reduction on the target audio signal after the fusion processing of the corresponding transmission channel information, insert a noise compensation audio signal in at least one target frequency band. Each target frequency band is a frequency band of an audio signal subjected to noise reduction in a target frequency range; and the noise compensation audio signal is used to compensate for the audio signal in the corresponding target frequency band.

[0163] In a possible implementation, the noise frequency band of the first audio signal and the noise frequency band of the second audio signal are obtained based on a target coherence coefficient between the first audio signal and the second audio signal.

[0164] In a possible implementation, the target coherence coefficient can include at least one of the following: an amplitude square coherence coefficient; a relative deviation coefficient; a relative intensity sensitivity coefficient; an amplitude square coherence coefficient of an amplitude spectrum; and an amplitude square coherence coefficient of a phase spectrum.

[0165] In the electronic device provided in the embodiments of the present application, because the electronic device can perform fusion processing on the transmission channel information based on the divided frequency bands and the transmission channel information corresponding to each audio signal before performing noise reduction processing on the audio signals collected by different microphones, and then perform noise reduction on the audio signals after the fusion processing on the corresponding transmission channel information, the electronic device can process the audio signals without needing to be based on the features of a single audio signal or all frequency points of multiple audio signals, but can combine the transmission channel information corresponding to different audio signals in different frequency bands, thereby improving the robustness of the electronic device in processing the audio signals.

[0166] The beneficial effects of the various implementations in the embodiments can refer to the beneficial effects of the corresponding implementations in the method embodiments described above. To avoid repetition, details are not described here.

[0167] It should be understood that, in the embodiments of the present application, the input unit 1004 can include a graphics processing unit (GPU) 10041 and a microphone 10042. The graphics processing unit 10041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1006 can include a display panel 10061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 can include two parts of a touch detection device and a touch controller. The other input devices 10072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which are not described here.

[0168] The memory 1009 can be used to store software programs and various data. The memory 1009 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 1009 can include a volatile memory or a non-volatile memory, or the memory 1009 can include both volatile and non-volatile memories. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0169] The processor 1010 can include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 1010.

[0170] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned audio signal processing method embodiments, and the same technical effects can be achieved. To avoid repetition, details are not described here.

[0171] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0172] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions to realize the processes of the above audio signal processing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0173] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system on chip (SoC), a system chip, a chip system or a system on chip (SoC), etc.

[0174] The embodiment of the present application provides a computer program product stored in a storage medium, which is executed by at least one processor to realize the processes of the above audio signal processing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0175] It should be noted that in this document, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in the opposite order, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0176] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and a necessary general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part that contributes to the prior art, which is stored in a storage medium (such as a ROM / RAM, a magnetic disc, an optical disc), and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0177] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. An audio signal processing method, characterized in that, The method includes: Based on the noise band of the first audio signal and the noise band of the second audio signal, the target frequency range is divided into a first frequency band and a second frequency band. The first audio signal is the audio signal obtained by the first microphone from the target sound source, and the second audio signal is the audio signal obtained by the second microphone from the target sound source. The first frequency band is the intersection frequency band of the noise band of the first audio signal and the noise band of the second audio signal, and the second frequency band is the extensional difference frequency band of the noise band of the first audio signal and the noise band of the second audio signal. Within the first frequency band, the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal are subjected to a first fusion process; Within the second frequency band, a second fusion process is performed on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal. The target audio signal, after being fused with the corresponding transmission channel information, is denoised. The target audio signal includes at least one of the first audio signal and the second audio signal.

2. The method according to claim 1, characterized in that, The first fusion process, performed within the first frequency band, on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal, includes: If the noise intensity of the first sub-audio signal is less than that of the second sub-audio signal, the transmission channel information corresponding to the first sub-audio signal is superimposed onto the transmission channel information corresponding to the second sub-audio signal with a first weight; or, If the noise intensity of the first sub-audio signal is greater than the noise intensity of the second sub-audio signal, the transmission channel information corresponding to the second sub-audio signal is superimposed on the transmission channel information corresponding to the first sub-audio signal with a second weight. Wherein, the first sub-audio signal is: the audio signal of the first audio signal within the first frequency band, and the second sub-audio signal is: the audio signal of the second audio signal within the first frequency band.

3. The method according to claim 1, characterized in that, The second fusion process, performed within the second frequency band, on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal, includes: If the third sub-audio signal is a noise-free audio signal, the transmission channel information corresponding to the third sub-audio signal is superimposed on the transmission channel information corresponding to the fourth sub-audio signal with a third weight; or... When the fourth sub-audio signal is a noise-free audio signal, the transmission channel information corresponding to the fourth sub-audio signal is superimposed on the transmission channel information corresponding to the third sub-audio signal with a fourth weight. The third sub-audio signal is the audio signal of the first audio signal within the second frequency band; the fourth sub-audio signal is the audio signal of the second audio signal within the second frequency band.

4. The method according to claim 1, characterized in that, The processing intensity of the first fusion process is less than that of the second fusion process.

5. The method according to claim 1, characterized in that, The noise reduction of the target audio signal after fusing the corresponding transmission channel information includes: When the signal-to-frequency ratio of the target audio signal is less than or equal to a preset threshold, the target audio signal is denoised using the target denoising method. The signal-to-wind ratio is determined based on the noise frequency band in the target audio signal. The signal-to-wind ratio is used to characterize the probability that there is a noise signal with a very large frequency band in the target audio signal. The target noise reduction method is a noise reduction method that performs a first noise reduction process on the target audio signal in a third frequency band and a second noise reduction process on the target audio signal in a fourth frequency band. The frequency of the third frequency band is less than or equal to a first frequency threshold, the frequency of the fourth frequency band is greater than or equal to a second frequency threshold, and the processing intensity of the first noise reduction process is less than the processing intensity of the second noise reduction process.

6. The method according to claim 1, characterized in that, After denoising the target audio signal obtained by fusing the corresponding transmission channel information, the method further includes: Insert a noise-compensated audio signal within at least one target frequency band; Each target frequency band is a frequency band within the target frequency range where the audio signal to be denoised is located; the noise-compensated audio signal is used to compensate for the audio signal within the corresponding target frequency band.

7. The method according to claim 1, characterized in that, The noise bands of the first audio signal and the second audio signal are obtained based on the target coherence coefficient between the first audio signal and the second audio signal.

8. The method according to claim 7, characterized in that, The target coherence coefficient includes at least one of the following: The amplitude-squared coherence coefficient of the power spectrum; Relative deviation coefficient; Relative intensity sensitivity coefficient; The amplitude squared coherence coefficient of the amplitude spectrum; The amplitude squared coherence coefficient of the phase spectrum.

9. An audio signal processing device, characterized in that, The device includes a segmentation module, a fusion module, and a noise reduction module; The segmentation module is used to divide the target frequency range into a first frequency band and a second frequency band based on the noise frequency band of the first audio signal and the noise frequency band of the second audio signal. The first audio signal is the audio signal obtained by the first microphone from the target sound source, and the second audio signal is the audio signal obtained by the second microphone from the target sound source. The first frequency band is the intersection frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal, and the second frequency band is the extensional difference frequency band of the noise frequency band of the first audio signal and the noise frequency band of the second audio signal. The fusion module is used to perform a first fusion process on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal within the first frequency band. The fusion module is further configured to perform a second fusion process on the transmission channel information corresponding to the first audio signal and the transmission channel information corresponding to the second audio signal within the second frequency band; The noise reduction module is used to reduce the noise of the target audio signal after the corresponding transmission channel information has been fused. The target audio signal includes at least one of the first audio signal and the second audio signal.

10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the audio signal processing method as described in any one of claims 1-8.

11. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio signal processing method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Adaptive mixing of sub-band signals

    CN107409255A

  • Sound collecting apparatus and stereophonic computing method

    JP2003299183A