Audio signal processing methods, apparatus, electronic devices and readable storage media

By using an audio signal processing method across multiple pickup channels, blocked pickup channels can be identified and fused, thus solving the problem of poor audio quality caused by microphone hole blockage and improving audio playback quality.

CN115696113BActive Publication Date: 2026-03-13VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

When the microphone hole is blocked, the quality of the acquired voice signal decreases significantly, resulting in poor audio quality.

Method used

By acquiring audio signals from multiple pickup channels, determining the empirical cumulative distribution of their frequency energy, identifying blocked pickup channels, and fusing the audio signals from blocked channels, detection accuracy is improved and audio playback quality is enhanced.

Benefits of technology

It improves the accuracy of sound pickup channel blockage detection, reduces false alarms, and improves audio playback quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115696113B_ABST
    Figure CN115696113B_ABST
Patent Text Reader

Abstract

This application discloses an audio signal processing method, apparatus, electronic device, and readable storage medium, belonging to the field of audio processing technology. The method includes: acquiring a first audio signal collected through a first pickup channel and a second audio signal collected through a second pickup channel; determining a first empirical cumulative distribution of a first frequency energy based on the first audio signal; determining a second empirical cumulative distribution of a second frequency energy based on the second audio signal; determining a target pickup channel based on the first and second empirical cumulative distributions, wherein the target pickup channel is a blocked pickup channel among the first and second pickup channels; and performing fusion processing on the audio signal of the target pickup channel to obtain a target audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio processing technology, specifically relating to an audio signal processing method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] Microphones pick up and record audio signals, such as speech and other sound signals. The microphone hole of a microphone can become blocked during use. For example, when a user holds a mobile phone in a handheld position for a call, their ear may block the microphone hole at the top of the phone; when the phone is placed upright on a table, the bottom microphone hole may be blocked by the tabletop; when the phone is in landscape mode, such as when using applications like games or video conferencing, the user's fingers, other body parts, or objects may also block the microphone hole. Furthermore, the microphone hole can become blocked by dust or other debris over a long period. When the microphone hole is blocked, the quality of the captured voice signal deteriorates significantly, resulting in poor audio quality. Summary of the Invention

[0003] The purpose of this application is to provide an audio signal processing method, apparatus, electronic device, and readable storage medium that can solve the problem of poor audio quality when the pickup channel is blocked.

[0004] In a first aspect, embodiments of this application provide an audio signal processing method applied to an electronic device, the electronic device including a first pickup channel and a second pickup channel, comprising:

[0005] Acquire a first audio signal acquired through the first pickup channel and a second audio signal acquired through the second pickup channel;

[0006] A first empirical cumulative distribution of energy at the first frequency is determined based on the first audio signal;

[0007] Determine the second empirical cumulative distribution of the second frequency energy based on the second audio signal;

[0008] Based on the first experience cumulative distribution and the second experience cumulative distribution, a target pickup channel is determined, wherein the target pickup channel is the blocked pickup channel among the first pickup channel and the second pickup channel;

[0009] The audio signal of the target pickup channel is fused to obtain the target audio signal.

[0010] Secondly, embodiments of this application provide an audio signal processing device applied to an electronic device, the electronic device including a first pickup channel and a second pickup channel, comprising:

[0011] The first acquisition module is used to acquire a first audio signal collected through the first pickup channel and a second audio signal collected through the second pickup channel;

[0012] The first determining module is used to determine a first empirical cumulative distribution of the energy of the first frequency based on the first audio signal;

[0013] The second determining module is used to determine the second empirical cumulative distribution of the second frequency energy based on the second audio signal;

[0014] The third determining module is used to determine the target pickup channel based on the first experience cumulative distribution and the second experience cumulative distribution, wherein the target pickup channel is the blocked pickup channel among the first pickup channel and the second pickup channel;

[0015] The fusion module is used to fuse the audio signal of the target pickup channel to obtain the target audio signal.

[0016] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0017] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0018] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0019] In a sixth aspect, embodiments of this application provide a program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0020] In this embodiment, a first audio signal acquired through the first pickup channel and a second audio signal acquired through the second pickup channel are obtained; a first empirical cumulative distribution of first frequency energy is determined based on the first audio signal; a second empirical cumulative distribution of second frequency energy is determined based on the second audio signal; a target pickup channel is determined based on the first and second empirical cumulative distributions, wherein the target pickup channel is the blocked pickup channel among the first and second pickup channels; the audio signal of the target pickup channel is fused to obtain the target audio signal. Through this method, whether a pickup channel is blocked can be determined based on the empirical cumulative distribution of the frequency energy of the audio signal, which can improve the accuracy of detecting whether a pickup channel is blocked, reduce the degree of false detection, and improve the audio playback effect by fusing the blocked channel. Attached Figure Description

[0021] Figure 1 This is a flowchart of an audio signal processing method provided in an embodiment of this application;

[0022] Figure 2 This is another flowchart of the audio signal processing method provided in the embodiments of this application;

[0023] Figures 3a-3c This is an empirical cumulative distribution curve provided in the embodiments of this application;

[0024] Figure 3d This is an audio signal spectrum diagram provided in an embodiment of this application;

[0025] Figure 3e These are the mean absolute error curve and phase variance curve provided in the embodiments of this application;

[0026] Figure 3f This is a spectrum diagram of the output audio signal provided in an embodiment of this application;

[0027] Figure 3g This is an audio signal spectrum diagram provided in an embodiment of this application;

[0028] Figure 3h These are the mean absolute error curve and phase variance curve provided in the embodiments of this application;

[0029] Figure 3i This is a spectrum diagram of the output audio signal provided in an embodiment of this application;

[0030] Figure 4 This is a structural diagram of the audio signal processing apparatus provided in the embodiments of this application;

[0031] Figure 5 This is one of the structural diagrams of the electronic device provided in the embodiments of this application;

[0032] Figure 6 This is the second structural diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0034] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0035] The audio signal processing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0036] Figure 1 This is a flowchart of an audio signal processing method provided in an embodiment of this application. The audio signal processing method in this embodiment is applied to an electronic device, which includes a first pickup channel and a second pickup channel, and includes the following steps:

[0037] Step 101: Acquire the first audio signal collected through the first pickup channel and the second audio signal collected through the second pickup channel.

[0038] The first and second pickup channels can be microphone holes. The electronic device can have dual microphones or multiple microphones. In the case of multiple microphones, audio signals from two or more microphones can be obtained from multiple microphones for subsequent operations.

[0039] Step 102: Determine the first empirical cumulative distribution of the first frequency energy based on the first audio signal.

[0040] Multiple frequencies can be determined from the first audio signal. For each frequency, the energy of that frequency (also called frequency energy or frequency point energy) can be determined. Based on the frequency energy, an empirical cumulative distribution is obtained. In this empirical cumulative distribution, the horizontal axis represents the frequency energy, and the vertical axis represents the cumulative value. The cumulative value ranges from 0 to 1, with a minimum of 0 and a maximum of 1.

[0041] Step 103: Determine the second empirical cumulative distribution of the second frequency energy based on the second audio signal.

[0042] Multiple frequencies can be determined from the second audio signal, and the energy of each frequency can be determined. Based on the energy of the frequency, an empirical cumulative distribution is obtained, where the horizontal axis represents the frequency energy and the vertical axis represents the cumulative value, which ranges from 0 to 1, with a minimum of 0 and a maximum of 1.

[0043] Step 104: Based on the first experience cumulative distribution and the second experience cumulative distribution, determine the target pickup channel, wherein the target pickup channel is the blocked pickup channel among the first pickup channel and the second pickup channel.

[0044] When a clogging phenomenon occurs, compared to an unclogging pickup channel (hereinafter also referred to as a channel), the frequency energy distribution of the clogging channel is more concentrated in the low-energy region. By observing the difference in the cumulative values ​​of the low-energy regions of the two channels, the clogging phenomenon can be effectively identified. Preferably, the low-energy regions of the first and second empirical cumulative distributions can be observed in detail. For example, the first and second empirical cumulative distributions have the same maximum and minimum frequency energies (in dB). The region between the maximum and minimum frequency energies is a low-energy region belonging to a preset energy range, which facilitates determining the channel clogging status based on the difference in the cumulative values ​​of the same low-energy regions in the two distributions.

[0045] For example, the process involves obtaining the cumulative values ​​corresponding to various frequencies within the same low-energy region; calculating the absolute value of the difference between the cumulative values ​​of the two distributions at each frequency; averaging the absolute values ​​corresponding to each frequency; and determining the channel congestion status based on the average value. For instance, if the cumulative values ​​corresponding to -80dB in the two distributions are 0.02 and 0.01 respectively, with an absolute difference of 0.01; and the cumulative values ​​corresponding to -60dB are 0.2 and 0.02 respectively, with an absolute difference of 0.18, averaging these absolute values ​​yields an average of 0.095. If the average value is greater than a preset threshold, it indicates that one of the two channels is congested while the other is not. If the average value is less than the preset threshold, it indicates that neither channel is congested.

[0046] Step 105: Perform fusion processing on the audio signal of the target pickup channel to obtain the target audio signal.

[0047] In this embodiment, a first audio signal acquired through the first pickup channel and a second audio signal acquired through the second pickup channel are obtained; a first empirical cumulative distribution of first frequency energy is determined based on the first audio signal; a second empirical cumulative distribution of second frequency energy is determined based on the second audio signal; a target pickup channel is determined based on the first and second empirical cumulative distributions, wherein the target pickup channel is the blocked pickup channel among the first and second pickup channels; the audio signal of the target pickup channel is fused to obtain the target audio signal. Through the above method, whether a pickup channel is blocked can be determined based on the empirical cumulative distribution of the frequency energy of the audio signal, which can improve the accuracy of detecting whether a pickup channel is blocked, reduce the degree of false detection, and improve the audio playback effect by fusing the blocked channel.

[0048] In one embodiment of this application, step 102, determining the first empirical cumulative distribution of the first frequency energy based on the first audio signal, includes steps 1021 and 1022:

[0049] Step 1021: Determine the first power spectrum based on the first audio signal;

[0050] Step 1022: Determine the first empirical cumulative distribution of the first frequency energy based on the first power spectrum;

[0051] Step 103, determining the second empirical cumulative distribution of the second frequency energy based on the second audio signal, includes steps 1031 and 1032:

[0052] Step 1031: Determine the second power spectrum based on the second audio signal;

[0053] Step 1032: Determine the second empirical cumulative distribution of the second frequency energy based on the second power spectrum;

[0054] The first frequency energy is determined based on the frequency power in the first power spectrum, the second frequency energy is determined based on the power in the second power spectrum, and the first empirical cumulative distribution and the second empirical cumulative distribution have the same minimum frequency energy and the same maximum frequency energy.

[0055] Specifically, a Short-Time Fourier Transform (STFT) can be performed on the first audio signal and the second audio signal respectively to obtain a first power spectrum and a second power spectrum. The horizontal axis of the power spectrum represents frequency, and the vertical axis represents power, which can also be converted into decibels.

[0056] For example, time-domain audio signals x1 (i.e., the first audio signal) and x2 (i.e., the second audio signal) are acquired through two pickup channels, respectively. After Short-Time Fourier Transform (STFT), the spectra X1(λ,μ) and X2(λ,μ) are obtained, where λ represents the time frame number, which corresponds to a time period. The length of this time period can be set according to actual conditions and is not limited here. The first and second audio signals can include an audio signal of a preset duration, which can include multiple time frames. A time frame can be understood as the smallest time unit for the electronic device to process audio signals in this application. The length of the preset duration can be determined according to actual conditions and is not limited here. μ represents the frequency point number, which corresponds to a frequency value. The power spectrum, which can be expressed in decibels (dB), is as follows:

[0057] Φ1(λ,μ)=20·log 10 |X1(λ,μ)| (1)

[0058] Φ2(λ,μ)=20·log 10 |X2(λ,μ)| (2)

[0059] The first frequency energy is determined based on the power in the first power spectrum. For example, in equation (1), the value of the first frequency energy Φ1(λ,μ) is the power corresponding to the frequency (μ is the corresponding frequency) in the first power spectrum (this power is expressed in decibels). The first frequency energy Φ1(λ,μ) is determined based on the spectrum X1(λ,μ), which is obtained by passing the first audio signal x1 through STFT.

[0060] Accordingly, the second frequency energy is determined based on the power in the second power spectrum. For example, in equation (2), the value of the second frequency energy Φ2(λ,μ) is the power corresponding to the frequency (μ is the corresponding frequency) in the second power spectrum (this power is expressed in decibels). The second frequency energy Φ2(λ,μ) is determined based on the spectrum X2(λ,μ), which is obtained by passing the second audio signal x2 through STFT.

[0061] It should be noted that the horizontal axes of the first and second power spectra have the same frequency boundaries, that is, they have the same minimum and maximum frequencies.

[0062] Within a time frame, the number of low-energy frequency points in the signal acquired by the blocked channel increases significantly. In this application, the Empirical Cumulative Distribution Function (eCDF) is used to statistically determine the frequency energy. eCDF is an important tool for displaying the statistical characteristics of random numbers, showing distribution properties, requiring no parameter estimation, and is easy to implement. The eCDF definitions for the frequency energy of each channel are as follows:

[0063]

[0064]

[0065] Where μ1 and μ2 represent the start and end numbers of the frequency point energy statistical samples, respectively; I{·} represents the indicator function, which maps to 1 when the condition within the parentheses is met, and to 0 otherwise; d l It is the segmented interval boundary of eCDF, d l ∈{d1,d2,…,d L Based on real-world recordings from multiple scenarios, it was determined that microphone occlusion has a weak effect on energy attenuation in frequency bands below 600Hz. That is, the characteristics of the occlusion channel are not obvious in the low-frequency band of 0-600Hz. In this embodiment, frequency bands above 600Hz can be observed. μ1 is set as the frequency point number that roughly corresponds to 600Hz, and μ2 corresponds to the maximum identifiable frequency (according to the Nyquist sampling theorem, it is 1 / 2 of the sampling frequency).

[0066] Figures 3a-3c Examples of frequency energy (eCDF) of recorded corpora under three conditions are given. Figure 3a The empirical cumulative distribution of frequency energy when channel 1 (i.e., the first pickup channel) is blocked and channel 2 (i.e., the second pickup channel) is not blocked; Figure 3b The empirical cumulative distribution of frequency energy when channel 1 and channel 2 are unblocked or without wind noise; Figure 3c The empirical cumulative distribution of frequency energy is given when channel 1 has strong wind noise and channel 2 has no wind noise.

[0067] like Figure 3a As shown, when channel blockage occurs, compared to unblocked channels, the frequency energy distribution of the blocked channel is more concentrated in the low-energy region, and the eCDF difference in this region is significant; for example... Figure 3c As shown, wind noise primarily increases the distribution of high-energy frequencies, and this difference is only slightly related to the background noise level. Based on this, channel blockage can be effectively identified by observing the difference in eCDF in the low-energy region of the dual channels.

[0068] To adapt to various background noise levels, it is necessary to define appropriate low-energy discrimination boundaries {d1,d2,…,d L}, the low-energy distinction boundary satisfies: (1) the rightmost boundary d L (That is, the preset maximum boundary) should ensure that the blocked audio signal picked up in an environment with high background noise levels has a suitable number of frequency points with energy less than d. L (2) The leftmost boundary d1 (i.e., the preset minimum boundary) should ensure that differences can still be captured in environments with low background noise; (3) The scale value (d l -d l-1 The captured eCDF differences should be significant. The first empirical cumulative distribution and the second empirical cumulative distribution have the same preset minimum boundary, i.e., minimum frequency energy, and the same preset maximum boundary, i.e., maximum frequency energy. The energy region between the preset minimum boundary and the preset maximum boundary is the low energy region. The low energy region can be the range of a low-pressure preset energy threshold, for example, the energy range below 0 dB. It can be set according to the actual situation and is not limited here.

[0069] Let {d1,d2,…,d L Once set to a fixed value, the dynamic range of background noise that can be adapted is very large, and no additional background noise estimation algorithm is needed to determine whether background noise exists.

[0070] In this embodiment, when channel blockage occurs, compared with the unblocked channel, the frequency energy distribution of the blocked channel is more concentrated in the low energy region, and the eCDF difference in this region is obvious. By observing the difference in the empirical cumulative distribution of the low energy region of the two channels, the channel blockage phenomenon can be effectively identified, and the identification accuracy can be improved.

[0071] In one embodiment of this application, step 104, determining the target pickup channel based on the first empirical cumulative distribution and the second empirical cumulative distribution, includes steps 1041-1044:

[0072] Step 1041: Determine the absolute difference of the frequency energy cumulative distribution based on the first empirical cumulative distribution and the second empirical cumulative distribution;

[0073] Step 1042: When the absolute difference of the cumulative frequency energy distributions is greater than the first threshold, obtain the first cumulative frequency energy and the second cumulative frequency energy, wherein the first cumulative frequency energy is the maximum cumulative frequency energy value in the first empirical cumulative distribution, and the second cumulative frequency energy is the maximum cumulative frequency energy value in the second empirical cumulative distribution.

[0074] Step 1043: Determine the first hole-closing probability of the pickup channel corresponding to the larger of the first frequency accumulated energy and the second frequency accumulated energy. The first hole-closing probability is determined based on the absolute difference of the frequency energy accumulated distribution.

[0075] Step 1044: Determine the target pickup channel based on the first hole blockage probability.

[0076] Specifically, the Mean Absolute Difference (MAD) of the cumulative frequency energy distribution measures the difference in the empirical cumulative distributions of the two pickup channels. The expression for calculating the mean absolute difference D(λ) of the frequency energy is as follows:

[0077]

[0078] In equation (5), when D(λ) is greater than the first threshold D lower At that time, it was considered that the low-energy frequency distributions of the two channels differed greatly, indicating that a hole blockage phenomenon had occurred, but the hole blockage probability of each channel needs further calculation. The rightmost value of eCDF is eCDF(λ,d). L Larger channels have more low-energy frequencies, suggesting that eCDF(λ,d) L The larger channel is considered blocked, i.e., the accumulated energy at the first frequency and the accumulated energy at the second frequency are compared, and the channel corresponding to the larger value is considered blocked, eCDF1(λ,d) L That is, the cumulative value corresponding to the maximum frequency energy in the first empirical cumulative distribution, eCDF2(λ,d) L That is, the cumulative value corresponding to the maximum frequency energy in the second empirical cumulative distribution.

[0079] Optionally, the probability of the first blockage in the blocked channel can be calculated, that is, D(λ) is linearly mapped to the probability of the first blockage p(λ). The following formula is used as an example of channel 1 being blocked:

[0080]

[0081] Based on the probability of the first blockage, determine whether the first and second pickup channels are blocked. For example, according to equation (6), D(λ) > D upper When the threshold (i.e., the third threshold) is reached, a "severe" blockage is considered to have occurred (probability 100%); D(λ) <D lower When the channel is considered to be blocked (probability 0%), it is assumed that no blockage has occurred. In other cases, the probability (degree) of blockage is positively correlated with D(λ). If the blocked channel is considered to be the first pickup channel, the second pickup channel is judged to be unblocked, with a blockage probability p2(λ) = 0.

[0082] In the above embodiments, the blockage probability of each of the two channels can be obtained, thereby determining the availability of each channel, which facilitates subsequent processing of the audio signals acquired from the two channels and improves the audio processing effect.

[0083] When a finger is used to block the aperture, the finger may slightly slide near the pickup channel aperture. This sliding momentarily forces a high-speed airflow into the pickup channel aperture, generating a large instantaneous energy that significantly interferes with the speech in that channel. At this point, D(λ) will still be greater than D. lower In this time frame, the channel is judged to be blocked. However, the output signal energy of the pickup channel that is being slid by a finger is extremely high, which may cause another pickup channel that is neither blocked nor slid to be judged as blocked. Because the airflow forced into the microphone by a finger sliding on it has extremely high bandwidth and energy, the phase of all frequency points within the corresponding time frame will converge towards the phase of this high-speed airflow, causing the phase variance within that frame to decrease rapidly. Therefore, to avoid misjudgment of channel blockage due to slight finger sliding near the pickup channel aperture, this embodiment of the application increases the phase variance to reduce misjudgment, specifically:

[0084] Step 1042, when the absolute difference of the cumulative frequency energy distribution is greater than the first threshold, obtain the first frequency cumulative energy and the second frequency cumulative energy, including the following steps:

[0085] Step 10421: If the absolute difference in frequency energy is greater than the first threshold, obtain the first phase variance of the first audio signal and the second phase variance of the second audio signal.

[0086] Step 10422: When both the first phase variance and the second phase variance are greater than the second threshold, obtain the first frequency accumulated energy and the second frequency accumulated energy.

[0087] When D(λ)>D lower In the case of the first audio signal and the second audio signal, the phase variance of each channel is calculated, that is, the first phase variance of the first audio signal and the second phase variance of the second audio signal are calculated. If the first phase variance is less than the second threshold and the second phase variance is greater than the second threshold, then it is determined that the first pickup channel is blocked and its blockage probability is set to p(λ) = 1, and the second pickup channel is not blocked and its blockage probability is set to p(λ) = 0.

[0088] If both the first phase variance and the second phase variance are greater than the second threshold, then the linear mapping of D(λ) to the occlusion probability is performed again, that is, the following steps are performed: obtaining the first frequency accumulated energy and the second frequency accumulated energy; obtaining the first occlusion probability of the pickup channel corresponding to the larger of the first frequency accumulated energy and the second frequency accumulated energy; and determining whether the first pickup channel and the second pickup channel are blocked based on the first occlusion probability.

[0089] In the above embodiments, the empirical cumulative distribution of frequency energy and phase variance are combined to detect channel blockage. This can identify wind noise interference without complex background noise estimation, thereby improving the robustness and accuracy of blockage detection.

[0090] In one embodiment of this application, step 104, determining the target pickup channel based on the first hole blockage probability, includes step 1041:

[0091] Step 1041: When the probability of the first hole blockage is greater than the third threshold, the first pickup channel is determined to be the target pickup channel, and the first pickup channel is the pickup channel corresponding to the larger of the first frequency accumulated energy and the second frequency accumulated energy.

[0092] Step 105, the fusion processing of the audio signal of the target pickup channel to obtain the target audio signal includes step 1051:

[0093] Step 1051: Based on the first hole blockage probability and the second hole blockage probability of the second pickup channel, perform amplitude fusion processing on the first audio signal to obtain a target audio signal, wherein the amplitude ratio of the second audio signal in the target audio signal is greater than the amplitude ratio of the first audio signal.

[0094] Specifically, according to equation (6), D(λ) > D upper When the threshold (i.e., the third threshold) is reached, a “serious” clogging phenomenon is considered to have occurred (probability 100%), meaning the first pickup channel is blocked and the second pickup channel is not blocked.

[0095] To improve the output performance of the blocked channel, the output signal needs to be processed. In this embodiment, amplitude fusion processing is performed on the output signal of the blocked channel. Specifically, based on the blocking probabilities p1(λ) and p2(λ) of the two channels, amplitude fusion is performed on the blocked channel to ensure that the amplitude of the channel with less blocking (which can also be understood as the unblocked channel) has a larger proportion while the phase remains unchanged, thus preserving effective information for the spatial filtering algorithm. The other unblocked channel is not processed. Taking channel 1 as an example with a blocked hole, the spectra of the dual-channel output signals are as follows:

[0096]

[0097] Y2(λ,μ)=X2(λ,μ) (8)

[0098] After inverse STFT, the spectral signal is finally converted into time-domain signals y1 and y2 for output.

[0099] In the above embodiments, the channel amplitude fusion processing is performed on the audio of the blocked channel by utilizing the blockage probability of each channel, thereby improving the listening experience of the user in the call or recording device, while maintaining a better user experience with dual-channel stereo output.

[0100] It should be noted that in the scheme of this application, the electronic device processes the audio signal in units of time frames. To avoid abrupt changes between adjacent time frames and to improve the continuity of the criteria and output signal, eCDF(λ), D(λ), and p(λ) in the above embodiments can all undergo inter-frame smoothing, and the smoothed data can be used in subsequent calculations. Regarding the hole-closing probability, after inter-frame smoothing, the hole-closing probability will change, and subsequent calculations can be performed based on the new hole-closing probability.

[0101] The following provides an example of the audio signal processing method provided in the embodiments of this application.

[0102] like Figure 2 As shown, the audio signal processing method includes the following steps:

[0103] Step 201, STFT. Time-domain audio signals x1 (i.e., the first audio signal) and x2 (i.e., the second audio signal) are acquired through two pickup channels respectively. After passing through STFT, the spectra X1(λ,μ) and X2(λ,μ) are obtained, where λ represents the time frame number, which corresponds to a time period. The length of this time period can be set according to the actual situation and is not limited here. The first audio signal and the second audio signal can include an audio signal of a preset duration. The preset duration can include multiple time frames. A time frame can be understood as the smallest time unit for the electronic device to process audio signals in this application. The length of the preset duration can be determined according to the actual situation and is not limited here. μ represents the frequency point number, which corresponds to a frequency value. The power spectrum, which can be expressed in decibels (dB), is shown in Equation (1) and Equation (2) respectively.

[0104] Step 202: Determine the power spectrum of each of the two pickup channels, as described above.

[0105] Step 203: Determine the eCDF for each of the two pickup channels. See the previous description for details.

[0106] Step 204: Calculate the mean absolute difference D. Within a time frame, the number of low-energy frequency points of the signal acquired by the blocked channel increases significantly. This application uses the Empirical Cumulative Distribution Function (eCDF) to statistically determine the frequency energy. eCDF is an important tool for displaying the statistical characteristics of random numbers, showing distribution characteristics, requiring no parameter estimation, and is easy to implement. The definitions of the frequency energy eCDF for each channel are shown in Equations (3) and (4), respectively.

[0107] Where μ1 and μ2 represent the start and end numbers of the frequency point energy statistical samples, respectively; I{·} represents the indicator function, which maps to 1 when the condition within the parentheses is met, and to 0 otherwise; d l It is the segmented interval boundary of eCDF, d l ∈{d1,d2,…,d L}

[0108] The difference in empirical cumulative distributions corresponding to the two pickup channels is measured by the MAD (abbreviated as D) of frequency energy. The mean absolute error D(λ) of frequency energy is calculated as shown in Equation (5).

[0109] Step 205: Determine if D is greater than the first threshold. If D(λ) is greater than the first threshold D... lower At that time, it was considered that the low-energy frequency distributions of the two channels differed greatly, indicating that a hole blockage phenomenon had occurred, but the hole blockage probability of each channel needs further calculation. The rightmost value of eCDF is eCDF(λ,d). L Larger channels have more low-energy frequencies, suggesting that eCDF(λ,d) L If a large channel is blocked, D(λ) is linearly mapped to its blocking probability p(λ), as shown in equation (6).

[0110] D(λ)>D upper When the threshold (i.e., the third threshold) is reached, a "severe" blockage is considered to have occurred (probability 100%); D(λ) <D lower When the first pickup channel is considered to be blocked (probability 0%), the probability of blockage in other cases is positively correlated with D(λ). If the blocked channel is considered the first pickup channel, the second pickup channel is judged to be unblocked, with a blockage probability p2(λ) = 0.

[0111] When a finger is used to block the aperture, the finger may slightly slide near the pickup channel aperture. This sliding momentarily forces a high-speed airflow into the pickup channel aperture, generating a large instantaneous energy that significantly interferes with the speech in that channel. At this point, D(λ) will still be greater than D. lowerIn this time frame, the channel is judged to be blocked. However, the output signal energy of the pickup channel that is being slid by a finger is extremely high, which may cause another pickup channel that is neither blocked nor slid to be judged as blocked. Because the airflow forced into the microphone by the finger sliding on it has extremely high bandwidth and energy, the phase of all frequency points within the corresponding time frame will move closer to the phase of this high-speed airflow, causing the phase variance within this frame to decrease rapidly.

[0112] Step 206, calculate the phase variance. When D(λ) > D lower In this case, calculate the phase variance of each of the two channels, that is, calculate the first phase variance of the first audio signal and the second phase variance of the second audio signal;

[0113] Step 207: Determine if there is a phase variance less than the second threshold. If the first phase variance is less than the second threshold and the second phase variance is greater than the second threshold, proceed to step 209; if both the first phase variance and the second phase variance are greater than the second threshold, proceed to step 208.

[0114] To avoid abrupt changes and improve the continuity of the criteria and output signals, eCDF(λ), D(λ), and p(λ) are all smoothed across frames.

[0115] Step 208: Map D(λ) to the probability of plugging the hole.

[0116] Step 209: Determine that the first pickup channel is blocked, and set its blockage probability to p(λ) = 1, and that the second pickup channel is not blocked, and set its blockage probability to p(λ) = 0.

[0117] Step 210: The channel with the larger value corresponding to the rightmost boundary of eCDF is the blocked channel, and the blocking probability is set to p; the other channel is considered not blocked, and the blocking probability is set to 0.

[0118] Step 211: In the time frame where the hole blockage occurs, for channels with a high probability of hole blockage exceeding the first threshold, amplitude fusion is performed on the blocked channels based on the hole blockage probabilities p1(λ) and p2(λ), so that the amplitude of the channel with a lower degree of hole blockage accounts for a larger proportion; the phase remains the same as the phase of the blocked channel, preserving effective information for the spatial filtering algorithm. The other unblocked channel is not processed. Taking channel 1 as an example of being blocked, the dual-channel output signal spectra are shown in Equations (7) and (8), respectively.

[0119] After inverse STFT, the spectral signal is finally converted into time-domain signals y1 and y2 for output.

[0120] The following examples use two audio recordings from indoor and outdoor environments with different background noise levels to demonstrate the estimation results of the MAD and phase variance of the eCDF signal, as well as the hole blockage probability and output signal. The output of the proposed algorithm in hole blockage and wind noise scenarios is compared to illustrate the robustness of the proposed hole blockage detection.

[0121] like Figure 3d-3f The data shown corresponds to speech recorded in a quiet indoor environment. Wind noise is present from 0 to 6 seconds, and the first pickup channel A is blocked from 8.5 to 15.5 seconds. Figure 3d The upper part of the image shows the audio signal 11 obtained from the first pickup channel, and the lower part shows the audio signal 12 obtained from the second pickup channel. Figure 3e The upper part of the image shows the MAD curve obtained according to the method provided in the embodiment of this application, and the dashed line in the image represents the first threshold 13; the middle and lower parts of the image show the phase variance curve obtained according to the method provided in the embodiment of this application, and the dashed line in the image represents the second threshold 14.

[0122] Figure 3f The upper part of the image shows the output signal obtained after processing the audio signal obtained from the first pickup channel. As can be seen from the figure, the probability of the first pickup channel being blocked is 1 between 8.5 seconds and 15.5 seconds, which is consistent with the actual situation. Figure 3f The lower part of the image shows the output signal obtained after processing the audio signal obtained from the second pickup channel. As can be seen from the figure, the probability of the second pickup channel being blocked is 0 from 0 seconds to 15.5 seconds, which is consistent with the actual situation.

[0123] By comparison Figure 3d and Figure 3f From the audio corresponding to the first pickup channel, it can be seen that compared to the amplitude of the input signal, the amplitude of the output signal obtained after amplitude fusion processing increases between 8.5 seconds and 15.5 seconds. Figure 3f The two images in the image show similar amplitudes between 8.5 seconds and 15.5 seconds. By performing amplitude fusion processing, the amplitude of the channel with less blockage is given a larger proportion, which can improve the output signal effect of the channel with the less blockage.

[0124] like Figure 3g-3i The data shown corresponds to real-life outdoor recordings. Wind noise is present from 0 to 9 seconds, and the first pickup channel is blocked from 9 to 16 seconds. Figure 3g The upper part of the image shows the audio signal obtained from the first pickup channel. The time period t in the image represents the period from 9 seconds to 16 seconds. The audio signal during this period shows a significant change compared to the unblocked period. The lower part of the image shows the audio signal obtained from the second pickup channel. Figure 3hThe upper part of the image is a MAD curve obtained according to the method provided in the embodiments of this application, and the dashed line in the image represents the first threshold; the middle part and the lower part of the image are phase variance curves obtained according to the method provided in the embodiments of this application, and the dashed line in the image represents the second threshold.

[0125] Figure 3i The upper part of the image shows the output signal obtained after processing the audio signal obtained from the first pickup channel. As can be seen from the figure, the probability of the first pickup channel being blocked is approximately 1 between 9 and 16 seconds, which is consistent with the actual situation. Figure 3i The lower part of the image shows the output signal obtained after processing the audio signal obtained from the second pickup channel. As can be seen from the figure, the probability of the second pickup channel being blocked is 0 from 0 to 16 seconds, which is consistent with the actual situation.

[0126] The audio signal sampling frequency is set to 16kHz, and the STFT frame length and Fourier transform length are both set to 512 points. lower =0.1, D upper =0.15, d l ={-90,-80,-70,-60} dB, Phase variance judgment threshold: 1 / 4·π 2 ≈2.47.

[0127] from Figure 3e and Figure 3h As can be seen, when the hole blockage occurs, the value of D increases significantly; when there is no hole blockage, regardless of whether there is wind noise, D is below the set threshold. Figure 3h In the figure, label 16 shows the phase variance curve corresponding to channel 2, and label 17 shows the phase variance curve corresponding to channel 1.

[0128] In environments with the aforementioned two background noise levels, the same fixed threshold and parameter settings can accurately detect hole blockage and distinguish wind noise, demonstrating good robustness. Combining the hole blockage probability based on D and phase variance, this application can output an effective dual-channel audio signal, i.e., a stereo output signal, for the blocked channel.

[0129] The method provided in this application can accurately detect channel blockage, effectively avoid wind noise interference, and prevent damage to the voice signal by the wind noise suppression module, thus improving the user's auditory experience for calls or recording devices. Simultaneously, it maintains a better user experience with dual-channel stereo output. Compared with existing methods for closing channel blockages, it retains more useful information and improves the quality of the output audio.

[0130] The audio signal processing method provided in this application can be executed by an audio signal processing device. This application uses an audio signal processing device executing the audio signal processing method as an example to illustrate the audio signal processing device provided in this application.

[0131] like Figure 4 As shown, an audio signal processing device 400 is applied to an electronic device, which includes a first pickup channel and a second pickup channel, comprising:

[0132] The first acquisition module 401 is used to acquire a first audio signal acquired through the first pickup channel and a second audio signal acquired through the second pickup channel;

[0133] The first determining module 402 is used to determine the first empirical cumulative distribution of the first frequency energy based on the first audio signal;

[0134] The second determining module 403 is used to determine the second empirical cumulative distribution of the second frequency energy based on the second audio signal;

[0135] The third determining module 404 is used to determine the target pickup channel based on the first experience cumulative distribution and the second experience cumulative distribution, wherein the target pickup channel is the blocked pickup channel among the first pickup channel and the second pickup channel;

[0136] The fusion module 405 is used to perform fusion processing on the audio signal of the target pickup channel to obtain the target audio signal.

[0137] Optionally, the first determining module 402 includes:

[0138] The first determining submodule is used to determine the first power spectrum based on the first audio signal;

[0139] The second determining submodule is used to determine the first empirical cumulative distribution of the energy at the first frequency based on the first power spectrum.

[0140] The second determining module 403 includes:

[0141] The third determining submodule is used to determine the second power spectrum based on the second audio signal;

[0142] The fourth determining submodule is used to determine the second empirical cumulative distribution of the second frequency energy based on the second power spectrum;

[0143] The first frequency energy is determined based on the frequency power in the first power spectrum, the second frequency energy is determined based on the power in the second power spectrum, and the first empirical cumulative distribution and the second empirical cumulative distribution have the same minimum frequency energy and the same maximum frequency energy.

[0144] Optionally, the third determining module 404 includes:

[0145] The fifth determining submodule is used to determine the absolute difference of the frequency energy cumulative distribution based on the first empirical cumulative distribution and the second empirical cumulative distribution;

[0146] The sixth determining submodule is used to obtain a first frequency accumulated energy and a second frequency accumulated energy when the absolute difference of the frequency energy accumulated distribution is greater than a first threshold. The first frequency accumulated energy is the maximum frequency energy accumulated value in the first empirical accumulated distribution, and the second frequency accumulated energy is the maximum frequency energy accumulated value in the second empirical accumulated distribution.

[0147] The seventh determining submodule is used to determine the first hole-closing probability of the pickup channel corresponding to the larger of the first frequency accumulated energy and the second frequency accumulated energy, wherein the first hole-closing probability is determined based on the absolute difference of the frequency energy accumulated distribution.

[0148] The eighth determining submodule is used to determine the target pickup channel based on the first hole blockage probability.

[0149] Optionally, the first acquisition submodule includes:

[0150] The first acquisition unit is used to obtain the first phase variance of the first audio signal and the second phase variance of the second audio signal when the absolute difference of the cumulative frequency energy distribution is greater than a first threshold.

[0151] The second acquisition unit is used to acquire the first frequency accumulated energy and the second frequency accumulated energy when both the first phase variance and the second phase variance are greater than the second threshold.

[0152] Optionally, the eighth determining submodule includes:

[0153] The determining unit is used to determine the first pickup channel as the target pickup channel when the first hole blockage probability is greater than the third threshold, wherein the first pickup channel is the pickup channel corresponding to the larger of the first frequency accumulated energy and the second frequency accumulated energy.

[0154] The fusion module 405 is used to perform amplitude fusion processing on the first audio signal based on the first hole blockage probability and the second hole blockage probability of the second pickup channel to obtain a target audio signal, wherein the amplitude of the second audio signal accounts for a greater proportion than the amplitude of the first audio signal in the amplitude of the target audio signal.

[0155] Optionally, the third determining module 404 includes:

[0156] The ninth determining submodule is used to determine the absolute difference of the frequency energy cumulative distribution based on the first empirical cumulative distribution and the second empirical cumulative distribution;

[0157] The tenth determining submodule is used to obtain the first phase variance of the first audio signal and the second phase variance of the second audio signal when the absolute difference of the cumulative frequency energy distribution is greater than the first threshold.

[0158] The eleventh determination submodule is used to determine the first pickup channel as the target pickup channel when the first phase variance is less than the second threshold and the second phase variance is greater than the second threshold.

[0159] The audio signal processing device 400 provided in this application embodiment can implement the various processes implemented in the aforementioned method embodiments, and will not be described again here to avoid repetition.

[0160] The audio signal processing device 400 in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific type of device.

[0161] The audio signal processing device 400 in this embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this embodiment does not specifically limit its use.

[0162] Optionally, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described audio signal processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0163] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0164] Figure 6 A hardware structure diagram of an electronic device to implement an embodiment of this application.

[0165] The electronic device 600 includes, but is not limited to, components such as: radio frequency unit 601, network module 602, audio output unit 603, input unit 604, sensor 605, display unit 606, user input unit 607, interface unit 608, memory 609, and processor 610.

[0166] Those skilled in the art will understand that the electronic device 600 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 610 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0167] The processor 610 is configured to acquire a first audio signal acquired through the first pickup channel and a second audio signal acquired through the second pickup channel; determine a first empirical cumulative distribution of first frequency energy based on the first audio signal; determine a second empirical cumulative distribution of second frequency energy based on the second audio signal; determine a target pickup channel based on the first empirical cumulative distribution and the second empirical cumulative distribution, wherein the target pickup channel is the blocked pickup channel among the first pickup channel and the second pickup channel; and perform fusion processing on the audio signal of the target pickup channel to obtain a target audio signal.

[0168] Optionally, the processor 610 is further configured to determine a first power spectrum based on the first audio signal;

[0169] Based on the first power spectrum, a first empirical cumulative distribution of the first frequency energy is determined; based on the second audio signal, a second power spectrum is determined; based on the second power spectrum, a second empirical cumulative distribution of the second frequency energy is determined; wherein, the first frequency energy is determined based on the frequency power in the first power spectrum, the second frequency energy is determined based on the power in the second power spectrum, and the first empirical cumulative distribution and the second empirical cumulative distribution have the same minimum frequency energy and the same maximum frequency energy.

[0170] Optionally, the processor 610 is further configured to: determine the mean absolute difference of frequency energy cumulative distributions based on the first empirical cumulative distribution and the second empirical cumulative distribution; if the mean absolute difference of frequency energy cumulative distributions is greater than a first threshold, acquire a first frequency cumulative energy and a second frequency cumulative energy, wherein the first frequency cumulative energy is the maximum frequency energy cumulative value in the first empirical cumulative distribution and the second frequency cumulative energy is the maximum frequency energy cumulative value in the second empirical cumulative distribution; determine a first occlusion probability of the pickup channel corresponding to the larger of the first frequency cumulative energy and the second frequency cumulative energy, wherein the first occlusion probability is determined based on the mean absolute difference of frequency energy cumulative distributions; and determine the target pickup channel based on the first occlusion probability.

[0171] Optionally, the processor 610 is further configured to, when the absolute difference of the cumulative frequency energy distribution is greater than a first threshold, obtain a first phase variance of the first audio signal and a second phase variance of the second audio signal; and when both the first phase variance and the second phase variance are greater than a second threshold, obtain a first frequency cumulative energy and a second frequency cumulative energy.

[0172] Optionally, the processor 610 is further configured to determine the first pickup channel as the target pickup channel when the first hole blockage probability is greater than a third threshold, wherein the first pickup channel is the pickup channel corresponding to the larger of the first frequency accumulated energy and the second frequency accumulated energy; and to perform amplitude fusion processing on the first audio signal according to the first hole blockage probability and the second hole blockage probability of the second pickup channel to obtain a target audio signal, wherein the amplitude of the second audio signal accounts for a greater proportion than the amplitude of the first audio signal in the amplitude of the target audio signal.

[0173] Optionally, the processor 610 is further configured to determine the absolute difference of the cumulative frequency energy distribution based on the first empirical cumulative distribution and the second empirical cumulative distribution; if the absolute difference of the cumulative frequency energy distribution is greater than a first threshold, obtain a first phase variance of the first audio signal and a second phase variance of the second audio signal; if the first phase variance is less than a second threshold and the second phase variance is greater than a second threshold, determine the first pickup channel as the target pickup channel.

[0174] The electronic device provided in this application embodiment can implement the various processes implemented in the foregoing method embodiments, and will not be described again here to avoid repetition.

[0175] It should be understood that, in this embodiment, the input unit 604 may include a graphics processing unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 606 may include a display panel 6061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0176] The memory 609 can be used to store software programs and various data. The memory 609 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 609 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0177] Processor 610 may include one or more processing units; optionally, processor 610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 610.

[0178] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio signal processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0179] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0180] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio signal processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0181] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0182] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio signal processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0183] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0184] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0185] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio signal processing method applied to an electronic device, the electronic device comprising a first pickup channel and a second pickup channel, characterized in that, The method comprises: obtaining a first audio signal collected through the first sound pickup channel and a second audio signal collected through the second sound pickup channel; determining a first empirical cumulative distribution of frequency energy according to the first audio signal; determining a second empirical cumulative distribution of frequency energy according to the second audio signal; determining a frequency energy cumulative distribution mean absolute difference according to the first empirical cumulative distribution and the second empirical cumulative distribution; determining a target sound pickup channel according to the frequency energy cumulative distribution mean absolute difference, the target sound pickup channel being the blocked sound pickup channel of the first sound pickup channel and the second sound pickup channel; performing fusion processing on the audio signal of the target sound pickup channel to obtain a target audio signal.

2. The method of claim 1, wherein, The method comprises: determining a first power spectrum according to the first audio signal; determining the first empirical cumulative distribution of frequency energy according to the first power spectrum; The method comprises: determining a second power spectrum according to the second audio signal; determining the second empirical cumulative distribution of frequency energy according to the second power spectrum; The first frequency energy is determined according to the frequency power in the first power spectrum, the second frequency energy is determined according to the power in the second power spectrum, the first empirical cumulative distribution and the second empirical cumulative distribution have the same minimum frequency energy and the same maximum frequency energy.

3. The method of claim 2, wherein, The method comprises: in the case that the frequency energy cumulative distribution mean absolute difference is greater than a first threshold, obtaining a first frequency cumulative energy and a second frequency cumulative energy, the first frequency cumulative energy being the maximum frequency energy cumulative value in the first empirical cumulative distribution, and the second frequency cumulative energy being the maximum frequency energy cumulative value in the second empirical cumulative distribution; determining a first hole blocking probability of the sound pickup channel corresponding to the greater one of the first frequency cumulative energy and the second frequency cumulative energy, the first hole blocking probability being determined according to the frequency energy cumulative distribution mean absolute difference; determining the target sound pickup channel according to the first hole blocking probability.

4. The method of claim 3, wherein, In the case that the frequency energy cumulative distribution mean absolute difference is greater than a first threshold, obtaining a first frequency cumulative energy and a second frequency cumulative energy, comprises: in the case that the frequency energy cumulative distribution mean absolute difference is greater than a first threshold, obtaining a first phase variance of the first audio signal and a second phase variance of the second audio signal; in the case that the first phase variance and the second phase variance are both greater than a second threshold, obtaining a first frequency cumulative energy and a second frequency cumulative energy.

5. The method of claim 3, wherein, The method comprises: determining that the first pickup channel is the target pickup channel in a case where the first hole blocking probability is greater than a third threshold value, the first pickup channel being a pickup channel corresponding to a greater one of the first frequency cumulative energy and the second frequency cumulative energy; the fusion processing of the audio signal of the target pickup channel to obtain a target audio signal, comprising: performing amplitude fusion processing on the first audio signal according to the first hole blocking probability and a second hole blocking probability of the second pickup channel to obtain a target audio signal, the proportion of the amplitude of the second audio signal in the amplitude of the target audio signal being greater than the proportion of the amplitude of the first audio signal.

6. The method of claim 1, wherein, the determination of the target pickup channel according to the frequency energy cumulative distribution mean absolute difference, comprising: obtaining a first phase variance of the first audio signal and a second phase variance of the second audio signal in a case where the frequency energy cumulative distribution mean absolute difference is greater than a first threshold value; determining that the first pickup channel is the target pickup channel in a case where the first phase variance is less than a second threshold value and the second phase variance is greater than the second threshold value.

7. An audio signal processing apparatus applied to an electronic device, the electronic device comprising a first pickup channel and a second pickup channel, characterized in that, comprising: a first acquisition module configured to acquire a first audio signal collected through the first pickup channel and a second audio signal collected through the second pickup channel; a first determination module configured to determine a first empirical cumulative distribution of first frequency energy according to the first audio signal; a second determination module configured to determine a second empirical cumulative distribution of second frequency energy according to the second audio signal; a third determination module configured to determine a frequency energy cumulative distribution mean absolute difference according to the first empirical cumulative distribution and the second empirical cumulative distribution, and determine a target pickup channel according to the frequency energy cumulative distribution mean absolute difference, the target pickup channel being a blocked pickup channel among the first pickup channel and the second pickup channel; a fusion module configured to perform fusion processing on an audio signal of the target pickup channel to obtain a target audio signal.

8. The apparatus of claim 7, wherein, the first determination module, comprising: a first determination submodule configured to determine a first power spectrum according to the first audio signal; a second determination submodule configured to determine the first empirical cumulative distribution of the first frequency energy according to the first power spectrum; the second determination module, comprising: a third determination submodule configured to determine a second power spectrum according to the second audio signal; a fourth determination submodule configured to determine the second empirical cumulative distribution of the second frequency energy according to the second power spectrum; wherein the first frequency energy is determined according to a frequency power in the first power spectrum, the second frequency energy is determined according to a power in the second power spectrum, the first empirical cumulative distribution and the second empirical cumulative distribution have the same minimum frequency energy and the same maximum frequency energy.

9. An electronic device, comprising: comprising a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the steps of the audio signal processing method according to any one of claims 1 to 6.

10. A readable storage medium, characterized by, The program or instruction is stored on the readable storage medium, and when executed by the processor, the program or instruction implements the steps of the audio signal processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech pickup method and related products

    CN109005272A

  • Terminal microphone test method and device, mobile terminal and storage medium

    CN113596700A