Audio signal compensation method and device, electronic equipment and readable storage medium
By determining the compensation mask and audio signal using the real-time power and phase information of the microphone array's audio signal, the problem of poor sound pickup quality of electronic devices when the target sound source moves is solved, the strength and spectral adaptability of the audio signal are improved, and the sound pickup effect is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2023-07-03
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies suffer from poor sound pickup quality of electronic devices when the target sound source moves, mainly due to spectral damage and reduced signal-to-noise ratio caused by tracking delay of the covariance matrix.
By using real-time power and phase information of audio signals acquired from a microphone array, a compensation mask and compensation audio signal are determined, and audio signal compensation processing is performed to improve the intensity and spectral adaptability of the audio signal.
It effectively reduces the tracking delay caused by spatial filtering, improves the sound pickup quality and signal-to-noise ratio of electronic devices, and enhances the overall quality of audio signals.
Smart Images

Figure CN116847246B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio technology, specifically relating to an audio signal compensation method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] Currently, when electronic devices use at least two microphones to pick up sound, spatial filtering algorithms are typically employed to process the acquired audio signals in order to capture target sound sources from specific directions and reduce noise interference. In this spatial filtering algorithm, a forgetting factor, used to control the proportion of historical and current values, is used to update the covariance matrix of the current frame.
[0003] However, updating the covariance matrix using the aforementioned forgetting factor will cause a tracking delay in the covariance matrix over a period of time. This tracking delay will cause spectral damage to the target sound source picked up during that period, resulting in poor sound pickup quality of the electronic device when the target sound source moves. Summary of the Invention
[0004] The purpose of this application is to provide an audio signal compensation method, apparatus, electronic device, and readable storage medium, which can solve the problem that the sound pickup quality of electronic devices is poor when the target sound source moves.
[0005] In a first aspect, embodiments of this application provide an audio signal compensation method, the method comprising: determining a compensation mask based on a first real-time power of a first audio signal and a second real-time power of a second audio signal, wherein the first audio signal includes audio signals collected by at least two microphones and the second audio signal is an audio signal obtained by spatial filtering the first audio signal; determining a compensation audio signal based on the average amplitude of the first audio signal and the phase of a third audio signal in the first audio signal, wherein the third audio signal is an audio signal collected by any one of the at least two microphones; and performing compensation processing on the second audio signal using the compensation mask and the compensation audio signal to obtain a compensated second audio signal.
[0006] Secondly, embodiments of this application provide an audio signal compensation device, which includes a determining module and a processing module. The determining module is used to determine a compensation mask based on a first real-time power of a first audio signal and a second real-time power of a second audio signal. The first audio signal includes audio signals collected by at least two microphones, and the second audio signal is an audio signal obtained by spatial filtering the first audio signal. The determining module is also used to determine a compensation audio signal based on the average amplitude of the first audio signal and the phase of a third audio signal in the first audio signal. The third audio signal is an audio signal collected by any one of the at least two microphones. The processing module is used to perform compensation processing on the second audio signal using the compensation mask and the compensation audio signal to obtain a compensated second audio signal.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0011] In this embodiment, a compensation mask can be determined based on the first real-time power of a first audio signal and the second real-time power of a second audio signal. The first audio signal includes audio signals collected by at least two microphones, and the second audio signal is an audio signal obtained after spatial filtering of the first audio signal. A compensation audio signal is determined based on the average amplitude of the first audio signal and the phase of a third audio signal within the first audio signal. The third audio signal is an audio signal collected by any one of the at least two microphones. The compensation mask and the compensation audio signal are then used to compensate the second audio signal, resulting in a compensated second audio signal. Through this scheme, the electronic device can first determine the compensation mask based on the real-time power of the audio signals collected by at least two microphones before and after spatial filtering; then, based on the characteristics of the audio signal, determine the compensation audio signal; and use the compensation mask and the compensation audio signal to compensate the spatially filtered audio signal. This allows for adjustment of the mixing ratio of the compensation audio signal through the compensation mask, thereby enhancing the intensity of the compensated audio signal and adaptively restoring its spectrum. This compensation process can improve the quality of the picked-up audio signal, reduce the impact of tracking delay caused by drop filtering, and thus improve the sound pickup quality of electronic devices. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of a microphone array picking up audio signals from a fixed sound source in traditional technology;
[0013] Figure 2 This is a schematic diagram of a microphone array picking up audio signals from a moving sound source in traditional technology;
[0014] Figure 3 This is one of the flowcharts of the audio signal compensation method provided in the embodiments of this application;
[0015] Figure 4 This is the second flowchart of the audio signal compensation method provided in the embodiments of this application;
[0016] Figure 5 This is the third flowchart of the audio signal compensation method provided in the embodiments of this application;
[0017] Figure 6 This is the fourth flowchart of the audio signal compensation method provided in the embodiments of this application;
[0018] Figure 7 This is the fifth flowchart of the audio signal compensation method provided in the embodiments of this application;
[0019] Figure 8 This is a spectrum diagram of the audio signal picked up by the dual-microphone array in the audio signal compensation method provided in the embodiments of this application;
[0020] Figure 9 This is a schematic diagram of the audio signal compensation method provided in the embodiments of this application;
[0021] Figure 10 This is a spectrum diagram of the microphone pickup enhancement signal in the audio signal compensation method provided in the embodiments of this application;
[0022] Figure 11 This is a schematic diagram of the audio signal compensation device provided in the embodiments of this application;
[0023] Figure 12 This is a schematic diagram of the electronic device provided in the embodiments of this application;
[0024] Figure 13 This is a hardware schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] The following section will first explain some of the terms or terms used in the specification and claims of this application.
[0028] Spatial filtering algorithms, also known as beamforming algorithms, are a core component of smart antenna research. They adjust antenna gain based on the different transmission paths of the user signal in space, applying appropriate gain to paths with better transmission quality to form a narrow beam aimed at the user signal. Conversely, they minimize sidelobes along paths with poorer transmission quality, employing directional reception to improve system capacity. This algorithm can also combine signals from multiple microphones (often omnidirectional) to suppress signals from non-target directions and enhance signals from the target direction. This allows for focused sound pickup in a specific direction, effectively improving the signal-to-noise ratio of the received signal and also contributing to noise reduction.
[0029] The basic principle of spatial filtering algorithms is wave interference. By adjusting the parameters between signals from different array units, signals at certain angles are enhanced, while signals at other angles cancel each other out. As the technology and performance of electronic devices become more mature and their costs decrease, the application of this algorithm is becoming more widespread, including but not limited to the following devices and application scenarios: call noise reduction in headphones / mobile phones, digital hearing aids, sound source localization, smart speakers, conference room microphones, phased array radar / antennas, vehicle-mounted sound pickup, reverberation removal, and far-field sound pickup.
[0030] Short Time Fourier Transform (STFT): STFT is a mathematical transform related to Fourier transform, used to determine the frequency and phase of a local sinusoidal wave in a time-varying signal.
[0031] The basic idea of STFT is to select a time-frequency localized window function g(t). Assuming that analysis shows g(t) is stationary (pseudo-stationary) over a short time interval, the window function g(t) is shifted to make f(t)g(t) a stationary signal over different finite time widths, thus calculating the power spectrum at different times. STFT uses a fixed window function; once determined, its shape remains unchanged, and the resolution of the short-time Fourier transform is also determined. To change the resolution, a new window function needs to be selected. STFT is suitable for analyzing piecewise stationary or approximately stationary signals, but for non-stationary signals, when signal changes are drastic, a high time resolution is required; while for moments with relatively gentle waveform changes, a high frequency resolution is required. STFT cannot simultaneously meet the requirements of both frequency and time resolution. The short-time Fourier transform window function is limited by the Heisenberg uncertainty principle, and the area of the time-frequency window is no less than 2. This also illustrates that the time and frequency resolutions of the short-time Fourier transform window function cannot be optimal simultaneously.
[0032] Signal-to-noise ratio (SNR): This refers to the ratio of the strength of the received useful signal to the strength of the received interference signal (including noise and interference). The SNR concept originated in multi-user detection. Assume there are two users, User 1 and User 2, and their transmitting antennas send two signals (code orthogonality is used in Code Division Multiple Access, and spectral orthogonality is used in Orthogonal Frequency Division Multiplexing, thus distinguishing the different data sent to the two users). User 1's receiver can receive both the data sent to User 1 (the useful signal) and the data sent to User 2 (the interference signal). Similarly, User 2's receiver can receive both the data sent to User 1 (the interference signal) and the data sent to User 2 (the useful signal).
[0033] The 3rd Generation Partnership Project (3GPP) proposals include Multiple-in Multiple-out (MIMO) technology, which requires a channel quality indicator to feed back channel characteristics to the transmitter in order to adjust the data rate of the transmit antenna and achieve adaptive modulation. Ideally, the complete channel characteristics, i.e., the channel matrix H, could be estimated and fed back. However, in practical systems, especially MIMO systems, accurately and timely estimating the channel matrix H is impractical, and the amount of feedback information is limited by the feedback channel. Therefore, most 3GPP proposals use the signal-to-noise ratio (SNR) as feedback information, serving as a control parameter for adaptive modulation.
[0034] Signal-to-noise ratio (SNR) can also serve as an important indicator for receivers, setting higher requirements for equipment sensitivity and anti-interference capabilities. Code division multiple access (CDMA) systems are interference-limited systems, but multi-user interference has a significant impact, necessitating SNR consideration in the design. This is because the spreading codes in CDMA systems are not perfectly orthogonal and possess a certain correlation value. When multiple user terminals are located close to each other, interference between terminals can be substantial. Furthermore, since CDMA base stations use the same frequencies, interference can also exist between different base stations.
[0035] Microphone: Also known as a microphone, microphone, or transducer; a microphone is an energy conversion device that converts sound signals into electrical signals; its working principle is to transmit sound vibrations to a polymer diaphragm with permanent charge isolation, which drives an internal magnet to form a changing current, and then transmits the changing current to the subsequent sound processing circuit for amplification.
[0036] The audio signal compensation method, apparatus, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0037] Currently, mobile terminal hardware devices (such as smartphones, tablets, and headphones) typically include microphone arrays composed of a certain number of microphones for sampling and filtering the spatial characteristics of the sound field. Microphone arrays can be used for specific tasks such as sound source localization (including angle and distance measurement), suppression of background noise, interference, reverberation, echo, signal extraction, and signal separation. For example, taking the microphone array in a smart speaker as an example, when the user is far away from the smart speaker, the smart speaker can use its microphone array to localize the sound source and suppress background noise, thereby accurately picking up the user's voice commands.
[0038] Commonly used microphone arrays can be categorized by their microphone layout shape into: linear arrays (such as two-microphone or three-microphone arrays), circular arrays, and rectangular arrays. Linear arrays can achieve 180-degree planar sound pickup but cannot distinguish between horizontal and vertical angles; circular arrays can achieve 360-degree planar sound pickup and can distinguish between horizontal and vertical angles; rectangular arrays can only distinguish between horizontal and vertical angles. The geometry of the microphone array is known by design, and all microphones in the array have the same frequency response and synchronized sampling clocks.
[0039] In scenarios such as calls, recordings, VoIP, and voice chat in games, microphone arrays can be combined with spatial filtering algorithms to pick up target sound sources from specific directions and attenuate interference signals (including interference sources from specific directions or environmental noise without a clear direction), thereby improving voice / audio quality.
[0040] Spatial filtering algorithms can be classified into many types according to different criteria.
[0041] Based on different objects of action, they can be divided into the following three categories:
[0042] 1. Adaptive Algorithm Based on Direction Estimation. This algorithm can be divided into two cases. In the first case, the direction of the reference user signal is known. In this case, adaptive weights can be calculated according to different criteria, such as the linearly constrained minimum variance criterion, the maximum likelihood criterion, or the maximum signal-to-noise ratio criterion. In the second case, the direction of the reference user signal is unknown. In this case, the direction of arrival (DOA) of the signal can be estimated using methods such as multi-signal classification and rotation-invariant signal parameter estimation. Although this algorithm is relatively convenient in analysis, it suffers from high computational complexity and high sensitivity to errors.
[0043] 2. Algorithms based on training or reference signals. This algorithm does not require estimation of the signal's direction of arrival and has few restrictions on the antenna's structure. However, the problem with this algorithm is that transmitting the training signal requires the recovery of the a priori carrier and symbols, which is difficult in situations with co-channel interference. Furthermore, transmitting the training signal reduces the utilization of the spectrum.
[0044] 3. Beamforming algorithm based on signal structure. This algorithm uses the time-domain characteristics of the signal to calculate the weights, and utilizes constant mode properties, wired code, cyclostationary properties, and higher-order statistics, etc. It is relatively robust to errors and does not require signal direction information. However, the problem with this algorithm is that the convergence speed is relatively slow.
[0045] II. Based on whether or not a reference signal needs to be transmitted, they can be divided into the following two categories:
[0046] 1. Non-blind algorithms. These algorithms determine the channel response by transmitting training sequences or pilot signals, and then adjust the weights according to certain criteria. Commonly used non-blind algorithms include the least mean square algorithm, the direct matrix inversion algorithm, and the recursive least squares algorithm.
[0047] 2. Blind Algorithms. These algorithms do not require transmitting training sequences or pilot signals. The receiver uses its own transmitted signal as a reference signal for estimation and then adjusts the weights. Typical blind algorithms include two types: one utilizes signal characteristics, such as constant modulus algorithms (which utilize the constant modulus property of the signal), periodic stationary algorithms (which utilize the cyclostationarity of the signal), or finite symbol set algorithms; the other utilizes DOA (Difference of Arrival), such as the MUSIC algorithm or the ESPRIT algorithm.
[0048] II. Based on different implementation methods, they can be divided into the following two categories:
[0049] 1. Analogous Beam Forming (ABF) Algorithm. The core idea of this algorithm is to down-convert the analog RF received signal to an intermediate frequency (IF) via an RF front-end, calculate weighting coefficients using a weight update algorithm, sum the weighted analog IF signals, and then convert them to a digital IF signal via an analog-to-digital converter for further processing. This algorithm has a relatively complex circuit and lower accuracy.
[0050] 2. Digital Beam Forming (DBF) Algorithm. The core idea of this algorithm is to perform beamforming in the digital domain. The received RF signal is down-converted to an intermediate frequency (IF) by the RF front-end. The analog IF signal is then converted to a digital signal by an analog-to-digital converter (ADC). Weighting coefficients are calculated using a weight update algorithm, and the digital IF signals are then summed using weighted averages. This algorithm offers good flexibility and supports parallel tracking of multiple targets.
[0051] It should be noted that the effectiveness and performance of spatial filtering algorithms are greatly influenced by the configuration of the microphone array; different microphone array configurations will result in different effects and performance of spatial filtering algorithms.
[0052] Figure 1 This diagram illustrates a conventional technique for a microphone array to pick up audio signals from a fixed sound source, such as... Figure 1 As shown, the microphone array 11 consists of multiple microphones 12. The microphone array 11 needs to collect the target audio signal 13. It can be seen that when the microphone array 11 collects the audio signal, it will also collect the noise audio signal 14. Therefore, the audio signal collected by the microphone array 11 can be processed by the spatial filtering algorithm to form a directional beam in space, so as to ensure that the target audio signal 13 is picked up while the interference of the noise audio signal 14 is attenuated.
[0053] For example, assuming the number of the aforementioned microphones 12 is m (m is an integer greater than or equal to 2), the frequency domain component (i.e., frequency domain index) is f, and the time axis component (i.e., frame number) is k; then the specific process of processing the audio signal acquired by the aforementioned microphone array 11 using the aforementioned spatial filtering algorithm is as follows (1-5):
[0054] 1. Perform STFT on the audio signal acquired by the microphone array 11 to obtain the spectral components X(f,k)=[x1(f,k),x2(f,k),…,x m [(f,k)];
[0055] 2. By utilizing the time-frequency statistical characteristics of the target audio signal 13 and the noise audio signal 14, activation detection of the target audio signal 13 is performed to obtain vad(k) information.
[0056] 3. Calculate the covariance matrix φ of the above audio signal. xx (f,k)=X(f,k)X H (f,k), and the covariance matrix φ of the target audio signal 13 is estimated by weighting using the above vad(k) information. ss (f,k), and the covariance matrix φ of the aforementioned noise audio signal 14. nn (f,k), and using a forgetting factor weighted by α, an exponential average is performed on the historical values of the covariance matrix; where φ ss (f,k) can be obtained by the following formula (1), φ nn (f,k) can be obtained by the following formula (2):
[0057] φ ss (f,k)=α·φ ss (f,k-1)+(1-α)·vad(k·φxx (f,k); (1)
[0058] φ nn (f,k)=α·φ nn (f,k-1)+(1-α)·(1-vad(k))·φ xx (f,k); (2)
[0059] 4. Based on the optimization objective of maximizing the signal-to-noise ratio, the optimal filter coefficients W(f,k)=[w1(f,k),w2(f,k),…,w m [f,k] is the generalized eigenvector that satisfies the constraints; where W(f,k) can be obtained through the following formulas (3) and (4):
[0060]
[0061] φ ss ·w i (f,k)=λ i ·φ nn ·w i (f,k){i∈[1,m]}; (4)
[0062] 5. The enhanced audio signal S(f,k) is obtained by weighting and summing the filter coefficients, as shown in the following formula (5):
[0063]
[0064] However, in the spatial filtering algorithm described above, when updating the covariance matrix of the target audio signal 13 and the covariance matrix of the noisy audio signal 14, a forgetting factor α is needed to control the proportion of historical values and current values. If the proportion of current values is too large, the filter may diverge due to numerical jitter; conversely, if the proportion of historical values is too large, the filter values will be stable, but a tracking delay problem will be introduced.
[0065] When the positions of the target audio signal 13 and the noise audio signal 14 are fixed, the tracking delay can be ignored; however, in the usage scenarios of mobile devices such as smartphones or smart tablets, the positions of the acquired audio signals are usually not fixed but frequently move, as shown in the following specific scenarios:
[0066] Scenario 1: A user is making a hands-free call while walking. Due to the swinging of the arm during walking, the user's voice (i.e., the target audio signal 13 mentioned above) is moving rapidly.
[0067] Scenario 2: When a user is holding a phone, the user's voice (i.e., the target audio signal 13 mentioned above) will move rapidly when the user's hand position changes or the head turns.
[0068] Scenario 3: When a user is making a call outdoors at a bus stop or subway station, the ambient noise (i.e., the noise audio signal 14 mentioned above) will move rapidly as vehicles pass by.
[0069] Figure 2 This diagram illustrates a conventional technique for a microphone array to pick up audio signals from a moving sound source, such as... Figure 2 As shown, when the target audio signal 13 or the noise audio signal 14 moves, especially when the moving speed is greater than the tracking speed, a tracking delay of the covariance matrix will inevitably occur, resulting in poor sound pickup quality.
[0070] For example, when the target audio signal 13 moves rapidly, it will cause tracking error in the covariance matrix of the target audio signal 13 for a period of time. In the audio signal picked up during that period of time, the target audio signal 13 will suffer from spectral damage, low volume, and intermittent sound.
[0071] For example, when the aforementioned noisy audio signal 14 moves rapidly, it will cause tracking error in the covariance matrix of the noisy audio signal 14 for a period of time. In the audio signals picked up during that period of time, the proportion of the target audio signal 13 will increase, which will reduce the signal-to-noise ratio of the target audio signal 13, resulting in the picked-up sound being very noisy and affecting the sound quality.
[0072] To address the aforementioned problems, embodiments of this application provide an audio signal compensation method, apparatus, electronic device, and readable storage medium. In this solution, a compensation mask is determined based on the first real-time power of a first audio signal collected by at least two microphones of a microphone array and the second real-time power of a second audio signal. The second audio signal is the audio signal obtained after spatial filtering using the aforementioned spatial filtering algorithm. A compensation audio signal is determined based on the average amplitude of the first audio signal and the phase of a third audio signal within the first audio signal. The third audio signal is the audio signal collected by any one of the at least two microphones. The compensation mask and the compensation audio signal are then used to compensate the second audio signal, resulting in a compensated second audio signal. Through this solution, the electronic device can first determine the compensation mask based on the real-time power of the audio signals collected by at least two microphones before and after spatial filtering; then determine the compensation audio signal based on the characteristics of the audio signal; and use the compensation mask and the compensation audio signal to compensate the spatially filtered audio signal. This allows for adjustment of the mixing ratio of the compensation audio signal through the compensation mask, thereby enhancing the intensity of the compensated audio signal and adaptively restoring its spectrum. This compensation process can improve the quality of the picked-up audio signal, reduce the impact of tracking delay caused by drop filtering, and thus improve the sound pickup quality of electronic devices.
[0073] It should be noted that the audio signal compensation method provided in this application can be executed by an audio signal compensation device, an electronic device, or a functional module within an electronic device. Some embodiments of this application use an electronic device executing the audio signal compensation method as an example to illustrate the audio signal compensation method provided in this application.
[0074] Figure 3 A flowchart of an audio signal compensation method provided in an embodiment of this application is shown. Figure 3 As shown, the audio signal compensation method provided in this application embodiment may include the following steps 301 to 303.
[0075] Step 301: The electronic device determines a compensation mask based on the first real-time power of the first audio signal and the second real-time power of the second audio signal.
[0076] The first audio signal includes audio signals collected by at least two microphones, and the second audio signal is an audio signal obtained by spatially filtering the first audio signal.
[0077] In this embodiment of the application, the compensation mask is used to indicate the degree of compensation for the second audio signal.
[0078] Optionally, in this embodiment of the application, the second audio signal may specifically be: an audio signal obtained by spatially filtering the first audio signal using the spatial filtering algorithm described above.
[0079] Optionally, in this embodiment of the application, the audio signals collected by the at least two microphones may include: the audio signal of the target sound source that the electronic device needs to collect, and the ambient noise audio signal.
[0080] Optionally, in this embodiment of the application, the above-mentioned at least two microphones may be microphones in a microphone array of an electronic device.
[0081] Optionally, in the embodiments of this application, any two of the above-mentioned at least two microphones may be the same or different.
[0082] Optionally, in the embodiments of this application, any one of the above-mentioned at least two microphones can be an electrodynamic microphone, a condenser microphone, a piezoelectric microphone, an electromagnetic microphone, a carbon microphone, or a semiconductor microphone, etc.
[0083] For example, taking any of the above microphones as an example, this microphone can be a dynamic microphone or a ribbon microphone, etc.
[0084] For example, taking any of the above microphones as a piezoelectric microphone, this microphone can be a crystal microphone or a ceramic microphone, etc.
[0085] Optionally, in this embodiment of the application, the target sound source can be any sound source that needs to be collected, such as the user's voice or the music being played.
[0086] Optionally, in this embodiment of the application, the aforementioned environmental noise audio signal may include: wind sound or bird call, etc.
[0087] Optionally, in this embodiment of the application, the first real-time power is used to indicate the energy of the first audio signal.
[0088] Optionally, in this embodiment of the application, the second real-time power is used to indicate the energy of the second audio signal.
[0089] The specific method for determining the aforementioned compensation mask in electronic devices will be explained in detail below.
[0090] Optionally, in the embodiments of this application, combined with Figure 3 ,like Figure 4 As shown, step 301 above can be implemented through step 301a below.
[0091] Step 301a: The electronic device determines the compensation mask as the quotient of the difference between the first real-time power and the second real-time power and the sum of the first real-time power and the second real-time power.
[0092] Optionally, in this embodiment of the application, before determining the compensation mask, the electronic device may first determine the average amplitude X of the first audio signal using the following formula (6). abs (f,k):
[0093]
[0094] Where m is the number of at least two microphones, f is the frequency domain component (i.e., frequency domain index), k is the time axis component (i.e., frame number), and x i (f,k) is the amplitude value of the audio signal collected by the i-th microphone among the at least two microphones at frequency f at time k.
[0095] It should be noted that the meanings of m, f, and k are the same in all formulas in the embodiments of this application. To avoid repetition, they will not be repeated in the following formulas.
[0096] Optionally, in this embodiment of the application, the electronic device determines the average amplitude X of the first audio signal. abs After (f,k), we can determine X based on this. abs (f,k), and the first real-time power PX(f,k) is determined by the following formula (7):
[0097] PX(f,k)=|X abs (f,k)| 2 (7)
[0098] Then, the electronic device can determine the second real-time power PS(f,k) based on the amplitude value S(f,k) of the second audio signal and using the following formula (8):
[0099] PS(f,k)=|S(f,k)| 2 (8)
[0100] Thus, the electronic device can determine the compensation mask R(f,k) by the quotient of the difference between the first real-time power PX(f,k) and the second real-time power PS(f,k) and the sum of the first real-time power PX(f,k) and the second real-time power PS(f,k) using the following formula (9):
[0101]
[0102] Optionally, in this embodiment of the application, if the value of PX(f,k)-PS(f,k) in the above formula (9) is negative, the electronic device can determine the above compensation mask R(f,k) as 0, and at this time the electronic device does not need to compensate the above second audio signal.
[0103] It can be seen that the range of the above compensation mask R(f,k) is (0, 1). The above formula (9) is actually a normalized representation of the difference between the above first real-time power PX(f,k) and the above second real-time power PS(f,k).
[0104] In this embodiment of the application, since the electronic device can determine the above-mentioned quotient as the above-mentioned compensation mask, that is, it can determine the compensation mask based on the energy difference of the audio signals before and after spatial filtering (i.e., the first audio signal and the second audio signal), the accuracy of determining the proportion of the second audio signal can be improved when the second audio signal is compensated by the compensation mask.
[0105] Step 302: The electronic device determines the compensation audio signal based on the average amplitude of the first audio signal and the phase of the third audio signal in the first audio signal.
[0106] The third audio signal is an audio signal collected by any one of the at least two microphones, and the compensation audio signal is used to compensate for the second audio signal.
[0107] Optionally, in the embodiments of this application, any of the above-mentioned microphones may be the microphone closest to the sound source of the first audio signal among the at least two microphones; or it may be the first microphone among the at least two microphones, etc.
[0108] The specific method for determining the aforementioned compensated audio signal in electronic devices will be explained in detail below.
[0109] Optionally, in this embodiment of the application, the electronic device may determine the above-mentioned compensation audio signal by means of method one or method two.
[0110] Method 1
[0111] Optionally, in the embodiments of this application, combined with Figure 3 ,like Figure 5 As shown, step 302 above can be implemented through step 302a below.
[0112] Step 302a: The electronic device uses the phase of the third audio signal as the phase part of Euler's formula and the average amplitude of the first audio signal as the amplitude part of Euler's formula to calculate the compensated audio signal.
[0113] It should be noted that Euler's formula is used to construct a frequency domain signal in complex form. Its input has two components, namely the amplitude and phase of the frequency domain signal (both of which are real numbers), and its output is a single component, namely the frequency domain signal in complex form.
[0114] The phase of the frequency domain signal in the complex form mentioned above is defined as follows: complex number z = x + iy, its phase is atan2(y,x), and its range is [-pi,pi]. atan2 is also called the arctangent of the four quadrants. With the phase of z obtained, the signal can be synthesized using Euler's formula: complex number z = x + jy = A*exp(j*theta), where A is the amplitude and theta is the phase angle.
[0115] For example, taking the third audio signal as the audio signal (i.e., x1(f,k), hereinafter referred to as audio signal 1) acquired by the first microphone among the at least two microphones, the electronic device, after obtaining the phase and the average amplitude of the audio signal 1, can determine the compensation audio signal S by the following formula (10). c (f,k):
[0116] S c (f,k)=X abs (f,k)·exp(j·∠x1(f,k)); (10)
[0117] Among them, X abs (f,k) is the average amplitude of the first audio signal mentioned above, exp(j·∠x1(f,k) is the phase of audio signal 1, and j is the imaginary unit.
[0118] It can be seen that the above-mentioned compensated audio signal S c (f,k) is a complex value at time k and frequency f, where the amplitude part is the average amplitude of the first audio signal and the phase part is the phase of the audio signal 1.
[0119] In this embodiment, since the electronic device can use the phase of the third audio signal as the phase part of Euler's formula and the average amplitude of the first audio signal as the amplitude part of Euler's formula to calculate the compensation audio signal, the compensation audio signal can be determined based on the characteristics of the audio signal before spatial filtering. Therefore, the determined compensation audio signal can compensate for the spectral damage caused by spatial filtering, thereby accurately compensating for the second audio signal.
[0120] Method 2
[0121] Optionally, in the embodiments of this application, combined with Figure 3 ,like Figure 6As shown, before step 302 above, the audio signal compensation method provided in this application embodiment may further include step A below, and step 302 above can be specifically implemented by step 302b below.
[0122] Step A: The electronic device determines the noise reduction gain based on the signal-to-noise ratio of the first audio signal.
[0123] The aforementioned noise reduction gain is used to perform noise reduction processing on the aforementioned compensated audio signal.
[0124] Optionally, in this embodiment of the application, noise reduction processing is performed on the above-mentioned compensated audio signal to improve the signal-to-noise ratio of the compensated audio signal.
[0125] Optionally, in this embodiment of the application, before determining the noise reduction gain, the electronic device may first take the diagonal mean of the covariance matrix of the environmental noise audio signal obtained in the beamforming algorithm to obtain the power spectrum mean of the environmental noise audio signal. Then, the posterior signal-to-noise ratio γ(f,k) of the first audio signal is calculated using the following formula (11):
[0126]
[0127] Wherein, PX(f,k) is the first real-time power mentioned above. This represents the power spectral mean of the aforementioned environmental noise audio signal.
[0128] Subsequently, the electronic device can employ a decision-directed approach, using the difference between the calculated posterior signal-to-noise ratio γ(f,k) and 1 as an estimate of the prior signal-to-noise ratio of the first audio signal, and smoothing it with the prior signal-to-noise ratio of the previous frame to obtain the prior signal-to-noise ratio ξ(f,k) of the first audio signal; the prior signal-to-noise ratio ξ(f,k) can be calculated using the following formula (12):
[0129] ξ(f,k)=α*ξ(f,k-1)+(1-α)*max(0,γ(f,k)-1); (12)
[0130] Where ξ(f,k-1) is the prior signal-to-noise ratio of the previous frame, γ(f,k) is the subsequent prior signal-to-noise ratio, and α is the smoothing coefficient, usually taken as 0.95.
[0131] Thus, after obtaining the prior signal-to-noise ratio of the first audio signal, the electronic device can determine the noise reduction gain G(f,k) based on the logarithmic amplitude spectrum estimation with minimum mean square error using the following formula (13):
[0132]
[0133] Where ξ(f,k) is the prior signal-to-noise ratio of the first audio signal mentioned above; γ(f,k) is the prior signal-to-noise ratio mentioned above.
[0134] Step 302b: The electronic device uses the phase of the third audio signal as the phase part of Euler's formula and the product of the noise reduction gain and the average amplitude of the first audio signal as the amplitude part of Euler's formula to calculate the compensated audio signal.
[0135] For example, taking the third audio signal as the audio signal (i.e., x2(f,k), hereinafter referred to as audio signal 2) acquired by the second microphone among the at least one microphone, the electronic device, after obtaining the noise reduction gain, the average amplitude, and the phase of the audio signal 2, can determine the compensated audio signal S by the following formula (14). c (f,k):
[0136] S c (f,k)=G(f,k)·X abs (f,k)·exp(j·∠x2(f,k)); (14)
[0137] Where G(f,k) is the noise reduction gain mentioned above, and X abs (f,k) is the average amplitude of the first audio signal, exp(j·∠x2(f,k) is the phase of the second audio signal, and j is the imaginary unit.
[0138] For a detailed description of step 302b above, please refer to the relevant description in step 302a above. To avoid repetition, it will not be repeated here.
[0139] In this embodiment, since the electronic device can first determine the noise reduction gain used to perform noise reduction processing on the above-mentioned compensated audio signal, and then use the phase of the above-mentioned third audio signal as the phase part of Euler's formula and the product of the noise reduction gain and the average amplitude of the above-mentioned first audio signal as the amplitude part of Euler's formula to calculate the above-mentioned compensated audio signal, the noise signal mixed in the compensated audio signal can be reduced, thereby improving the signal-to-noise ratio of the compensated audio signal and enhancing the compensation effect of the compensated audio signal.
[0140] Step 303: The electronic device uses a compensation mask and a compensation audio signal to perform compensation processing on the second audio signal to obtain the compensated second audio signal.
[0141] The specific method for obtaining the compensated second audio signal by the electronic device is explained in detail below.
[0142] Optionally, in the embodiments of this application, combined with Figure 3 ,like Figure 7 As shown, step 303 above can be specifically implemented through step 303a below.
[0143] Step 303a: The electronic device uses a compensation mask to perform weighted summation on the compensated audio signal and the second audio signal to obtain the compensated second audio signal.
[0144] Optionally, in this embodiment of the application, the electronic device uses the above-mentioned compensation mask to perform weighted summation processing on the above-mentioned compensation audio signal and the above-mentioned second audio signal. This can be understood as follows: the electronic device uses the compensation mask to weight the compensation audio signal, and uses the difference between 1 and the compensation mask to weight the second audio signal, and then adds the two parts to obtain the above-mentioned compensated second audio signal.
[0145] Optionally, in this embodiment of the application, the electronic device can determine the compensated second audio signal S using the following formula (15). enhance :
[0146] S enhance =R(f,k)·S c (f,k)+(1-R(f,k))·S(f,k); (15)
[0147] Where R(f,k) is the compensation mask mentioned above, S c (f,k) is the aforementioned compensated audio signal, and S(f,k) is the aforementioned second audio signal.
[0148] It can be seen that, through the above compensation mask R(f,k), the mixing ratio of the above-compensated audio signal and the above-mentioned second audio signal in the above-compensated second audio signal can be determined, so as to determine the degree of compensation of the second audio signal.
[0149] In this embodiment of the application, since the electronic device can use the above-mentioned compensation mask to perform weighted summation processing on the above-mentioned compensation audio signal and the above-mentioned second audio signal to obtain the compensated second audio signal, the mixing ratio of the compensation audio signal and the second audio signal can be adaptively adjusted through the compensation mask, so that the compensated second audio signal can accurately achieve the enhancement effect, thereby improving the sound pickup quality.
[0150] The audio signal compensation method provided in the embodiments of this application will be described exemplarily below with reference to the accompanying drawings.
[0151] For example, suppose the microphone array in the electronic device includes microphone 1 and microphone 2 (i.e., at least two microphones as described above). If the electronic device needs to pick up 15 seconds of audio signal from a user, where the user is moving rapidly between the 6th and 11th seconds, then as follows: Figure 8 As shown, between the 6th and 11th seconds, due to the tracking delay caused by the user's rapid movement, the quality of the audio signals picked up by both microphone 1 and microphone 2 is poor.
[0152] After that, as Figure 9 As shown, the electronic device can first determine a compensation mask R(f,k) based on the real-time power before and after spatial filtering by the adaptive spatial filtering module of the audio signal x1(f,k) picked up by the microphone 1 and the audio signal x2(f,k) picked up by the microphone 2 (i.e., microphone signal x(f,k), which is also the first audio signal). Then, based on the average amplitude of x(f,k) and the phase of x1(f,k) (i.e., the third audio signal), a compensation audio signal is determined to compensate for the spatially filtered x(f,k) (i.e., S(f,k), which is also the second audio signal). Afterwards, R(f,k) and the compensation audio signal are used to compensate S(f,k) to obtain the output array pickup enhancement signal (i.e., the compensated second audio signal).
[0153] Figure 10 The spectrum diagram of the microphone pickup enhancement signal described in the embodiment of this application is shown, as follows: Figure 10 As shown, compared to the audio signal 101 output by the existing audio output scheme between the 6th and 11th seconds, the audio signal 102 between the 6th and 11th seconds in the microphone pickup enhancement signal obtained by the audio signal compensation method provided in this application embodiment is significantly enhanced, thereby reducing the impact of the tracking delay and improving the quality of audio pickup.
[0154] In the audio signal compensation method provided in this application embodiment, the electronic device can first determine a compensation mask based on the real-time power of the audio signals collected by at least two microphones before and after spatial filtering; then, based on the characteristics of the audio signal, determine a compensation audio signal; and use the compensation mask and the compensation audio signal to perform compensation processing on the spatially filtered audio signal. Thus, the mixing ratio of the compensation audio signal can be adjusted through the compensation mask, so that the intensity of the audio signal obtained after compensation processing is improved and the spectrum is adaptively restored. In this way, the quality of the picked-up audio signal can be improved through this compensation processing, reducing the impact of tracking delay caused by airdrop filtering, thereby improving the sound pickup quality of the electronic device.
[0155] The audio signal compensation method provided in this application can be executed by an audio signal compensation device. This application uses an audio signal compensation device executing the audio signal compensation method as an example to illustrate the audio signal compensation device provided in this application.
[0156] like Figure 11 As shown in the figure, this application embodiment provides an audio signal compensation device 110, which may include a determination module 111 and a processing module 112.
[0157] The determining module 111 can be used to determine a compensation mask based on a first real-time power of a first audio signal and a second real-time power of a second audio signal. The first audio signal includes audio signals collected by at least two microphones, and the second audio signal is an audio signal obtained by spatially filtering the first audio signal. The determining module 111 can also be used to determine a compensation audio signal based on the average amplitude of the first audio signal and the phase of a third audio signal in the first audio signal. The third audio signal is an audio signal collected by any one of the at least two microphones. The processing module 112 can be used to perform compensation processing on the second audio signal using the compensation mask and the compensation audio signal to obtain a compensated second audio signal.
[0158] In one possible implementation, the determining module 111 can be specifically used to calculate the compensated audio signal by taking the phase of the third audio signal as the phase part of Euler's formula and taking the average amplitude of the first audio signal as the amplitude part of Euler's formula.
[0159] In one possible implementation, the determining module 111 can also be used to determine the noise reduction gain based on the signal-to-noise ratio of the first audio signal before determining the compensated audio signal based on the average amplitude of the first audio signal and the phase of the third audio signal. Specifically, the determining module 111 can be used to calculate the compensated audio signal by using the phase of the third audio signal as the phase part of Euler's formula and the product of the noise reduction gain and the average amplitude of the first audio signal as the amplitude part of Euler's formula.
[0160] In one possible implementation, the determining module 111 can be specifically used to determine the compensation mask as the quotient of the difference between the first real-time power and the second real-time power and the sum of the first real-time power and the second real-time power.
[0161] In one possible implementation, the processing module 112 can be used to perform weighted summation processing on the compensated audio signal and the second audio signal using the compensation mask to obtain the compensated second audio signal.
[0162] In the audio signal compensation device provided in this application embodiment, the device first determines a compensation mask based on the real-time power of audio signals collected by at least two microphones before and after spatial filtering; then, based on the characteristics of the audio signal, it determines a compensation audio signal; and uses the compensation mask and the compensation audio signal to perform compensation processing on the spatially filtered audio signal. Thus, the mixing ratio of the compensation audio signal can be adjusted through the compensation mask, thereby improving the intensity of the audio signal obtained after compensation processing and adaptively restoring its spectrum. This compensation processing can improve the quality of the picked-up audio signal, reducing the impact of tracking delay caused by airdrop filtering, and thus improving the quality of sound pickup.
[0163] The audio signal compensation device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0164] The audio signal compensation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0165] The audio signal compensation device provided in this application embodiment can realize the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.
[0166] like Figure 12As shown, this application embodiment also provides an electronic device 1200, including a processor 1201 and a memory 1202. The memory 1202 stores a program or instructions that can run on the processor 1201. When the program or instructions are executed by the processor 1201, they implement the various steps of the audio signal compensation method embodiment described above and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0167] It should be noted that the electronic devices in the embodiments of this application include mobile electronic devices and non-mobile electronic devices.
[0168] Figure 13 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0169] like Figure 13 As shown, the electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.
[0170] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 13 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0171] The processor 1010 can be used to determine a compensation mask based on a first real-time power of a first audio signal and a second real-time power of a second audio signal. The first audio signal includes audio signals collected by at least two microphones, and the second audio signal is an audio signal obtained by spatially filtering the first audio signal. The processor 1010 can also be used to determine a compensation audio signal based on the average amplitude of the first audio signal and the phase of a third audio signal in the first audio signal. The third audio signal is an audio signal collected by any one of the at least two microphones. The processor 1010 can also be used to perform compensation processing on the second audio signal using the compensation mask and the compensation audio signal to obtain a compensated second audio signal.
[0172] In one possible implementation, the processor 1010 can be specifically used to calculate the compensated audio signal by taking the phase of the third audio signal as the phase part of Euler's formula and taking the average amplitude of the first audio signal as the amplitude part of Euler's formula.
[0173] In one possible implementation, the processor 1010 can further be used to determine the noise reduction gain based on the signal-to-noise ratio of the first audio signal before determining the compensated audio signal based on the average amplitude of the first audio signal and the phase of the third audio signal. Specifically, the processor 1010 can be used to calculate the compensated audio signal by using the phase of the third audio signal as the phase part of Euler's formula and the product of the noise reduction gain and the average amplitude of the first audio signal as the amplitude part of Euler's formula.
[0174] In one possible implementation, the processor 1010 can specifically be used to determine the compensation mask as the quotient of the difference between the first real-time power and the second real-time power and the sum of the first real-time power and the second real-time power.
[0175] In one possible implementation, the processor 1010 can be used to perform a weighted summation of the compensated audio signal and the second audio signal using the aforementioned compensation mask to obtain the compensated second audio signal.
[0176] In the electronic device provided in this application embodiment, the electronic device can first determine a compensation mask based on the real-time power of audio signals collected by at least two microphones before and after spatial filtering; then, based on the characteristics of the audio signal, determine a compensation audio signal; and use the compensation mask and the compensation audio signal to perform compensation processing on the spatially filtered audio signal. Thus, the mixing ratio of the compensation audio signal can be adjusted through the compensation mask, so that the intensity of the audio signal obtained after compensation processing is improved and the spectrum is adaptively restored. In this way, the quality of the picked-up audio signal can be improved through compensation processing to reduce the impact of tracking delay caused by airdrop filtering, thereby improving the sound pickup quality of the electronic device.
[0177] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0178] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0179] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.
[0180] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio signal compensation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0181] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0182] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio signal compensation method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0183] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0184] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio signal compensation method embodiment described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0185] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0186] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0187] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An audio signal compensation method, characterized in that, The method includes: A compensation mask is determined based on the first real-time power of the first audio signal and the second real-time power of the second audio signal. The first audio signal includes audio signals collected by at least two microphones, and the second audio signal is an audio signal obtained by spatial filtering the first audio signal. Based on the average amplitude of the first audio signal and the phase of the third audio signal in the first audio signal, a compensation audio signal is determined, wherein the third audio signal is an audio signal collected by any one of the at least two microphones. The second audio signal is compensated using the compensation mask and the compensation audio signal to obtain the compensated second audio signal.
2. The method according to claim 1, characterized in that, Determining the compensated audio signal based on the average amplitude of the first audio signal and the phase of the third audio signal in the first audio signal includes: The compensated audio signal is calculated by taking the phase of the third audio signal as the phase part of Euler's formula and taking the average amplitude of the first audio signal as the amplitude part of Euler's formula.
3. The method according to claim 1, characterized in that, Before determining the compensated audio signal based on the average amplitude of the first audio signal and the phase of the third audio signal in the first audio signal, the method further includes: The noise reduction gain is determined based on the signal-to-noise ratio of the first audio signal; Determining the compensated audio signal based on the average amplitude of the first audio signal and the phase of the third audio signal in the first audio signal includes: The compensated audio signal is calculated by taking the phase of the third audio signal as the phase part of Euler's formula and taking the product of the noise reduction gain and the average amplitude of the first audio signal as the amplitude part of Euler's formula.
4. The method according to claim 1, characterized in that, Determining the compensation mask based on the first real-time power of the first audio signal and the second real-time power of the second audio signal includes: The quotient of the difference between the first real-time power and the second real-time power and the sum of the first real-time power and the second real-time power is determined as the compensation mask.
5. The method according to any one of claims 1 to 4, characterized in that, The step of using the compensation mask and the compensation audio signal to perform compensation processing on the second audio signal to obtain the compensated second audio signal includes: Using the compensation mask, the compensated audio signal and the second audio signal are weighted and summed to obtain the compensated second audio signal.
6. An audio signal compensation device, characterized in that, The device includes a determining module and a processing module; The determining module is used to determine a compensation mask based on the first real-time power of the first audio signal and the second real-time power of the second audio signal. The first audio signal includes audio signals collected by at least two microphones, and the second audio signal is an audio signal obtained by spatial filtering the first audio signal. The determining module is further configured to determine a compensation audio signal based on the average amplitude of the first audio signal and the phase of the third audio signal in the first audio signal, wherein the third audio signal is an audio signal collected by any one of the at least two microphones. The processing module is used to perform compensation processing on the second audio signal using the compensation mask and the compensation audio signal to obtain the compensated second audio signal.
7. The apparatus according to claim 6, characterized in that, The determining module is specifically used to calculate the compensated audio signal by taking the phase of the third audio signal as the phase part of Euler's formula and the average amplitude of the first audio signal as the amplitude part of Euler's formula.
8. The apparatus according to claim 6, characterized in that, The determining module is further configured to determine the noise reduction gain based on the signal-to-noise ratio of the first audio signal before determining the compensated audio signal based on the average amplitude and the phase of the third audio signal; The determining module is specifically used to calculate the compensated audio signal by taking the phase of the third audio signal as the phase part of Euler's formula and the product of the noise reduction gain and the average amplitude of the first audio signal as the amplitude part of Euler's formula.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the audio signal compensation method as described in any one of claims 1-5.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio signal compensation method as described in any one of claims 1-5.