Audio processing method, device, equipment, medium and program product
By acquiring the fundamental frequency and harmonic energy of the target audio frame, constructing energy compensation parameters, and repairing the damaged audio frame after wind noise suppression, the problem of wind noise suppression mistakenly suppressing speech components is solved, and the clarity and intelligibility of the audio signal are improved.
Patent Information
- Application Number
- CN202410720761.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-12-05
AI Technical Summary
In outdoor audio recording scenarios, wind noise interference masks the low-frequency components of the speech signal. Existing deep learning networks mistakenly suppress speech components when suppressing wind noise, resulting in a decrease in audio quality and a reduction in the clarity and intelligibility of the speech signal.
By acquiring the fundamental frequency and harmonic energy of the target audio frame, determining the peak and valley energy parameters, and constructing energy compensation parameters using preset mapping information, the audio frame damaged after wind noise suppression is repaired, restoring the clarity and intelligibility of the speech signal.
It effectively repairs the energy loss of audio frames after wind noise suppression, improves the clarity and intelligibility of the final output sound source signal, and enhances audio quality.
Smart Images

Figure CN121075352A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of artificial intelligence, and particularly relates to an audio processing method and device, equipment, medium and program product. BACKGROUND
[0002] With the continuous development and innovation of technology, mobile terminal devices are more and more widely applied in people's daily life, and people's requirements for audio quality are also getting higher and higher. In an outdoor sound receiving scene, such as outdoor audio and video call, live broadcast, recording, voice message recording, etc., the microphone will capture the wind noise and the sound of the terminal object. Since the wind noise energy is usually large, the terminal object sound accounts for a low proportion in the sound collected by the microphone, and the wind noise signal accounts for a high proportion, so the obtained audio quality is poor, such as the audio receiving object cannot hear the audio content. Therefore, it is necessary to suppress the wind noise signal in the audio.
[0003] However, since the wind noise interference component in the noisy signal is much larger than the speech component, when the wind noise component in the audio is suppressed, some frequency band speech signals with extremely low signal-to-noise ratio are often mis-suppressed, so that the low-frequency speech component in the wind noise suppressed audio signal is severely damaged. On the hearing, only the high-frequency speech component is left, so the sound hearing is relatively thin and the volume is small, and it is not easy to hear clearly. SUMMARY
[0004] To solve the above technical problems, the present application provides an audio processing method, device, equipment, medium and program product.
[0005] In one aspect, the present application provides an audio processing method, which comprises:
[0006] obtaining a target audio frame and a target fundamental frequency of the target audio frame; the target audio frame is obtained after wind noise suppression processing;
[0007] determining a target peak-valley energy parameter according to the harmonic energy of at least two target post-harmonics corresponding to the target fundamental frequency; the target peak-valley energy parameter is determined according to the ratio of the target energy peak value to the target energy valley value in the harmonic energy of the at least two target post-harmonics;
[0008] determining an energy compensation parameter corresponding to the target fundamental frequency and the target peak-valley energy parameter based on preset mapping information; the preset mapping information is used to represent the mapping relationship between the frequency interval where the fundamental frequency is located and the peak-valley energy parameter interval where the peak-valley energy parameter is located and the energy compensation parameter; the energy compensation parameter is obtained based on statistical processing of the harmonic energy of the harmonics in a normal audio frame; the normal audio frame is an audio frame not containing wind noise signal;
[0009] The modified audio frame corresponding to the target audio frame is determined according to the energy compensation parameter and the harmonic energy of the at least two target post-harmonics.
[0010] In another aspect, the embodiments of the present application further provide an audio processing device, which comprises:
[0011] The target audio frame is obtained through wind noise suppression processing.
[0012] The target peak-valley energy parameter is determined according to the ratio of the target energy peak value to the target energy valley value in the harmonic energy of the at least two target post-harmonics.
[0013] The energy compensation parameter corresponding to the target fundamental frequency and the target peak-valley energy parameter is determined based on preset mapping information. The energy compensation parameter is obtained based on statistical processing of the harmonic energy of harmonics in a normal audio frame.
[0014] The modified audio frame corresponding to the target audio frame is determined according to the energy compensation parameter and the harmonic energy of the at least two target post-harmonics.
[0015] In another aspect, the embodiments of the present application further provide an electronic device for audio processing, which comprises a processor and a memory.
[0016] In another aspect, the embodiments of the present application further provide a computer readable storage medium.
[0017] In another aspect, the embodiments of the present application further provide a computer program product.
[0018] The audio processing method, device, equipment, medium and program product provided in the embodiments of the present application determine a target peak-valley energy parameter according to the harmonic energy of at least two target later-stage harmonics corresponding to a target fundamental frequency, then determine an energy compensation parameter corresponding to the target fundamental frequency and the target peak-valley energy parameter based on preset mapping information, and determine a corrected audio frame corresponding to the target audio frame according to the energy compensation parameter and the harmonic energy of the at least two target later-stage harmonics. Since the energy compensation parameter is obtained based on the statistical processing of the harmonic energy of the harmonics in a normal audio frame, the energy compensation parameter can accurately reflect the harmonic energy of the harmonics in the normal audio frame. The energy-compensated target audio frame after wind noise suppression processing is repaired by using the energy compensation parameter, so that the repaired audio frame can be restored to a relatively normal level, thereby improving the intelligibility and clarity of the final output sound source signal. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 is a schematic diagram of an implementation environment of an audio processing method according to an example embodiment.
[0021] Figure 2 is a schematic diagram of a flow of an audio processing method according to an example embodiment Figure One .
[0022] Figure 3 is a schematic diagram of the structure of a wind noise suppression model according to an example embodiment.
[0023] Figure 4 is a schematic diagram of the energy before and after the repair of a target audio frame according to an example embodiment.
[0024] Figure 5 is a schematic diagram of a flow of an audio processing method according to an example embodiment Figure Two .
[0025] Figure 6 is a schematic diagram of a flow of an audio processing method according to an example embodiment Figure Three .
[0026] Figure 7 is a block diagram of an audio processing device according to an example embodiment.
[0027] Figure 8is a hardware structure block diagram of a server of an audio processing method according to an exemplary embodiment. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0029] It should be noted that the terms "first", "second" and the like in the specification and claims of the embodiments of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in other than the order illustrated or described herein. Therefore, the features defined as "first", "second" can explicitly or implicitly include one or more features. In the description of the embodiments, unless otherwise specified, the meaning of "a plurality of" is two or more. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0030] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, and do not limit the embodiments of the present application.
[0031] The embodiments of the present application relate to artificial intelligence (AI) and machine learning (ML) technology.
[0032] Artificial intelligence is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machine has the functions of perception, reasoning and decision-making.
[0033] Artificial intelligence is a comprehensive discipline involving a wide range of fields, both hardware and software technologies. The basic technologies of artificial intelligence generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation interaction system, mechatronics and other technologies; the software technologies of artificial intelligence generally include computer vision technology, natural language processing technology, and machine learning / deep learning and other major directions. With the development and progress of artificial intelligence, artificial intelligence is researched and applied in multiple fields, such as common smart home, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, unmanned vehicles, autonomous vehicles, robots, intelligent medical treatment and the like. It is believed that with the further development of future technologies, artificial intelligence will be applied in more fields and play an increasingly important role.
[0034] Machine learning is a multi-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance.
[0035] Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Deep learning is the core of machine learning and a technology for implementing machine learning. Machine learning generally includes deep learning, reinforcement learning, transfer learning, inductive learning and other technologies, and deep learning includes mobile visual neural network (Mobilenet), convolutional neural network (CNN), deep belief network, recurrent neural network, autoencoder, generative adversarial network and other technologies.
[0036] In the outdoor sound receiving scenario, wind noise problem is often encountered. The wind noise problem is directed to the process of voice communication, in which the wind sound is collected by the microphone and interferes with the communication. Since the microphone is sensitive, the microphone will capture the wind noise and the terminal object sound, so that the audio receiving party cannot hear clearly, hears very hard or misunderstands the communication content.
[0037] In a sound signal containing wind noise, the wind noise causes the sound signal to be disturbed, especially in the low frequency part (approximately below 1000 Hz) where the wind noise energy is relatively concentrated, the low frequency harmonic components of the normal speech signal are basically masked by the wind noise signal and cannot be recognized. The deep learning network can better suppress wind noise according to a large number of samples. After the wind noise suppression processing of the deep learning network is performed on the sound signal containing wind noise, the wind noise can be well suppressed and cleaned. However, since the wind noise interference component in the noisy signal is much larger than the speech component, the deep learning network tends to suppress the wind noise component, and the probability of mis-suppression of the speech signal in some frequency band with extremely low signal-to-noise ratio is high, thereby causing the low frequency speech component to be severely damaged, resulting in a decrease in audio quality.
[0038] Therefore, embodiments of the present application provide an audio processing method, device, equipment, medium and program product. The target peak valley energy parameter is determined according to the harmonic energy of at least two target later-stage harmonics corresponding to the target fundamental frequency. Then, the energy compensation parameter corresponding to the target fundamental frequency and the target peak valley energy parameter is determined based on the preset mapping information. The corrected audio frame corresponding to the target audio frame is determined according to the energy compensation parameter and the harmonic energy of the at least two target later-stage harmonics. Since the energy compensation parameter is obtained based on the statistical processing of the harmonic energy of the harmonics in the normal audio frame, the energy compensation parameter can accurately reflect the harmonic energy of the harmonics in the normal audio frame. The energy-damaged target audio frame after wind noise suppression processing is repaired using the energy compensation parameter, which can restore the repaired audio frame to a relatively normal level, thereby improving the intelligibility and clarity of the final output sound source signal.
[0039] Figure 1 Fig. 1 is an implementation environment schematic diagram of an audio processing method according to an example embodiment. As shown in Figure 1 Fig. 1, the implementation environment can at least include a terminal device 01.
[0040] Specifically, the terminal device 01 can include a smart phone, a desktop computer, a tablet computer, a notebook computer, a vehicle-mounted terminal, a digital assistant, a smart wearable device, a voice interaction device, and the like. It can also include software running in the device, such as web pages provided by some service providers to users, and applications provided by the service providers to users. Specifically, the terminal device 01 can have a microphone for receiving an audio signal. The terminal device 01 can perform wind noise suppression on the received audio signal. The terminal device 01 can also repair the damaged audio signal after wind noise suppression to restore the audio signal to a relatively normal level, thereby improving the intelligibility and clarity of the terminal object collected sound signal.
[0041] It should be noted that, Figure 1This is merely an example. Other implementation environments can also be included in other scenarios.
[0042] Figure 2 is a flowchart of an audio processing method according to an example embodiment Figure One . The method can be used in the implementation environment in Figure 1 . This specification provides method operation steps as described in the embodiments or flowcharts, but can include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiments is only one of the many execution orders of the steps, and does not represent the only execution order. In actual system or server product execution, the method order shown in the embodiments or drawings can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically as shown in Figure 2 , the method can include:
[0043] S101: Obtain a target audio frame and a target fundamental frequency of the target audio frame.
[0044] In the outdoor sound collection scenario, such as outdoor audio and video call, outdoor live broadcast and the like, the microphone of the terminal device captures the terminal object sound and also captures the wind noise at the same time. Since the wind noise seriously interferes with the terminal object sound, it affects the understanding of the voice content at the audio receiving end. Therefore, in order to improve the audio quality, after the terminal object obtains the audio signal, the terminal object can process the audio signal to suppress the wind noise component in the audio signal.
[0045] Specifically, the terminal device acquires an original audio signal. The original audio signal can be a signal collected by a microphone of the terminal device, or a signal acquired by the terminal device through an audio acquisition channel, such as a signal acquired through a communication application, a website download channel, and the like. The original audio signal can be a signal containing wind noise or a signal not containing wind noise. The terminal device can determine whether the original audio signal contains wind noise by processing the original audio signal. After acquiring the original audio signal, the terminal device can obtain a first preset number of original audio frames by performing frame windowing processing on the original audio signal. The first preset number is different according to the length of the original audio signal. Generally, an original audio signal with a fixed length of 20 ms can be taken as a frame of signal. For example, the length of the original audio signal is 1 min, and after frame windowing processing, 3000 original audio frames can be obtained. After obtaining the original audio frames, each original audio frame is converted from the time domain to the frequency domain, thereby obtaining a first preset number of preset audio frames. The original audio frame obtained by frame windowing processing on the original audio signal is a time domain signal, and therefore needs to be converted into a frequency domain signal for subsequent processing. The original audio frame can be converted from the time domain to the frequency domain by Fourier transform. Optionally, the Fourier transform used can be continuous Fourier transform, discrete Fourier transform, or fast Fourier transform, etc. The original audio frame in the time domain is converted into a preset audio frame in the frequency domain by Fourier transform. After obtaining the preset audio frame, a time-frequency domain feature corresponding to each preset audio frame is obtained by performing feature extraction on each preset audio frame. The time-frequency domain feature corresponding to the preset audio frame can include a spectral energy feature, a pitch period, etc. The spectral energy feature refers to the energy distribution of different frequency components of a signal in the frequency domain, including a frequency point energy and a spectral energy, etc. The spectral energy feature can be obtained by a power spectral density (PSD) of the signal. The power spectral density describes the distribution of each frequency component of the signal relative to the total energy. The power spectral density can be expressed as the square of the modulus of the complex spectrum divided by the frequency resolution (frequency interval). After obtaining the power spectral density, the frequency point energy corresponding to each frequency point can be obtained by integrating the power spectral density at each frequency point. In addition, the cumulative energy, i.e., the spectral energy of the entire spectrum, can be obtained by integrating the power spectral density in the frequency band of the signal. The spectral energy feature corresponding to the preset audio frame can be obtained by performing feature extraction on the preset audio frame in the above manner. The pitch period refers to the period of the strongest periodic component in a signal. In speech signal processing, the pitch period is an important parameter that describes the frequency of the periodic signal generated by the vibration of the vocal cords. When pronouncing, whether the vocal cords vibrate can divide the speech signal into two types: clear and dull. The dull sound shows obvious periodicity in the time domain, which is caused by the vibration of the vocal cords due to the airflow through the glottis.The frequency of such vocal cord vibration is called the fundamental frequency, and the corresponding period is called the fundamental period. The fundamental period can be estimated using autocorrelation method, parallel processing method, average amplitude difference method, data reduction method, etc., so the fundamental period of the preset audio frame can be calculated using these methods. In some embodiments, an artificial intelligence model can also be used to extract spectral energy features, fundamental period and other time-frequency domain features.
[0046] By performing frame windowing, time-frequency domain conversion, and feature extraction on the original audio signal, the key features of the audio signal can be extracted, the understanding and description ability of the local characteristics of the audio signal can be enhanced, and the quality and accuracy of signal analysis can be improved.
[0047] After obtaining the time-frequency domain features corresponding to the preset audio frame, wind noise suppression processing can be performed based on the time-frequency domain features, and whether the original audio signal is a signal containing wind noise can be determined according to the wind noise suppression processing result. Specifically, for any one preset audio frame, the time-frequency domain features of the preset audio frame are obtained. The time-frequency domain features of the preset audio frame include the initial energy of each frequency point in the preset audio frame. Then the time-frequency domain features are input to the wind noise suppression model for wind noise suppression processing to obtain the attenuation gain parameters of each frequency point in the preset audio frame. According to the attenuation gain parameters of each frequency point in the preset audio frame and the initial energy of each frequency point in the preset audio frame, the gain energy of each frequency point in the preset audio frame is determined. By comparing the gain energy of each frequency point in the preset audio frame and the initial energy of each frequency point in the preset audio frame, the comparison result is obtained. By performing wind noise suppression processing on the energy of each frequency point in the preset audio frame and comparing the energy of each frequency point in the preset audio frame before and after wind noise suppression, the comparison result can be used to determine whether the obtained audio signal contains wind noise, so that wind noise suppression and determination of whether the audio signal is a normal audio signal are combined together, which can improve the processing efficiency of the terminal device for the audio signal, and at the same time can save the system resources of the terminal device.
[0048] In the embodiments of the present application, the wind noise suppression model can be a deep neural network model, such as a recurrent neural network for audio noise reduction (RNNoise), a deep complex convolution recurrent network (DCCRN), etc. The time-frequency domain features input to the wind noise suppression model can be determined according to the type of the wind noise model. Taking the RNNoise network as an example, Figure 3 is a structural diagram of a wind noise suppression model according to an example embodiment, as Figure 3As shown, the RNNoise network mainly includes three Gate Recurrent Units (GRUs) and three dense layers, the number of neurons contained in the three GRUs is 24 / 48 / 96 respectively, and the number of neurons contained in the three dense layers is 24 / 1 / 22 respectively. The input features of RNNoise include 22 Bark-Frequency Cepstral Coefficients (BFCCs), 12 features obtained by taking the first and second order derivatives of the first 6 BFCCs in the time domain, the first 6 coefficients of the discrete cosine transform value of the pitch correlation in the entire frequency band, and the pitch period and spectral non-stationarity metric value, a total of 42 input features (22+6*2+6+1+1). The output of the RNNoise network is 22 attenuation gains in the Bark frequency domain. Through Bark domain to linear domain processing, the attenuation gain value of each frequency point of the spectrum can be obtained. For example, the Fourier transform used in the time-frequency conversion process is 256 points, so there are 257 frequency points and 257 frequency point attenuation gain values corresponding to them. The RNNoise model structure is relatively simple, the effect is relatively practical, and wind noise in the audio signal can be effectively suppressed.
[0049] In the embodiments of the present application, after obtaining the attenuation gain parameters of each frequency point in the preset audio frame, the attenuation gain parameters of each frequency point in the preset audio frame can be multiplied by the initial energy of each frequency point in the preset audio frame, so as to obtain the gain energy of each frequency point in the preset audio frame.
[0050] After obtaining the gain energy of each frequency point in the preset audio frame, the gain energy of each frequency point in the preset audio frame can be compared with the initial energy of each frequency point in the preset audio frame. In the comparison process, the initial spectral energy absolute value of the preset audio frame can be determined based on the initial energy of each frequency point in the preset audio frame, and the spectral energy absolute value of the preset audio frame after wind noise suppression processing can be determined according to the gain energy of each frequency point in the preset audio frame. Then, the spectral energy absolute value of the preset audio frame after wind noise suppression processing is compared with the initial spectral energy absolute value of the preset audio frame to obtain a comparison result. The spectral energy absolute value is the sum of the energy of each frequency point. In determining the initial spectral energy absolute value of the preset audio frame, the initial energy of each frequency point in the preset audio frame can be summed to obtain the initial spectral energy absolute value of the preset audio frame. In some embodiments, the initial spectral energy absolute value of the preset audio frame can also be obtained when the feature extraction is performed on the preset audio frame. In determining the spectral energy absolute value of the preset audio frame after wind noise suppression processing, the gain energy of each frequency point in the preset audio frame can be summed to obtain the spectral energy absolute value of the preset audio frame after wind noise suppression processing. By comparing the spectral energy absolute values of the preset audio frame before and after wind noise suppression, whether the wind noise is contained in the preset audio frame can be accurately measured, so as to improve the accuracy of the judgment result, and at the same time, the complexity of the judgment process can be reduced.
[0051] In the embodiments of the present application, it can be determined whether the wind noise is contained in the preset audio frame according to the comparison result. Specifically, in the case that the comparison result satisfies the first preset condition, it can be determined that the wind noise is contained in the preset audio frame, and in this case, the target audio frame after wind noise suppression processing can be determined according to the gain energy of each frequency point in the preset audio frame.
[0052] In the embodiments of the present application, the comparison result is the quotient of the absolute value of the spectrum energy after wind noise suppression processing of the preset audio frame and the absolute value of the initial spectrum energy of the preset audio frame. The first preset condition can be that the quotient of the absolute value of the spectrum energy after wind noise suppression processing of the preset audio frame and the absolute value of the initial spectrum energy of the preset audio frame is less than a preset value, and the absolute value of the initial spectrum energy of the preset audio frame is greater than a preset energy value. That is, in the case that the quotient of the absolute value of the spectrum energy after wind noise suppression processing of the preset audio frame and the absolute value of the initial spectrum energy of the preset audio frame is less than a preset value, and the absolute value of the initial spectrum energy of the preset audio frame is greater than a preset energy value, it can be determined that the wind noise is contained in the preset audio frame, and in this case, the target audio frame is determined according to the gain energy of each frequency point in the preset audio frame. In the audio signal containing wind noise, the energy of the wind noise often occupies most of the energy of the audio signal. The quotient of the absolute value of the spectrum energy after wind noise suppression processing of the preset audio frame and the absolute value of the initial spectrum energy of the preset audio frame is less than a preset value, such as less than 0.2, which can determine that the spectrum energy of the preset audio frame is greatly weakened after wind noise suppression processing. And the absolute value of the initial spectrum energy of the preset audio frame is greater than a preset energy value, which is used to determine whether the preset audio frame has enough energy, and the combination of the two can improve the accuracy of the wind noise judgment.
[0053] In the embodiments of the present application, in the case that the comparison result does not satisfy the first preset condition, it can be determined that the wind noise is contained in the preset audio frame, and in this case, the preset audio frame can be buffered. Specifically, in the case that the comparison result does not satisfy the first preset condition, the preset audio frame is buffered as an original normal audio frame. Not satisfying the first preset condition means that the quotient of the absolute value of the spectrum energy after wind noise suppression processing of the preset audio frame and the absolute value of the initial spectrum energy of the preset audio frame is greater than or equal to a preset value, or the absolute value of the initial spectrum energy of the preset audio frame is less than or equal to a preset energy value. In this case, it can be considered that the preset audio frame does not contain wind noise. For the preset audio frame which does not contain wind noise, it can be buffered as an original normal audio frame, so that the normal audio frame can be determined according to the original normal audio frame subsequently, and then the energy analysis is performed on the normal audio frame to determine the energy compensation parameter. In the case that the comparison result determines that the preset audio frame does not contain wind noise, buffering the preset audio frame as an original normal audio frame can increase the availability of the preset audio frame and provide protection for repairing damaged audio.
[0054] In the embodiments of the present application, after obtaining the attenuation gain gain value of each frequency point of the preset audio frame output by the wind noise suppression model, the energy of each frequency point of the preset audio frame is multiplied by the attenuation gain gain and summed, and the absolute value of the spectrum energy after wind noise suppression processing (Eout) can be obtained. Next, the sum of the low frequency energies before and after wind noise suppression of the preset audio frame is compared. The low frequency refers to a frequency lower than 1000 Hz, and the frequency of wind noise is usually within this range. Then, according to the comparison result, it is determined whether the preset audio frame has wind noise. When the sum of the low frequency energies output by the wind noise suppression Eout divided by the sum of the low frequency energies of the preset audio before wind noise suppression Ein is less than a preset value, such as 0.1, and Ein is greater than a preset energy value, it is determined that the preset audio frame has wind noise. On the contrary, if the above conditions are not met, it is determined that the current input frame does not have wind noise.
[0055] In the embodiments of the present application, when the preset audio frame has wind noise, and the target audio frame after wind noise suppression processing is determined according to the gain energy of each frequency point in the preset audio frame. Since the target audio frame is obtained after wind noise suppression processing, during the wind noise suppression processing, the low frequency components close to the wind noise frequency are often suppressed together as wind noise components, which may cause the speech components of the low frequency in the obtained target audio frame to be severely damaged. Therefore, in order to improve the audio quality, the target audio frame can be repaired. When repairing the target audio frame, the target fundamental frequency of the target audio frame can be determined first.
[0056] The fundamental frequency can be determined according to the fundamental period. The fundamental period is the longest repeating period in a periodic signal. The fundamental period can be determined by a fundamental period estimation algorithm. Optionally, the fundamental period estimation algorithm includes but is not limited to autocorrelation method, cepstrum method, average magnitude difference function method, linear prediction method, wavelet-autocorrelation function method, spectral subtraction-autocorrelation function method, etc. As an example, the autocorrelation method can be used to estimate the fundamental period of the target audio frame. The autocorrelation method is to find the fundamental period by calculating the similarity between the signal and its delayed version. The calculation process is as follows: first, calculate the autocorrelation function of the input signal (target audio frame): for a discrete signal x[n], its autocorrelation function Rxx[k] is defined as: Rxx[k] = Σx[n]*x[n-k], where k is the delay. Then, find the maximum value of the autocorrelation function: remove the point k=0 from the autocorrelation function, find the maximum value of the remaining autocorrelation function corresponding to the delay k, here k is the fundamental period.
[0057] After calculating the fundamental period of the target audio frame, the reciprocal of the fundamental period of the target audio frame is the target fundamental frequency of the target audio frame.
[0058] S103: Determine the target peak valley energy parameter according to the harmonic energy of at least two target post-harmonics corresponding to the target fundamental frequency.
[0059] A speech signal is a complex signal generated by the joint action of a sound source (such as vocal cord vibration) and a sound channel (such as the mouth, nose, etc.). In a speech signal, a harmonic is a frequency component that is an integer multiple of the fundamental frequency, and the frequency of the fundamental frequency is the pitch frequency. The fundamental frequency is usually related to the pitch of the sounder, and it determines the basic period of the speech signal. The harmonics affect the tone and quality of the speech. In a speech signal, the relative strength, phase relationship, and distribution characteristics of the fundamental frequency and the harmonics jointly determine the characteristics of the sound. The harmonic frequency is an integer multiple of the fundamental frequency, such as 2f0, 3f0, 4f0, etc. In a speech signal, the harmonics usually decrease with increasing frequency, but the specific attenuation characteristics may vary depending on the sounder, the way of pronunciation, and the shape of the sound channel.
[0060] In the embodiments of the present application, when repairing the target audio frame, the target peak-valley energy parameter corresponding to the target audio frame can be determined first. The target pitch frequency can correspond to multiple harmonics, which can be divided into front harmonics and rear harmonics according to the frequency from small to large. That is, the frequency of the front harmonics is smaller than that of the rear harmonics. As an example, the pitch frequency is f0, and the harmonics corresponding to f0 include 2f0, 3f0, 4f0, 5f0, 6f0, 7f0, 8f0, 9f0, and 10f0. The front harmonics can be 2f0, 3f0, 4f0, and 5f0, and the rear harmonics can be 6f0, 7f0, 8f0, 9f0, and 10f0. Optionally, the front harmonics and the rear harmonics corresponding to the target pitch frequency can each have multiple harmonics. Since the harmonic energy usually decreases gradually with the increase of the frequency, the energy of the front harmonics can be repaired when repairing the target audio frame. On the one hand, the repaired audio signal can be restored to a normal level, and on the other hand, the complexity of the repair can be reduced, thereby improving the processing efficiency of the terminal device for the audio signal.
[0061] In the embodiments of the present application, the target peak-valley energy parameter is determined according to the ratio of the target energy peak value to the target energy valley value in the harmonic energy of at least two target rear harmonics. When determining the target peak-valley energy parameter, the harmonic energy of at least two target rear harmonics corresponding to the target pitch frequency can be obtained first, and then the target energy peak value and the target energy valley value are determined in the harmonic energy of the at least two target rear harmonics. According to the target energy peak value and the target energy valley value, the target peak-valley energy ratio is determined, and then the target peak-valley energy parameter is obtained by logarithmic calculation on the target peak-valley energy ratio.
[0062] For the target audio frame, in the wind noise suppression process, the severely damaged is the harmonic of the low frequency band, that is, the energy of the target front harmonic is severely lost, and the energy of the target rear harmonic is relatively small. Therefore, the peak valley energy parameter of the target rear harmonic can accurately represent the energy characteristics of the target audio frame to a certain extent. By obtaining the harmonic energy of the target rear harmonic, the target energy peak and the target energy valley are determined, and the target peak valley energy parameter is determined, so that the harmonic energy range corresponding to the target audio frame can be determined according to the target peak valley energy parameter, to improve the accuracy of subsequent repair of the target audio frame.
[0063] S105: Determine the energy compensation parameter corresponding to the target fundamental frequency and the target peak valley energy parameter based on the preset mapping information.
[0064] In the embodiment of the application, after obtaining the target peak valley energy parameter corresponding to the target audio frame, the energy compensation parameter can be determined according to the target fundamental frequency and the target peak valley energy parameter of the target audio frame by using the preset mapping information. The preset mapping information is used to represent the mapping relationship between the frequency interval where the fundamental frequency is located and the peak valley energy parameter interval where the peak valley energy parameter is located and the energy compensation parameter. The energy compensation parameter is obtained based on the statistical processing of the harmonic energy of the harmonic in the normal audio frame. The energy compensation parameter is used to compensate the energy of the target front harmonic corresponding to the target audio frame, so that the energy of the target front harmonic is restored to a relatively normal level. The normal audio frame is an audio frame that does not contain wind noise signal. By analyzing the harmonic energy envelope of the normal audio frame, the energy characteristics of the normal audio signal of the current terminal device and the current terminal object can be obtained, and the energy characteristics of the audio signal are used to repair the harmonic energy in the target audio frame.
[0065] Because the harmonic energy envelope of normal speech is related to many factors, in addition to the sound cavity structure and pronunciation characteristics of the terminal object, it is also related to the acoustic frequency response of the recording device, so it is difficult to accurately predict and restore the harmonic energy characteristics of the damaged speech signal using a general model. In addition, if a neural network or a deep learning network method is used to learn the personalized sound characteristics of the terminal object online, it needs a long learning training period, has high computational complexity, and may not have good results. Therefore, the application embodiment adopts a simple and low complexity method to statistically analyze the harmonic characteristics of the normal audio signal of the terminal object, and uses the statistical analysis result to predict the damaged harmonic energy of the target audio frame.
[0066] Specifically, the harmonic characteristics of the normal audio signal are obtained by statistical analysis on the harmonic characteristics of the normal audio signal, and the harmonic characteristics of the normal audio signal are taken as the energy compensation parameter. The preset mapping information is obtained by constructing the mapping relationship between the fundamental frequency, the peak-valley energy parameter of the post-stage harmonic, and the energy compensation parameter. When repairing the target audio frame, the energy compensation parameter can be obtained conveniently, and then the target audio frame is repaired.
[0067] In the embodiments of the present application, the preset mapping information can be constructed by the following method. First, the normal audio frame and the preset fundamental frequency of the normal audio frame are obtained. The frequency interval in which the preset fundamental frequency is located is determined in at least two frequency intervals. Then, the preset peak-valley energy parameter is determined according to the harmonic energy of the at least two preset post-stage harmonics corresponding to the preset fundamental frequency. When determining the preset peak-valley energy parameter, the preset peak-valley energy ratio is determined according to the preset energy peak value and the preset energy valley value in the harmonic energy of the at least two preset post-stage harmonics, and then the preset peak-valley energy parameter is obtained by logarithmic calculation of the preset peak-valley energy ratio. The preset peak-valley energy parameter obtained by logarithmic calculation of the preset peak-valley energy ratio can be used to improve the distribution characteristics of the harmonic energy, so that it is more consistent with the normal distribution, thereby simplifying the process of data analysis and model construction. Then, the peak-valley energy parameter interval in which the preset peak-valley energy parameter is located is determined in at least two peak-valley energy parameter intervals. Then, the energy compensation parameter is determined according to the harmonic energy of the preset pre-stage harmonic corresponding to the preset fundamental frequency and the preset energy peak value. The preset mapping information is obtained by constructing the mapping relationship between the frequency interval, the peak-valley energy parameter interval, and the energy compensation parameter.
[0068] By determining the frequency interval in which the preset fundamental frequency is located, the peak-valley energy parameter interval in which the preset peak-valley energy parameter is located, and the energy compensation parameter, the statistical analysis of the harmonic characteristics of the normal audio signal is realized, the envelope of the harmonic energy can be accurately analyzed, and the mapping relationship between the frequency interval, the peak-valley energy parameter interval, and the energy compensation parameter is constructed, which can improve the convenience of subsequent repair of the target audio frame.
[0069] In the embodiments of the present application, the normal audio frame refers to an audio frame that meets certain conditions and does not contain wind noise signals. The normal audio frame can be obtained by detecting and screening the original acupuncture audio frame. Specifically, a second preset number of original normal audio frames are obtained. The original normal audio frames are used to be screened to obtain normal audio frames. The second preset number can be selected according to the statistical analysis needs of the candidate, for example, the second preset number can be 100, that is, 100 frames of original normal audio frames are obtained for screening processing to obtain normal audio frames. Then, each original normal audio frame is detected based on a voice activity detection algorithm to obtain a first normal audio frame containing a target sound signal. Then, the first normal audio frame is detected based on a clear and hoarse sound detection algorithm to obtain a second normal audio frame containing a hoarse sound signal. The second normal audio frame is taken as a normal audio frame.
[0070] In the embodiments of the present application, when analyzing the harmonic energy envelope of the normal audio signal, the normal audio signal of the non-wind noise speech segment needs to be obtained first. The normal audio signal can be an audio frame not containing wind noise in the current voice call, or can be the latest cached original normal audio frame in the cache. In actual application, when analyzing the harmonic energy envelope of the normal audio signal, the normal audio signal without wind noise can be obtained from the current recording (for example, a call application) first. If the current recording process has been disturbed by wind noise all the time, the harmonic feature statistical result of the normal audio signal in the last non-wind noise scene in the voice call application can be searched and applied to the current wind noise scene for voice harmonic repair processing. For the non-wind noise scene, the pitch period value of the normal audio signal is estimated first. It is worth noting that not all input sound signals have a pitch period, for example, the silent frame signal and the clear tone frame signal do not have a pitch period. Since the hoarse sound signal is the sound signal most affected by wind noise suppression, the method provided in the embodiments of the present application is mainly for low-frequency repair processing of the hoarse sound signal with a pitch period. Therefore, after obtaining the original normal audio frame, it needs to be filtered by a voice activity detection (VAD) algorithm and a clear and hoarse sound detection algorithm. Filtering the original normal audio frame by the voice activity detection algorithm and the clear and hoarse sound detection algorithm can ensure that the obtained normal audio frame contains the target sound signal, that is, the hoarse sound signal of the terminal object sound, so as to ensure that the harmonic energy envelope obtained in the subsequent analysis can accurately reflect the harmonic energy characteristics of the terminal object voice signal.
[0071] In the embodiments of the present application, after obtaining the normal audio frame, a preset pitch period of the normal audio frame can be determined based on a pitch period estimation algorithm, and then a preset pitch frequency of the normal audio frame can be determined according to the preset pitch period. Optionally, the pitch period estimation algorithm includes, but is not limited to, autocorrelation method, cepstrum method, average magnitude difference function method, linear prediction method, wavelet-autocorrelation function method, spectral subtraction-autocorrelation function method, etc. The preset pitch frequency is the reciprocal of the preset pitch period.
[0072] In the embodiments of the present application, since the severely damaged part in the target audio frame is the harmonic wave in the low frequency band, the energy of the front harmonic wave of the target audio frame is mainly repaired when the target audio frame is repaired. Therefore, when the harmonic energy envelope of the normal audio frame is analyzed, the energy characteristics of the front harmonic wave of the normal audio frame can be counted, and the energy compensation parameter is obtained.
[0073] Specifically, when the energy compensation parameter is determined, the front harmonic wave energy normalization value can be determined according to the preset front harmonic wave energy corresponding to the preset pitch frequency and the preset energy peak value. Then, the front harmonic wave energy normalization values corresponding to the third preset number of normal audio frames are obtained in the peak-valley energy parameter interval. The front harmonic wave energy normalization values corresponding to the third preset number of normal audio frames are used for statistical processing, so that the finally obtained energy compensation parameter has higher accuracy and representativeness. The third preset number can be selected according to the statistical needs, such as 10, 20, 30, etc. The energy compensation parameter is obtained by weighted average processing of the front harmonic wave energy normalization values corresponding to the third preset number of normal audio frames. By determining the front harmonic wave energy normalization value, the energy of the front harmonic wave is adjusted to a unified standard, so that different harmonics can be statistically combined. By weighted average of the front harmonic wave energy normalization values corresponding to multiple normal audio frames, the obtained energy compensation parameter can more accurately reflect the actual situation of the front harmonic wave energy in the current frequency interval and the current peak-valley energy parameter interval, so that the energy compensation parameter has higher credibility and representativeness, thereby improving the accuracy of repairing the target audio frame.
[0074] When the harmonic energy envelope of the normal audio signal is analyzed, the pitch frequency can be divided into a limited number of intervals according to the range of the pitch frequency value. For example, if the range of the pitch frequency value is 100 Hz to 400 Hz, a frequency interval can be divided every 30 Hz from 100 Hz to 400 Hz, and there are 11 frequency intervals. The pitch frequency of the normal audio frame can be obtained according to the pitch period conversion of the normal audio frame. For any normal audio frame, the frequency interval in which it is located can be determined according to its pitch frequency.
[0075] For any normal audio frame, according to the frequency from small to large, the corresponding first N harmonics in its corresponding harmonics can be selected for statistical analysis. In the N harmonics, the first M are called the front harmonics, and the M+1th to the Nth harmonics are called the rear harmonics. Taking N=10 and M=4 as an example, first, the maximum value PeakVal and the minimum value ValleyVal of the harmonic energy are determined in the rear harmonics, i.e., the 5th to the 10th of the 10 harmonics, a total of 6 harmonics, and the logarithm value of the ratio of the maximum value PeakVal and the minimum value ValleyVal is calculated, i.e., the peak-valley energy parameter P2V=10log10(PeakVal / ValleyVal). For the front harmonics, i.e., the 1st to the 4th of the 10 harmonics, a total of 4 harmonics, the 4 harmonic energies are divided by PeakVal for normalization processing to obtain the front harmonic energy normalized value corresponding to each front harmonic.
[0076] For the value range of P2V, it can also be divided into several peak-valley energy parameter intervals to refine the analysis granularity. For example, the actual value of P2V is divided into an interval every 2dB, for example, it can be divided into 10 intervals, from (-2], (-2, 0], (0, 2], (2, 4], (4, 6], (6, 8], (8, 10], (10, 12], (12, 14], (14,].
[0077] For any normal audio frame, after obtaining its corresponding P2V value, the P2V interval corresponding to the normal audio frame can be determined. Correspondingly, for a P2V interval, a plurality of normal audio frames can correspond to the P2V interval. For any P2V interval, the weighted average processing can be performed on the normalized values of the front segment harmonic energy of the audio frames corresponding to the P2V interval, so as to obtain the corresponding weighted average result. As an example, 10 normal audio frames belong to the interval with P2V of (4, 6], the 10 normal audio frames correspond to 10 normalized values of the front segment harmonic energy of the first harmonic, 10 normalized values of the front segment harmonic energy of the second harmonic, 10 normalized values of the front segment harmonic energy of the third harmonic, and 10 normalized values of the front segment harmonic energy of the fourth harmonic. The weighted average processing is performed on the 10 normalized values of the front segment harmonic energy of the first harmonic, so as to obtain the weighted average result corresponding to the first harmonic. The weighted average processing is performed on the 10 normalized values of the front segment harmonic energy of the second harmonic, so as to obtain the weighted average result corresponding to the second harmonic. The weighted average processing is performed on the 10 normalized values of the front segment harmonic energy of the third harmonic, so as to obtain the weighted average result corresponding to the third harmonic. The weighted average processing is performed on the 10 normalized values of the front segment harmonic energy of the fourth harmonic, so as to obtain the weighted average result corresponding to the fourth harmonic. Then, the weighted average result corresponding to the first harmonic, the weighted average result corresponding to the second harmonic, the weighted average result corresponding to the third harmonic, and the weighted average result corresponding to the fourth harmonic are taken as energy compensation sub-parameters respectively, and the four energy compensation sub-parameters are taken as an array, so as to obtain the energy compensation parameter.
[0078] It should be noted that the number of values included in the energy compensation parameter corresponds to the number of front segment harmonics.
[0079] As described in the above example, the fundamental frequency corresponds to 11 frequency intervals, and the peak-valley energy parameter P2V corresponds to 10 peak-valley energy parameter intervals, so that the preset mapping information including 110 groups of mapping relationships can be constructed. Alternatively, the preset mapping information can also be represented by a three-dimensional data table. In the three-dimensional data table, the three-dimensional information of the table is the frequency interval, the peak-valley energy parameter interval, and the energy compensation parameter corresponding to the weighted average result of a group of front segment harmonic energy normalization values.
[0080] In the embodiment of the present application, the statistical analysis on the harmonic energy envelope of the non-wind-noise normal recording sound signal is performed in intervals, the statistical characteristics of the harmonic statistics obtained by the current terminal device of the current terminal object are counted, and the damaged low-frequency harmonics of the low-frequency speech in the wind noise are repaired according to the statistical characteristics in intervals.
[0081] S107: determining a corrected audio frame corresponding to the target audio frame according to the energy compensation parameter and the harmonic energy of the at least two target rear segment harmonics.
[0082] In the embodiments of the present application, after obtaining the target fundamental frequency and the target peak-valley energy parameter corresponding to the target audio frame, the energy compensation parameter corresponding to the target fundamental frequency can be obtained by querying the preset mapping information, and then the energy compensation parameter is used to repair the target audio frame.
[0083] Specifically, the target fundamental frequency corresponds to at least two target pre-harmonics, and the energy compensation parameter includes a fourth preset number of energy compensation sub-parameters. The number of energy compensation sub-parameters is equal to the number of target pre-harmonics, and one-to-one correspondence. The energy compensation sub-parameters are obtained by statistically analyzing and processing the pre-harmonic energy of the normal sound signal. Therefore, the number of energy compensation sub-parameters in the energy compensation parameter depends on the number of pre-harmonics set in the statistical analysis and processing of the pre-harmonic energy of the normal sound signal. For example, when four pre-harmonics are selected in the statistical analysis and processing of the pre-harmonic energy of the normal sound signal, the number of energy compensation sub-parameters in the energy compensation parameter obtained is four, that is, the fourth preset number is four.
[0084] In determining the corrected audio frame corresponding to the target audio frame, the target post-harmonic energy satisfying the second preset condition can be determined in the harmonic energy of the at least two target post-harmonics. Then, according to each energy compensation sub-parameter and the target post-harmonic energy, the correction harmonic energy corresponding to each target pre-harmonic is determined. Finally, according to the correction harmonic energy corresponding to each target pre-harmonic and the harmonic energy of the at least two target post-harmonics, the corrected audio frame corresponding to the target audio frame is determined. The energy compensation sub-parameters are obtained based on the harmonic energy envelope analysis of the normal audio frame. Using the energy compensation sub-parameters to repair the pre-harmonic in the target audio frame can restore the energy of the pre-harmonic in the target audio frame to a normal level, thereby realizing the repair of the target audio frame and improving the quality of the audio signal.
[0085] In repairing the target audio frame, the target post-harmonic energy satisfying the second preset condition can be determined in the harmonic energy of the at least two target post-harmonics. The second preset condition can be the maximum value of the harmonic energy in the target post-harmonic, or the average value of the harmonic energy in the target post-harmonic, etc. The embodiments of the present application do not make too many limitations.
[0086] In actual application, in the wind noise stage, when repairing the target audio frame, according to the target fundamental frequency of the target audio frame, the frequency interval in which the target fundamental frequency is located is determined, and according to the target peak-valley energy parameter, the peak-valley energy parameter interval in which the target peak-valley energy parameter is located is determined. Then, by searching the preset mapping information obtained in the non-wind noise stage, the energy compensation parameter corresponding to the frequency interval in which the target fundamental frequency is located and the peak-valley energy parameter interval in which the target peak-valley energy parameter is located is obtained, and then each energy compensation sub-parameter in the energy compensation parameter is used to repair the corresponding front harmonic in the target audio frame. As an example, the energy compensation parameter obtained by querying the preset mapping information includes four values of energy compensation sub-parameters: [0.8, 1.6, 2.5, 0.7], and the four values are multiplied by the maximum energy value of the rear harmonics (i.e. the 5th to 10th harmonics) of the target audio frame, so as to obtain the energy values of the four low-frequency target front harmonics, that is, the energy results of the repaired four front harmonic signals in the target audio frame.
[0087] Figure 4 is a target audio frame before and after repair energy diagram according to an example embodiment, as shown in Figure 4 each wavelet peak position represents the position of the harmonic, after wind noise suppression by the wind noise suppression model, the front harmonics of the target audio frame are obviously damaged, and the energy of the front harmonic signals of the target audio frame is repaired by the above-mentioned audio processing method, so that the sound can be restored.
[0088] Figure 5 is a flowchart of an audio processing method according to an example embodiment, as shown in Figure Two , the method can include: Figure 5
[0089] S201: The microphone collects a sound signal.
[0090] S203: Deep learning wind noise suppression.
[0091] S205: Wind noise existence decision.
[0092] S207: Non-wind noise speech envelope feature analysis estimation.
[0093] S209: Wind noise suppression after speech harmonic component extraction and repair.
[0094] S211: The repaired sound signal.
[0095] In the steps S201 to S211, the microphone of the terminal device collects a sound signal. The sound signal collected by the terminal device can be a sound signal containing wind noise or a sound signal not containing wind noise. Then, a deep network learning model, such as a wind noise suppression model such as RNNoise, DCCRN, etc., is used to suppress wind noise of the sound signal. Then, wind noise judgment is performed according to the sound signal after wind noise suppression to determine whether the sound signal collected by the terminal device contains wind noise. For the sound signal not containing wind noise, envelope feature analysis estimation can be performed to obtain the envelope feature of the normal speech signal. For the sound signal containing wind noise, after obtaining the sound signal after wind noise suppression, the envelope feature of the normal speech signal is used to extract and repair the harmonic component, so as to obtain the repaired sound signal.
[0096] It should be noted that in actual application, the sound signal will be processed by the steps S201 to S211 in the form of an audio frame.
[0097] In the embodiment of the present application, after obtaining the repaired frequency spectrum energy, the modified audio frame is restored to a time domain sound signal through inverse Fourier transform processing, and a repaired wind noise suppression result is obtained.
[0098] Figure 6 Fig. 1 shows a flowchart of an audio processing method according to an example embodiment. Figure Three As shown in Fig. 1, the method can include: Figure 6
[0099] S301: Collecting a signal by a microphone.
[0100] S303: Time-frequency conversion and feature extraction.
[0101] S305: Deep learning wind noise suppression processing.
[0102] S307: Whether there is wind noise.
[0103] S309: Fundamental frequency estimation and harmonic energy envelope analysis.
[0104] S311: Fundamental frequency estimation and harmonic correction.
[0105] S313: Frequency domain to time domain.
[0106] S315: Outputting a sound signal.
[0107] In the steps S301 to S315, the microphone of the terminal device collects a sound signal, and then converts the sound signal from the time domain to the frequency domain through time-frequency conversion, and extracts the time-frequency domain features. Then, the extracted time-frequency domain features are input into the wind noise suppression model based on deep learning for wind noise suppression processing. Then, wind noise judgment is performed based on the sound signal after wind noise suppression to determine whether the sound signal collected by the terminal device contains wind noise. For the sound signal without wind noise, the fundamental frequency estimation and the harmonic energy envelope analysis can be performed to obtain the envelope features of the normal speech signal. For the sound signal containing wind noise, after obtaining the sound signal after wind noise suppression, the fundamental frequency estimation is performed, and the harmonic correction is performed using the envelope features of the normal speech signal to restore the harmonic energy. Then, the sound signal after energy restoration is converted from the frequency domain to the time domain. Finally, the sound signal is output.
[0108] The audio processing method provided in the embodiments of the present application is used for the outdoor sound collection scene, and the wind noise signal interference is very serious. The noise features mainly include high low-frequency energy, which can seriously mask the original normal sound signal. Therefore, even if the wind noise is suppressed by the commonly used wind noise suppression algorithm, the original normal low-frequency speech component is difficult to restore, the sound hearing is relatively thin, and the volume is small. If the sound is used in a call, the other party will hear it with great effort or make mistakes in understanding the content. In view of the above defects, the low-frequency harmonic energy envelope of the terminal object is estimated to repair the damaged speech signal. The harmonic component feature information of the terminal object is collected in the current call process or in the historical call. When wind noise interference occurs and is suppressed, the damaged harmonic can be repaired, so that the sound is restored to a relatively normal level, thereby improving the intelligibility and clarity of the sound signal collected by the terminal object.
[0109] The embodiments of the present application also provide an audio processing device, Figure 7 According to an exemplary embodiment, an audio processing device is shown in a block diagram as shown in Figure 7 The audio processing device can at least include:
[0110] The acquisition module 301 is configured to acquire a target audio frame and a target fundamental frequency of the target audio frame. The target audio frame is obtained through wind noise suppression processing.
[0111] The target peak-valley energy parameter determination module 303 is configured to determine a target peak-valley energy parameter according to the harmonic energy of at least two target later-stage harmonics corresponding to the target fundamental frequency. The target peak-valley energy parameter is determined according to the ratio of the target energy peak value to the target energy valley value in the harmonic energy of the at least two target later-stage harmonics.
[0112] The energy compensation parameter determination module 305 is configured to determine an energy compensation parameter corresponding to the target pitch frequency and the target peak-valley energy parameter based on preset mapping information. The preset mapping information is used to represent a mapping relationship between a frequency interval in which the pitch frequency is located and a peak-valley energy parameter interval in which the peak-valley energy parameter is located and the energy compensation parameter. The energy compensation parameter is obtained based on harmonic energy statistical processing of harmonics in a normal audio frame. The normal audio frame is an audio frame that does not contain wind noise signals.
[0113] The corrected audio frame determination module 307 is configured to determine a corrected audio frame corresponding to the target audio frame according to the energy compensation parameter and harmonic energy of at least two target later-stage harmonics.
[0114] In some optional embodiments, the apparatus further includes:
[0115] The time-frequency domain feature acquisition module is configured to acquire time-frequency domain features of the preset audio frame. The time-frequency domain features of the preset audio frame include initial energies of each frequency point in the preset audio frame.
[0116] The wind noise suppression module is configured to input the time-frequency domain features to a wind noise suppression model for wind noise suppression processing, to obtain attenuation gain parameters of each frequency point in the preset audio frame.
[0117] The gain energy determination module is configured to determine gain energies of each frequency point in the preset audio frame according to the attenuation gain parameters of each frequency point in the preset audio frame and the initial energies of each frequency point in the preset audio frame.
[0118] The comparison module is configured to perform comparison processing on the gain energies of each frequency point in the preset audio frame and the initial energies of each frequency point in the preset audio frame, to obtain a comparison result.
[0119] The target audio frame determination module is configured to determine the target audio frame according to the gain energies of each frequency point in the preset audio frame in a case where the comparison result satisfies a first preset condition.
[0120] In some optional embodiments, the time-frequency domain feature acquisition module includes:
[0121] The original audio signal acquisition submodule is configured to acquire an original audio signal.
[0122] The framing and windowing submodule is configured to perform framing and windowing processing on the audio signal, to obtain a first preset number of original audio frames.
[0123] The conversion submodule is configured to convert each original audio frame from the time domain to the frequency domain, to obtain a first preset number of preset audio frames.
[0124] The feature extraction submodule is configured to perform feature extraction on each preset audio frame, to obtain time-frequency domain features corresponding to each preset audio frame.
[0125] In some alternative embodiments, the comparison module comprises:
[0126] An initial spectral energy absolute value determination submodule is configured to determine an initial spectral energy absolute value of the preset audio frame based on initial energy of each frequency point in the preset audio frame.
[0127] A wind noise suppression spectral energy absolute value determination submodule is configured to determine a spectral energy absolute value of the preset audio frame after wind noise suppression processing based on gain energy of each frequency point in the preset audio frame.
[0128] A comparison submodule is configured to compare the spectral energy absolute value of the preset audio frame after wind noise suppression processing with the initial spectral energy absolute value of the preset audio frame to obtain a comparison result.
[0129] In some alternative embodiments, the comparison result is a quotient of the spectral energy absolute value of the preset audio frame after wind noise suppression processing and the initial spectral energy absolute value of the preset audio frame; and the target audio frame determination module comprises:
[0130] A target audio frame determination submodule is configured to determine the target audio frame based on the gain energy of each frequency point in the preset audio frame when the quotient of the spectral energy absolute value of the preset audio frame after wind noise suppression processing and the initial spectral energy absolute value of the preset audio frame is less than a preset value and the initial spectral energy absolute value of the preset audio frame is greater than a preset energy value.
[0131] In some alternative embodiments, the apparatus further comprises:
[0132] An original normal audio frame determination module is configured to cache the preset audio frame as an original normal audio frame when the comparison result does not satisfy a first preset condition; and the original normal audio frame is used to determine a normal audio frame.
[0133] In some alternative embodiments, the target peak-valley energy parameter determination module comprises:
[0134] A harmonic energy acquisition determination submodule is configured to acquire harmonic energy of at least two target post-periodic waves corresponding to a target fundamental frequency.
[0135] A target energy value determination submodule is configured to determine a target energy peak value and a target energy valley value from the harmonic energy of the at least two target post-periodic waves.
[0136] A target peak-valley energy parameter determination submodule is configured to determine a target peak-valley energy parameter based on the target energy peak value and the target energy valley value.
[0137] In some alternative embodiments, the apparatus further comprises a preset mapping information determination module, and the preset mapping information determination module comprises:
[0138] an obtaining sub-module, configured to obtain a normal audio frame and a preset fundamental frequency of the normal audio frame;
[0139] a frequency interval determination sub-module, configured to determine a frequency interval in which the preset fundamental frequency is located, from at least two frequency intervals;
[0140] a preset peak-valley energy parameter determination sub-module, configured to determine a preset peak-valley energy parameter according to harmonic energies of at least two preset post-harmonics corresponding to the preset fundamental frequency; the preset peak-valley energy parameter is determined according to a ratio of a preset energy peak value to a preset energy valley value in the harmonic energies of the at least two preset post-harmonics;
[0141] a peak-valley energy parameter interval determination sub-module, configured to determine a peak-valley energy parameter interval in which the preset peak-valley energy parameter is located, from at least two peak-valley energy parameter intervals;
[0142] an energy compensation parameter determination sub-module, configured to determine an energy compensation parameter according to a harmonic energy of a preset pre-harmonic corresponding to the preset fundamental frequency and the preset energy peak value;
[0143] a preset mapping information determination sub-module, configured to construct a mapping relationship between the frequency interval, the peak-valley energy parameter interval and the energy compensation parameter, and obtain preset mapping information.
[0144] In some optional embodiments, the obtaining sub-module comprises:
[0145] an original normal audio frame obtaining unit, configured to obtain a second preset number of original normal audio frames;
[0146] a first normal audio frame determination unit, configured to detect each original normal audio frame based on a voice activity detection algorithm, and obtain a first normal audio frame containing a target sound signal;
[0147] a second normal audio frame determination unit, configured to detect the first normal audio frame based on a voiced-unvoiced sound detection algorithm, and obtain a second normal audio frame containing a voiced sound signal;
[0148] a normal audio frame determination unit, configured to take the second normal audio frame as the normal audio frame;
[0149] a preset fundamental period determination unit, configured to determine a preset fundamental period of the normal audio frame based on a fundamental period estimation algorithm;
[0150] a preset fundamental frequency determination unit, configured to determine a preset fundamental frequency of the normal audio frame according to the preset fundamental period.
[0151] In some optional embodiments, the energy compensation parameter determination sub-module comprises:
[0152] The pre-harmonic energy normalization value determination unit is configured to determine a pre-harmonic energy normalization value according to a preset harmonic energy of a preset pre-harmonic corresponding to a preset fundamental frequency and a preset energy peak value;
[0153] The pre-harmonic energy normalization value acquisition unit is configured to acquire pre-harmonic energy normalization values corresponding to a third preset number of normal audio frames in the peak-valley energy parameter interval.
[0154] The energy compensation parameter determination unit is configured to perform weighted average processing on the pre-harmonic energy normalization values corresponding to the third preset number of normal audio frames to obtain an energy compensation parameter.
[0155] In some optional embodiments, the target fundamental frequency corresponds to at least two target pre-harmonics; the energy compensation parameter includes a fourth preset number of energy compensation sub-parameters, the number of the energy compensation sub-parameters is equal to the number of the target pre-harmonics, and the energy compensation sub-parameters are in one-to-one correspondence; and the corrected audio frame determination module includes:
[0156] The target post-harmonic energy determination sub-module is configured to determine target post-harmonic energy satisfying a second preset condition from the harmonic energies of the at least two target post-harmonics.
[0157] The corrected harmonic energy determination sub-module is configured to determine, according to each energy compensation sub-parameter and the target post-harmonic energy, a corrected harmonic energy corresponding to each target pre-harmonic.
[0158] The corrected audio frame determination sub-module is configured to determine, according to the corrected harmonic energy corresponding to each target pre-harmonic and the harmonic energies of the at least two target post-harmonics, a corrected audio frame corresponding to the target audio frame.
[0159] It should be noted that the audio processing device embodiments provided by the embodiments of the present application are based on the same inventive concept as the above-mentioned audio processing method embodiments.
[0160] The embodiments of the present application also provide an electronic device for audio processing, which includes a processor and a memory, the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program is loaded and executed by the processor to implement the audio processing method provided in any of the above embodiments.
[0161] The embodiments of the present application also provide a computer readable storage medium, which can be arranged in a terminal to save at least one instruction or at least one program for implementing an audio processing method in the method embodiments, the at least one instruction or the at least one program is loaded and executed by the processor to implement the audio processing method provided in the above method embodiments.
[0162] Optionally, in the embodiments of the present application, the storage medium can be located in at least one of the plurality of network servers of the computer network. Optionally, in the embodiments, the storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0163] The memory of the embodiments of the present application can be used to store software programs and modules, and the processor executes various function applications and data processing by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by functions, etc.; and the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory can also include a memory controller to provide access for the processor to the memory.
[0164] The embodiments of the present application also provide a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio processing method provided by the method embodiments.
[0165] The method embodiments provided by the embodiments of the present application can be executed in a terminal, a computer terminal, a server or a similar computing device. Taking the case of running on a server as an example, Figure 8 is a hardware structure block diagram of a server according to an exemplary embodiment of an audio processing method. As shown in Figure 8As shown, the server 700 can vary in configuration and performance based on the needs of the user. The server 700 can include one or more Central Processing Units (CPU) 710 (the CPU 710 can include, but is not limited to, a microprocessor, an Application Specific Integrated Circuit (ASIC), a programmable logic device (PLD), a processor, a controller, a microcontroller, a microprocessor, other processing units, or one or more processors of any kind), a memory 730 for storing data, one or more storage media 720 (such as one or more mass storage devices) for storing applications 723 or data 722. The memory 730 and the storage media 720 can be of any type generally known or used in the art including non-volatile memory, volatile memory, disk storage, magnetic storage, optical storage, or any other suitable type of storage. The applications 723 and the data 722 stored in the storage media 720 can include one or more modules that can include a series of instructions for operating the server. Further, the CPU 710 can be configured to communicate with the storage media 720 to execute the series of instructions in the storage media 720 to operate the server 700. The server 700 can also include one or more power supplies 760, one or more wired or wireless network interfaces 750, one or more input / output interfaces 740, and / or one or more operating systems 721, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0166] The input / output interface 740 can be configured to receive or transmit data via a network. Examples of the network can include a wireless network provided by a communication provider of the server 700. In one example, the input / output interface 740 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the input / output interface 740 can be a radio frequency (RF) module that is configured to communicate with the Internet through a wireless manner.
[0167] Those skilled in the art can understand that Figure 8 The structure shown is merely illustrative and does not limit the structure of the electronic device described above. For example, the server 700 can include more or less components than those shown, or have a different configuration of components than those shown. Figure 8 For example, the server 700 can include more or less components than those shown, or have a different configuration of components than those shown. Figure 8 For example, the server 700 can include more or less components than those shown, or have a different configuration of components than those shown.
[0168] It should be noted that the above-mentioned order of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. And the above describes the specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0169] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device and server embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0170] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk.
[0171] The above is only the preferred embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An audio processing method, characterized by, The method comprises: obtaining a target audio frame and a target fundamental frequency of the target audio frame; the target audio frame is obtained through wind noise suppression processing; determining a target peak-valley energy parameter according to harmonic energy of at least two target latter harmonics corresponding to the target fundamental frequency; the target peak-valley energy parameter is determined according to a ratio of a target energy peak value to a target energy valley value in the harmonic energy of the at least two target latter harmonics; determining an energy compensation parameter corresponding to the target fundamental frequency and the target peak-valley energy parameter based on preset mapping information; the preset mapping information is used to represent a mapping relationship between a frequency interval where the fundamental frequency is located and a peak-valley energy parameter interval where the peak-valley energy parameter is located and the energy compensation parameter; the energy compensation parameter is obtained based on statistical processing of harmonic energy of harmonics in a normal audio frame; the normal audio frame is an audio frame not containing wind noise signals; determining a corrected audio frame corresponding to the target audio frame according to the energy compensation parameter and the harmonic energy of the at least two target latter harmonics.
2. The method of claim 1, wherein, Before the obtaining of the target audio frame and the target fundamental frequency of the target audio frame, the method further comprises: obtaining a time-frequency domain feature of a preset audio frame; the time-frequency domain feature of the preset audio frame comprises initial energy of each frequency point in the preset audio frame; inputting the time-frequency domain feature into a wind noise suppression model for wind noise suppression processing to obtain an attenuation gain parameter of each frequency point in the preset audio frame; determining gain energy of each frequency point in the preset audio frame according to the attenuation gain parameter of each frequency point in the preset audio frame and the initial energy of each frequency point in the preset audio frame; performing comparison processing on the gain energy of each frequency point in the preset audio frame and the initial energy of each frequency point in the preset audio frame to obtain a comparison result; in a case where the comparison result satisfies a first preset condition, determining the target audio frame according to the gain energy of each frequency point in the preset audio frame.
3. The method of claim 2, wherein, The obtaining of the time-frequency domain feature of the preset audio frame comprises: obtaining an original audio signal; performing frame windowing processing on the audio signal to obtain a first preset number of original audio frames; converting each original audio frame from a time domain to a frequency domain to obtain a first preset number of preset audio frames; performing feature extraction on each preset audio frame to obtain a time-frequency domain feature corresponding to each preset audio frame.
4. The method according to claim 2 or 3, characterized in that, The comparison processing on the gain energy of each frequency point in the preset audio frame and the initial energy of each frequency point in the preset audio frame to obtain a comparison result comprises: determining an initial spectral energy absolute value of the preset audio frame based on the initial energy of each frequency point in the preset audio frame; determining a spectral energy absolute value of the preset audio frame after wind noise suppression processing according to the gain energy of each frequency point in the preset audio frame; comparing the spectral energy absolute value of the preset audio frame after wind noise suppression processing with the initial spectral energy absolute value of the preset audio frame to obtain a comparison result.
5. The method of claim 4, wherein, The comparison result is a quotient of the spectral energy absolute value of the preset audio frame after wind noise suppression processing and the initial spectral energy absolute value of the preset audio frame. The method further comprises: In a case where the comparison result does not satisfy the first preset condition, the preset audio frame is cached as an original normal audio frame; the original normal audio frame is used to determine the normal audio frame.
6. The method of claim 2, wherein, The method further comprises: The method further comprises:
7. The method of claim 1, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises:
8. The method of claim 1, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises:
9. The method of claim 8, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises:
10. The method according to claim 8 or 9, characterized in that, The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The determine a pre-harmonic energy normalization value according to the pre-harmonic energy corresponding to the preset fundamental frequency and the preset energy peak value; obtain a pre-harmonic energy normalization value corresponding to each of the third preset number of normal audio frames in the peak-valley energy parameter interval; perform weighted average processing on the pre-harmonic energy normalization values corresponding to the third preset number of normal audio frames to obtain an energy compensation parameter.
11. The method of claim 1, wherein, The target fundamental frequency corresponds to at least two target pre-harmonics, and the energy compensation parameter includes a fourth preset number of energy compensation sub-parameters, the number of the energy compensation sub-parameters is equal to the number of the target pre-harmonics, and each energy compensation sub-parameter corresponds to one target pre-harmonic; The method of determining the corrected audio frame corresponding to the target audio frame according to the energy compensation parameter and the harmonic energy of the at least two target post-harmonics includes: determining a target post-harmonic energy that satisfies a second preset condition from the harmonic energy of the at least two target post-harmonics; determining a corrected harmonic energy corresponding to each target pre-harmonic according to each energy compensation sub-parameter and the target post-harmonic energy; determining a corrected audio frame corresponding to the target audio frame according to the corrected harmonic energy corresponding to each target pre-harmonic and the harmonic energy of the at least two target post-harmonics.
12. An audio processing apparatus, characterized by comprising: The device includes: An acquisition module is configured to acquire a target audio frame and a target fundamental frequency of the target audio frame, wherein the target audio frame is obtained after wind noise suppression processing; A target peak-valley energy parameter determination module is configured to determine a target peak-valley energy parameter according to the harmonic energy of at least two target post-harmonics corresponding to the target fundamental frequency, wherein the target peak-valley energy parameter is determined according to a ratio of a target energy peak value to a target energy valley value in the harmonic energy of the at least two target post-harmonics; An energy compensation parameter determination module is configured to determine an energy compensation parameter corresponding to the target fundamental frequency and the target peak-valley energy parameter based on preset mapping information, wherein the preset mapping information is used to represent a mapping relationship between a frequency interval in which a fundamental frequency is located and a peak-valley energy parameter interval in which a peak-valley energy parameter is located and an energy compensation parameter, the energy compensation parameter is obtained based on statistical processing of harmonic energy of harmonics in a normal audio frame, and the normal audio frame is an audio frame that does not contain wind noise signals; A corrected audio frame determination module is configured to determine a corrected audio frame corresponding to the target audio frame according to the energy compensation parameter and the harmonic energy of the at least two target post-harmonics.
13. An electronic device for audio processing, the electronic device comprising: The device includes a processor and a memory, and the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program is loaded and executed by the processor, and the audio processing method of any one of claims 1-11 is implemented.
14. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program, the at least one instruction or the at least one program is loaded and executed by the processor, and the audio processing method of any one of claims 1-11 is implemented.
15. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the audio processing method of any one of claims 1-11.