Audio processing method, electronic device, chip system, storage medium and program product
By adjusting the gain by combining the frequency band sharpness distribution and the sound pressure level distribution, sibilance suppression is achieved, which solves the problem of tone damage during sibilance suppression and improves the user experience.
Patent Information
- Application Number
- CN202411297305.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-09-14
AI Technical Summary
Existing technologies can easily lead to tone damage during sibilance suppression, affecting the user experience.
Sibilance suppression is achieved by combining the sharpness distribution and sound pressure level distribution of the frequency band with the sharpness and amplitude of the sibilance signal, and adjusting the gain to reduce timbre damage.
It significantly reduces the sharpness of high-sharp sibilance signals while reducing the suppression of low-sharp sibilance signals, thus minimizing the damage to timbre caused by sibilance suppression.
Smart Images

Figure CN119418716B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminals, and in particular to an audio processing method, an electronic device, a chip system, a storage medium and a program product. BACKGROUND
[0002] Sibilance is the sound produced when air flow and teeth rub together when a person is speaking, and the energy thereof is mainly concentrated in 4kHz-12kHz, which is in a frequency band that is more sensitive to human ears, and is easy to cause a user to have a bad subjective listening experience. Therefore, currently, before an electronic device plays audio data, the audio data is subjected to sibilance recognition and sibilance suppression. Exemplarily, after the electronic device recognizes one or more frames of sibilance in the audio data, the electronic device can perform sibilance suppression through an equalizer (EQ) or a multiband dynamic range compressor (MBDRC) or the like.
[0003] However, such a method can cause timbre damage, so that the timbre of the audio after sibilance suppression is distorted. SUMMARY
[0004] Embodiments of the present application provide an audio processing method, an electronic device, a chip system, a storage medium and a program product, which are applied to the technical field of terminals. The electronic device can perform suppression based on a sharpness distribution and a frequency band sound pressure level distribution corresponding to the frequency band, and the timbre damage to the audio data in the sibilance suppression process is small, thereby being capable of improving user experience.
[0005] In a first aspect, an embodiment of the present application provides an audio processing method, which comprises: displaying a first interface of a first application, the first interface comprising a first control; based on a first operation of a user on the first control, obtaining a first audio, the first audio comprising a first sibilance signal and a second sibilance signal, the amplitude of the first sibilance signal being consistent with the amplitude of the second sibilance signal, and the first sharpness of the first sibilance signal being less than the second sharpness of the second sibilance signal; and playing a second audio, the second audio being obtained after sibilance suppression is performed on the first audio, the second audio comprising a third sibilance signal and a fourth sibilance signal, the third sibilance signal being the first sibilance signal after processing, and the fourth sibilance signal being the second sibilance signal after processing, and a first difference between the amplitude of the first sibilance signal and the amplitude of the third sibilance signal being less than a second difference between the amplitude of the second sibilance signal and the amplitude of the fourth sibilance signal.
[0006] The first application can be an application capable of playing audio, and the first control can be a control for triggering the playing of audio. The first application can be, for example, a music application, and the first control can be a control for playing music; the first application can also be a video application, and the first control can be a control for triggering the playing of a video; the first application can also be a call application, and the first control can be a control for making or receiving a call; or, the first application can also be any other application capable of playing audio, and the first control can be any other control capable of triggering the playing of audio. The first operation can be, but is not limited to, a click operation, a sliding operation, or the like. The first audio can be one frame or multiple frames of audio data, for example, the xth frame of audio data in the following. The first tooth signal can be one frame of tooth, or a part of the audio signal in one frame of tooth, for example, the kth frequency band in the following, that is, the first tooth signal is an audio signal corresponding to one frequency band in the xth frame of audio data (tooth).
[0007] The first sharpness can reflect the sharpness of the first tooth signal heard by the user, and the second sharpness can reflect the sharpness of the second tooth signal heard by the user. The magnitude of the first tooth signal being consistent with the magnitude of the second tooth signal can mean that the magnitude of the first tooth signal is exactly equal to the magnitude of the second tooth signal, or the difference between the magnitude of the first tooth signal and the magnitude of the second tooth signal is less than a threshold value.
[0008] The third tooth signal is the first tooth signal after tooth suppression, and the terminal device can suppress the tooth of the first tooth signal based on the gain corresponding to the first tooth signal to obtain the third tooth signal. The greater the gain corresponding to the first tooth signal, the greater the degree of suppression of the first tooth signal, and the greater the difference between the magnitude of the first tooth signal and the magnitude of the third tooth signal. The fourth tooth signal is the second tooth signal after tooth suppression, and the terminal device can suppress the tooth of the second tooth signal based on the gain corresponding to the second tooth signal to obtain the fourth tooth signal. The greater the gain corresponding to the second tooth signal, the greater the degree of suppression of the second tooth signal, and the greater the difference between the magnitude of the second tooth signal and the magnitude of the fourth tooth signal.
[0009] It can be understood that the magnitude and the frequency band sound pressure level are positively correlated. The magnitude of the first tooth signal being consistent with the magnitude of the second tooth signal can also mean that the sum of the frequency band sound pressure level distributions corresponding to the first tooth signal is consistent with the sum of the frequency band sound pressure level distributions corresponding to the second tooth signal.
[0010] The audio processing method of the present application, the terminal device suppresses the sibilance signal in combination with the sharpness and amplitude of the sibilance signal. For sibilance signals with similar amplitudes and different sharpnesses, the greater the sharpness of the sibilance signal, the greater the degree of suppression of the sibilance signal, so that the sharpness of the suppressed sibilance signal is significantly reduced; the smaller the sharpness of the sibilance signal, the smaller the degree of suppression of the sibilance signal. In this way, the sibilance signal is suppressed based on the sharpness and amplitude, which not only significantly reduces the sharpness of the sibilance signal with high sharpness, but also has a lower degree of suppression for the sibilance signal with low sharpness, so that the hearing of the suppressed sibilance signal is not too dull, and the damage to the timbre of the sibilance suppression is small.
[0011] In combination with the first aspect, in some implementations of the first aspect, the first sharpness is the sum of the sharpness distributions corresponding to the first sibilance signal, and the second sharpness is the sum of the sharpness distributions corresponding to the second sibilance signal.
[0012] The sharpness distribution corresponding to the first sibilance signal can be understood as the sharpness distribution of the Brak frequency band corresponding to the first sibilance signal, for example, the sharpness of one or more Bark frequency bands corresponding to the kth frequency band in the following, and the first sharpness is, for example, S1 in the following; the sharpness distribution corresponding to the second sibilance signal can be understood as the sharpness distribution of the Brak frequency band corresponding to the second sibilance signal, for example, the sharpness of one or more Bark frequency bands corresponding to the wth frequency band in the following, and the second sharpness is, for example, S2 in the following. The sum of the sharpness distributions corresponding to the first sibilance signal can be calculated, for example, by the part of calculating the sum of the sharpness distributions in formula 13 in the following; the sum of the sharpness distributions corresponding to the second sibilance signal can be calculated, for example, by the part of calculating the sum of the sharpness distributions in formula 13 in the following.
[0013] In this way, the Brak frequency band is based on the frequency scale of human auditory perception, so that the first sharpness of the first sibilance signal and the second sharpness of the second sibilance signal can more truly reflect human hearing, so that the sibilance signal with a higher degree of sharpness can be suppressed.
[0014] In combination with the first aspect, in some implementations of the first aspect, the third sibilance signal is a signal obtained by processing the first sibilance signal based on a gain of the first sibilance signal, and the fourth sibilance signal is a signal obtained by processing the second sibilance signal based on a gain of the second sibilance signal, and the gain of the first sibilance signal is smaller than the gain of the second sibilance signal.
[0015] In this way, since the first sharpness of the first sibilance signal is less than the second sharpness of the second sibilance signal, the gain of the first sibilance signal is less than the gain of the second sibilance signal. That is, when the amplitudes are similar, the smaller the sharpness is, the smaller the degree of suppression of the sibilance signal is, so that the sibilance signal with smaller sharpness is suppressed to a smaller degree or is not suppressed, and the timbre damage caused in the sibilance suppression process is smaller.
[0016] In combination with the first aspect, in some implementations of the first aspect, the gain of the first sibilance signal is determined based on a sum of the first sharpness and a frequency band sound pressure level distribution corresponding to the first sibilance signal, and the gain of the second sibilance signal is determined based on a sum of the second sharpness and a frequency band sound pressure level distribution corresponding to the second sibilance signal, the sum of the frequency band sound pressure level distribution corresponding to the first sibilance signal is positively correlated with the amplitude of the first sibilance signal, and the sum of the frequency band sound pressure level distribution corresponding to the second sibilance signal is positively correlated with the amplitude of the second sibilance signal.
[0017] In this way, since the first sharpness of the first sibilance signal is less than the second sharpness of the second sibilance signal, the gain of the first sibilance signal is less than the gain of the second sibilance signal. That is, when the amplitudes are similar, the smaller the sharpness is, the smaller the degree of suppression of the sibilance signal is, so that the sibilance signal with smaller sharpness is suppressed to a smaller degree or is not suppressed, and the timbre damage caused in the sibilance suppression process is smaller. In this way, since the first sharpness of the first sibilance signal is less than the second sharpness of the second sibilance signal, the gain of the first sibilance signal is less than the gain of the second sibilance signal. That is, when the amplitudes are similar, the smaller the sharpness is, the smaller the degree of suppression of the sibilance signal is, so that the sibilance signal with smaller sharpness is suppressed to a smaller degree or is not suppressed, and the timbre damage caused in the sibilance suppression process is smaller. In this way, since the first sharpness of the first sibilance signal is less than the second sharpness of the second sibilance signal, the gain of the first sibilance signal is less than the gain of the second sibilance signal. That is, when the amplitudes are similar, the smaller the sharpness is, the smaller the degree of suppression of the sibilance signal is, so that the sibilance signal with smaller sharpness is suppressed to a smaller degree or is not suppressed, and the timbre damage caused in the sibilance suppression process is smaller.
[0018] In this way, since the first sharpness of the first sibilance signal is less than the second sharpness of the second sibilance signal, the gain of the first sibilance signal is less than the gain of the second sibilance signal. That is, when the amplitudes are similar, the smaller the sharpness is, the smaller the degree of suppression of the sibilance signal is, so that the sibilance signal with smaller sharpness is suppressed to a smaller degree or is not suppressed, and the timbre damage caused in the sibilance suppression process is smaller.
[0019] In some implementations of the first aspect, the first sibilance signal belongs to a c1th frame of audio data, the second sibilance signal belongs to a c2th frame of audio data, the c1th frame of audio data and the c2th frame of audio data are sibilance, the c1th frame of audio data is identified by the terminal device based on a sharpness distribution corresponding to the c1th frame of audio data and a spectral centroid of the c1th frame of audio data, and the c2th frame of audio data is identified by the terminal device based on a sharpness distribution corresponding to the c2th frame of audio data and a spectral centroid of the c2th frame of audio data.
[0020] The c1th frame of audio data and the c2th frame of audio data can be one frame of sibilance (c1 is the same as c2) or different frames of sibilance (c1 is different from c2). For example, when the c1th frame of audio data and the c2th frame of audio data are different frames of sibilance, the first sibilance signal can be a kth frequency band in the c1th frame of audio data, and the second sibilance signal can be a kth frequency band in the c2th frame of audio data, that is, the frequency ranges of the first sibilance signal and the second sibilance signal are the same.
[0021] The sharpness distribution corresponding to the c1th frame of audio data and the sharpness distribution corresponding to the c2th frame of audio data can each be a set of sharpnesses including one or more Bark frequency bands.
[0022] In this way, the sharpness distribution can reflect the sharpness of the audio data, and the spectral centroid can reflect the center position of the energy distribution in the spectrum, so that the terminal device can identify sibilance with a higher energy part distributed at a high frequency position and a higher sharpness. Audio data with a higher energy part distributed at a low frequency position is less likely to be identified as sibilance. This makes the accuracy of sibilance identification higher.
[0023] In some implementations of the first aspect, a product of a sum of the sharpness distribution corresponding to the c1th frame of audio data and a normalized spectral centroid of the c1th frame of audio data is greater than a fourth threshold value, and a product of a sum of the sharpness distribution corresponding to the c2th frame of audio data and a normalized spectral centroid of the c2th frame of audio data is greater than the fourth threshold value.
[0024] The product of the sum of the sharpness distribution corresponding to the c1th frame of audio data and the normalized spectral centroid of the c1th frame of audio data and the product of the sum of the sharpness distribution corresponding to the c2th frame of audio data and the normalized spectral centroid of the c2th frame of audio data can each be understood as the identified sharpness in the following description and can be calculated according to Formula 9. The fourth threshold value can be the threshold value 1 in the following description.
[0025] In this way, the value of the normalized spectral centroid can be distributed in a relatively small range, for example, between 0 and 1, so that the fourth threshold can be easily determined; and the normalized spectral centroid can reduce the influence of different audio signals due to different sampling rates and frequency ranges, so that the identification of the sharpness can more accurately reflect the center position of the energy distribution in the spectrum, which helps to improve the accuracy of sibilance identification.
[0026] In combination with the first aspect, in some implementations of the first aspect, the sharpness distribution corresponding to the c1th frame of audio data and the sharpness distribution corresponding to the c2th frame of audio data respectively include the sharpness of a plurality of Bark frequency bands, and the plurality of Bark frequency bands are Bark frequency bands with a frequency greater than or equal to a preset frequency.
[0027] The plurality of Bark frequency bands can be a plurality of specific Bark frequency bands as follows.
[0028] In this way, the sharpness distribution corresponding to the c1th frame of audio data and the sharpness distribution corresponding to the c2th frame of audio data can reflect the sharpness of the high-frequency part of the audio data, which facilitates more accurate identification of sibilance.
[0029] In the second aspect, an audio processing method is provided. The method includes: determining that an xth frame of audio data in a plurality of frames of audio data is sibilance; dividing a preset frequency range of the xth frame of audio data into y frequency bands, the preset frequency range including a sibilance frequency band, or the preset frequency range partially overlapping with the sibilance frequency band, y being a positive integer; calculating a gain of each frequency band in the y frequency bands based on a sharpness distribution and a frequency band sound pressure level distribution of each frequency band in the y frequency bands, the sharpness distribution including the sharpness of at least one frequency band, and the frequency band sound pressure level distribution including the frequency band sound pressure level of at least one frequency band; and performing suppression on the y frequency bands based on the gain of each frequency band in the y frequency bands to obtain sibilance-suppressed audio data.
[0030] The sharpness distribution corresponding to each frequency band in the y frequency bands can include the sharpness of at least one frequency band corresponding to the each frequency band; and the frequency band sound pressure level distribution corresponding to each frequency band in the y frequency bands can include the frequency band sound pressure level of at least one frequency band corresponding to the each frequency band.
[0031] The audio processing method provided by the present application can perform frequency-division sibilance suppression on a preset frequency unit of sibilance, and such fine-grained sibilance suppression can reduce timbre damage. In addition, the terminal device calculates the gain corresponding to each frequency band based on the sharpness distribution and the frequency band sound pressure level distribution, where the sharpness distribution can reflect the sharpness of the frequency band, and the frequency band sound pressure level distribution can reflect the energy of the frequency band. In this way, the terminal device can suppress the frequency band with high sharpness and large energy. The probability of suppressing the frequency band with low sharpness can be reduced, so that the timbre damage caused in the sibilance suppression process is small.
[0032] In some implementations of the second aspect, the calculating the gain of each of the y frequency bands includes: calculating a sum of the first frequency band sound pressure level distribution corresponding to a kth frequency band of the y frequency bands to obtain a first value; calculating a sum of the first sharpness distribution corresponding to the kth frequency band of the y frequency bands to obtain a second value; calculating a suppression sharpness of the kth frequency band based on the first value and the second value, the suppression sharpness of the kth frequency band being used to represent a size of the first value and the second value; and determining the first gain of the kth frequency band based on the suppression sharpness.
[0033] The first value can be a value calculated by formula 12 below. The second value can be a value calculated by formula 13 below. The first gain can be a band gain g k of the kth frequency band calculated by formula 14 below.
[0034] In this way, the suppression sharpness of the kth frequency band can reflect the sharpness and the energy of the kth frequency band. The determination of the first gain based on the suppression sharpness can reduce the probability of suppressing the frequency band with low sharpness, thereby reducing the timbre damage of the sibilance suppression.
[0035] In some implementations of the second aspect, the determining the first gain of the kth frequency band based on the suppression sharpness includes: in a case where the suppression sharpness is less than or equal to a first threshold value, the first gain is 0; and / or in a case where the suppression sharpness is greater than the first threshold value, the first gain is calculated based on the suppression sharpness, the first threshold value, and a second threshold value, the second threshold value being used to represent a suppression capability.
[0036] The first threshold value can be threshold value 4 below. The second threshold value can be threshold value 5 below. Threshold value 5 can also be referred to as a suppression capability or a suppression gain.
[0037] In this way, when the suppression sharpness is small, it indicates that the sharpness of the kth frequency band is small and / or the energy is low, and the terminal device can not suppress the kth frequency band, thereby reducing the timbre damage.
[0038] In some implementations of the second aspect, the first gain g k satisfies the following formula:
[0039]
[0040] In the formula, sharp k is the suppression sharpness, thres sharpnes is the first threshold value, and progain i is the second threshold value.
[0041] In this way, the first gain increases with the increase of the suppression sharpness of the kth frequency band, so that the greater the sharpness of the kth frequency band, the greater the amplitude (sum of the frequency band sound pressure level distribution), the greater the suppression sharpness of the kth frequency band, the greater the first gain, and the greater the suppression of the kth frequency band by the terminal device based on the first gain, so that the sharpness of the kth frequency band can be reduced. Conversely, if the suppression sharpness of the kth frequency band is small, the first gain is smaller, and the suppression of the kth frequency band by the terminal device based on the first gain is smaller, so that the hearing of the suppressed kth frequency band will not be dull.
[0042] In addition, the second threshold is a preset value that can reflect the suppression degree or suppression ability, which can also be referred to as a suppression gain. In the case of a certain suppression sharpness, the greater the second threshold, the greater the gain calculated by the terminal device, and the greater the suppression ability of the frequency band; the smaller the second threshold, the smaller the gain calculated by the terminal device, and the smaller the suppression ability of the frequency band. Therefore, the second threshold can be a value that can reasonably suppress the sibilance by the terminal device according to experience or experiments, etc.
[0043] In combination with the second aspect, in some implementations of the second aspect, the suppression sharpness of the kth frequency band is calculated based on the first value and the second value, including: calculating a ratio of the first value and a third threshold to obtain a third value; converting the third value into a fourth value based on a preset function, the fourth value being less than or equal to the third value; and calculating a product of the fourth value and the second value, the suppression sharpness being the product of the fourth value and the second value.
[0044] The third threshold can be threshold 3 in the following. The third value can be p' in the following. The fourth value can be f(p') in the following.
[0045] In this way, the larger first value can be converted into a smaller fourth value, so as to facilitate the control of the suppression sharpness of the frequency band within a smaller fluctuation range, thereby facilitating the setting of the first threshold.
[0046] In combination with the second aspect, in some implementations of the second aspect, the first value is obtained, including: dividing the amplitude spectrum of the xth frame of audio data into E frequency bands, E being a positive integer; calculating a second frequency band sound pressure level distribution of the xth frame of audio data, the second frequency band sound pressure level distribution including the frequency band sound pressure level of each of the E frequency bands; and calculating a sum of the first frequency band sound pressure level distribution based on the second frequency band sound pressure level distribution, the first frequency band sound pressure level distribution including the frequency band sound pressure level of at least one of the E frequency bands.
[0047] The E frequency bands can be e1 frequency bands in the following. The second frequency band sound pressure level distribution can be the frequency band sound pressure level p x The E frequency bands may, for example, be divided based on 1 / 3 octave filtering.
[0048] With reference to the second aspect, in some implementations of the second aspect, in the order from low to high frequency, the at least one of the E frequency bands includes a first frequency band and a second frequency band, the first frequency band is one of the E frequency bands to which a lower limit frequency of the kth frequency band belongs, and the second frequency band is one of the E frequency bands to which an upper limit frequency of the kth frequency band belongs.
[0049] wherein the first frequency band is, for example, one of the E frequency bands indicated by h=k_fl below; and the second frequency band is, for example, one of the E frequency bands indicated by h=k_fh below. h=k_fl and h=k_fh can be the same or different.
[0050] In this way, the frequency band SPL distribution corresponding to the kth frequency band can be determined based on the frequency band SPL distribution of the E frequency bands, and then the sum of the first frequency band SPL distributions is calculated. The sum of the first frequency band SPL distributions is related to the amplitude of the kth frequency band, and can reflect the amplitude or energy of the kth frequency band.
[0051] Exemplarily, the sum Psum of the first frequency band SPL distributions satisfies the following formula:
[0052]
[0053] wherein k_fl is the lower limit frequency of the kth frequency band, h=k_fl is the index of one of the E frequency bands to which the lower limit frequency belongs, k_fh is the upper limit frequency of the kth frequency band, h=k_fh is the index of one of the E frequency bands to which the upper limit frequency belongs, h is the index of a frequency band in the E frequency bands, p(h) is the frequency band SPL of the frequency band with index h in the second frequency band SPL distribution, and the index of the frequency band in the E frequency bands increases as the frequency of the frequency band increases.
[0054] With reference to the second aspect, in some implementations of the second aspect, obtaining the second value includes: calculating a second sharpness distribution of the xth frame of audio data, the second sharpness distribution including a sharpness of each Bark frequency band; and calculating a sum of first sharpness distributions based on the second sharpness distribution, the first sharpness distribution including a sharpness of at least one Bark frequency band.
[0055] wherein the second sharpness distribution can be a Bark domain sharpness distribution as described below. The first sharpness distribution belongs to the second sharpness distribution.
[0056] In this way, the first sharpness distribution conforms to the human ear perception characteristics, and the sibilance suppression based on the first sharpness distribution makes the user's listening experience better.
[0057] Exemplarily, the sum Ssum of the first sharpness distributions satisfies the following formula:
[0058]
[0059] wherein k_fl is a lower limit frequency of the kth frequency band, k_fh is an upper limit frequency of the kth frequency band, z is an index of a Bark frequency band, z = bark() is a function of determining the index of the Bark frequency band based on the lower limit frequency or the upper limit frequency, sp(z) is a sharpness of the Bark frequency band with the index z in the second sharpness distribution.
[0060] wherein z = bark() can be a product of bark1 and 10 in formula 12.
[0061] In this way, the sum of the first sharpness distribution can reflect the sharpness of the kth frequency band heard by the human ear, and the timbre damage of the sibilance suppression based on the sum of the first sharpness distribution is smaller.
[0062] With reference to the second aspect, in some implementations of the second aspect, determining that the xth frame of audio data in the plurality of frames of audio data is sibilance includes: determining that the xth frame of audio data is sibilance based on a sum of third sharpness distributions of the xth frame of audio data and a spectral centroid of the xth frame of audio data.
[0063] wherein the sum of the third sharpness distributions can be a sum of the sharpnesses of a plurality of specific Bark frequency bands in the following.
[0064] In this way, the sum of the third sharpness distributions can reflect the sharpness of the xth frame of audio data, and the spectral centroid of the xth frame of audio data can reflect the center position of the energy distribution of the xth frame of audio data. Thus, whether the xth frame of audio data is sibilance is determined in combination with the sharpness and the energy distribution of the xth frame of audio data. In this way, the audio data with the energy distribution concentrated in the high frequency and the high sharpness can be identified as sibilance, so that the accuracy of sibilance identification is higher.
[0065] With reference to the second aspect, in some implementations of the second aspect, the third sharpness distribution includes the sharpness of each Bark frequency band in a plurality of Bark frequency bands, and the plurality of Bark frequency bands are Bark frequency bands with frequencies greater than or equal to a preset frequency.
[0066] wherein the Bark frequency bands with frequencies greater than or equal to the preset frequency can be a plurality of specific Bark frequency bands in the following.
[0067] Alternatively, the plurality of Bark frequency bands can also be Bark frequency bands with indexes greater than or equal to a preset index (bark). Since the index of the Bark frequency band increases as the frequency increases, the Bark frequency band with the index greater than or equal to the preset index can also be a Bark frequency band with a frequency higher than a certain threshold.
[0068] In this way, the third sharpness distribution can also reflect the sharpness of the high frequency part in the xth frame of audio data.
[0069] With reference to the second aspect, in some implementations of the second aspect, determining that the xth frame of audio data is a sibilance includes: determining that the xth frame of audio data is a sibilance based on the sum of the third sharpness distribution and the normalized spectral centroid of the xth frame of audio data.
[0070] In this way, the spectral centroid of the xth frame of audio data can be within a standardized range, and the influence of different sampling rates of the audio signal can be reduced. The accuracy of audio recognition is higher.
[0071] With reference to the second aspect, in some implementations of the second aspect, determining that the xth frame of audio data is a sibilance includes: calculating a product of the sum of the third sharpness distribution and the normalized spectral centroid to obtain a recognition sharpness of the xth frame of audio data; and determining that the xth frame of audio data is a sibilance in a case where the recognition sharpness is greater than a fourth threshold value.
[0072] The fourth threshold value can be threshold value 1 described below.
[0073] In this way, the greater the sum of the third sharpness distribution and the normalized spectral centroid, the greater the recognition sharpness, and the audio with high sharpness and the center position of the energy distribution at a high frequency position can be recognized as a sibilance based on the recognition sharpness, so that the accuracy of sibilance recognition is higher.
[0074] With reference to the second aspect, in some implementations of the second aspect, the method further includes: in a case where the recognition sharpness is less than or equal to the fourth threshold value, determining whether the (x-1)th frame of audio data is a sibilance; in a case where the (x-1)th frame of audio data is a sibilance, determining whether the recognition sharpness is greater than a fifth threshold value, the fifth threshold value being less than the fourth threshold value; and determining that the xth frame of audio data is a sibilance in a case where the recognition sharpness is greater than the fifth threshold value.
[0075] The fifth threshold value can be threshold value 2 described below.
[0076] In this way, for a plurality of continuous sibilances, a sibilance with a smaller amplitude (smaller recognition sharpness) can be recognized in a sibilance amplitude decline phase, which helps to further improve the accuracy of sibilance recognition.
[0077] With reference to the second aspect, in some implementations of the second aspect, the method further includes: in a case where the (x-1)th frame of audio data is not a sibilance, determining that the xth frame of audio data is not a sibilance; or, in a case where the recognition sharpness is less than or equal to the fifth threshold value, determining that the xth frame of audio data is not a sibilance.
[0078] The x-1th frame of audio data is not a sibilance, the xth frame of audio data is not a sibilance in an amplitude decreasing phase, the sharpness is less than the fourth threshold value, and it is determined that the xth frame of audio data is not a sibilance. The sharpness is less than or equal to the fifth threshold value, the sharpness of the xth frame of audio data is small or the high frequency position energy is small, and it is determined that the xth frame of audio data is not a sibilance. In this way, the non-sibilance can be accurately identified.
[0079] With reference to the second aspect, in some implementations of the second aspect, the method further includes:
[0080] The sibilance-suppressed audio data is filtered by a filter, and a first parameter of the filter is determined based on a frequency response curve of the audio playing module. The first parameter includes one or more of the following: a center frequency, a quality factor, or a gain.
[0081] The first parameter is a center frequency (FC), a Q value, and a gain of the filter processing module.
[0082] In this way, the filter processing of the filter processing module can reduce the amplification effect of the sibilance frequency range by the audio playing module such as a loudspeaker. This helps to further improve the sound quality of the audio data and improve the user experience.
[0083] In a third aspect, an audio processing apparatus is provided. The audio processing apparatus can be an electronic device, or a chip or chip system in the electronic device. The audio processing apparatus can include an audio playing unit and a processing unit. When the audio processing apparatus is an electronic device, the audio playing unit can be a loudspeaker, and the processing unit can be a processor. The audio processing apparatus can further include a storage unit, which can be a memory. The storage unit is configured to store instructions, and the processing unit is configured to execute the instructions stored in the storage unit, so that the electronic device implements the audio processing method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect. When the audio processing apparatus is a chip or chip system in the electronic device, the processing unit can be a processor. The processing unit executes the instructions stored in the storage unit, so that the electronic device implements the audio processing method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect. The storage unit can be a storage unit (e.g., a register, a cache, etc.) in the chip, or a storage unit (e.g., a read-only memory, a random access memory, etc.) in the electronic device and located outside the chip.
[0084] In a fourth aspect, an electronic device is provided, which includes a processor and a memory. The memory is configured to store code instructions, and the processor is configured to execute the code instructions to perform the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.
[0085] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program or instructions. When the computer program or instructions are executed on a computer, the computer is caused to perform the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.
[0086] In a sixth aspect, a computer program product including a computer program is provided. When the computer program is executed on a computer, the computer is caused to perform the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.
[0087] In a seventh aspect, a chip or chip system is provided, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is configured to execute a computer program or instructions to perform the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect. The communication interface in the chip can be an input / output interface, a pin, or a circuit, etc.
[0088] In a possible implementation, the chip or chip system described above in the present application further includes at least one memory, which stores instructions. The memory can be a storage unit inside the chip, such as a register, a cache, etc., or a storage unit of the chip (such as a read-only memory, a random access memory, etc.).
[0089] It should be understood that the third aspect to the seventh aspect of the present application correspond to the technical solution of the first aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding possible implementation are similar, which will not be repeated. BRIEF DESCRIPTION OF DRAWINGS
[0090] Figure 1 is a schematic diagram of an interface for an application scenario;
[0091] Figure 2A is a schematic diagram of a time-domain waveform and a frequency spectrum of audio data;
[0092] Figure 2B is a time-domain waveform diagram and a frequency spectrum diagram of audio data;
[0093] Figure 3A A standard sharpness profile and a spectrogram of an audio data;
[0094] Figure 3B A time-domain waveform, a standard sharpness profile, and a spectrogram of an audio data;
[0095] Figure 4 A schematic block diagram of a hardware architecture of a terminal device provided by an embodiment of the present application;
[0096] Figure 5 A schematic block diagram of a software architecture of a terminal device provided by an embodiment of the present application;
[0097] Figure 6 A flowchart of an audio processing method provided by an embodiment of the present application;
[0098] Figure 7A A standard sharpness profile and a recognized sharpness profile of an audio data provided by an embodiment of the present application;
[0099] Figure 7B A standard sharpness profile and a recognized sharpness profile of an audio data provided by an embodiment of the present application;
[0100] Figure 8 A process diagram of a correction and combination of a frequency band sound pressure level provided by an embodiment of the present application;
[0101] Figure 9 A diagram of a determination process of a Bark domain loudness profile provided by an embodiment of the present application;
[0102] Figure 10 A diagram of a frequency band characteristic loudness variation curve of an audio data provided by an embodiment of the present application;
[0103] Figure 11 A flowchart of a sibilance recognition method provided by an embodiment of the present application;
[0104] Figure 12A A diagram of an original sequence spectrogram and a sibilance-suppressed spectrogram provided by an embodiment of the present application;
[0105] Figure 12B An original sequence spectrogram and a sibilance-suppressed spectrogram provided by an embodiment of the present application;
[0106] Figure 13A A diagram for showing a sibilance-suppressed part provided by an embodiment of the present application;
[0107] Figure 13B A spectrogram for showing a sibilance-suppressed part provided by an embodiment of the present application. Detailed Implementation
[0108] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:
[0109] 1. Frame
[0110] In the field of audio signal processing, a frame refers to a short segment of continuous audio data. Each frame typically contains a certain number of sampling points, which represent the amplitude changes of the audio signal within that time period. A frame can be a segment of audio data, such as 10ms, 20ms, or 30ms. This segment of time is called the frame length, or the duration of each frame.
[0111] 2. Sibilance
[0112] The sound produced when airflow passes through the narrow passage between the teeth and tongue during pronunciation, such as the pronunciation of initials like "x", "s", "q" and "c".
[0113] 3. Sharpness
[0114] Sharpness typically refers to the harshness of a sound or the intensity of high-frequency components in the sound spectrum. High-sharpness sounds can be harsh and unpleasant to listen to.
[0115] 4. Clarity
[0116] Sharpness usually refers to the brightness of a sound, which is generally related to the distribution of energy in the sound's spectrum. A sound with more energy in the high-frequency range is considered brighter.
[0117] 5. Bandwidth energy
[0118] Audio signals contain energy within a specific frequency band (i.e., a certain range of frequencies), measured in dB.
[0119] For example, an audio signal can be decomposed into different frequency components using Fourier transform or other spectral analysis methods. Band energy is the energy of these frequency components within a certain band, which is then integrated or summed. For instance, an audio signal can be decomposed into several bands (such as 20Hz to 100Hz, 100Hz to 200Hz, etc.), and then the energy within each band can be calculated.
[0120] 6. Sound pressure level (SPL)
[0121] It refers to the logarithmic measure of sound pressure relative to a reference sound pressure.
[0122] Frequency Band Sound Pressure Level is the sound pressure level calculated within a certain frequency range (i.e. frequency band).
[0123] 7. Frequency Bin Amplitude
[0124] Frequency Bin Amplitude refers to the amplitude value of an audio signal at a specific frequency bin.
[0125] Relationship between Frequency Bin Amplitude and Sound Pressure Level: Sound Pressure Level can be calculated based on Frequency Bin Amplitude. Specifically, there is a logarithmic relationship between Frequency Bin Amplitude and Sound Pressure Level, i.e. Sound Pressure Level is a logarithmic representation of Frequency Bin Amplitude.
[0126] 8. Bark Scale
[0127] Also known as Bark Spectrum, Bark Scale or Bark Frequency Scale. It is a frequency scale based on human auditory perception. The Bark Frequency Scale aims to better reflect the human ear's ability to perceive different frequencies. The traditional Hertz (Hz) frequency scale is linear, while the Bark frequency scale is nonlinear, more in line with human auditory characteristics.
[0128] Bark Band is a discrete frequency interval in the Bark domain. The entire auditory frequency range (about 20 Hz to 20 kHz) is usually divided into 24 Bark bands. Each Bark band corresponds to a specific frequency range. These Bark bands are divided based on the resolution ability of the human auditory system to frequency.
[0129] Furthermore, each of the 24 Bark bands corresponds to a Bark value. The Bark value can also be understood as the index or serial number of the Bark band.
[0130] Exemplarily, the 24 Bark bands and the corresponding Bark values of each Bark band can be shown in Table 1. The frequency range of the Bark band corresponding to Bark value 1 is 20-100 Hz; the frequency range of the Bark band corresponding to Bark value 2 is 100-200 Hz, and so on.
[0131] Table 1
[0132]
[0133]
[0134] It should be noted that when the frequency ranges of two Bark bands overlap, the overlapping frequency can belong to either of the two Bark bands. For example, the frequency range of the Bark band corresponding to Bark value 1 is 20-100 Hz; the frequency range of the Bark band corresponding to Bark value 2 is 100-200 Hz, and the frequency 100 Hz can belong to the Bark band corresponding to Bark value 1 or the Bark band corresponding to Bark value 2. The present application does not make specific limitations on this.
[0135] 9、Equalizer (EQ)
[0136] An audio data processing tool used to adjust the gain of different frequency bands in audio data. It can enhance or reduce the volume of a specific frequency range to improve sound quality or achieve a specific sound effect.
[0137] 10、Multiband Dynamic Range Compressor (MBDRC)
[0138] An advanced dynamic processing tool that can compress different frequency bands of audio data separately.
[0139] 11、Threshold of Hearing
[0140] Also known as the hearing threshold, it refers to the minimum sound intensity that the human ear can hear in a very quiet environment. This minimum sound intensity varies with frequency and is usually represented by the threshold of hearing. The threshold of hearing is a frequency-dependent curve called the threshold of hearing curve, which describes the minimum audible sound pressure level at different frequencies.
[0141] 12、Excitation Level
[0142] It refers to the sound energy or loudness in a certain frequency band. It reflects the excitation degree of the sound in the frequency band to the human ear. The excitation level is usually expressed in decibels (dB).
[0143] 13、Spectral Centroid
[0144] It can be understood as the "center of gravity" of the spectrum, i.e. the center position of the energy distribution in the spectrum. The spectral centroid is obtained by calculating the weighted average of the frequency components of the spectrum.
[0145] 14、Normalized Spectral Centroid
[0146] Normalization of spectral centroid is to normalize the value of spectral centroid in a standardized range (usually between 0 and 1). This normalization can eliminate the influence of different sampling rates and frequency ranges between different audio signals, making it easier to compare.
[0147] 15、Frequency response
[0148] Also known as frequency response. It refers to the relationship between the output sound pressure level of an audio playback device (such as a speaker) at different frequencies and the frequency of the input audio signal. It is usually used to describe the performance of a speaker in the entire audio frequency spectrum.
[0149] Frequency response is usually represented in the form of a frequency response curve (or frequency response curve), with the horizontal axis (x-axis) representing frequency or frequency point (usually in hertz), and the vertical axis (y-axis) representing sound pressure level (usually in decibels).
[0150] The frequency response curve shows the output sound pressure level of an audio playback device (such as a speaker) at different frequencies, which can be used to evaluate the uniformity and accuracy of an audio playback device (such as a speaker) in the entire frequency range.
[0151] 16、Center frequency of filter
[0152] It is the main frequency parameter of the filter, indicating the position of the filter on the frequency response curve that reaches the maximum gain or minimum attenuation.
[0153] 17、Quality factor (Q value) of filter
[0154] It describes the frequency selectivity or bandwidth of the filter. The higher the Q value, the narrower the bandwidth of the filter and the stronger the selectivity; the lower the Q value, the wider the bandwidth and the weaker the selectivity.
[0155] 18、Gain of filter
[0156] It represents the degree of amplification or attenuation of the audio signal by the filter at a specific frequency.
[0157] 19、Other terms
[0158] In the embodiments of the present application, "first", "second" and the like are used to distinguish the same or similar items with basically the same function and effect. For example, the first chip and the second chip are only used to distinguish different chips, and do not limit the sequence. Those skilled in the art can understand that "first", "second" and the like do not limit the quantity and execution sequence, and "first", "second" and the like do not necessarily mean different.
[0159] It should be noted that the terms "exemplary" and "for example" are used herein to mean "serving as an example, instance, or illustration," and not "preferred" or "advantageous over other examples." The usage of these terms in this application is not intended to convey any preference or advantage for the embodiments or examples described with such terms.
[0160] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, wherein a, b, and c can be single or multiple.
[0161] 20. An electronic device
[0162] The electronic device of the embodiments of the present application can include a handheld device, a vehicle-mounted device, etc. with a face recognition function. For example, some electronic devices are: a mobile phone, a tablet computer, a palm computer, a notebook computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with a wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a wearable device, a terminal device in a 5G network, or a terminal device in a future evolved public land mobile network (PLMN), etc. The embodiments of the present application are not limited thereto.
[0163] By way of example and not limitation, in the embodiments of the present application, the electronic device can also be a wearable device. The wearable device can also be referred to as a wearable smart device, which is a general term for devices that are designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing, and shoes, etc. The wearable device is a portable device that is directly worn on the body or integrated into the clothes or accessories of the user. The wearable device is not only a hardware device, but also a device that realizes powerful functions through software support and data interaction and cloud interaction. The general wearable smart device includes a device with full functions and large size, which can realize complete or partial functions without relying on a smart phone, such as a smart watch or smart glasses, etc., and a device that focuses on a certain application function and needs to be used in cooperation with other devices, such as a smart phone, such as various smart wristbands and smart jewelry for monitoring vital signs, etc.
[0164] Furthermore, in this embodiment of the application, the electronic device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.
[0165] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0166] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.
[0167] With the development of terminal technology, terminal devices are offering increasingly richer functions. For example, users can listen to music, record and play recordings, watch videos, play games, and make calls using terminal devices. However, the audio played by terminal devices may contain sibilance, resulting in a poor subjective listening experience for users. Therefore, before playing audio data, the acquired audio data can be processed, such as sibilance suppression, to improve the sound quality of the played audio data and provide users with a better audio-visual experience.
[0168] For example, Figure 1 This is a schematic diagram of an interface for an application scenario provided in an embodiment of this application. For example... Figure 1 As shown in interface (a), interface (a) can be understood as the music playback interface of a music application. Interface (a) includes control 101. In response to the user's input operation on control 101, such as clicking control 101, the terminal device can switch from interface (a) to interface (b), and the terminal device will start playing the audio data of the song.
[0169] The audio data of the song played when the terminal device displays interface (b) is processed audio data obtained through sibilance suppression and the like. That is, in response to the input operation of the user on the control 101, the terminal device can obtain the audio data of the song. And the terminal device can perform sibilance suppression and the like on the obtained audio data, thereby obtaining the processed audio data.
[0170] Sibilance can be understood as part of the pronunciation in the human voice. For example, when the initial consonant is "x", "s", "q" and "c", the tip of the tongue is pressed against the upper incisors, and the airflow and the teeth rub to produce sibilance. Exemplarily, in combination with Figure 2A With Figure 2B , wherein Figure 2B is a time-domain waveform diagram and a spectrum diagram of a piece of audio data. Figure 2A is a time-domain waveform and a spectrum diagram. Figure 2B
[0171] , wherein Figure 2A The selected part of the block 201 is the time-domain waveform and the spectrum diagram corresponding to the "x" pronunciation. Figure 2B The selected part of the block 205 is the time-domain waveform and the spectrum diagram corresponding to the "x" pronunciation. Figure 2A The selected part of the block 202 is the time-domain waveform and the spectrum diagram corresponding to the "c" pronunciation. Figure 2B The selected part of the block 206 is the time-domain waveform and the spectrum diagram corresponding to the "c" pronunciation. Figure 2A The selected part of the block 203 is the time-domain waveform and the spectrum diagram corresponding to the "x" pronunciation. Figure 2B The selected part of the block 207 is the time-domain waveform and the spectrum diagram corresponding to the "x" pronunciation. Figure 2A The selected part of the block 204 is the time-domain waveform and the spectrum diagram corresponding to the "s" pronunciation. Figure 2B The selected part of the block 208 is the time-domain waveform and the spectrum diagram corresponding to the "s" pronunciation.
[0172] In combination with Figure 2A With Figure 2B , in the spectrum diagram or the spectrum diagram, the vertical coordinate is the frequency, and the frequency uniformly increases from bottom to top. And, in the spectrum diagram of Figure 2A , the lighter the color (white part) represents the higher energy; in the spectrum diagram of Figure 2B , the more yellow the color represents the higher energy.
[0173] From the spectrum diagram of the selected part of the block 201, the block 202, the block 203 and the block 204, and the spectrum of the selected part of the block 205, the block 206, the block 207 and the block 208, it can be seen that in the spectrum of "x", "c" and "s", the energy is higher at the position of higher frequency, then the pronunciation of "x", "c" and "s" heard by the user is more sharp.
[0174] Furthermore, a frame of audio data is typically a segment of tens of milliseconds, such as 10ms, 20ms, or 30ms, while a sibilance is usually a segment of 40-400ms. Therefore, sibilances in an audio data segment usually appear in consecutive frames. For example, combining... Figure 2A The portions circled in boxes 201, 202, 203, and 204 are all multi-frame audio data, each of which may include multiple frames of sibilance. Furthermore, combining the schematic diagrams of the time-domain waveforms of the multi-frame sibilances circled in boxes 201, 202, 203, and 204, and the time-domain waveforms of the multi-frame sibilances circled in boxes 205, 206, 207, and 208, it can be seen that the amplitude of consecutive multi-frame sibilances typically exhibits a slow increasing trend initially, followed by a slow decreasing trend.
[0175] exist Figure 2A The schematic diagram of the spectrum shown is consistent with Figure 2B The spectrum diagram shown compares the spectra of dental consonants (the sounds of "x", "c", and "s") with those of non-dental consonants (e.g., the uncircled portions). It can be seen that for dental consonants, higher frequencies have higher energy. For non-dental consonants, higher frequencies have lower energy.
[0176] Typically, sibilance has a frequency range of 4kHz-12kHz, which is in the mid-to-high frequency band and falls within the range where the human ear is most sensitive. Therefore, when audio is played on a terminal device, if the audio playback device (such as a speaker) amplifies the audio energy, the user will hear a high-pitched sound, potentially resulting in a poor user experience.
[0177] Therefore, to provide good sound quality, terminal devices can suppress sibilance in audio data before playing it. Suppressing sibilance in audio data requires two processes: sibilance recognition and sibilance suppression.
[0178] Taking the a-th frame of audio data in a segment of audio data as an example, the terminal device can currently identify whether each frame of audio data in the segment of audio data is sibilant in the following way.
[0179] S1. The terminal device calculates the loudness distribution of the Bark domain of the audio data in frame a.
[0180] The loudness distribution in the Bark domain can also be understood as the loudness of each Bark frequency band corresponding to the audio data of frame a, that is, the characteristic loudness of each Bark frequency band. For example, the 24 Bark frequency bands shown in Table 1 above can be referred to. Each of the 24 Bark frequency bands corresponds to a loudness level.
[0181] S2, the terminal device calculates a sum of Bark domain loudness distributions of the a-th frame of audio data, to obtain a total loudness N of the Bark domain loudness distributions of the a-th frame of audio data a .
[0182] Exemplarily, the total loudness N of the Bark domain loudness distributions can satisfy the following formula: a
[0183] N a =∫1 24 n'(z)dz (formula 1);
[0184] wherein z is a Bark value, that is, an index of a Bark band; n'(z) is a loudness of the Bark band with the index z.
[0185] S3, based on the Bark domain loudness distributions and the total loudness N a , the terminal device calculates a standard sharpness (S) of the a-th frame of audio data.
[0186] Exemplarily, the standard sharpness (S) can satisfy the following formula:
[0187]
[0188] wherein N a is the total loudness of the Bark domain loudness distributions; n'(z) is the loudness of the Bark band with the index z; g(z) is an excitation function, which can satisfy the following formula 3.
[0189]
[0190] S4, the terminal device determines whether the standard sharpness is greater than or equal to a preset threshold 1.
[0191] In a case where the standard sharpness is greater than or equal to the preset threshold 1, the terminal device determines that the a-th frame of audio data is a sibilance;
[0192] In a case where the standard sharpness is less than the preset threshold 1, the terminal device determines that the a-th frame of audio data is not a sibilance.
[0193] Based on the above manner, the terminal device can identify a sibilance in a piece of audio data. Further, the terminal device can perform sibilance suppression through an equalizer (EQ) or a multiband dynamic range compressor (MBDRC).
[0194] However, the above manner of processing audio data can have the following problems.
[0195] Problem one, the accuracy of sibilance identification can be poor.
[0196] In combination with the above formula 2, the standard sharpness is the weighted sum of the sharpness of each Bark band of the a-th frame of audio data. Thus, when the sharpness of the Bark band with lower frequency is larger in the a-th frame of audio data, the standard sharpness calculated by the terminal device can also be larger. Thus, the terminal device judges that the a-th frame of audio data is a sibilance based on the standard sharpness being larger than the preset threshold 1. However, sibilance usually refers to audio data with medium-high frequency and large energy. Therefore, such a manner can cause the accuracy of sibilance recognition to be poor.
[0197] In addition, when the audio data has a complex sound source, the accuracy of sibilance recognition based on the preset threshold 1 can be poor. For example, in a complex scene such as music, the audio data contains a large number of musical instruments and vocals. When the loudness of the low-frequency musical instrument is large, the standard sharpness corresponding to the musical instrument sound calculated is also large. If sibilance recognition is performed based on a higher preset threshold 1, the situation of identifying the musical instrument sound as sibilance can be reduced, but it can cause some sibilance in other audio (for example, vocals) to be misrecognized or not recognized completely; if sibilance recognition is performed based on a lower preset threshold 1, because the standard sharpness corresponding to the musical instrument sound is large, more musical instrument sounds are also identified as sibilance.
[0198] Exemplarily, in combination with Figure 3A With Figure 3B . Figure 3B The time-domain waveform diagram, the standard sharpness distribution, and the frequency spectrum diagram of a piece of audio data are shown, Figure 3A is Figure 3B a schematic diagram of the standard sharpness distribution and the frequency spectrum diagram in
[0199] As shown in Figure 3A and Figure 3B , the piece of audio data includes audio data 1 of a piano, audio data 2 of a mixed sound source (vocal and instrument sound), and audio data 3 of a drum set.
[0200] In Figure 3A the frequency spectrum diagram shown, the white part represents a part with high energy. In Figure 3B the frequency spectrum diagram shown, the closer to yellow, the higher the energy. That is, Figure 3A the white part in Figure 3B is a schematic of the part with high energy in
[0201] As shown in Figure 3A and Figure 3BAs shown, in combination with the spectrum of the audio data 1, it can be seen that the energy of the audio data 1 is generally low; in combination with the spectrum of the audio data 2, it can be seen that in the audio data 2, the energy of part of the audio data is high, and the energy of part of the audio data is low; in combination with the spectrum of the audio data 3, it can be seen that in the audio data 3, the energy of the low-frequency audio data is high.
[0202] Based on this, the standard sharpness of each frame of audio data in the audio data 1 is generally low; the standard sharpness of each frame of audio data in the audio data 2 is uneven; the standard sharpness of each frame of audio data in the audio data 2 is generally high.
[0203] When the terminal device identifies the sibilance based on the preset threshold 1, it is assumed that the preset threshold 1 is large, for example, x1. Since the standard sharpness distribution of the audio data 1 is less than x1, the terminal device cannot identify the sibilance in the audio data 1. Moreover, for the audio data 2, the terminal device can only identify part of the sibilance with high standard sharpness in the audio data 2, and may miss part of the sibilance with relatively small standard sharpness.
[0204] That is, since the sibilance usually appears continuously in multiple frames, and the amplitude of the multiple frames of continuous sibilance usually changes in a trend of slow growth and then slow decrease, the standard sharpness distribution of the multiple frames of continuous sibilance is also in a trend of low to high and then high to low. In this way, when the terminal device identifies the sibilance of the audio data 2 based on the large preset threshold 1 (x1), the sibilance identified by the terminal device may be several frames of sibilance with large standard sharpness in the multiple frames of continuous sibilance of the audio data 2, and the multiple frames of sibilance with small standard sharpness are not identified. Therefore, when the terminal device performs sibilance suppression subsequently, it is also to suppress several frames of sibilance with large standard sharpness in the multiple frames of continuous sibilance. Therefore, the user may hear that the audio data has a mutation, resulting in poor user experience.
[0205] In addition, it is assumed that the preset threshold 1 is small, for example, x2. Since the standard sharpness distribution of the audio data 3 is mostly greater than x3, the terminal device may identify most of the audio data in the audio data 3 as sibilance. Therefore, the terminal device suppresses most of the audio data in the audio data 3, resulting in distortion of the timbre of the audio data 3, so that the user hears distorted drum sounds, and the user experience is poor.
[0206] Therefore, for the audio data shown in the figure, if the preset threshold 1 is high, the terminal device may not be able to identify all the sibilance in the audio data 2; if the preset threshold 2 is low, the terminal device may identify the sound of the musical instrument as sibilance. Therefore, the accuracy of sibilance identification based on the preset threshold 1 is poor. Figure 3A And Figure 3B The audio data shown in the figure. If the preset threshold 1 is high, the terminal device may not be able to identify all the sibilance in the audio data 2; if the preset threshold 2 is low, the terminal device may identify the sound of the musical instrument as sibilance. Therefore, the accuracy of sibilance identification based on the preset threshold 1 is poor.
[0207] Question 2: sibilance suppression may cause timbre distortion.
[0208] When the terminal device performs sibilance suppression through EQ or MBDRC, the measurement index is the energy (also referred to as amplitude, in dB) of each frequency band.
[0209] However, for different frequency bands, the frequency ranges are different, and the sensitivity of human ears to perception is different. Using only the amplitude as the judgment standard may cause the terminal device to suppress the frequency band with low sensitivity of human ears to perception or the frequency band with low sharpness, so that the audio timbre heard by the user is distorted and the listening experience is dull.
[0210] For example, it is assumed that the amplitudes of sibilance 1 and sibilance 2 are similar, but the sharpness of sibilance 2 is higher, and the listening experience of sibilance 2 is also more sharp; the sharpness of sibilance 1 is lower, and sibilance 1 does not sound sharp.
[0211] In this way, if only the amplitude is used as the judgment basis, the amplitudes of sibilance 1 and sibilance 2 are similar, and the terminal device may suppress sibilance 1 and sibilance 2. Since the sharpness of sibilance 2 is higher, sibilance 2 is no longer sharp after being suppressed, so that the listening experience of the user is better; but since the sharpness of sibilance 1 is lower, sibilance 1 may be suppressed after being suppressed, causing timbre damage, so that the suppressed sibilance 1 is not clear and bright enough, and the listening experience of the user is poor.
[0212] For example, for the sibilance in wind sound or flat friction sound “v” and “f”, the energy of each frequency band is uniformly distributed, and in general, the wind sound or flat friction sound itself is not sharp for the user. If the terminal device suppresses the sibilance in the wind sound or flat friction sound based on the amplitude, timbre damage may be caused, so that the timbre of the wind sound or flat friction sound heard by the user is distorted and the listening experience is dull.
[0213] Therefore, the present application provides an audio processing method. After the terminal device identifies at least one frame of sibilance, for each frame of sibilance, a preset frequency range (for example, sibilance frequency band 4 kHz-12 kHz) in each frame of sibilance can be divided into y frequency bands (y is an integer greater than 1). The terminal device can perform sibilance suppression on each frequency band based on the sharpness distribution and the frequency band sound pressure level distribution corresponding to each frequency band in the y frequency bands.
[0214] The sharpness distribution can reflect the sharpness of the frequency band, and the frequency band sound pressure level distribution can reflect the energy level of the frequency band. In this way, by combining the sharpness distribution and the frequency band sound pressure level distribution, the terminal device can identify the frequency band with high sharpness and high energy in the y frequency bands, and then suppress the frequency band, so that the probability of the terminal device performing sibilance suppression on the frequency band with low sharpness is small, and the timbre damage of the audio data caused by the terminal device performing sibilance suppression is small.
[0215] The following will be described in combination with Figure 4 to Figure 13BThe technical solutions of this application and how they solve the aforementioned technical problems are described in detail with specific embodiments. The following specific embodiments can be implemented independently or in combination with each other. Identical or similar concepts or processes may not be described again in some embodiments.
[0216] The embodiments shown in this application can be executed by a terminal device. This terminal device can be the terminal device itself, a chip, chip system, or processor that supports the implementation of audio processing methods in the terminal device, or a logic module or software capable of implementing all or part of the terminal device's functions.
[0217] It should be understood that the specific form and quantity of each device shown in the embodiments of this application are merely examples and should not constitute any limitation on the implementation of the methods provided in this application.
[0218] To facilitate understanding this solution, we will first combine... Figure 4 The hardware structure of terminal device 400 is described.
[0219] Figure 4 This is a schematic block diagram of the hardware architecture of the terminal device 400 provided in an embodiment of this application. Figure 4 As shown, the terminal device 400 may include a processor 410, an external memory interface 420, an internal memory 421, a universal serial bus (USB) interface 430, a charging management module 440, a power management module 441, a battery 442, an antenna 1, an antenna 2, a mobile communication module 450, a wireless communication module 460, an audio module 470, a sensor module 480, a button 490, an indicator 492, a camera 493, and a display screen 494, etc.
[0220] The sensor module 480 may include, but is not limited to, one or more of the following sensors: pressure sensor, gyroscope sensor, barometric pressure sensor, magnetic sensor, accelerometer, distance sensor, proximity sensor, fingerprint sensor, temperature sensor, touch sensor, ambient light sensor, and bone conduction sensor, etc.
[0221] The audio module 470 may include a speaker 470A, a receiver 470B, a microphone 470C, and a headphone jack 470D. The speaker can be used to play audio data from the processor 410.
[0222] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the terminal device 400. In other embodiments of the present application, the terminal device 400 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0223] The processor 410 can include one or more processing units, for example: the processor 410 can include an application processor (AP), a modem, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Optionally, a memory can also be provided in the processor 410 for storing instructions and data. Different processing units can be independent devices or can be integrated into one or more processors.
[0224] Optionally, the processor 410 can obtain audio data and perform sibilance identification on the audio data to determine at least one frame of sibilance in the audio data; the processor 410 can also perform sibilance suppression on the at least one frame of sibilance to obtain sibilance-suppressed audio data; and the processor 410 can be preconfigured with a peak filter, parameters (such as center frequency, Q value, and gain, etc.) of the peak filter can be determined based on a frequency response curve of the loudspeaker 470A, and the processor 410 can filter the sibilance-suppressed audio data through the peak filter to obtain filtered audio data. Then, the processor 410 can transmit the filtered audio data to the audio module 470 so that the audio module 470 (such as the loudspeaker 470A) plays the filtered audio data.
[0225] The wireless communication function of the terminal device 400 can be implemented through the antenna 1, the antenna 2, the mobile communication module 450, the wireless communication module 460, the modem, and the baseband processor, etc.
[0226] The terminal device 400 implements the display function through the GPU, the display screen 494, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 494 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 410 can include one or more GPUs that execute program instructions to generate or change display information.
[0227] The display screen 494 is configured to display images, videos, and the like. In some embodiments, the terminal device 400 can include one or N display screens 494, where N is a positive integer greater than 1.
[0228] The external memory interface 420 can be configured to connect an external memory card, such as a Micro SD card, to extend the storage capability of the terminal device 400. The internal memory 421 can be configured to store computer executable program code, including instructions.
[0229] The following describes the software architecture of the terminal device 400, taking the Android system as an example. Figure 5
[0230] It should be understood that the software system of the terminal device 400 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture.
[0231] The layered architecture divides the software into several layers, each of which has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom, the application layer, the Java framework layer, the system service layer, the system runtime and system library, the hardware abstraction layer, and the kernel layer.
[0232] 1. Application layer
[0233] The application layer can include a series of application packages. The application packages may, for example, include a first application, through which a user can interact with the terminal device 400, and through which the terminal device 400 can play audio. The first application may, for example but not limited to, be a call application, a music application, a game application, a video application, a social application, and the like. For example, the first application is a music application, and the terminal device 400 can play a song through the first application. This process can be described with reference to the description of the music application in the foregoing. Figure 1
[0234] During the process of playing audio through the first application, the first application can instruct the audio processing module in the FWK layer to obtain and process audio data.
[0235] 2. Java framework layer, also referred to as application framework (FWK) layer
[0236] The Java framework layer can provide an application programming interface (API) and a programming framework for applications in the application layer. The Java framework layer includes some pre-defined functions. The Java framework layer can include, but is not limited to, an activity manager, an input manager, a phone manager, a window manager, and an audio processing module.
[0237] The window manager is configured to manage window programs. The window manager can acquire a display screen size, determine whether there is a status bar, lock a screen, and capture a screen.
[0238] The content provider is configured to store and acquire data, and make the data accessible to applications. The data can include videos, images, audios, dialed and received calls, browsing history and bookmarks, and a phonebook.
[0239] The view system includes visual controls, such as a control for displaying text and a control for displaying pictures. The view system can be used to build an application. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying pictures.
[0240] The phone manager is configured to provide a communication function of the terminal device 400. For example, the phone manager is configured to manage a call state (including call connection and call hang-up).
[0241] The audio processing module is configured to acquire audio data from the application layer or other layers, and sequentially perform tooth sound identification, tooth sound suppression, and filtering processing on the audio data to obtain filtered audio data. The audio processing module is further configured to write the filtered audio data into a buffer.
[0242] It should be understood that the audio processing module can be a software module, or can include a plurality of software modules. For example, the audio processing module can include a tooth sound identification and suppression module, and a filtering processing module.
[0243] The tooth sound identification and suppression module is configured to perform tooth sound identification and tooth sound suppression on the audio data, and transmit the tooth sound suppressed audio data to the filtering processing module.
[0244] The filtering processing module, which can also be referred to as a peak filter, is configured to perform filtering processing on the tooth sound suppressed audio data to obtain filtered audio data.
[0245] It should be noted that the audio processing module (or sibilance recognition and suppression module) can perform sibilance recognition based on the sharpness distribution 1 and the normalized spectral centroid of the audio data, and / or can perform sibilance suppression based on the sharpness distribution 2 and the frequency band sound pressure level distribution corresponding to each frequency band in the sibilance, which can be referred to the description below.
[0246] 3、System services layer
[0247] The system services layer contains various system services that interact with the application framework layer through the Binder IPC mechanism. The system services layer manages and coordinates system resources and provides core functionalities for applications.
[0248] The system services layer includes an audio flinger. The audio flinger can read the filtered audio data from the buffer and perform mixing, sound effect processing, etc. on the filtered audio data to obtain the audio data a.
[0249] 4、Hardware abstraction layer
[0250] The hardware abstraction layer (HAL) provides hardware-related abstract interfaces, so that the system services and libraries of the upper layer can interact with different hardware.
[0251] The HAL includes a Bluetooth HAL, a camera HAL, an audio HAL, and a sensor HAL, etc. Each library module can provide an interface for the corresponding hardware component implementation.
[0252] The audio HAL can obtain the audio data a from the audio flinger. The audio HAL can further process the audio data a based on the hardware requirements for playing the audio data, such as format conversion, sampling rate adjustment, volume control, etc., to obtain the audio data b. The audio HAL can transmit the audio data b to the hardware (such as a speaker, an earpiece, or a headset, etc.) in the terminal device through the kernel layer, so that the hardware in the terminal device plays the audio.
[0253] 6、Kernel, also known as Linux Kernel
[0254] The kernel layer is the layer between hardware and software. The kernel layer is used to drive hardware to work. The kernel layer includes display drivers, screen drivers, graphics processing unit (GPU) drivers, camera drivers, and audio drivers, etc., which are not limited by the embodiments of the present application.
[0255] The audio driver can obtain audio data b from the audio HAL, and can drive a digital-to-analog converter to play the audio data b. Illustratively, the audio driver can convert the audio data b from an electrical signal (or a digital signal) to an acoustic signal (or an analog signal) through the digital-to-analog converter, and play the acoustic signal through an output device such as a speaker or a headset.
[0256] It should be understood that Figure 5 The layered architecture shown is merely an example, and the operating system of the terminal device 400 can also include more or fewer layers. For example, the operating system also includes an Android Runtime and a system library, etc. And each layer can include more or fewer software modules, for example, the FWK layer can also include an activity manager, etc.
[0257] In combination with Figure 5 In combination with the software architecture shown, the terminal device can Figure 6 The method 600 shown performs audio processing. This is described in detail below.
[0258] Figure 6 A flowchart of an audio processing method 600 provided by an embodiment of the present application is shown. As shown in Figure 6 The method 600 includes the following steps:
[0259] S601, the sibilance identification and suppression module obtains M frames of audio data.
[0260] Wherein, M is a positive integer. The M frames of audio data can be a time domain signal.
[0261] S602, the sibilance identification and suppression module converts the M frames of audio data from a time domain signal to a frequency domain signal.
[0262] Wherein, the frequency domain signal of the M frames of audio data can also be understood as the frequency spectrum or amplitude spectrum of the M frames of audio data.
[0263] Illustratively, the sibilance identification and suppression module can convert the M frames of audio data from a time domain signal to a frequency domain signal through fast Fourier transform (FFT) or the like.
[0264] The sibilance identification and suppression module can sequentially identify the sibilance of each frame of audio data in the M frames of audio data. Hereinafter, taking the xth frame of audio data as an example, the process (S603 to S608) of sibilance identification by the sibilance identification and suppression module is described in detail.
[0265] As x takes all integers between 1 and M, the sibilance identification and suppression module can identify the sibilance in the M frames of audio data.
[0266] S603, the sibilance recognition and suppression module calculates the frequency band sound pressure level p of the xth frame of audio data (frequency domain signal). x .
[0267] Among them, the frequency band sound pressure level p x This can be understood as a sequence of sound pressure levels across multiple frequency bands. That is, the entire frequency band of the x-th frame of audio data can be divided into e1 frequency bands, each of which corresponds to a sound pressure level, for a total of e1 frequency band sound pressure levels. These e1 frequency band sound pressure levels constitute the frequency band sound pressure level p of the x-th frame of audio data. x e1 is a positive integer.
[0268] For example, the sibilance recognition and suppression module can filter the spectrum of the x-th frame of audio data using a 1 / 3 octave band filter to obtain e1 frequency bands, where e1 is, for example, 28. The width of each of the e1 frequency bands is 1 / 3 times the center frequency of the 1 / 3 octave band filter.
[0269] Taking the sound pressure level of one frequency band out of the e1 frequency bands as an example, the frequency band sound pressure level can be calculated using the following formula:
[0270]
[0271] Where p0 is the sound pressure level of this frequency band; id1 is the lower limit frequency of this frequency band; id2 is the upper limit frequency of this frequency band; mag i p represents the frequency amplitude. re1f The standard sound pressure level, also known as the reference sound pressure level, can be 20 micropascals (μPa).
[0272] It should be understood that for each of the e1 frequency bands, the band sound pressure level can be calculated in the above manner, thus obtaining the e1 frequency band sound pressure levels. For example, if e1 can be 28, then through the above process, the sibilance recognition and suppression module can obtain the band sound pressure level of each of the 28 frequency bands.
[0273] It should be noted that the frequency band division shown above using 1 / 3 octave band filtering is only an example. The sibilance recognition and suppression module can also divide the frequency band using 1 / 1 octave band filtering, etc., so that the divided frequency band can more accurately describe the spectral characteristics of the noise.
[0274] S604, based on frequency band sound pressure level p x The sibilance recognition and suppression module calculates the Bark domain loudness distribution of the x-th frame of audio data.
[0275] The Bark domain loudness distribution can include a characteristic loudness of each of the Bark bands, i.e., the Bark domain loudness distribution can be understood as including a set or sequence of characteristic loudnesses.
[0276] S605, based on the Bark domain loudness distribution of the xth frame of audio data, the tooth identification and suppression module calculates a sharpness distribution 1 of the xth frame of audio data.
[0277] The sharpness distribution 1 can include a sharpness of each of the 24 Bark bands.
[0278] The tooth identification and suppression module may, for example, calculate the sharpness distribution 1 of the xth frame of audio data by the following way:
[0279] First, the tooth identification and suppression module calculates a sum of the Bark domain loudness distribution of the xth frame of audio data, i.e., a total loudness (N sum ).
[0280] The total loudness N sum of the Bark domain loudness distribution may, for example, satisfy the following formula:
[0281]
[0282] Wherein n(z) is the characteristic loudness of the Bark band with index z.
[0283] Then, the tooth identification and suppression module calculates the sharpness distribution 1 based on the total loudness and the Bark domain loudness distribution.
[0284] For the xth frame of audio data, the sharpness of the Bark band with index z in the Bark bands may satisfy the following formula:
[0285]
[0286] Wherein sp(z) is the sharpness of the Bark band with index z, n(z) is the characteristic loudness of the Bark band with index z, N sum is the total loudness, and g(z) is an excitation function.
[0287] The excitation function may satisfy the following formula:
[0288]
[0289] In this way, when z traverses 1 to 24, the tooth identification and suppression module can determine the sharpness of each of the 24 Bark bands, i.e., the sharpness distribution 1 of the xth frame of audio data.
[0290] S606, the tooth identification and suppression module calculates a normalized spectral centroid NSC of the xth frame of audio data.
[0291] Exemplarily, the normalized spectral centroid NSC of the xth frame of audio data satisfies the following formula:
[0292]
[0293] wherein n is the nth frequency point, N is the total number of frequency points included in the spectrum of the xth frame of audio data, f(n) is the frequency corresponding to the nth frequency point, E(n) is the frequency point energy of the nth frequency point, and fs is the sampling rate.
[0294] S607, the sibilance recognition and suppression module calculates the recognition sharpness of the xth frame of audio data based on the normalized spectral centroid and the sharpness distribution 1 of the xth frame of audio data.
[0295] wherein the recognition sharpness can be understood as the product of the normalized spectral centroid and the sum of the sharpnesses of a plurality of specific Bark frequency bands. The plurality of specific Bark frequency bands may, for example, be a plurality of Bark frequency bands with high frequencies, for example, a plurality of Bark frequency bands with z being 17 to 24, and the like.
[0296] Exemplarily, when the plurality of specific Bark frequency bands are a plurality of Bark frequency bands with z being 17 to 24, the recognition sharpness of the xth frame of audio data satisfies the following formula:
[0297]
[0298] wherein NSC is the normalized spectral centroid of the xth frame of audio data, and sp(z) is the sharpness of the Bark frequency band with index z.
[0299] It should be understood that in formula 9, 17 can also be replaced by other values, for example, 16 or 18, and the like. This does not constitute a limitation on the embodiments of the present application.
[0300] By combining the sum of the sharpness distribution and the normalized spectral centroid to calculate the recognition sharpness, and performing sibilance recognition based on the recognition sharpness, in the audio recognition process, not only the sharpness of the audio data is considered, but also the high or low frequency of the position with high energy in the audio data is considered, so that the terminal device can recognize the sibilance with high energy concentrated in the medium-high frequency position and high sharpness. Thus, the discrimination degree of sibilance and non-sibilance spectral features is improved, which is beneficial to improving the recognition effect in complex scenes such as music and video.
[0301] S608, the sibilance recognition and suppression module judges whether the xth frame of audio data is sibilance based on the recognition sharpness of the xth frame of audio data.
[0302] Exemplarily, when the recognition sharpness of the xth frame of audio data is greater than a threshold 1, it can be judged that the xth frame of audio data is sibilance; when the recognition sharpness of the xth frame of audio data is less than or equal to the threshold 1, it can be judged that the xth frame of audio data is not sibilance.
[0303] The threshold 1 can be a preset value.
[0304] It should be understood that, since the multiple specific Bark frequency bands are high frequency bands, the identification sharpness of the xth frame of audio data can reflect the sharpness of the high frequency bands; in addition, by normalizing the spectral centroid, it can also reflect whether the position with high energy in the xth frame of audio data is high in frequency. Moreover, since the normalized spectral centroid is normalized, the influence of different audio signals due to different sampling rates and frequency ranges can be eliminated, thereby facilitating comparison.
[0305] Therefore, by identifying the sibilance through the identification sharpness, the sharpness of the high frequency bands in the audio data and the frequency of the position with high energy are comprehensively considered. When the identification sharpness of the xth frame of audio data is large, it indicates that the xth frame of audio data is high in sharpness, and the position with high energy is high in frequency, which means that the xth frame of audio data has a high probability of being sibilance, so that the xth frame of audio data can be determined to be sibilance.
[0306] Compared with identifying sibilance through the standard sharpness, the identification sharpness considers the sharpness of the high frequency bands and the frequency of the position with high energy, so that the accuracy of sibilance identification is higher. In addition, since the normalized spectral centroid is normalized, the influence of different frequency ranges of different audio data can be eliminated, so that the identification sharpness of different audio data can be distributed at a relatively stable level, thereby making the sibilance identification based on the identification sharpness applicable to various types of audio data, such as instrument sound and human voice.
[0307] Exemplarily, in combination with Figure 7A and Figure 7B wherein Figure 7B shows the time domain waveform diagram, the standard sharpness distribution, the identification sharpness distribution and the spectrum diagram of a piece of audio data. Compared with Figure 3B , Figure 7B the identification sharpness distribution is added in the standard sharpness distribution. Figure 7A for the Figure 7B shows the standard sharpness distribution and the identification sharpness distribution. Figure 7A the standard sharpness distribution in Figure 3A is the same as the standard sharpness distribution shown in
[0308] in combination with Figure 7A or Figure 7BComparing the sharpness distribution with the standard sharpness distribution, it can be seen that the sharpness distribution of this audio data is relatively more concentrated, making it easier to identify sibilance based on a preset value (threshold 1). Furthermore, referring to the standard sharpness distribution and the sharpness distribution of audio data 1 and audio data 3, it can be seen that the sharpness distribution of audio data 1 and audio data 3 is almost stable at around 0. Thus, when identifying sibilance based on sharpness, the probability of identifying audio data emitted by musical instruments as sibilance is low, resulting in a higher accuracy rate for sibilance recognition.
[0309] Based on the above embodiments, S604 can be implemented in the following manner, namely, based on the frequency band sound pressure level p x The sibilance recognition and suppression module can determine the Bark domain loudness distribution of the x-th frame audio data in the following way.
[0310] S01, the sibilance recognition and suppression module simulates the filtering characteristics of the human ear, correcting and merging the band sound pressure level of the low-frequency bands in the e1 frequency bands to obtain the band sound pressure level of each band in the e2 frequency bands, where e2 is a positive integer less than e1. The band sound pressure level of each band in the e2 frequency bands constitutes a new band sound pressure level p. x New frequency band sound pressure level p x It includes sound pressure levels in two frequency bands.
[0311] like Figure 8 As shown, assuming that the frequency bands are arranged from low to high, the sound pressure level of the first frequency band is p1, the sound pressure level of the second frequency band is p2, and so on, with the sound pressure level of the eth frequency band being pe.
[0312] The sibilance recognition and suppression module merges the first to sixth frequency bands to obtain a new first frequency band, and the frequency range of the new first frequency band can be 25-80Hz. Correspondingly, the band sound pressure level (p1') of the new first frequency band is the sum of the band sound pressure levels of the first to sixth frequency bands, that is, p1' = p1 + p2 + p3 + p4 + p5 + p6.
[0313] The sibilance recognition and suppression module merges frequency bands 7 through 9 to obtain a new second frequency band, and the frequency range of the new second frequency band can be 100-160Hz. Correspondingly, the band sound pressure level (p2') of the new second frequency band is the sum of the band sound pressure levels of the 7th frequency band and the band sound pressure levels of the 9th frequency band, that is, p2' = p7 + p8 + p9.
[0314] The sibilance recognition and suppression module combines the 10th frequency band and the 11th frequency band to obtain a new 3rd frequency band, and the frequency range of the new 3rd frequency band can be 200-250 Hz. Correspondingly, the band sound pressure level (p3') of the new 3rd frequency band is the sum of the band sound pressure level of the 10th frequency band and the band sound pressure level of the 11th frequency band, i.e., p3' = p10 + p11.
[0315] The band sound pressure levels of the 12th frequency band to the e1th frequency band are not changed. In this way, the e1 frequency bands are combined into e2 frequency bands, and the band sound pressure levels p x are also updated to the new band sound pressure levels p x . Wherein, e1 can be 20 for example, and e2 can be 20 for example.
[0316] S02, based on the new band sound pressure levels p x of the e2 frequency bands, the sibilance recognition and suppression module calculates the characteristic loudness of each frequency band in the e2 frequency bands.
[0317] Taking one frequency band in the e2 frequency bands as an example, based on the band sound pressure level p0 of the one frequency band, the sibilance recognition and suppression module can calculate the characteristic loudness N0 of the one frequency band through the following formula. p0 is one of the new band sound pressure levels.
[0318]
[0319] Wherein, E TQ is the quiet listening area excitation level, E0 is the reference sound intensity I0 is the excitation level when the sound intensity is 10 -12 watts per square meter (W / m 2 ), and E is the excitation level of the one frequency band. In calculating the characteristic loudness N1, E can be determined based on the band sound pressure level p0 of the one frequency band, for example, E = 10 0.1×p0 and the like.
[0320] Wherein, the unit of the characteristic loudness can be song per Bark (Bark), i.e., song / Bark.
[0321] It should be understood that for each frequency band in the e2 frequency bands, the characteristic loudness of the frequency band can be calculated in the above manner, thereby obtaining e2 characteristic loudnesses corresponding to the e2 frequency bands.
[0322] S03, the sibilance recognition and suppression module converts the e2 characteristic loudnesses into a Bark domain loudness distribution.
[0323] Exemplarily, as Figure 9As shown, the characteristic loudness of the new first frequency band in the e2 frequency bands is N1; the characteristic loudness of the new second frequency band is N2, and so on, and the characteristic loudness of the new e2th frequency band is Ne2. Each of the e2 frequency bands is a frequency range, and corresponds to an upper frequency limit (the maximum frequency in the frequency range) and a lower frequency limit (the minimum frequency in the frequency range).
[0324] Process 1: For each of the e2 frequency bands, the index of the Bark subband to which the lower frequency limit corresponding to the frequency band belongs can be calculated based on the lower frequency limit, and the index of the Bark subband to which the upper frequency limit corresponding to the frequency band belongs can be calculated based on the upper frequency limit, and then the plurality of Bark subbands corresponding to the frequency band can be determined.
[0325] The number of Bark subbands can be 240, and the indices of the 240 Bark subbands can be integers between 1 and 240, for example, in ascending order of frequency. The Bark subbands can correspond to Bark frequency bands, for example, the first Bark frequency band can correspond to the first Bark subband to the tenth Bark subband, and so on.
[0326] Exemplarily, taking the new first frequency band in the e2 frequency bands as an example, the lower frequency limit corresponding to the new first frequency band is hz1, and the lower frequency limit corresponding to the new first frequency band is hz2. Based on hz1 (or hz2) and the following formula, the initial Bark subband index (bark0) corresponding to hz1 (or hz2) can be calculated.
[0327]
[0328] The hz can be the upper frequency limit or the lower frequency limit.
[0329] Then, the initial Bark subband index (bark0) can be corrected by the following formula to obtain the Bark subband index (bark1).
[0330]
[0331] For example, substituting hz1 into formula 11, bark0 is calculated to be equal to r1', and then substituting r1' into formula 12, bark1 is calculated to be equal to the product of r1' and 10; substituting hz2 into formula 11, bark0 is calculated to be equal to r2', and then substituting r2' into formula 12, bark1 is calculated to be equal to 0.1xr2, and r2 is the product of bark1 and 10. Thus, the Bark subbands corresponding to the new first frequency band are the r1th Bark subband (the Bark subband with index r1) to the r2th Bark subband (the Bark subband with index r2). In the same way, the Bark subbands corresponding to each of the e2 frequency bands can be determined.
[0332] Process 2: The sibilance identification and suppression module determines the characteristic loudness corresponding to each Bark subband.
[0333] For example, in order of frequency from low to high, for the e2 frequency bands, when the characteristic loudness corresponding to the new ith frequency band is greater than or equal to the characteristic loudness of the new (i-1)th frequency band, the characteristic loudness of the Bark subbands corresponding to the new ith frequency band is equal to the characteristic loudness of the new ith frequency band; when the characteristic loudness corresponding to the new ith frequency band is less than the characteristic loudness of the new (i-1)th frequency band, the Bark subbands corresponding to the new ith frequency band introduce a masking effect, and the new characteristic loudness of the Bark subbands corresponding to the new ith frequency band is obtained. i takes an integer from 1 to e2. When i is 1, if the characteristic loudness N1 of the new first frequency band is greater than 0, the r1th Bark subband to the r2th Bark subband corresponding to the new first frequency band all correspond to the characteristic loudness N1.
[0334] For example, if the characteristic loudness N2 of the new second frequency band is greater than the characteristic loudness of the new first frequency band, the r3th Bark subband to the r4th Bark subband corresponding to the new second frequency band all correspond to the characteristic loudness N2.
[0335] If the characteristic loudness N3 of the new third frequency band is less than the characteristic loudness N2 of the new second frequency band, the r5th Bark subband to the r6th Bark subband corresponding to the new third frequency band introduce a masking effect.
[0336] It can be understood that, for a piece of audio data, when the characteristic loudness of the audio data decreases, the characteristic loudness shows a slow downward trend. For example, in combination with Figure 10 After point A, the characteristic loudness decreases, and the characteristic loudness decreases slowly.
[0337] Therefore, when N3 is less than N2, the characteristic loudness corresponding to the r5th Bark subband to the r6th Bark subband is slowly reduced from N2 to N3. That is, for the r5th Bark subband to the r6th Bark subband, in order of frequency from low to high, the characteristic loudness of each Bark subband is sequentially reduced, that is, the characteristic loudness of the r5th Bark subband is the highest, close to N2, and the characteristic loudness of the r6th Bark subband is the lowest, close to N3. For example, the characteristic loudness of the r5th Bark subband can be N2*0.9+N3*0.1; the characteristic loudness of the r5+1th Bark subband can be N2*0.8+N3*0.2; and so on, the characteristic loudness of the r6th Bark subband can be N2*0.1+N3*0.9.
[0338] When the sibilance identification and suppression module traverses the characteristic loudness of the e2 frequency bands, the characteristic loudness corresponding to each Bark subband in the plurality of Bark subbands corresponding to the e2 frequency bands can be determined.
[0339] Process 3, the sibilance identification and suppression module can determine the characteristic loudness of each Bark frequency band in the 24 Bark frequency bands based on the characteristic loudness corresponding to each Bark subband in the plurality of Bark subbands corresponding to the e2 frequency bands, that is, the Bark domain loudness distribution.
[0340] It should be understood that the Bark subband corresponding to each Bark frequency band is determined, for example, the Bark frequency band with z being 1 can correspond to the 1st to 10th Bark subbands; the Bark frequency band with z being 2 can correspond to the 11th to 20th Bark subbands, and so on. The characteristic loudness of each Bark frequency band is the sum of the characteristic loudness of the plurality of Bark subbands corresponding to the Bark frequency band. For example, the characteristic loudness of the Bark frequency band with z being 1 (the 1st Bark frequency band) is the sum of the characteristic loudness of the 1st to 10th Bark subbands, and so on.
[0341] Therefore, in the case of determining the characteristic loudness corresponding to each Bark subband in the plurality of Bark subbands corresponding to the e2 frequency bands, the characteristic loudness of each Bark frequency band can be calculated, that is, the Bark domain loudness distribution.
[0342] On the basis of the above-mentioned embodiments, in order to further improve the accuracy of sibilance identification, S608 can also be implemented by Figure 11 The process is implemented as follows:
[0343] S1101, the sibilance identification and suppression module determines whether the identification sharpness of the xth frame of audio data is greater than or equal to a threshold value 1.
[0344] If the identification sharpness of the xth frame of audio data is greater than the threshold value 1, it is determined that the xth frame of audio data is sibilance.
[0345] If the identification sharpness of the xth frame of audio data is less than or equal to threshold 1, S1102 is performed.
[0346] S1102, the sibilance identification and suppression module determines whether the (x-1)th frame of audio data is sibilance.
[0347] If the (x-1)th frame of audio data is not sibilance, it is determined that the xth frame of audio data is not sibilance.
[0348] If the (x-1)th frame of audio data is sibilance, S1103 is performed.
[0349] S1103, the sibilance identification and suppression module determines whether the identification sharpness of the xth frame of audio data is greater than or equal to threshold 2.
[0350] If the identification sharpness of the xth frame of audio data is greater than threshold 2, it is determined that the xth frame of audio data is sibilance.
[0351] If the identification sharpness of the xth frame of audio data is less than or equal to threshold 2, it is determined that the xth frame of audio data is not sibilance.
[0352] It should be understood that threshold 1 and threshold 2 are preset values, and threshold 2 is less than threshold 1.
[0353] In combination Figure 2A With Figure 2B As can be seen from the time-domain waveform of sibilance shown, for a plurality of consecutive frames of sibilance, the amplitude of sibilance first presents a slow growth trend and then presents a slow decline trend. Therefore, in the amplitude decline phase, the amplitude of the latter frame of sibilance is generally lower than that of the former frame of sibilance, that is, the energy of the latter frame of sibilance is generally lower than that of the former frame of sibilance, thereby causing the identification sharpness of the latter frame of sibilance to be generally lower than that of the former frame of sibilance. Therefore, in the case where the identification sharpness of the xth frame of audio data is less than or equal to threshold 1 and the (x-1)th frame of audio data is sibilance, whether the xth frame of audio data is sibilance can be determined based on the identification sharpness of the xth frame of audio data and the smaller threshold 2. In this way, for a plurality of consecutive frames of sibilance, in the sibilance amplitude decline phase, the sibilance with smaller amplitude can be identified, which helps to further improve the accuracy of sibilance identification.
[0354] The above describes the process of sibilance identification of the sibilance identification and suppression module on the xth frame of audio data. When the sibilance identification and suppression module traverses each frame of audio data (x takes an integer in 1 to M) in the M frames of audio data, the sibilance identification and suppression module can identify all sibilance in the M frames of audio data.
[0355] Further, in the case where the xth frame of audio data is sibilance, the sibilance identification and suppression module can perform sibilance suppression through S610 to S615.
[0356] In case that the xth frame of audio data is not a sibilance, the sibilance identification and suppression module performs S616, i.e., converting the xth frame of audio data from a frequency domain signal to a time domain signal.
[0357] The process of sibilance suppression is described in detail as follows.
[0358] S610, the sibilance identification and suppression module divides a preset frequency range in the xth frame of audio data into y frequency bands.
[0359] The preset frequency range can include a sibilance frequency band, i.e., the preset frequency range includes 4-12 kHz, so that the sibilance identification and suppression module can perform sibilance suppression on the sibilance frequency band in the xth frame of audio data. Alternatively, the preset frequency range can also partially coincide with the sibilance frequency band. For example, the preset frequency range is 5-13 kHz, etc. So that the sibilance identification and suppression module can perform sibilance suppression on other frequency bands in the xth frame of audio data.
[0360] Y can be a preset integer greater than 1, such as 3, 9, or 10, etc.
[0361] S611, based on the sharpness distribution 1, the sibilance identification and suppression module calculates the suppression sharpness of each frequency band in the y frequency bands.
[0362] The sharpness distribution 1 is the sharpness of the xth frame of audio data calculated in S605, i.e., the sharpness of each Bark frequency band. The suppression sharpness of each frequency band in the y frequency bands corresponding to the xth frame of audio data is related to the sharpness distribution 1, and is also related to the frequency band sound pressure level p x of the xth frame of audio data. x The frequency band sound pressure level p
[0363] Exemplarily, taking the kth frequency band in the y frequency bands as an example, the suppression sharpness (sharp k ) of the kth frequency band satisfies the following formula:
[0364]
[0365] Wherein, z is the index of the bark frequency band, k_fl is the lower limit frequency of the kth frequency band (the minimum frequency in the frequency range), z=bark(k_fl) is the index of the Bark frequency band to which the lower limit frequency of the kth frequency band belongs, k_fh is the upper limit frequency of the kth frequency band (the maximum frequency in the frequency range), z=bark(k_fh) is the index of the Bark frequency band to which the upper limit frequency of the kth frequency band belongs, z=bark() is a function of determining the index of the Bark frequency band based on the lower limit frequency or the upper limit frequency, and f(p') is a function of the frequency band sound pressure level p xThe frequency band pressure level score corresponding to the determined k frequency bands.
[0366] The suppression sharpness of the kth frequency band can be understood as the product of f(p') and the sum of the sharpness of one or more Bark frequency bands corresponding to the kth frequency band. The one or more Bark frequency bands corresponding to the kth frequency band are from the Bark frequency band to which the lower limit frequency of the kth frequency band belongs to the Bark frequency band to which the upper limit frequency of the kth frequency band belongs.
[0367] In formula 13, z = bark() is similar to formula 11 and formula 12 above, and z in z = bark() is the product of bark1 and 10 in formula 12. That is, the way of determining z based on the lower limit frequency or the upper limit frequency is similar to the way of determining r1 and r2 above, and the description above can be referred to, which will not be repeated here.
[0368] f(p') can be determined based on the sum of the sound pressure levels of one or more frequency bands of e1 frequency bands corresponding to the kth frequency band, for example. f(p') can satisfy the following formula:
[0369]
[0370] Wherein, h is the index of the frequency band of e1 frequency bands, k_fl is the lower limit frequency (the minimum frequency in the frequency range) of the kth frequency band, h = k_fl is the index of one frequency band of e1 frequency bands to which the lower limit frequency of the kth frequency band belongs, k_fh is the upper limit frequency (the maximum frequency in the frequency range) of the kth frequency band, h = k_fh is the index of one frequency band of e1 frequency bands to which the upper limit frequency of the kth frequency band belongs, p(h) is one frequency band sound pressure level of the hth frequency band of e1 frequency bands), and thres_p is a preset threshold 3. x
[0371] That is, the determination of one or more frequency bands of e1 frequency bands corresponding to the kth frequency band is similar to the determination of one or more Bark frequency bands corresponding to the kth frequency band. The lower limit frequency of the kth frequency band is within the frequency range of one frequency band of e1 frequency bands, and the one frequency band is the frequency band corresponding to the lower limit frequency of the kth frequency band; the upper limit frequency of the kth frequency band is within the frequency range of one frequency band of e1 frequency bands, and the one frequency band is the frequency band corresponding to the upper limit frequency of the kth frequency band. The one or more frequency bands of e1 frequency bands corresponding to the kth frequency band are from the frequency band corresponding to the lower limit frequency of the kth frequency band to the frequency band corresponding to the upper limit frequency of the kth frequency band.
[0372] For one or more frequency bands of e1 frequency bands corresponding to the kth frequency band, each frequency band corresponds to a frequency band sound pressure level, That is, the sum of the sound pressure levels of one or more frequency bands corresponding to the one or more frequency bands, where the sound pressure levels of one or more frequency bands corresponding to the one or more frequency bands can also be understood as the sound pressure level distribution of the k-th frequency band.
[0373] The larger the value, the greater the sharpness of the k-th frequency band, based on... The calculated suppression sharpness of the k-th frequency band k The probability is relatively high. The smaller the value, the lower the sharpness of the k-th frequency band, based on... The calculated suppression sharpness of the k-th frequency band k (It may be relatively small.)
[0374] f(p') can be determined based on p', for example, f(p') can be determined in any of the following ways.
[0375] For example, f(p') can satisfy the following formula:
[0376]
[0377] f(p') can also satisfy the following formula:
[0378]
[0379] Combining formula 15 or formula 16, it can be seen that when p' is less than the threshold 3 (thres_p), f(p') can be set to 0. That is, in If the value is less than the threshold 3 (thres_p), it means The intensity is relatively low. In this case, even if the sharpness corresponding to the k-th frequency band is relatively high, the energy of the k-th frequency band is relatively low, so the subjective listening experience of the user for the k-th frequency band is not sharp, making it unnecessary to suppress the k-th frequency band. Therefore, f(p') can be set to 0 at this time. The threshold 3 (thres_p) can be a value determined empirically and preset in the terminal device, for example.
[0380] Furthermore, f(p') can also be equal to p'. This eliminates the need for p' conversion, thus reducing the computational load on the terminal device.
[0381] In other words, f(p') and p' can have the same trend: the larger p' is, the larger f(p') is; the smaller p' is, the smaller f(p') is. When p' is small, it reflects that the energy of the k-th frequency band is low, so f(p') can be small, for example, it can be 0, or equal to p'. This ensures that the suppression sharpness of the k-th frequency band calculated based on f(p') is... kA smaller value indicates a lower value. Conversely, a larger value indicates a higher energy level in the k-th frequency band, allowing for a larger value for f(p') to improve the suppression sharpness of the k-th frequency band calculated based on f(p'). k (It is relatively large.)
[0382] Therefore, based on f(p') and The calculated suppression sharpness of the k-th frequency band k It can reflect the sharpness and energy level of the k-th frequency band.
[0383] In this way, the sibilance recognition and suppression module can identify the frequency bands with higher sharpness and energy among the y frequency bands based on the suppression sharpness of each frequency band, and thus suppress these high-sharpness and high-energy frequency bands. Compared to suppressing high-energy frequency bands, this embodiment also considers sharpness, making it less likely that the sibilance recognition and suppression module will suppress frequency bands with high energy but low sharpness, thereby helping to reduce the damage to timbre caused by sibilance suppression. This can improve the sound quality of the processed audio data and enhance the user experience.
[0384] S612, the sibilance recognition and suppression module determines whether the suppression sharpness of the k-th frequency band is greater than the threshold 4. The threshold 4 is a preset value.
[0385] If the suppression sharpness of the k-th frequency band is less than or equal to the threshold 4, it indicates that the sharpness and / or energy of the k-th frequency band is low, and S613 can be executed, that is, the frequency band gain of the k-th frequency band is set to 0.
[0386] If the suppression sharpness of the k-th frequency band is greater than the threshold 4, it indicates that the sharpness and energy of the k-th frequency band are relatively high, and S614 can be executed, that is, the band gain g of the k-th frequency band can be calculated. k .
[0387] Among them, the bandwidth gain g k It can be determined based on the suppression sharpness of the k-th frequency band, threshold 4, and threshold 5. For example, the band gain g... k The following formula can be satisfied:
[0388]
[0389] Among them, sharp k Let thres be the suppression sharpness of the k-th frequency band. sharpnes With a threshold of 4, progress i The threshold is 5.
[0390] The threshold 5 is a preset value less than 0, which can be used to represent the suppression ability. The smaller the threshold 5 is, the smaller the suppression strength of the kth frequency band by the sibilance recognition and suppression module is; the larger the threshold 5 is, the larger the suppression strength of the kth frequency band by the sibilance recognition and suppression module is.
[0391] It can be understood that, in combination with formula 17, it can be determined that the gain of the kth frequency band is positively correlated with the suppression sharpness of the kth frequency band; in combination with formula 13 to formula 16, it can be determined that the suppression sharpness of the kth frequency band is positively correlated with the sum of the sharpness of one or more Bark frequency bands corresponding to the kth frequency band. That is, for two frequency bands, assuming that f(p') corresponding to the two frequency bands is equal or very close (for example, the difference is less than a threshold), the suppression sharpness of the frequency band with a larger sum of the sharpness of one or more Bark frequency bands corresponding to the frequency band is larger, and the gain corresponding to the frequency band is also larger.
[0392] Exemplarily, for the kth frequency band and the wth frequency band, k and w can be different, and the kth frequency band and the wth frequency band can belong to one frame of audio data (recognized as sibilance); or k and w are the same, and the kth frequency band and the wth frequency band respectively belong to different frames of audio data (recognized as sibilance). For example, the kth frequency band belongs to the xth frame of audio data (recognized as sibilance), the wth frequency band belongs to the x+hth frame of audio data (recognized as sibilance), and so on, and k and w are equal.
[0393] Assuming that the amplitudes of the kth frequency band and the wth frequency band are the same or similar, and f(p') corresponding to the kth frequency band and f(p') corresponding to the wth frequency band are equal or very close. And the sum of the sharpness of one or more Bark frequency bands corresponding to the kth frequency band is S1, then the suppression sharpness of the kth frequency band is the product of S1 and f(p'), which is, for example, sharp1; the sum of the sharpness of one or more Bark frequency bands corresponding to the wth frequency band is S2, then the suppression sharpness of the wth frequency band is the product of S2 and f(p'), which is, for example, sharp2. In the case of S1 being less than S2, sharp1 is less than sharp2. By substituting sharp1 and sharp2 into formula 17 respectively, the gain of the kth frequency band calculated is also less than the gain of the wth frequency band. Or, if sharp1 is less than threshold 4, the gain of the kth frequency band is 0.
[0394] Therefore, based on the determination of the gain according to the suppression sharpness, for the frequency band with higher sharpness, the gain calculated by the terminal device is relatively larger; for the frequency band with smaller sharpness, the gain calculated by the terminal device is relatively smaller. So that the terminal device can have a higher suppression degree for the sibilance signal with higher sharpness, and a lower suppression degree for the sibilance signal with lower sharpness. Thus, the sibilance suppression has less damage to the timbre.
[0395] Based on the above manner (S611 to S614), the sibilance identification and suppression module can calculate the band gain of each of the y frequency bands corresponding to the xth frame of audio data in the case that k takes all integers between 1 and y.
[0396] S615, the sibilance identification and suppression module performs time-frequency domain gain smoothing on the y frequency bands based on the band gain of the y frequency bands.
[0397] That is, the band gain of the y frequency bands can be smoothed from two dimensions of time domain and frequency domain, so that the processed xth frame of audio data sounds smoother and more natural.
[0398] Exemplarily, the sibilance identification and suppression module can perform time-frequency domain gain smoothing in the following manner.
[0399] Suppose y is 3, and the band gain of the first frequency band in the three frequency bands is g1, the band gain of the second frequency band is g2, and the band gain of the third frequency band is g3 in order of frequency from low to high. There may be 0 or no 0 in g1, g2 and g3.
[0400] Then in order of frequency from low to high, the sibilance identification and suppression module can perform smoothing processing on the next frequency band based on the band gain of the previous frequency band and the band gain of the next frequency band in the y frequency bands, and the band gain g2 of the second frequency band can be replaced by g1*0.9+g2*0.1; the band gain g2 of the third frequency band can be replaced by g2*0.9+g3*0.1, etc. In this way, after the y frequency bands are suppressed, the suppressed audio data is smoother and more natural.
[0401] It should be understood that the y frequency bands shown above are the y frequency bands corresponding to the xth frame of audio data.
[0402] In addition, the sibilance identification and suppression module can also perform smoothing processing on each of the y frequency bands corresponding to the xth frame of audio data based on the band gain of each of the y frequency bands corresponding to the x-1th frame of audio data.
[0403] Suppose according to the frequency from low to high, the band gain of the first frequency band in the three frequency bands corresponding to the x-1th frame of audio data is g1', the band gain of the second frequency band is g2', and the band gain of the third frequency band is g3'. There may be 0 or no 0 in g1', g2' and g3'.
[0404] According to the order from low to high of the frequencies, the sibilance identification and suppression module can perform smoothing processing on the frequency band gain of the fth frequency band corresponding to the xth frame of audio data based on the frequency band gain of the fth frequency band corresponding to the x-1th frame of audio data, where f is a positive integer. For example, the frequency band gain g1 of the 1st frequency band in the xth frame of audio data can be replaced by g1’*0.9+g1*0.1; the frequency band gain g2 of the 2nd frequency band in the xth frame of audio data can be replaced by g2’*0.9+g2*0.1; the frequency band gain g3 of the 3rd frequency band in the xth frame of audio data can be replaced by g3’*0.9+g3*0.1, and so on.
[0405] In this way, the sibilance identification and suppression module can determine the smoothed gain of each frequency band in the y frequency bands (corresponding to the xth frame of audio data).
[0406] Then, the sibilance identification and suppression module can suppress the y frequency bands corresponding to the xth frame of audio data based on the smoothed gain.
[0407] S616, convert the xth frame of audio data from a frequency domain signal to a time domain signal.
[0408] It should be understood that whether the xth frame of audio data is sibilance or not, the xth frame of audio data is converted from a frequency domain signal to a time domain signal.
[0409] In the case that the xth frame of audio data is not sibilance, the sibilance identification and suppression module does not perform S610 to S615; in the case that the xth frame of audio data is sibilance, the sibilance identification and suppression module performs S610 to S615 to suppress sibilance of the xth frame of audio data.
[0410] Through S610 to S616, when x takes all integers between 1 and M, the sibilance identification and suppression module can suppress sibilance of all identified sibilance in M frames of audio data, so that the processed M frames of audio data are smoother.
[0411] Further, when the terminal device plays the sibilance-suppressed audio data through the audio playing module such as a loudspeaker, the sibilance-suppressed audio data may still sound sharp due to the characteristics of the audio playing module itself.
[0412] For example, when the audio playing module such as a loudspeaker is small, such as the loudspeaker of a mobile phone, the energy of the sibilance frequency band in the audio data may be amplified when playing the audio data, so that even if the played audio data is sibilance-suppressed audio data, the audio data played by the terminal device may still sound sharp.
[0413] Therefore, the terminal device can further perform filtering processing on the audio data processed by the sibilance identification and suppression module through a filtering processing module. Specifically as follows.
[0414] In S617, the sibilance recognition and suppression module transmits the converted time domain signal to the filtering processing module.
[0415] Taking the time domain signal of the xth frame of audio data as an example, the filtering processing module can perform S618, i.e., filtering processing on the time domain signal of the xth frame of audio data to obtain the filtered audio data of the xth frame of audio data.
[0416] It can be understood that the center frequency (FC), Q value, and gain of the filtering processing module and the like can be preset. The center frequency (FC), Q value, and gain of the filtering processing module and the like are determined based on the frequency response curve of the loudspeaker or the like audio playing module.
[0417] Exemplarily, for the peak value in the frequency response curve, the gain is higher, so that the filtering processing module can amplify the energy of the frequency band corresponding to the frequency point (the abscissa of the peak value) of the peak value. If the amplified frequency band coincides with the sibilance in the audio data, the energy of the sibilance will be amplified. Therefore, the center frequency (FC), Q value, and gain of the filtering processing module and the like need to be set to make the frequency response curve of the filtering processing module more gentle.
[0418] For example, the center frequency (FC) can be set as the frequency point of the peak value in the initial frequency response curve. The Q value can be set as the ratio of the center frequency (FC) to the effective bandwidth. The effective bandwidth can be set based on the width of the peak in the initial frequency response curve. The wider the width of the peak in the initial frequency response curve (the difference between the abscissas), the larger the effective bandwidth can be. The smaller the width of the peak in the initial frequency response curve, the smaller the effective bandwidth can be. The gain of the filtering processing module can be determined based on the ordinate of the peak value in the initial frequency response curve. For example, the gain of the filtering processing module can be the difference between the ordinate of the peak value and the ordinate of the horizontal line. The horizontal line is a horizontal straight line determined based on the relatively stable part in the initial frequency response curve.
[0419] In this way, through the filtering processing of the filtering processing module, the amplification effect of the sibilance frequency band by the loudspeaker or the like audio playing module can be reduced. Thus, it is helpful to further improve the sound quality of the audio data and improve the user experience.
[0420] The audio data (filtered audio data) subjected to the sibilance suppression and filtering processing can be played in the following manner (S619 to S625). The following manner is only an example. In some possible implementation, the terminal device can also play the audio data (filtered audio data) subjected to the sibilance suppression and filtering processing in other manners, and does not constitute a limitation on the embodiments of the present application.
[0421] S619, the filter processing module writes the filtered audio data of the xth frame of audio data into the cache.
[0422] It can be understood that when x takes all integers between 1 and M, the filter processing module can write filtered audio data of M frames of audio data into the cache.
[0423] S620, the audio mixer obtains filtered audio data of one or more frames of audio data from the cache. The filtered audio data of one or more frames of audio data belongs to part or all of the filtered audio data of M frames of audio data.
[0424] S621, the audio mixer performs mixing, sound effect, etc. processing on the filtered audio data of one or more frames of audio data to obtain audio data a.
[0425] S622, the audio mixer transmits the audio data a to the audio HAL.
[0426] S623, the audio HAL performs further processing such as format conversion, sampling rate adjustment, volume control, etc. on the audio data a to obtain audio data b.
[0427] S624, the audio HAL transmits the audio data b to the audio driver.
[0428] S625, the audio driver plays the audio data b through a speaker or other audio playing module.
[0429] It can be understood that S620 to S625 can be executed multiple times to enable the terminal device to play all audio data in M frames of audio data.
[0430] Through the method of the present application, the sibilance in the audio data can be accurately identified, and the identified sibilance can be effectively suppressed with little damage to the tone color. Exemplarily, in combination with Figure 12A and Figure 12B wherein Figure 12B is a spectrogram of audio data before and after sibilance suppression provided by an embodiment of the present application, Figure 12A is a spectrogram of Figure 12B is a schematic diagram of the spectrogram.
[0431] Figure 12B (a) in (a) shows an original sequence spectrogram, i.e. a spectrogram before sibilance suppression, and correspondingly, Figure 12A (a) in (a) shows a schematic diagram of the original sequence spectrogram; Figure 12B (b) in (b) shows a spectrogram after sibilance identification and sibilance suppression by the method of the present application, and correspondingly, Figure 12B (b) in (b) shows a schematic diagram of the spectrogram after sibilance suppression.
[0432] In Figure 12A , the white part represents a part with higher energy, and the closer to white, the higher the energy. Figure 12B In Figure 12A , the part circled in the box 1201 in (a) can be understood as a part of Figure 12B , and the part circled in the box 1205 in (a) can be understood as a part of Figure 12A , the part circled in the box 1202 in (b) can be understood as a part of Figure 12B , and the part circled in the box 1206 in (b) can be understood as a part of
[0433] In combination with Figure 12A , the spectrograms circled in the box 1201 and the box 1202 are compared, or in combination with Figure 12B , the spectrograms circled in the box 1205 and the box 1206 are compared, and the spectrograms circled in the box 1207 and the box 1208 are compared, it can be seen that after the audio data is subjected to sibilance recognition and sibilance suppression by the method of the present application, the sibilance with higher frequency and higher energy is significantly suppressed, so that in Figure 12A or Figure 12A , the energy of the sibilance at the position with higher frequency is reduced in (b).
[0434] In addition, in combination with the part circled in the box 1203 in (a) of Figure 12B and the part circled in the box 1204 in (a) of Figure 13A , or in combination with the spectrograms of the low-frequency parts in (a) and (b) of Figure 13B , it can be seen that the method of the present application does not suppress the audio data with lower frequency and higher energy, so that the timbre damage of sibilance suppression is smaller.
[0435] In addition, by the method of the present application for sibilance recognition, compared with sibilance recognition by the standard sharpness and the preset threshold 1, the accuracy of sibilance recognition is higher. And by the method of the present application for sibilance suppression, compared with sibilance suppression by the EQ or MBDRC, the timbre damage can be reduced.
[0436] Exemplarily, in combination with Figure 13B and Figure 13A , wherein Figure 13B is a spectrogram provided by an embodiment of the present application for displaying a sibilance suppression part, Figure 13A is a schematic of Figure 13B . Figure 13A , (a) and Figure 13BThe (a) in FIG. 6 shows the original sequence of the same audio data, which includes the sound of musical instruments and human voice. If the sibilance identification is based on the amplitude of the audio data and the suppression is based on the way of the EQ, the suppressed audio part can be shown as Figure 13A (b) in FIG. 6 (where the white part) or Figure 13B (b) in FIG. 6 (where the red and yellow parts). It can be seen that this audio processing method also suppresses the sound of musical instruments (piano sound and drum sound), and also suppresses the low frequency part in the voice and accompaniment sound, resulting in timbre loss, audio timbre distortion heard by the user, and poor listening experience.
[0437] And when the sibilance identification and sibilance suppression are performed by the method of the present application, the suppressed audio part can be shown as Figure 1 (c) in FIG. 6 (where the white part) or Figure 1 (c) in FIG. 6 (where the red or yellow part). It can be seen that the method of the present application has high accuracy in audio identification, and is not easy to identify the sound of musical instruments as sibilance, so that almost no suppression is performed on the sound of musical instruments; and also almost no suppression is performed on the low frequency part in the audio data. By comparison, it can be seen that the method of the present application can reduce timbre damage.
[0438] It should be understood that the size of the serial number in each of the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic.
[0439] The audio processing method of the embodiments of the present application has been described above, and the device provided by the embodiments of the present application for executing the above method will be described below. Those skilled in the art can understand that the method and the device can be combined and referenced with each other, and the related device provided by the embodiments of the present application can execute the steps in the above listed method.
[0440] The embodiments of the present application also provide an audio processing device. The audio processing device can include a processor, a communication interface and a memory. Wherein the processor, the communication interface and the memory communicate with each other through an internal connection path, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory. The communication interface can be used to send signals to other devices (such as the touch screen of the processor or the terminal equipment), and can also be used to receive signals from other devices (such as the memory). Illustratively, the communication interface reads the instructions stored in the memory and sends the instructions to the processor.
[0441] It should be understood that the audio processing apparatus can be specifically a terminal device in the above-described embodiments, and can be used to perform each step and / or process in the above-described method embodiments corresponding to the terminal device. Optionally, the memory can include a read-only memory and a random access memory, and provide instructions and data for the processor. A part of the memory can also include a non-volatile random access memory. For example, the memory can also store device type information. The processor can be used to execute the instructions stored in the memory, and when the processor executes the instructions stored in the memory, the processor is used to perform each step and / or process of the above-described method embodiments.
[0442] It should be understood that in the embodiments of the present application, the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0443] In the implementation process, each step of the above-described method can be completed by the integrated logic circuit of hardware in the processor or the instruction in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as hardware processor execution completion, or executed by the combination of hardware and software modules in the processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory, and the processor executes the instructions in the memory to complete the steps of the above-described method in combination with the hardware thereof. To avoid repetition, it will not be described in detail here.
[0444] The application starting method provided by the embodiments of the present application can be applied to an electronic device with communication function. The electronic device includes a terminal device, and the specific device form of the terminal device can refer to the above-mentioned related description, which will not be described here.
[0445] The embodiments of the present application provide an electronic device, which includes: a processor and a memory; the memory stores computer execution instructions; and the processor executes the computer execution instructions stored in the memory, so that the terminal device executes the above-described method.
[0446] The embodiments of the present application provide a chip. The chip includes a processor, and the processor is used to call a computer program in a memory to execute the technical solutions in the above-described embodiments. The implementation principles and technical effects are similar to those of the above-mentioned related embodiments, which will not be described here.
[0447] The embodiment of the present application further provides a computer readable storage medium. The computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the method described above. The method described in the above embodiment can be implemented by software, hardware, firmware or any combination thereof, in whole or in part. If implemented in software, the functions can be stored in or transmitted as one or more instructions or code on a computer-readable medium. The computer-readable medium can include a computer storage medium and a communication medium, and can further include any medium that can carry the computer program from one place to another. The storage medium can be any target medium that can be accessed by a computer.
[0448] In a possible implementation, the computer readable medium can include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that is targeted to carry desired program codes in the form of instructions or data structures and can be accessed by a computer. Moreover, any connection is appropriately referred to as a computer readable medium. For example, if software is transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave), the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave) is included in the definition of the medium. As used herein, a disk and a disc include compact discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, wherein a disk usually reproduces data magnetically, while a disc reproduces data optically with a laser. Combinations of the above should also be included in the scope of the computer readable medium.
[0449] The embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed, the computer program causes the computer to execute the method described above.
[0450] The embodiment of the present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general purpose computer, a special purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the computer or other programmable data processing device produce a device that implements the functions described in the flowcharts and / or block diagrams. one or more processes and / or blocks an apparatus for performing functions specified in one or more blocks.
[0451] The above detailed description has disclosed, for the purpose of the present application, the purpose, technical solutions and beneficial effects. It should be understood that the above is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.
Claims
1. An audio processing method, characterized in that, Applied to a terminal device, the method includes: Displaying a first interface of a first application, the first interface including a first control; Based on the user's first operation on the first control, a first audio is obtained. The first audio includes a first sibilant signal and a second sibilant signal. The amplitude of the first sibilant signal is the same as the amplitude of the second sibilant signal. The first sharpness of the first sibilant signal is less than the second sharpness of the second sibilant signal. Play a second audio, which is the audio obtained after suppressing sibilance in the first audio. The second audio includes a third sibilance signal and a fourth sibilance signal. The third sibilance signal is the processed first sibilance signal, and the fourth sibilance signal is the processed second sibilance signal. The first difference between the amplitude of the first sibilance signal and the amplitude of the third sibilance signal is less than the second difference between the amplitude of the second sibilance signal and the amplitude of the fourth sibilance signal.
2. The method according to claim 1, characterized in that, The first sharpness is the sum of the sharpness distributions corresponding to the first dentate sound signal, and the second sharpness is the sum of the sharpness distributions corresponding to the second dentate sound signal.
3. The method according to claim 1 or 2, characterized in that, The third dentate sound signal is obtained by processing the first dentate sound signal based on the gain of the first dentate sound signal, and the fourth dentate sound signal is obtained by processing the second dentate sound signal based on the gain of the second dentate sound signal, wherein the gain of the first dentate sound signal is less than the gain of the second dentate sound signal.
4. The method according to claim 3, characterized in that, The gain of the first dentate sound signal is determined based on the sum of the first sharpness and the frequency band sound pressure level distribution corresponding to the first dentate sound signal. The gain of the second dentate sound signal is determined based on the sum of the second sharpness and the frequency band sound pressure level distribution corresponding to the second dentate sound signal. The sum of the frequency band sound pressure level distribution corresponding to the first dentate sound signal is positively correlated with the amplitude of the first dentate sound signal, and the sum of the frequency band sound pressure level distribution corresponding to the second dentate sound signal is positively correlated with the amplitude of the second dentate sound signal.
5. The method according to any one of claims 1 to 4, characterized in that, The first sibilant signal belongs to the c1 frame audio data, and the second sibilant signal belongs to the c2 frame audio data. The c1 frame audio data and the c2 frame audio data are sibilant sounds. The c1 frame audio data is identified by the terminal device based on the sharpness distribution and the spectral centroid of the c1 frame audio data. The c2 frame audio data is identified by the terminal device based on the sharpness distribution and the spectral centroid of the c2 frame audio data.
6. The method according to claim 5, characterized in that, The product of the sum of the sharpness distributions corresponding to the audio data of frame c1 and the normalized spectral centroid of the audio data of frame c1 is greater than the fourth threshold, and the product of the sum of the sharpness distributions corresponding to the audio data of frame c2 and the normalized spectral centroid of the audio data of frame c2 is greater than the fourth threshold.
7. The method according to claim 5 or 6, characterized in that, The sharpness distribution corresponding to the audio data of the c1th frame and the sharpness distribution corresponding to the audio data of the c2th frame respectively include the sharpness of multiple Bark frequency bands, and the multiple Bark frequency bands are Bark frequency bands with frequencies greater than or equal to preset frequencies.
8. An audio processing method, characterized in that, include: Determine that the x-th frame of audio data in a multi-frame audio dataset is sibilant; The preset frequency range of the xth frame audio data is divided into y frequency bands, the preset frequency range includes the serrated audio band, or the preset frequency range partially overlaps with the serrated audio band, where y is a positive integer; Based on the sharpness distribution and frequency band sound pressure level distribution corresponding to each of the y frequency bands, the gain of each of the y frequency bands is calculated. The sharpness distribution includes the sharpness of at least one frequency band, and the frequency band sound pressure level distribution includes the frequency band sound pressure level of at least one frequency band. Based on the gain of each of the y frequency bands, the y frequency bands are suppressed to obtain sibilance-suppressed audio data.
9. The method according to claim 8, characterized in that, The calculation of the gain of each of the y frequency bands includes: The sum of the sound pressure level distributions of the first frequency band corresponding to the kth frequency band in the y frequency bands is calculated to obtain the first value; The sum of the first sharpness distributions corresponding to the kth frequency band among the y frequency bands is calculated to obtain the second value; The suppression sharpness of the kth frequency band is calculated based on the first value and the second value, and the suppression sharpness of the kth frequency band is used to represent the magnitude of the first value and the second value; Based on the suppression sharpness, the first gain of the k-th frequency band is determined.
10. The method according to claim 9, characterized in that, Determining the first gain of the k-th frequency band based on the suppression sharpness includes: When the suppression sharpness is less than or equal to a first threshold, the first gain is 0; and / or, When the suppression sharpness is greater than the first threshold, the first gain is calculated based on the suppression sharpness, the first threshold, and the second threshold, where the second threshold is used to represent the suppression capability.
11. The method according to claim 10, characterized in that, The first gain g k Satisfy the following formula: Among them, sharp k For the suppression sharpness, thres sharpnes For the first threshold, progress i This is the second threshold.
12. The method according to any one of claims 9 to 11, characterized in that, The calculation of the suppression sharpness of the k-th frequency band based on the first value and the second value includes: Calculate the ratio of the first value to the third threshold to obtain the third value; Based on a preset function, the third value is converted into a fourth value, wherein the fourth value is less than or equal to the third value; The product of the fourth value and the second value is calculated, and the suppression sharpness is the product of the fourth value and the second value.
13. The method according to any one of claims 9 to 12, characterized in that, The process of obtaining the first value includes: The amplitude spectrum of the x-th frame audio data is divided into E frequency bands, where E is a positive integer; Calculate the second frequency band sound pressure level distribution of the x-th frame audio data, whereby the second frequency band sound pressure level distribution includes the frequency band sound pressure level of each of the E frequency bands; Based on the second frequency band sound pressure level distribution, the sum of the first frequency band sound pressure level distributions is calculated, wherein the first frequency band sound pressure level distribution includes the frequency band sound pressure level of at least one of the E frequency bands.
14. The method according to claim 13, characterized in that, In ascending order of frequency, at least one of the E frequency bands includes a first frequency band to a second frequency band, wherein the first frequency band is a frequency band to which the lower limit frequency of the k-th frequency band in the E frequency bands belongs, and the second frequency band is a frequency band to which the upper limit frequency of the k-th frequency band in the E frequency bands belongs.
15. The method according to any one of claims 9 to 14, characterized in that, The process of obtaining the second value includes: Calculate the second sharpness distribution of the x-th frame audio data, the second sharpness distribution including the sharpness of each Bark band; Based on the second sharpness distribution, the sum of the first sharpness distribution is calculated, wherein the first sharpness distribution includes the sharpness of at least one Bark frequency band.
16. The method according to any one of claims 8 to 15, characterized in that, The determination that the x-th frame of audio data in a multi-frame audio data is a sibilant includes: Based on the sum of the third sharpness distribution of the x-th frame audio data and the spectral centroid of the x-th frame audio data, the x-th frame audio data is determined to be sibilant.
17. The method according to claim 16, characterized in that, The third sharpness distribution includes the sharpness of each of the multiple Bark frequency bands, wherein the multiple Bark frequency bands are Bark frequency bands with frequencies greater than or equal to a preset frequency.
18. The method according to claim 16 or 17, characterized in that, Determining that the xth frame of audio data is sibilant includes: Based on the sum of the third sharpness distribution and the normalized spectral centroid of the x-th frame audio data, the x-th frame audio data is determined to be sibilant.
19. The method according to claim 18, characterized in that, Determining that the xth frame of audio data is sibilant includes: The recognition sharpness of the x-th frame audio data is obtained by calculating the product of the sum of the third sharpness distribution and the centroid of the normalized spectrum. If the sharpness of the identification is greater than the fourth threshold, the xth frame of audio data is determined to be sibilant.
20. The method according to claim 19, characterized in that, The method further includes: If the recognition sharpness is less than or equal to the fourth threshold, determine whether the audio data of the (x-1)th frame is sibilant; If the audio data of the (x-1)th frame contains sibilance, it is determined whether the recognition sharpness is greater than the fifth threshold, and the fifth threshold is less than the fourth threshold; If the sharpness of the identification is greater than the fifth threshold, the xth frame of audio data is determined to be sibilant.
21. The method according to claim 20, characterized in that, The method further includes: If the audio data of the (x-1)th frame is not sibilant, then the audio data of the xth frame is determined to be not sibilant; or, If the sharpness of the identification is less than or equal to the fifth threshold, it is determined that the xth frame audio data is not sibilant.
22. The method according to any one of claims 8 to 21, characterized in that, The method further includes: The audio data after sibilance suppression is filtered by a filter. The first parameter of the filter is determined based on the frequency response curve of the audio playback module. The first parameter includes one or more of the following: center frequency, quality factor, or gain.
23. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 22.
24. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 22.
25. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 22.
26. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 22.
Citation Information
Patent Citations
Tooth tone adjustment method and device, electronic equipment and computer readable storage medium
CN112951266A
Audio detection method and device, electronic equipment and storage medium
CN113611330A