A method, apparatus and electronic device for sibilance processing
By identifying the core and adaptive suppression frequency bands in the audio frame and gradually increasing the gain to process sibilance, the problem of sound quality loss during sibilance suppression is solved, achieving a balanced improvement in sound quality and suppression effect.
Patent Information
- Application Number
- CN202311345722.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-17
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2043-10-17
AI Technical Summary
Existing technologies damage sound quality and are not ideal in suppressing sibilance, making it difficult to achieve a good balance between sound quality and suppression effect.
By determining the core suppression band and adaptive suppression band of the audio frame, expanding the suppression band and gradually increasing the gain, sibilance processing is performed to ensure that the sound quality loss is minimized.
It significantly improves sibilance suppression while reducing sound quality degradation, achieving a balance between sound quality and suppression effect.
Smart Images

Figure CN119889336B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio, more particularly, to a sibilance processing method, device and electronic equipment. BACKGROUND
[0002] Sibilance refers to all the fricative sounds emitted by people, which are strong and heavily emphasized fricative consonants. Sibilance usually ranges from 4 kHz to 8 kHz, corresponding to high sharpness, which is generally not suitable for human ears to listen to. Therefore, electronic equipment will suppress sibilance in the audio signal to reduce sibilance energy and try to keep each frame of audio within a suitable sharpness range to avoid sibilance with high sharpness from causing damage to human hearing.
[0003] When sibilance suppression is performed, the audio quality of the audio signal will be damaged to some extent. Considering the audio quality, the frequency range of the sibilance segment of the suppressed audio signal is currently small, resulting in unsatisfactory sibilance suppression effect.
[0004] Therefore, it is necessary to provide a sibilance processing method that can improve the sibilance suppression effect while taking into account the audio quality. SUMMARY
[0005] The embodiments of the present application provide a sibilance processing method that balances the audio quality and sibilance suppression, and effectively improves the sibilance suppression effect while minimizing the damage to the audio quality.
[0006] In a first aspect, a sibilance processing method is provided, including: determining that a current audio frame is a sibilance frame; determining a core suppression frequency band and an adaptive suppression frequency band of the current audio frame; determining a first suppression gain of the core suppression frequency band; in a case where the frequency band range of the adaptive suppression frequency band includes the frequency band range of the core suppression frequency band, or in a case where the frequency band range of the adaptive suppression frequency band and the frequency band range of the core suppression frequency band have a partial intersection, expanding the suppression frequency band of the sibilance through the adaptive suppression frequency band to determine at least one extended suppression frequency band other than the core suppression frequency band; determining a second suppression gain of each extended suppression frequency band, and in accordance with a frequency change direction from a first boundary frequency closest to the core suppression frequency band to a second boundary frequency farthest from the core suppression frequency band in each extended suppression frequency band, the second suppression gain of each extended suppression frequency band increases from the first suppression gain to 1; and applying the first suppression gain to the energy of the current audio frame in the core suppression frequency band, and applying the second suppression gain of each extended suppression frequency band to the energy of the current audio frame in each extended suppression frequency band.
[0007] In the above embodiment, the determined suppression frequency band of the audio frame includes a core suppression frequency band and an adaptive suppression frequency band used to expand the core suppression frequency band, and in a case where a frequency band range of the adaptive suppression frequency band includes a frequency band range of the core suppression frequency band or there is a partial intersection between the frequency band range of the adaptive suppression frequency band and the frequency band range of the core suppression frequency band, the suppression frequency band of the sibilance is expanded by the adaptive suppression frequency band, that is, an expanded suppression frequency band other than the core suppression frequency band is obtained, and the frequency band range of the suppression frequency band is increased, and the sibilance processing is performed on the suppression frequency band with a larger range, so that a better sibilance suppression effect can be obtained. Moreover, since the adaptive suppression frequency band expands the core suppression frequency band, the gain of the core suppression frequency band is the first suppression gain, and for the expanded expanded suppression frequency band, the second suppression gain of the expanded suppression frequency band is increased from the first suppression gain to 1, so that the problem of quality damage caused by still using a smaller gain to suppress the energy of the expanded suppression frequency band can be avoided, and therefore, the gradually increasing second suppression gain can gently damage the quality, and can improve the sibilance suppression effect while minimizing the damage to the quality. In summary, the sibilance processing method provided in the embodiments of the present application balances the quality and sibilance suppression in two aspects, and can effectively improve the sibilance suppression effect while minimizing the damage to the quality.
[0008] Optionally, before the steps of determining the core suppression frequency band and the adaptive suppression frequency band of the current audio frame in the above embodiment, the method further includes: determining a first bark frequency band in a bark domain corresponding to a maximum sharpness according to a sharpness distribution of the current audio frame; and in the steps of determining the core suppression frequency band and the adaptive suppression frequency band of the current audio frame in the above embodiment, the core suppression frequency band and the adaptive suppression frequency band are determined according to the first bark frequency band, and the core suppression frequency band and the adaptive suppression frequency band both include the first bark frequency band.
[0009] Since the first bark frequency band is a frequency band corresponding to the maximum sharpness, the sibilance energy of the first bark frequency band is relatively high, and the core suppression frequency band determined based on the first bark frequency band can better determine a frequency band with relatively high sibilance energy, so as to improve the sibilance suppression effect.
[0010] Optionally, in the step of determining the core suppression frequency band according to the first bark frequency band in the above embodiment, the first bark frequency band is taken as a center frequency band, M1 bark frequency bands are expanded to both sides of the first bark frequency band respectively, and the core suppression frequency band is obtained, where M1 is an integer greater than 1 or equal to 1.
[0011] Generally, the region with high sibilance energy is concentrated on both sides of the first bark band, and therefore, the core suppression band obtained by extending M1 bark bands to both sides of the first bark band can well and comprehensively obtain the frequency band with relatively high sibilance energy, so as to further improve the sibilance suppression effect. In addition, the frequency band is extended by the number M1 of bark bands, and the implementation process is simple and convenient to implement.
[0012] Optionally, in the step of determining the adaptive suppression band according to the first bark band in the above embodiment, the frequency band is extended to both sides of the first bark band according to the sharpness distribution of the current audio frame, so as to determine the adaptive suppression band, wherein the frequency band sharpness of the adaptive suppression band satisfies a first preset condition for limiting the extension of the adaptive suppression band.
[0013] In the above embodiment, the frequency band is extended to both sides of the first bark band according to the sharpness distribution of the current audio frame, and the extension of the frequency band is stopped when the frequency band sharpness of the extended frequency band satisfies the first preset condition, so as to obtain the adaptive suppression band. Since the adaptive suppression band is extended based on the sharpness distribution of the current audio frame, the adaptive suppression band obtained for each audio frame is greatly related to the characteristics of the audio frame itself, and therefore, the adaptive suppression band obtained for each audio frame is more in line with the characteristics of each audio frame, so as to accurately determine the suppression band of each audio frame, thereby better improving the sibilance suppression effect, and improving the flexibility of sibilance processing.
[0014] Optionally, in the step of extending the frequency band to both sides of the first bark band according to the sharpness distribution of the current audio frame to determine the adaptive suppression band in the above embodiment, the sharpness corresponding to each bark band in the plurality of bark bands on both sides of the first bark band is compared K times according to the sharpness distribution of the current audio frame, the frequency band corresponding to the candidate suppression interval obtained in the Kth comparison process is determined as the adaptive suppression band when the frequency band sharpness of the candidate suppression interval obtained in the Kth comparison process satisfies the first preset condition, K is an integer greater than 1; wherein,
[0015] In the first comparison process, the sharpness corresponding to the two bark bands on both sides of the first bark band is compared, and the bark band corresponding to the maximum sharpness in the first comparison process is determined as the initial candidate bark band; the interval between the first bark band and the initial candidate bark band including the first bark band and the initial candidate bark band is determined as the candidate suppression interval in the first comparison process, and the frequency band sharpness of the candidate suppression interval in the first comparison process does not satisfy the first preset condition.
[0016] In the i-th comparison process, the sharpness corresponding to the two bark bands on both sides of the candidate suppression interval obtained in the (i-1)-th comparison process is compared, the bark band corresponding to the maximum sharpness in the i-th comparison process is determined as the intermediate candidate bark band, i is greater than 1 and less than K; the interval between the intermediate candidate bark band and the boundary bark band is determined as the candidate suppression interval in the i-th comparison process, the boundary bark band is the bark band far away from the intermediate candidate bark band in the candidate suppression interval obtained in the (i-1)-th comparison process, and the frequency band sharpness of the candidate suppression interval in the i-th comparison process does not satisfy the first preset condition;
[0017] K-i comparison processes are performed after the i-th comparison process until the adaptive suppression frequency band is determined in the K-th comparison process.
[0018] In the above embodiment, according to the sharpness distribution of the current audio frame, the sharpness corresponding to each bark band in the multiple bark bands on both sides of the first bark band is compared for K times, the maximum sharpness in the sharpness corresponding to the two bark bands on both sides of the first bark band is determined in each comparison process, the candidate suppression interval of each comparison process is determined based on the maximum sharpness obtained in each comparison process, the purpose of expanding the suppression interval by a part of interval in each comparison process is achieved, and the expansion of the suppression interval is finally completed through multiple comparison processes to obtain the adaptive suppression frequency band. Based on this way of cyclically traversing the sharpness of the current audio frame, the adaptive suppression frequency band can be obtained as much as possible based on the characteristics of the current audio frame to search for the bark band with a larger sharpness value to more accurately obtain the adaptive suppression frequency band conforming to the current audio frame.
[0019] Optionally, the first preset condition is that the frequency band sharpness of the candidate suppression interval obtained in each comparison process is greater than a first sharpness threshold.
[0020] Optionally, the first sharpness threshold is the product of the frame sharpness of the current audio frame and a coefficient, and the coefficient is less than 1.
[0021] Since the frame sharpness of the current audio frame is related to the current audio frame, the first sharpness threshold is designed as the product of the frame sharpness of the current audio frame and a coefficient, so that the first preset condition is related to each audio frame, and the adaptive suppression frequency band obtained is more in line with each audio frame, which is conducive to more accurately determining the suppression frequency band of each audio frame, so that the effect of sibilance suppression can be better improved.
[0022] Optionally, in the step of determining the first suppression gain of the core suppression frequency band in the above embodiment, the first suppression gain is determined according to the gain coefficient and the initial suppression gain, the gain coefficient being related to the current audio frame.
[0023] Since the gain coefficient is related to the current audio frame, the gain coefficient varies dynamically for different audio frames, i.e., the gain coefficient can be different for different audio frames. In this way, the first suppression gain obtained based on the gain coefficient and the initial suppression gain can make the first suppression gain more consistent with and adaptive to each audio frame, improve the effect of sibilance suppression, and improve the flexibility of sibilance processing.
[0024] Optionally, the gain coefficient is related to the frame sharpness of the current audio frame and a second sharpness threshold.
[0025] By associating the gain coefficient with the frame sharpness of the audio frame and the second sharpness threshold, the gain coefficient can be conveniently dynamically changed for different audio frames.
[0026] Optionally, in the step of determining the second suppression gain of each extended suppression frequency band in the above embodiment, the second suppression gain of each extended suppression frequency band is determined according to the first suppression gain as a reference, in a manner of increasing a gain step for each incremental frequency band width, in a frequency variation direction from the first boundary frequency to the second boundary frequency of each extended suppression frequency band.
[0027] Optionally, the gain step is obtained according to the following formula:
[0028]
[0029] Wherein, step is the gain step, P is the number of frequency points in each extended suppression frequency band, and gain is the first suppression gain.
[0030] Optionally, the method further comprises: in a case where the frequency band range of the core suppression frequency band includes the frequency band range of the adaptive suppression frequency band, the first suppression gain acts on the energy of the current audio frame in the core suppression frequency band.
[0031] Optionally, in the step of determining the current audio frame as a sibilance frame in the above embodiment, a target feature value of the current audio frame is determined; and in a case where the target feature value satisfies a second preset condition, the current audio frame is determined as a sibilance frame.
[0032] Optionally, the target feature value is the frame sharpness of the current audio frame; and the second preset condition is that the frame sharpness of the current audio frame is greater than a second sharpness threshold.
[0033] In a second aspect, an electronic device is provided, which is configured to execute the method provided in the first aspect. Specifically, the electronic device can include modules configured to perform any of the possible implementation manners of the first aspect.
[0034] In a third aspect, an electronic device is provided, which includes a processor. The processor is coupled with a memory and is configured to execute instructions in the memory to implement the method in any of the possible implementation manners of the first aspect. Optionally, the electronic device further includes the memory. Optionally, the electronic device further includes a communication interface, and the processor is coupled with the communication interface.
[0035] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program. When the computer program is executed by an apparatus, the apparatus is caused to implement the method in any of the possible implementation manners of the first aspect.
[0036] In a fifth aspect, a computer program product is provided, which includes instructions. When the instructions are executed by a computer, the apparatus is caused to implement the method in any of the possible implementation manners of the first aspect.
[0037] In a sixth aspect, a chip is provided, which includes an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through internal connection paths. The processor is configured to execute code in the memory. When the code is executed, the processor is configured to execute the method in any of the possible implementation manners of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 FIG. 1 is a schematic architecture diagram of an electronic device provided by an embodiment of the present application.
[0039] Figure 2 FIG. 2 is a software structure block diagram of an electronic device provided by an embodiment of the present application.
[0040] Figure 3 FIG. 3 is a schematic flow chart of a sibilance processing method provided by an embodiment of the present application.
[0041] Figure 4 FIG. 4 is a schematic diagram of a suppression frequency band and a suppression gain provided by an embodiment of the present application.
[0042] Figure 5 FIG. 5 is another schematic diagram of a suppression frequency band and a suppression gain provided by an embodiment of the present application.
[0043] Figure 6 FIG. 6 is still another schematic diagram of a suppression frequency band and a suppression gain provided by an embodiment of the present application.
[0044] Figure 7Fig. 2 is another schematic flowchart of the method for sibilance processing provided by the embodiments of the present application.
[0045] Figure 8 Fig. 3 is a schematic diagram of the sharpness distribution of an audio frame represented in a bark domain provided by the embodiments of the present application.
[0046] Figure 9 Fig. 4 is a schematic diagram of a candidate suppression interval determined in a multiple comparison process provided by the embodiments of the present application.
[0047] Figure 10 Fig. 5 is a comparative diagram of the energy of an audio frame before and after sibilance processing of the audio frame provided by the embodiments of the present application.
[0048] Figure 11 Fig. 6 is a schematic flowchart of the method for determining the frame sharpness of an audio frame provided by the embodiments of the present application.
[0049] Figure 12 Fig. 7 is still another schematic flowchart of the method for sibilance processing provided by the embodiments of the present application.
[0050] Figure 13 Fig. 8 is an exemplary block diagram of the apparatus for sibilance processing provided by the embodiments of the present application.
[0051] Figure 14 Fig. 9 is a schematic structural diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0052] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0053] In order to facilitate the understanding of the solutions, first, the related terms involved in the embodiments of the present application are introduced.
[0054] Sibilance
[0055] Sibilance refers to all the fricative sounds emitted by a person, with strong and heavy friction consonants, corresponding to a higher sharpness, which is generally not suitable for human ears to listen to. Sibilance is usually concentrated in the frequency range of 4kHz to 8kHz.
[0056] For the received audio signal, the electronic device will process the sibilance in the audio signal to suppress the sibilance, so that each frame of audio in the audio signal is within a suitable sharpness range, avoiding the damage to the hearing of human ears caused by the sibilance with higher sharpness.
[0057] Sharpness
[0058] Sharpness is a perceptual measure of how pleasing a sound is, and is used to distinguish between sounds that are sharp or dull. Sharpness perception is mainly related to the spectral content and center frequency of narrowband sounds, and is independent of the loudness level and the details of the spectral content. The unit of sharpness is acum, and the reference sound for 1 acum is a narrowband noise with a center frequency of 1 kHz, a sound pressure level of 60 dB, and a bandwidth equal to one critical band.
[0059] Loudness
[0060] Loudness is the strength of a sound, and describes the loudness of a sound, which is the subjective perception of a sound by a human ear. The unit of loudness is sone, and the loudness of a pure tone with a frequency of 1 kHz and a sound pressure level of 40 dB is defined as 1 sone.
[0061] The perception of a sound by a human ear is not only related to sound pressure, but also related to frequency. Sounds with the same sound pressure level but different frequencies sound different in loudness. For example, an air compressor and an electric saw, both with a sound pressure level of 100 decibels, sound very different in loudness. According to the characteristics of the perception of a sound by a human ear, a subjective auditory perception quantity of a sound by a human is defined according to sound pressure and frequency, and is called loudness level, with a unit of phon.
[0062] Bark scale
[0063] Because the perception of a sound by a human ear (such as frequency and pitch) is nonlinear, a series of scales have been developed to measure the perception of a sound, such as the mel scale, the bark scale, and various other scales. The bark scale is an early proposed psychoacoustic scale of sound. The bark scale is a scale in Hz that maps physical frequencies to 24 critical bands in psychoacoustics. In simple terms, it is a conversion of physical frequencies to psychoacoustic frequencies.
[0064] For example, the range of 0 Hz to 16 kHz in the audible range of a human ear is usually divided into 24 critical bands, and the width of a critical band (referred to as critical bandwidth) is represented by a critical band level, with a unit of bark: 1 bark = 1 critical bandwidth. When the frequency f < 500 Hz, 1 bark = f / 100, and the critical bandwidth is almost constant at 100 Hz; when the frequency f > 500 Hz, 1 bark = 4log(f / 100), and the critical bandwidth increases with the increase of the center frequency, and is about 20% of the center frequency. Therefore, the frequency domain represented by the bark scale can be understood as a bark domain, and the unit of the bark domain is bark. The bandwidth of different positions in the bark domain is not necessarily the same.
[0065] Exemplarily, taking an example of dividing 0Hz-16kHz in the audible range of human ear into 24 critical bands, the critical bandwidth, frequency boundary and critical band level of different center frequencies are shown in Table 1.
[0066] Table 1
[0067]
[0068]
[0069] Regarding the critical band, the critical band refers to the frequency bandwidth of the hearing filter generated due to the cochlea structure. In the hearing system, the cochlea plays a role of frequency spectrum analysis, and a specific position point on the basilar membrane is the maximum response to a certain characteristic frequency (CF), and the response of the point decreases when the sound wave deviates from the CF, so each point on the basilar membrane can be equivalent to a band-pass filter with a CF, and the entire hearing system can be equivalent to a series of band-pass filters with continuous CFs, which are called "hearing filters". The critical band is a reflection of the band-pass filtering function of the hearing system, and the bandwidth of the hearing filter is the critical bandwidth.
[0070] Therefore, in acoustic research, people use hearing filters to simulate different critical bands, and the audio signal can present 24 critical bands on the frequency band, which are 1 to 24, which is the bark domain. Eberhanrd Zwicker proposed that the 24 critical bands of hearing can be roughly simulated using hearing filters, that is, the signal method is described in the bark domain.
[0071] It should be understood that although the above example is to divide the frequency band of 0Hz-16kHz into 24 critical bands (i.e., 24 barks), in the implementation, the frequency band of 0Hz-16kHz can also be subdivided into more critical bands, for example, the frequency band of 0Hz-16kHz is divided into 240 critical bands (i.e., 240 barks). In the following embodiments, the related processing process of the audio frame is described taking 240 barks as an example.
[0072] For the sibilance, the electronic device will perform suppression processing on the sibilance in the audio signal to reduce the sibilance energy, so as to try to make each frame of audio in the audio signal be in a suitable sharpness range, and avoid the sibilance with high sharpness from causing damage to the hearing of the human ear. Taking a call scene as an example, if the sibilance energy is high, the speech will have an unnatural harsh feeling, which will reduce the quality of the audio signal and make the user feel annoyed. By suppressing the energy of the sibilance segment through the related algorithm, the human voice can be clearer and more smooth, and less harsh.
[0073] However, the quality of the audio signal is damaged to some extent when the sibilance suppression is performed, and the bandwidth of the suppression frequency band is small in view of the quality, which results in an unsatisfactory sibilance suppression effect. If the bandwidth of the suppression frequency band is blindly expanded, the quality of the audio signal will be damaged to a large extent.
[0074] To solve the above problems, the embodiment of the present application provides a sibilance processing method. In the case that an audio frame is determined to be a sibilance frame, a core suppression frequency band and an adaptive suppression frequency band of the audio frame are determined. The first suppression gain (less than 1) of the core suppression frequency band is a fixed value. In the case that the frequency band range of the adaptive suppression frequency band includes the frequency band range of the core suppression frequency band or there is a partial intersection between the frequency band range of the adaptive suppression frequency band and the frequency band range of the core suppression frequency band, the suppression gain of at least one extended suppression frequency band other than the core suppression frequency band in the adaptive suppression frequency band is smoothed according to a gain step based on the first suppression gain of the core suppression frequency band, so that the suppression gain of each extended frequency band increases from the first suppression gain to 1. The first suppression gain and the second suppression gain are used to process the sibilance of the audio frame. In this way, the suppression frequency band is expanded by the adaptive suppression frequency band in addition to the core suppression frequency band, which effectively improves the sibilance suppression effect. Moreover, the suppression gain of at least one extended suppression frequency band other than the core suppression frequency band in the adaptive suppression frequency band is smoothed according to a gain step, so that the suppression gain of the extended suppression frequency band can be smoothly transitioned from the first suppression gain to 1, which reduces the damage to the quality of the audio frame. In summary, the embodiment of the present application balances the quality and sibilance suppression, and effectively improves the sibilance suppression effect while minimizing the damage to the quality.
[0075] It can be understood that the core suppression frequency band and the adaptive suppression frequency band are both frequency bands used for sibilance suppression in the audio frame.
[0076] The core suppression frequency band is a frequency band with a small frequency band range, and the sibilance energy is high and concentrated, which is the core frequency band for sibilance suppression and can suppress more sibilance energy.
[0077] The adaptive suppression frequency band is used to expand the core suppression frequency band to additionally increase the frequency band for sibilance suppression. Moreover, the amplitude of sibilance suppression on the extended suppression frequency band expanded based on the adaptive suppression frequency band is small, and the sibilance energy of the extended suppression frequency band is smoothly processed, which can maintain the quality as much as possible.
[0078] The sibilance processing method provided by the embodiment of the present application can be applied to various electronic devices capable of playing and processing sound, such as mobile phones, tablet computers, wearable devices, notebook computers, netbooks, sound boxes, multimedia consoles, vehicle terminal devices, and the like. The embodiment of the present application does not make any limitation on the specific type of the electronic device.
[0079] Figure 1 A structural diagram of the electronic device 100 is shown. The electronic device 100 can include a processor 110, a mobile communication module 120, a wireless communication module 130, an audio module 140, a display screen 151, a camera 152, and a subscriber identification module (SIM) card interface 153, etc. Among them, the audio module 140 includes a speaker 140A, a receiver 140B, a microphone 140C, and a headset interface 140D.
[0080] In the embodiments of the present application, the processor 110 is configured to determine whether an audio frame is a tooth frame, and in a case where it is determined that the audio frame is a tooth frame, determine a suppression frequency band and a suppression gain, so as to perform tooth processing on the audio frame through the suppression frequency band and the suppression gain.
[0081] For example, the processor 110 is configured to determine whether an audio frame is a tooth frame according to a frame sharpness of the audio frame.
[0082] For example, the processor 110 is further configured to perform the following steps: determining a core suppression frequency band and dynamically determining an adaptive suppression frequency band of the audio frame to obtain the suppression frequency band; determining a first suppression gain (less than 1) of the core suppression frequency band; taking the first suppression gain of the core suppression frequency band as a reference, performing smoothing processing on a gain of at least one extended suppression frequency band other than the core suppression frequency band in the adaptive suppression frequency band according to a gain step to obtain a second suppression gain (less than 1) of each extended suppression frequency band; and performing processing on the audio frame according to the first suppression gain and the second suppression gain to suppress the tooth.
[0083] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.
[0084] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of instruction fetching and instruction execution.
[0085] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.
[0086] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0087] The wireless communication function of the electronic device 100 can be realized through the antenna 1, the antenna 2, the mobile communication module 120, the wireless communication module 130, the modem processor, and the baseband processor, etc.
[0088] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antennas can be used in combination with tuning switches.
[0089] The mobile communication module 120 can provide a solution including 2G / 3G / 4G / 5G, etc. wireless communication applied to the electronic device 100. The mobile communication module 120 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 120 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the same to the modem processor to be demodulated. The mobile communication module 120 can also amplify the signal modulated by the modem processor, and radiate the same as electromagnetic waves through the antenna 1. In some embodiments, at least part of the function modules of the mobile communication module 120 can be disposed in the processor 110. In some embodiments, at least part of the function modules of the mobile communication module 120 can be disposed in the same device as at least part of the modules of the processor 110.
[0090] The modem processor can include a modulator and a demodulator. The modulator is used to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the microphone 170B, etc.), or displays an image or a video through the display screen 194. In some embodiments, the modem processor can be an independent device. In other embodiments, the modem processor can be independent of the processor 110, and disposed in the same device as the mobile communication module 120 or other function modules.
[0091] The wireless communication module 130 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module 130 can be one or more devices that integrate at least one communication processing module. The wireless communication module 130 receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 130 can also receive signals to be transmitted from the processor 110, frequency-modulate them, amplify them, and radiate them as electromagnetic waves via the antenna 2.
[0092] In some embodiments, the antenna 1 and the mobile communication module 120 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 130 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include a global positioning system (GPS), a global navigation satellite system (GLONASS), a beidu navigation satellite system (BDS), a quasi-zenith satellite system (QZSS), and / or a satellite based augmentation systems (SBAS).
[0093] The electronic device 100 implements a display function through a GPU, a display 151, and an application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display 151 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.
[0094] The electronic device 100 can implement a photographing function through an ISP, a camera 152, a video codec, a GPU, a display 151, and an application processor, etc.
[0095] The electronic device 100 can implement an audio function through an audio module 140, a speaker 140A, a receiver 140B, a microphone 140C, an earphone interface 140D, and an application processor, etc. For example, music playing, recording, etc.
[0096] The audio module 140 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 140 can also be configured to encode and decode audio signals. In some embodiments, the audio module 140 can be disposed in the processor 110, or some functional modules of the audio module 140 can be disposed in the processor 110.
[0097] The speaker 140A, also known as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or take a hands-free call through the speaker 140A.
[0098] The receiver 140B, also known as a "earpiece", is configured to convert an audio electrical signal into a sound signal. When the electronic device 100 receives a call or a voice message, the user can take the voice message by holding the receiver 140B close to the ear.
[0099] The microphone 140C, also known as a "microphone", "voice microphone", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can make a sound by holding the mouth close to the microphone 140C, and input the sound signal into the microphone 140C. The electronic device 100 can be provided with at least one microphone 140C. In other embodiments, the electronic device 100 can be provided with two microphones 140C, in addition to collecting sound signals, the noise reduction function can also be realized. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 140C, in addition to collecting sound signals, noise reduction, it can also identify the source of the sound, realize the function of directional recording, etc.
[0100] The earphone interface 140D is configured to connect a wired earphone.
[0101] The SIM card interface 153 is configured to connect a SIM card. The SIM card can be inserted into or pulled out of the SIM card interface 153, so as to realize contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, N is a positive integer greater than 1. The electronic device 100 interacts with the network through the SIM card, and realizes the functions of call and data communication, etc. In some embodiments, the electronic device 100 adopts eSIM, i.e. embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0102] It should be understood that the structure illustrated by the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware. For example, the electronic device 100 can further include an external memory interface, a charge management module, a power management module, a battery, an internal memory, a sensor module, a key, a motor, an indicator, and the like.
[0103] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. The embodiments of the present application take the Android system with a layered architecture as an example to illustrate the software structure of the electronic device 100.
[0104] Figure 2 is a software structure block diagram of the electronic device 100 provided by the embodiments of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through a software interface. In some embodiments, the Android system is divided into five layers, from top to bottom, the application layer, the application framework layer, the Android runtime and system library, the hardware abstraction layer, and the kernel layer.
[0105] The application layer can include a series of application packages.
[0106] Illustratively, the application packages can include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, and the like.
[0107] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions.
[0108] Illustratively, the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0109] The window manager is used to manage the window program. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and intercept the screen, and the like.
[0110] The content provider is used to store and obtain data, and make the data accessible to the applications. The data can include video, image, audio, dialed and received calls, browsing history and bookmarks, phonebook, and the like.
[0111] The view system includes visual controls, such as controls that display text, controls that display pictures, and the like. The view system can be used to build an application. A display interface can be composed of one or more views. For example, a display interface that includes a short message notification icon can include a view that displays text and a view that displays a picture.
[0112] The telephony manager is used to provide the communication function of the electronic device 100. For example, the management of the call state (including the connection, hang-up, and the like).
[0113] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, and the like.
[0114] The notification manager enables the application to display notification information in the status bar, which can be used to convey a message of the notification type, can automatically disappear after a short stay, and does not require user interaction. For example, the notification manager is used to notify the completion of the download, the message reminder, and the like. The notification manager can also be a notification that appears in the form of a chart or a scroll bar text in the top status bar of the system, such as a notification of an application running in the background, and can also be a notification that appears in the form of a dialogue window on the screen. For example, the text information is prompted in the status bar, a prompt sound is emitted, the electronic device is vibrated, the indicator light is blinked, and the like.
[0115] The Android runtime includes the core library and the virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0116] The core library contains two parts: one part is the function function that the java language needs to call, and the other part is the core library of Android.
[0117] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application layer and the application framework layer into a binary file. The virtual machine is used to perform the management of the object life cycle, the stack management, the thread management, the security and exception management, and the garbage collection, and the like.
[0118] The system library can include multiple functional modules. For example: the surface manager, the media library, the audio and video processing library, and the like.
[0119] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0120] The media library supports the playback and recording of multiple commonly used audio, video formats, and the like, as well as static image files. The media library can support multiple audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, and the like.
[0121] The hardware abstraction layer is an abstract layer between the kernel layer and the system library. The hardware abstraction layer can be a package of hardware drivers, and provides a unified interface for the call of the upper layer application.
[0122] Exemplarily, the hardware abstraction layer can include an audio module, a video module, a hardware configuration module, and the like.
[0123] The kernel layer is a layer between hardware and software, and includes various drivers.
[0124] Exemplarily, the kernel layer can include a display driver, a camera driver, an audio driver, a sensor driver, and the like.
[0125] It should be understood that the embodiments of the present application are only exemplified by using the Android system, and the scheme of the present application can also be implemented in other operating systems (for example, Windows system, IOS system, and the like) as long as the functions of the various functional modules are similar to those of the embodiments of the present application.
[0126] It should be noted that the processing of determining whether an audio frame is a tooth frame, determining a suppression frequency band and a suppression gain, and tooth suppression according to the suppression frequency band and the suppression gain can be implemented in multiple software architecture layers of the electronic device. For example, the application layer of the electronic device can perform the above processing, and the application layer involves audio / video related application programs (for example, calls, music, videos, and the like). In addition, the audio / video processing library in the system library of the electronic device, the audio / video module of the hardware abstraction layer, and the audio / video driver of the driver layer can all perform the above processing, thereby implementing tooth processing of the audio frame. The software layer level of the specific implementation of the tooth processing is not limited in the embodiments of the present application.
[0127] The method of tooth processing of the embodiments of the present application can be applied to various application scenarios of processing audio signals, for example, a call scenario, a scenario of playing music or video, and the like.
[0128] It should also be understood that the audio signal includes multiple audio frames, and the electronic device performs the method of tooth processing of the embodiments of the present application on each audio frame, that is, in the case of determining that an audio frame is a tooth frame, tooth processing is performed on the audio frame.
[0129] In the following, the method of tooth processing will be described in combination with Figures 3 to 12 The process of the method of tooth processing will be described.
[0130] It should be noted that the electronic device can also be a processor or a chip in the electronic device, and the like, which executes the method of tooth processing of the embodiments of the present application. In order to facilitate description, the electronic device is taken as the execution subject, and the embodiments of the present application are described.
[0131] Figure 3FIG. 3 is a schematic flowchart of a method 300 for sibilance processing according to an embodiment of the present application. Figures 4 to 6 FIG. 4 is a schematic diagram of a suppressed frequency band and a suppressed gain in various cases according to an embodiment of the present application.
[0132] In step S310, the electronic device acquires a current audio frame.
[0133] In this step, the electronic device performs frame processing on the collected audio signal to acquire each audio frame, and the current audio frame is any audio frame.
[0134] It should be understood that the audio signal can be real-time data (for example, real-time collected data in a call scene) or recorded data, and the present application does not make any limitation.
[0135] In step S320, the electronic device determines a target feature value of the current audio frame.
[0136] The target feature value can be any audio parameter that can be used to measure or evaluate the audio.
[0137] In an example, the target feature value can include sharpness for measuring the sharpness of the sound, and for the current audio frame, the target feature value can be the frame sharpness of the current audio frame.
[0138] In another example, the target feature value can also include other audio parameters such as loudness.
[0139] In some embodiments, the current audio frame is a frequency domain signal. In implementation, the electronic device converts the current audio frame from a time domain signal to a frequency domain signal, and then determines the target feature value of the frequency domain signal.
[0140] For the process of the electronic device determining the target feature value of the current audio frame, reference can be made to the specific description of this step in method 500 below, which will not be repeated here.
[0141] In step S330, the electronic device determines whether the current audio frame is a sibilance frame according to the target feature value of the current audio frame.
[0142] In the case where it is determined that the current audio frame is a sibilance frame, the electronic device performs step S340; in the case where it is determined that the current audio frame is not a sibilance frame, the electronic device does not perform sibilance processing on the current audio frame, and outputs the current audio frame in a normal manner.
[0143] In some embodiments, in the case where the target feature value of the current audio frame meets a preset condition, the electronic device determines that the current audio frame is a sibilance frame; in the case where the target feature value of the current audio frame does not meet the preset condition, the electronic device determines that the current audio frame is not a sibilance frame.
[0144] In an example, the preset condition is that the target feature value of the current audio frame is greater than a threshold value.
[0145] That is, in a case where the target feature value of the current audio frame is greater than the threshold value, the electronic device determines that the current audio frame is an abrasive frame; in a case where the target feature value of the current audio frame is less than the threshold value, the electronic device determines that the current audio frame is not an abrasive frame.
[0146] It should be understood that, for a case where the target feature value is equal to the threshold value, it can be divided into a case that satisfies the determination of the current audio frame as an abrasive frame, or a case that does not satisfy the current audio frame as an abrasive frame, which is not limited here.
[0147] In step S340, the electronic device determines the suppression frequency band.
[0148] It should be understood that the suppression frequency band is a frequency band that needs to be processed by the abrasive processing, and is a part of the frequency band occupied by the current audio frame.
[0149] In some embodiments, the suppression frequency band can include a full set of two frequency bands, which are respectively a core suppression frequency band and an adaptive suppression frequency band. The electronic device needs to determine the core suppression frequency band and the adaptive suppression frequency band to obtain the suppression frequency band and determine the suppression gain.
[0150] In the algorithm design process, theoretically, the core suppression frequency band is a frequency band with a smaller frequency band range, and the purpose of determining the adaptive suppression frequency band in the embodiments of the present application is to expand the range of the suppression frequency band. Therefore, based on the design idea, in most cases, the frequency band range of the adaptive suppression frequency band includes the frequency band range of the core suppression frequency band or the frequency band range of the adaptive suppression frequency band and the frequency band range of the core suppression frequency band exist partial intersection, in a few cases, the frequency band range of the core suppression frequency band includes the frequency band range of the adaptive suppression frequency band. Especially for the call scene, the case that the frequency band range of the core suppression frequency band includes the frequency band range of the adaptive suppression frequency band rarely occurs, so the case that the frequency band range of the core suppression frequency band includes the frequency band range of the adaptive suppression frequency band can be regarded as an extreme scene of the embodiments of the present application.
[0151] In the following, the core suppression frequency band and the adaptive suppression frequency band are combined Figures 4 to 6 The suppression frequency band formed based on the core suppression frequency band and the adaptive suppression frequency band is illustratively described.
[0152] Case 1, the frequency band range of the adaptive suppression frequency band includes the frequency band range of the core suppression frequency band
[0153] Reference is made to Figure 4In case 1, the maximum frequency point of the adaptive suppression frequency band is greater than the maximum frequency point of the core suppression frequency band, and the minimum frequency point of the adaptive suppression frequency band is less than the minimum frequency point of the core suppression frequency band. In this case, the suppression frequency band is the adaptive suppression frequency band with a larger frequency range, that is, the union of the adaptive suppression frequency band and the core suppression frequency band is the adaptive suppression frequency band.
[0154] For ease of description, the frequency band in the adaptive suppression frequency band other than the core suppression frequency band is referred to as an extended suppression frequency band. In case 1, the suppression frequency band includes two extended suppression frequency bands, which are located on the two sides of the core suppression frequency band.
[0155] Case 2, the frequency range of the adaptive suppression frequency band partially intersects with the frequency range of the core suppression frequency band
[0156] In case 2, both of the extreme frequency points of the adaptive suppression frequency band are greater than or less than the extreme frequency points of the core suppression frequency band.
[0157] In an example, referring to (a) in Figure 5 , the maximum frequency point of the adaptive suppression frequency band is greater than the maximum frequency point of the core suppression frequency band, and the minimum frequency point of the adaptive suppression frequency band is between the minimum frequency point and the maximum frequency point of the core suppression frequency band. In this case, the suppression frequency band includes the core suppression frequency band and one extended suppression frequency band in the adaptive suppression frequency band other than the intersection of the two frequency bands, and the extended suppression frequency band is located on one side of the maximum frequency point of the core suppression frequency band.
[0158] In another example, referring to (b) in Figure 5 , the minimum frequency point of the adaptive suppression frequency band is less than the minimum frequency point of the core suppression frequency band, and the maximum frequency point of the adaptive suppression frequency band is between the minimum frequency point and the maximum frequency point of the core suppression frequency band. In this case, the suppression frequency band includes the core suppression frequency band and one extended suppression frequency band in the adaptive suppression frequency band other than the intersection of the two frequency bands, and the extended suppression frequency band is located on one side of the minimum frequency point of the core suppression frequency band.
[0159] It can be seen that in case 2, the suppression frequency band includes the core suppression frequency band and one extended suppression frequency band, or in other words, the union of the adaptive suppression frequency band and the core suppression frequency band includes the core suppression frequency band and the extended suppression frequency band.
[0160] Case 3, the frequency range of the core suppression frequency band includes the frequency range of the adaptive suppression frequency band
[0161] Referring to Figure 6 , in case 3, the maximum frequency point of the adaptive suppression frequency band is less than the maximum frequency point of the core suppression frequency band, and the minimum frequency point of the adaptive suppression frequency band is greater than the minimum frequency point of the core suppression frequency band.
[0162] In this case, in an example, the suppression frequency band is a core suppression frequency band, that is, the full set of the adaptive suppression frequency band and the core suppression frequency band is the core suppression frequency band, and there is no extended suppression frequency band. It can be understood that, in the design process, the frequency range of the core suppression frequency band is small in itself, and the suppression of the energy of the core suppression frequency band has little effect on the sound quality or is within an acceptable range, so even if the frequency range of the core suppression frequency band is smaller than the frequency range of the adaptive suppression frequency band, in order to facilitate algorithm implementation and avoid excessive judgment logic, the core suppression frequency band can be taken as the suppression frequency band, and the energy of the core suppression frequency band can be treated as the sibilance.
[0163] Of course, in other examples, for case 3, the adaptive suppression frequency band can also be taken as the suppression frequency band, and the energy of the small-range adaptive suppression frequency band can be suppressed, so as to better reduce the effect on the sound quality while achieving the effect of sibilance suppression.
[0164] In step S350, the electronic device determines the suppression gain.
[0165] The suppression gain is the suppression gain corresponding to the suppression frequency band, and is a value less than 1.
[0166] In the embodiments of the application, the frequency band that does not need to be treated as sibilance is referred to as a non-suppression frequency band, and the suppression gain of the non-suppression frequency band is 1, that is, the energy of the non-suppression frequency band does not need to be suppressed.
[0167] In the embodiments of the application, the suppression gain of the core suppression frequency band is always a fixed value.
[0168] In the above case 1 and case 2, the suppression gain of the suppression frequency band includes two parts of gain, one part of gain is the suppression gain of the core suppression frequency band (denoted as the first suppression gain), and the other part of gain is the suppression gain of the extended suppression frequency band in the adaptive suppression frequency band except the core suppression frequency band (denoted as the second suppression gain), wherein the second suppression gain is greater than the first suppression gain, and the second suppression gain is incremental, and in the extended suppression frequency band, the second suppression gain is increased from the first suppression gain to the gain of the non-suppression frequency band, that is, the second suppression gain is increased from the first suppression gain to 1.
[0169] For the second suppression gain of the extended suppression frequency band, the second suppression gain is gradually increased, which can be understood as that the second suppression gain includes a plurality of gain values different from each other, the extended suppression frequency band is divided into a plurality of sub-frequency bands, one gain corresponds to one sub-frequency band of the extended suppression frequency band, and in the extended suppression frequency band, the plurality of gains are increased from the first suppression gain to 1 according to the plurality of corresponding sub-frequency bands. In an example, the frequency band widths of the plurality of sub-frequency bands are the same. In other examples, the frequency band widths of the plurality of sub-frequency bands are not completely the same, which is not limited here.
[0170] For example, the extended suppression frequency band includes five sub-frequency bands with the same frequency band width, and the five sub-frequency bands are denoted as sub-frequency band 1, sub-frequency band 2, sub-frequency band 3, sub-frequency band 4 and sub-frequency band 5 in order of frequency from small to large. The second suppression gain includes five gains, and the five gains are denoted as gain 1, gain 2, gain 3, gain 4 and gain 5 in order of gain from small to large. Then, the corresponding relationship between the five sub-frequency bands and the five gains is: sub-frequency band 1-gain 1, sub-frequency band 2-gain 2, sub-frequency band 3-gain 3, sub-frequency band 4-gain 4 and sub-frequency band 5-gain 5.
[0171] In the following, the process of determining the core suppression frequency band and the adaptive suppression frequency band and determining the two suppression gains by the electronic device will be described in detail. Figures 4 to 6 The suppression gain of the suppression frequency band is schematically described.
[0172] For case 1, continuing to refer to Figure 4 , the suppression frequency band includes the core suppression frequency band and two extended suppression frequency bands, and it is assumed that the first suppression gain is 0.6. For each extended suppression frequency band, the second suppression gain of each extended suppression frequency band is increased from 0.6 to 1 starting from the boundary frequency close to the core suppression frequency band, and the gain of the non-suppression frequency band is 1.
[0173] For case 2, continuing to refer to Figure 5 , the suppression frequency band includes the core suppression frequency band and one extended suppression frequency band, and it is assumed that the first suppression gain is 0.6. For the extended suppression frequency band, the second suppression gain of the extended suppression frequency band is increased from 0.6 to 1 starting from the boundary frequency close to the core suppression frequency band.
[0174] For case 3, continuing to refer to Figure 6 Since the suppression frequency band is the core suppression frequency band, the suppression gain of the suppression frequency band is the first suppression gain of the core suppression frequency band, which is a fixed value of 0.6.
[0175] It should be noted that Figure 4 and Figure 5 The second suppression gain of the extended suppression frequency band shown as a smooth straight line is only a schematic description, only showing that the second suppression gain is a kind of increasing trend, and does not represent that the second suppression gain of the extended suppression frequency band in the actual curve is also formed as a straight line. In the actual curve, the second suppression gain can be a wavy line composed of multiple points, and only the overall trend is increasing.
[0176] The process of determining the core suppression frequency band and the adaptive suppression frequency band and determining the two suppression gains by the electronic device will be described in detail below, which will not be described here.
[0177] In step S360, the electronic device performs the fricative processing on the current audio frame according to the suppression frequency band and the suppression gain.
[0178] In this step, the electronic device applies the suppression gain to the energy of the current audio frame in the suppression frequency band, to achieve the sibilance processing of the current audio frame, to suppress the sibilance.
[0179] For the above case 1 and case 2, the suppression frequency band includes a core suppression frequency band and an extended suppression frequency band, and the suppression gain includes a first suppression gain and a second suppression gain. The first suppression gain is applied to the energy of the current audio frame in the core suppression frequency band, and the second suppression gain is applied to the energy of the current audio frame in the extended suppression frequency band.
[0180] Since the first suppression gain and the second suppression gain are both less than 1, applying the two suppression gains to the corresponding suppression frequency band is equivalent to multiplying the original energy of the current audio frame by a coefficient less than 1, which reduces the energy of the current audio frame in the suppression frequency band and achieves the purpose of sibilance suppression. In addition, since the extended suppression frequency band is determined based on the adaptive suppression frequency band and the core suppression frequency band, it is an extended suppression frequency band of the core suppression frequency band, which increases the frequency range of the suppression frequency band. Processing the sibilance in a wide range of suppression frequency band can achieve better sibilance suppression effect.
[0181] In addition, since the core suppression frequency band is extended, the second suppression gain of the extended suppression frequency band gradually increases from the first suppression gain, which can avoid the problem of audio quality damage caused by still using a small gain to suppress the energy of the extended suppression frequency band. Therefore, the gradually increasing second suppression gain can gently damage the audio quality, improve the sibilance suppression effect, and reduce the damage to the audio quality.
[0182] For case 3, although the frequency range of the adaptive suppression frequency band is smaller than that of the core suppression frequency band, the frequency range of the core suppression frequency band itself is relatively small. Processing the sibilance in the core suppression frequency band as the final suppression frequency band has little effect on the audio quality or is within an acceptable range while achieving sibilance suppression. Moreover, since the first suppression gain is the gain of the core suppression frequency band, directly using the first suppression gain to process the sibilance in the core suppression frequency band can avoid the increase in algorithm difficulty caused by additional steps, facilitating algorithm implementation.
[0183] Figure 7 Fig. 4 is a schematic flowchart of a sibilance processing method 400 provided by an embodiment of the present application. Compared with the method 300 shown in Fig. 3, the method 400 focuses on how the electronic device determines the suppression frequency band and the suppression gain when an audio frame is determined to be a sibilance frame, and the process of sibilance processing according to the suppression frequency band and the suppression gain. Figure 3
[0184] In the method 400, the electronic device determines the core suppression frequency band and the adaptive suppression frequency band according to the sharpness distribution of an audio frame, to obtain a suppression frequency band, determines a suppression gain according to different situations, and finally performs the sibilance processing on the audio frame according to the suppression frequency band and the suppression gain.
[0185] As described above, since the perception of sound (e.g., frequency, pitch) by the human ear is nonlinear, it is more accurate to measure the perception of sound by the human ear by using a scale different from the linear frequency domain. Exemplarily, the bark scale is used to measure the sharpness distribution of the audio frame in the embodiments of the present application, and after the core suppression interval and the adaptive suppression interval in the bark domain are determined, the core suppression interval and the adaptive suppression interval are converted to the linear frequency domain to obtain the corresponding core suppression frequency band and adaptive suppression frequency band.
[0186] It should be understood that the sharpness of the audio frame measured by the bark scale and the determination of the core suppression frequency band and the adaptive suppression frequency band according to the sharpness distribution listed below are only illustrative, and other scales such as the mel spectrum scale can also be used to measure the perception of sound by the human ear. When other scales such as the mel spectrum scale are used, it is not limited to determining the core suppression frequency band and the adaptive suppression frequency band according to the target characteristic value of the sharpness of the current audio frame, and the core suppression frequency band and the adaptive suppression frequency band can be determined according to other target characteristic values for representing audio characteristics.
[0187] In step S410, the electronic device determines a first bark frequency band corresponding to the maximum sharpness according to the sharpness distribution of the current audio frame.
[0188] For the sharpness measured by the bark scale, it represents the sharpness distribution of the current audio frame in each bark frequency band in the bark domain. Here, a bark frequency band is a frequency band represented by a bark label in the bark domain. Different bark labels represent different frequency bands, and the frequency band widths of the frequency bands represented by different bark labels are not necessarily the same.
[0189] The maximum sharpness represents the maximum sharpness in each sharpness corresponding to each bark frequency band of the current audio frame. In order to facilitate description, the bark frequency band corresponding to the maximum sharpness of the current audio frame is defined as the first bark frequency band.
[0190] Figure 8 is a schematic diagram of the sharpness distribution of the audio frame represented by the bark domain provided by the embodiments of the present application.
[0191] Reference Figure 8, the horizontal axis is bark domain, 0Hz-16kHz in the audible range of human ear is divided into 240 critical frequency bands, a number represents a bark label, and a bark label represents a bark frequency band. For example, the center frequency of the 50th bark is 500Hz, and the frequency band width is 10Hz, so the bark frequency band corresponding to the 50th bark is 495Hz-505Hz; the center frequency of the 150th bark is 2718Hz, and the frequency band width is 31.25Hz, so the bark frequency band corresponding to the 150th bark is 2702.375Hz-2733.625Hz. In addition, the frequency band width corresponding to different bark labels is not necessarily the same, for example, the frequency band width corresponding to the 50th bark is 10Hz, and the frequency band width corresponding to the 150th bark is 31.25Hz. The vertical axis is sharpness, and the bark frequency bands represented by different bark labels have different sharpness. Among the sharpnesses corresponding to the respective bark frequency bands, the peak A is the maximum sharpness, and the bark frequency band corresponding to the bark label represented by the peak A is the first bark frequency band.
[0192] In step S420, the electronic device determines the core suppression interval according to the first bark frequency band. In some embodiments, the electronic device determines the core suppression interval according to the first bark frequency band and the bandwidth parameter.
[0193] It can be understood that the bandwidth parameter is used to expand the frequency band, and can be used to determine the bark frequency band that needs to be expanded. The bandwidth parameter can be a predefined parameter, or a parameter that dynamically changes according to different audio frames, which is not limited here.
[0194] Exemplarily, the bandwidth parameter is used to indicate the number M1 of bark frequency bands. In implementation, the electronic device takes the first bark frequency band as the center frequency band, expands M1 bark frequency bands to both sides of the first bark frequency band respectively, and obtains the core suppression interval. The core suppression interval can be expressed in the following manner: [(index max -M1)bark, (index max +M1)bark], unit: bark, wherein index max represents the bark label of the first bark frequency band corresponding to the maximum sharpness.
[0195] For example, referring to Figure 8 , it is assumed that the bark label of the first bark frequency band is 165, that is, the first bark frequency band is the 165th bark, and M1=35, then index max -M1=165-35=130, index max+M1= 165 + 35 = 200, thus, the core suppression interval is [130bark, 200bark], i.e., the core suppression interval is greater than or equal to the 130th bark and less than or equal to the 200th bark.
[0196] In step S430, the electronic device determines the adaptive suppression interval according to the first bark band.
[0197] In this step, the electronic device takes the first bark band as a reference, gradually expands the suppression interval according to the sharpness corresponding to the bark bands on both sides of the first bark band, and finally obtains the adaptive suppression interval.
[0198] In some embodiments, the electronic device takes the first bark band as a reference, performs a plurality of comparison processes, determines the maximum sharpness of the sharpness corresponding to the two bark bands on both sides of the first bark band through each comparison process, determines the candidate suppression interval of each comparison process based on the maximum sharpness, realizes the purpose of expanding the suppression interval by a part of interval in each comparison process, and finally completes the expansion of the suppression interval through the plurality of comparison processes to obtain the adaptive suppression interval.
[0199] It should be understood that the two bark bands on both sides of the first bark band in any two comparison processes are not completely the same, wherein the two bark bands on both sides of the first bark band in the current comparison process are the two bark bands on both sides of the candidate suppression interval obtained in the last comparison process.
[0200] It should also be understood that the two bark bands on both sides of the candidate suppression interval represent two bark bands located on both sides of the candidate suppression interval with the same interval as the boundary frequency band of the candidate suppression interval. For example, the two bark bands on both sides of the candidate suppression interval represent the two bark bands located on both sides of the candidate suppression interval closest to the two boundary frequency bands of the candidate suppression interval. For example, the candidate suppression interval is [165bark, 166bark], and the two bark bands on both sides of the candidate suppression interval are the 164th bark and the 167th bark, respectively.
[0201] For determining the candidate compression interval of each comparison process based on the maximum sharpness, the electronic device may, for example, take the bark band corresponding to the maximum sharpness as one boundary band of the current candidate compression interval, and take the boundary band of the candidate compression interval obtained in the last comparison process that is far away from the bark band corresponding to the maximum sharpness as the other boundary band of the current candidate compression interval. For example, the candidate compression interval obtained in the last comparison process is [165 bark, 166 bark], the two bark bands on both sides of the candidate compression interval are the 164th bark and the 167th bark respectively, and the bark band corresponding to the maximum sharpness is the 164th bark. Therefore, for the current candidate compression interval, the bark band (the 164th bark) corresponding to the maximum sharpness is one boundary band of the current candidate compression interval, and the boundary band (the 166th bark) of the candidate compression interval [165 bark, 166 bark] that is far away from the bark band (the 164th bark) corresponding to the maximum sharpness. Therefore, the current candidate compression interval is [164 bark, 166 bark].
[0202] In the process of expanding the compression interval to determine the adaptive compression interval, in some embodiments, for the candidate compression interval obtained in each comparison process, the sharpness of each bark band in the candidate compression interval is accumulated and summed to obtain the band sharpness of the candidate compression interval. If the band sharpness of the candidate compression interval of a certain comparison process meets a preset condition (referred to as a first preset condition), the expansion of the interval is stopped, and the candidate compression interval obtained in the certain comparison process is determined as the adaptive compression interval. If the band sharpness of the candidate compression interval of a certain comparison process does not meet the first preset condition, the next comparison process is continued to continue expanding the compression interval until a certain candidate compression interval meets the first preset condition.
[0203] It should be understood that the band sharpness of the candidate compression interval represents the sum of the sharpness of each bark band of the candidate compression interval.
[0204] In an example, the first preset condition is that the band sharpness of the candidate compression interval obtained in each comparison process is greater than a first sharpness threshold.
[0205] That is, when the frequency band sharpness of the candidate suppression interval obtained in a certain comparison process is greater than the first sharpness threshold, it means that the frequency band sharpness of the candidate suppression interval is already large enough to meet the first preset condition, and there is no need to expand it any more. Therefore, the electronic device stops expanding the interval, and determines the candidate suppression interval obtained in the certain comparison process as the adaptive suppression interval. If the frequency band sharpness of the candidate suppression interval obtained in a certain comparison process is less than the first sharpness threshold, it means that the frequency band sharpness of the candidate suppression interval is too small to meet the first preset condition, and the suppression interval can still be expanded. Therefore, the next comparison process is continued to continue expanding the suppression interval.
[0206] It should be noted that, for the case that the frequency band sharpness of the candidate suppression interval is equal to the first sharpness threshold, it can be divided into the case of meeting the first preset condition, or the case of not meeting the first preset condition, which is not limited here.
[0207] Exemplarily, the first sharpness threshold can be the product of the frame sharpness of the current audio frame and a coefficient coef, where the coefficient coef is less than 1.
[0208] Exemplarily, the coefficient coef is greater than or equal to 0.5 and less than or equal to 0.9. For example, the coefficient coef is 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.9, etc.
[0209] In the following, the process of determining the adaptive suppression interval by the electronic device according to the first bark frequency band is described in detail. Figure 8 and Figure 9
[0210] Figure 9 is a schematic diagram of the candidate suppression interval determined in each comparison process provided by the embodiments of the present application. As an example, Figure 9 only three candidate suppression intervals determined in three comparison processes are shown in the above table, which should not be construed as a limitation on the embodiments of the present application.
[0211] First comparison process
[0212] The electronic device takes the first bark band as a reference, compares the sharpness corresponding to the two bark bands on both sides of the first bark band, determines the maximum sharpness in the sharpness corresponding to the two bark bands, retains the bark band corresponding to the maximum sharpness (denoted as candidate bark band 1), takes the interval between the first bark band and the candidate bark band 1 including the first bark band and the candidate bark band 1 as a candidate suppression interval 1, accumulates and sums the sharpness corresponding to each bark band of the candidate suppression interval 1, and if the frequency band sharpness obtained by the summing result is less than the first sharpness threshold, the expansion of the suppression interval is continued.
[0213] For example, taking the first bark band in Figure 8 as the 165th bark band as an example, in Figure 9 , the two bark bands on both sides of the 165th bark band are the 164th bark band and the 166th bark band. Comparing the sharpness corresponding to the 164th bark band and the 166th bark band respectively, the sharpness corresponding to the 166th bark band is greater than the sharpness corresponding to the 164th bark band. The 166th bark band is retained. The interval between the 166th bark band and the 165th bark band (i.e., [165 bark-166 bark]) including the 166th bark band and the 165th bark band is taken as a candidate suppression interval 1, and the sharpness corresponding to the 166th bark band and the 165th bark band in the candidate suppression interval 1 is accumulated and summed. The frequency band sharpness of the candidate suppression interval 1 obtained by the summing result is less than the first sharpness threshold, and the expansion of the suppression interval is continued.
[0214] Second comparison process
[0215] Taking the candidate suppression interval 1 obtained by the first comparison process as a reference, the electronic device compares the sharpness corresponding to the two bark bands on both sides of the candidate suppression interval 1, determines the maximum sharpness in the sharpness corresponding to the two bark bands, retains the bark band of the maximum sharpness (denoted as candidate bark band 2), takes the interval between the boundary bark band 1 away from the candidate bark band 2 of the candidate suppression interval 1 and the candidate bark band 2 including the boundary bark band 1 and the candidate bark band 2 as a candidate suppression interval 2, accumulates and sums the sharpness corresponding to each bark band of the candidate suppression interval 2, and if the frequency band sharpness of the candidate suppression interval 2 obtained by the summing result is less than the first sharpness threshold, the expansion of the suppression interval is continued.
[0216] Continuing to take Figure 9For example, the candidate suppression interval 1 on the left side of the candidate suppression interval 1 is the 164th bark, and the candidate suppression interval 1 on the right side of the candidate suppression interval 1 is the 167th bark. The sharpness corresponding to the 164th bark is greater than the sharpness corresponding to the 167th bark. The 164th bark is retained, and the interval between the 164th bark and the 166th bark (that is, [164 bark-166 bark]) is recorded as a candidate suppression interval 2. The sharpness corresponding to the 164th bark, the 165th bark, and the 166th bark in the candidate suppression interval 2 is accumulated and summed, and the frequency band sharpness of the candidate suppression interval 2 obtained by the summation is less than the first sharpness threshold. The expansion of the suppression interval is continued.
[0217] The third comparison process
[0218] The candidate suppression interval 2 obtained by the second comparison process is taken as a reference. The electronic device compares the sharpness corresponding to the two bark frequencies on the left and right sides of the candidate suppression interval 2, determines the maximum sharpness in the sharpness corresponding to the two bark frequencies, retains the bark frequency of the maximum sharpness (recorded as a candidate bark frequency 3), records the interval between the boundary bark frequency 2 away from the candidate bark frequency 3 of the candidate suppression interval 2 and the candidate bark frequency 3 as a candidate suppression interval 3, accumulates and sums the sharpness corresponding to each bark frequency of the candidate suppression interval 3, and the frequency band sharpness of the candidate suppression interval 3 obtained by the summation result is less than the first sharpness threshold. The expansion of the suppression interval is continued.
[0219] The candidate suppression interval 2 obtained by the second comparison process is taken as a reference. The electronic device compares the sharpness corresponding to the two bark frequencies on the left and right sides of the candidate suppression interval 2, determines the maximum sharpness in the sharpness corresponding to the two bark frequencies, retains the bark frequency of the maximum sharpness (recorded as a candidate bark frequency 3), records the interval between the boundary bark frequency 2 away from the candidate bark frequency 3 of the candidate suppression interval 2 and the candidate bark frequency 3 as a candidate suppression interval 3, accumulates and sums the sharpness corresponding to each bark frequency of the candidate suppression interval 3, and the frequency band sharpness of the candidate suppression interval 3 obtained by the summation result is less than the first sharpness threshold. The expansion of the suppression interval is continued. Figure 9 For example, the candidate suppression interval 2 obtained by the second comparison process is taken as a reference. The electronic device compares the sharpness corresponding to the two bark frequencies on the left and right sides of the candidate suppression interval 2, determines the maximum sharpness in the sharpness corresponding to the two bark frequencies, retains the bark frequency of the maximum sharpness (recorded as a candidate bark frequency 3), records the interval between the boundary bark frequency 2 away from the candidate bark frequency 3 of the candidate suppression interval 2 and the candidate bark frequency 3 as a candidate suppression interval 3, accumulates and sums the sharpness corresponding to each bark frequency of the candidate suppression interval 3, and the frequency band sharpness of the candidate suppression interval 3 obtained by the summation result is less than the first sharpness threshold. The expansion of the suppression interval is continued.
[0220] In the same way, the comparison process is continued in the same way, and the comparison process is continued for multiple times until the sum of the sharpness corresponding to each bark band in the extended candidate suppression band (i.e., the band sharpness) is greater than the first sharpness threshold, and then the expansion of the band is stopped, and the candidate suppression band obtained finally is determined as the adaptive suppression band.
[0221] It should be understood that Figure 8 It should be understood that the schematic diagram is a schematic diagram of the case that the band range of the adaptive suppression band includes the band range of the core suppression band, and should not be regarded as a limitation on the embodiments of the present application. Based on the above description of the cases 1 to 3 of the suppression band, in the implementation, the band range of the adaptive suppression band and the band range of the core suppression band can also have a partial intersection, or the band range of the core suppression band can also include the band range of the adaptive suppression band.
[0222] In step S440, the electronic device converts the core suppression band and the adaptive suppression band from the bark domain to the linear frequency domain to obtain a core suppression frequency band and an adaptive suppression frequency band.
[0223] The linear frequency domain here is a frequency domain with Hz as the unit.
[0224] The electronic device converts the core suppression band and the adaptive suppression band from the bark domain to the linear frequency domain to obtain a core suppression frequency band and an adaptive suppression frequency band.
[0225] The relationship between the core suppression frequency band and the adaptive suppression frequency band can be as shown in the above Figures 4 to 6 three cases.
[0226] In step S450, the electronic device determines a first suppression gain of the core suppression frequency band.
[0227] The first suppression gain of the core suppression frequency band is a fixed value, which can be a predefined value or a dynamically changing value.
[0228] Since there is a difference in the tooth energy of each audio frame, the first suppression gain is designed as a dynamically changing value, which can make the first suppression gain more suitable and adaptive to each audio frame, improve the effect of tooth suppression, and improve the flexibility of tooth processing.
[0229] In some embodiments, the first suppression gain can be obtained by the following first formula:
[0230] gain = ucgain x coef-prob
[0231] Wherein, gain is a first compression gain, ucgain is an initial gain, coef-prob is a gain coefficient, sharpness is a frame sharpness of the current audio frame, and threold is a sharpness threshold.
[0232] In step S460, the electronic device determines whether there is an extended compression frequency band between the core compression frequency band and the adaptive compression frequency band.
[0233] The extended compression frequency band is a frequency band other than the core compression frequency band in the adaptive compression frequency band.
[0234] If the electronic device determines that there is an extended compression frequency band between the core compression frequency band and the adaptive compression frequency band, the electronic device performs step S470; if the electronic device determines that there is no extended compression frequency band between the core compression frequency band and the adaptive compression frequency band, the electronic device performs step S492.
[0235] As described above, in the above case 1 (the frequency band range of the adaptive compression frequency band includes the frequency band range of the core compression frequency band) and case 2 (the frequency band range of the adaptive compression frequency band partially intersects with the frequency band range of the core compression frequency band), there is an extended compression frequency band between the core compression frequency band and the adaptive compression frequency band; in the above case 3 (the frequency band range of the core compression frequency band includes the frequency band range of the adaptive compression frequency band), there is no extended compression frequency band between the core compression frequency band and the adaptive compression frequency band.
[0236] Since the core compression frequency band is small in the frequency band range itself, and the purpose of the adaptive compression frequency band is to expand the range of the compression frequency band. Therefore, based on the design idea, the relationship between the core compression frequency band and the adaptive compression frequency band is mostly case 1 and case 2, and case 3 occurs in a very small number of cases.
[0237] In step S470, the electronic device determines at least one extended compression frequency band.
[0238] In this step, the electronic device determines the frequency band other than the core compression frequency band in the adaptive compression frequency band as the extended compression frequency band. Wherein, the extended compression frequency band can be one or two.
[0239] In case 1, referring to Figure 4 , the frequency band range of the adaptive compression frequency band includes the frequency band range of the core compression frequency band, two extended compression frequency bands are determined, and the two extended compression frequency bands are located on both sides of the core compression frequency band.
[0240] In case 2, referring to Figure 5 , the frequency band range of the adaptive compression frequency band partially intersects with the frequency band range of the core compression frequency band, one extended compression frequency band is determined, and the extended compression frequency band is located on one side of the core compression frequency band.
[0241] In step S480, the electronic device determines the second suppression gain of each extended suppression frequency band.
[0242] In some embodiments, the electronic device determines the second suppression gain of each extended suppression frequency band in a frequency variation direction from a first boundary frequency close to the core suppression frequency band to a second boundary frequency away from the core suppression frequency band, based on the first suppression gain, in a manner that the second suppression gain of each extended suppression frequency band is increased by one gain step for each increment of one frequency band width. The second suppression gain of each extended suppression frequency band is increased from the first suppression gain to 1.
[0243] It can be understood that the first boundary frequency is the frequency close to the core suppression frequency band in the extended suppression frequency band, and the second boundary frequency is the frequency away from the core suppression frequency band in the extended suppression frequency band. For example, the first boundary frequency of the one extended suppression frequency band is the boundary frequency on the left side of the one extended suppression frequency band, and the second boundary frequency of the one extended suppression frequency band is the boundary frequency on the right side of the one extended suppression frequency band. Figure 4 For example, the first boundary frequency of the one extended suppression frequency band is the boundary frequency on the left side of the one extended suppression frequency band, and the second boundary frequency of the one extended suppression frequency band is the boundary frequency on the right side of the one extended suppression frequency band.
[0244] The frequency band width can be of any length, which represents the granularity of the frequency variation when the gain of the extended suppression frequency band varies.
[0245] For example, the frequency band width can be 15.625 Hz. Such a small frequency band width makes the gain variation relatively smooth, which can better reduce the impact on the sound quality while suppressing the sibilance.
[0246] The gain step represents the granularity of the gain variation, and the gain of the extended suppression frequency band is increased in a manner that the gain is increased by one gain step for each increment of one frequency band width.
[0247] For example, the gain step can be obtained by the following second formula: Wherein, step is the gain step, P is the number of frequency points in the extended suppression frequency band, and gain is the first suppression gain. It should be understood that 1 actually represents the gain of the non-suppression frequency band.
[0248] For example, the first suppression gain gain = 0.6, one extended suppression frequency band is [5 kHz, 5.2 kHz], the frequency band width is 10 Hz, therefore, the number of frequency points P = 20, and the gain step For the second suppression gain of the extended suppression frequency band, the gain of 5 kHz is 0.6, the gain of 5.01 kHz is 0.62, the gain of 5.02 kHz is 0.64, and so on, and finally the second suppression gain of the extended suppression frequency band is obtained.
[0249] In step S491, the electronic device applies the first suppression gain to the energy of the current audio frame in the core suppression frequency band, and applies the second suppression gain of each extended suppression frequency band to the energy of the current audio frame in each extended suppression frequency band, to perform sibilance suppression.
[0250] In step S492, the electronic device applies the first suppression gain to the energy of the current audio frame in the core suppression frequency band, to perform sibilance suppression.
[0251] For specific descriptions of steps S491 and S492, refer to the related descriptions of step S360, which will not be repeated.
[0252] In the method 300 and the method 400, steps S420 to S440 can be understood as specific descriptions of step S340, steps S450 to S480 can be understood as specific descriptions of step S350, and steps S491 and S492 can be understood as specific descriptions of step S360.
[0253] It should be understood that the size of the serial number of each process of the above-mentioned embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0254] For example, steps S420 and S430 can be exchanged in order. For another example, step S450 can be executed at any time before step S491 or step S1492.
[0255] Figure 10 FIG. 1 is a comparison chart of the energy of an audio frame before sibilance processing and the energy of the audio frame after sibilance processing provided by an embodiment of the present application. Wherein, Figure 10 (a) in FIG. 1 shows the energy of the audio frame before sibilance processing. (b) in FIG. 1 shows the energy of the audio frame after sibilance processing.
[0256] Figure 10 The area S1 in (a) is the energy of the sibilance segment before sibilance processing, which is generally high, Figure 10 The area S2 in (b) is the energy of the sibilance segment after sibilance processing, which is obviously different from the energy of the area S1, the color is lighter, which means that the energy of the sibilance segment after sibilance processing is reduced, and the effect of sibilance suppression is achieved.
[0257] As described above, for an audio frame, the electronic device determines whether the audio frame is a sibilance frame according to a target characteristic value of the audio frame, which can be an audio parameter such as a frame sharpness or a loudness of the audio frame. In the following, the process in which the electronic device determines the frame sharpness of the audio frame is described by taking the frame sharpness of the audio frame as an example. Figure 11
[0258] Figure 11 FIG. 5 is a schematic flowchart of a method 500 for determining the frame sharpness of an audio frame according to an embodiment of the present application.
[0259] It should be noted that the method 500 describes the process of determining the frame sharpness of the audio frame based on the loudness model specified in the DIN45631 standard, and the embodiments of the present application only briefly introduce the process, and the specific description can refer to the related description of the loudness model.
[0260] It should be further noted that before the method 500 is executed, the electronic device converts the audio frame from a time domain signal to a frequency domain signal, and processes the frequency domain signal of the audio frame.
[0261] In step S510, the electronic device calculates the frequency sound pressure level of the current audio frame.
[0262] In step S520, the electronic device obtains the frequency band sound pressure level of the current audio frame by 1 / 3 octave filtering.
[0263] In step S530, the electronic device determines the frequency band characteristic loudness of the current audio frame.
[0264] In this step, in order to simulate human ear filtering, it is necessary to correct and combine the low frequency band, and the frequency band characteristic loudness of the current audio frame in the bark domain can be obtained according to the following formula:
[0265]
[0266] wherein E TQ is the quiet hearing threshold sound pressure level, E0 is the frequency band sound pressure level corresponding to the reference sound pressure level, N' is the frequency band loudness, the unit of the frequency band characteristic loudness is sone / bark, and it is a loudness value determined in the bark domain.
[0267] For the frequency band characteristic loudness in the bark domain, the frequency band is moved in a unit step (for example, the step is 1 bark frequency band) in the bark domain, and by comparing the frequency band characteristic loudness of adjacent bark frequency bands, it is determined whether there is a masking effect, and if there is, a slope absorption loudness is introduced to simulate the masking effect to cover some components.
[0268] In step S540, the electronic device calculates the total loudness of the current audio frame according to the band characteristic loudness of the current audio frame.
[0269] It should be understood that the total loudness of the audio frame represents the sum of the band characteristic loudness of each bark band in the bark domain.
[0270] Taking the bark domain including 240 bands as an example, the electronic device integrates the band characteristic loudness of the bark domain to obtain the total loudness:
[0271] In step S550, the electronic device determines the frame sharpness of the current audio frame according to the band characteristic loudness and the total loudness of the current audio frame.
[0272] Exemplarily, the electronic device can determine the frame sharpness of the current audio frame according to the following formula:
[0273]
[0274] wherein S is the frame sharpness of the audio frame, N is the total loudness of the audio frame, z is the critical band, N' is the band characteristic loudness of the audio frame in the bark domain, and g(z) is the weighting coefficient.
[0275] After determining the frame sharpness of the current audio frame, in a case where the frame sharpness of the current audio frame satisfies a preset condition (denoted as a second preset condition), the current audio frame is determined as a sibilance frame; in a case where the frame sharpness of the current audio frame does not satisfy the second preset condition, the current audio frame is determined as not a sibilance frame.
[0276] Exemplarily, the second preset condition is that the frame sharpness of the current audio frame is greater than a second sharpness threshold.
[0277] That is, in a case where the frame sharpness of the current audio frame is greater than the second sharpness threshold, the electronic device determines the current audio frame as a sibilance frame. In a case where the frame sharpness of the current audio frame is less than the second sharpness threshold, the electronic device determines the current audio frame as not a sibilance frame.
[0278] It should be understood that, for a case where the frame sharpness of the current audio frame is equal to the second sharpness threshold, it can be divided into a case satisfying that the current audio frame is determined as a sibilance frame, or a case not satisfying that the current audio frame is a sibilance frame, which is not limited herein.
[0279] It should be understood that the process of determining the frame sharpness of the audio frame described in the method 500 is only illustrative, and should not constitute a limitation on the embodiments of the present application. Any method capable of determining the frame sharpness of the audio frame is within the protection scope of the embodiments of the present application.
[0280] Figure 12 is a schematic flowchart of a method 600 of sibilance processing provided by an embodiment of the present application.
[0281] In step S610, the electronic device determines that the current audio frame is a sibilance frame.
[0282] As described previously, the electronic device can determine whether the current audio frame is a sibilance frame according to the target feature value of the current audio frame: in a case where the target feature value satisfies a second preset condition, the electronic device determines that the current audio frame is a sibilance frame, and in a case where the target feature value does not satisfy the second preset condition, the electronic device determines that the current audio frame is not a sibilance frame.
[0283] In some embodiments, the target feature value can be a frame sharpness of the current audio frame; and the second preset condition can be that the frame sharpness of the current audio frame is greater than a second sharpness threshold.
[0284] For a specific description of determining that the current audio frame is a sibilance frame in a case where the frame sharpness of the current audio frame satisfies the second preset condition, reference can be made to the related description hereinabove, and no further elaboration is made.
[0285] In step S620, the electronic device determines a core suppression frequency band and an adaptive suppression frequency band of the current audio frame.
[0286] In some embodiments, the electronic device can measure the sharpness distribution of the audio frame in a bark scale, and determine the core suppression frequency band and the adaptive suppression frequency band of the current audio frame.
[0287] In other embodiments, the electronic device can measure the perception of the human ear to sound in other scales such as a mel-spectral scale, and is not limited to determining the core suppression frequency band and the adaptive suppression frequency band according to the sharpness of the current audio frame, but can also determine the core suppression frequency band and the adaptive suppression frequency band according to other target feature values for representing audio characteristics.
[0288] In the embodiment of measuring the sharpness distribution of the audio frame in the bark scale, optionally, before determining the core suppression frequency band and the adaptive suppression frequency band of the current audio frame, the method comprises: the electronic device determining a first bark frequency band in a bark domain corresponding to a maximum sharpness according to the sharpness distribution of the current audio frame; and in the step of determining the core suppression frequency band and the adaptive suppression frequency band of the current audio frame, the electronic device determines the core suppression frequency band and the adaptive suppression frequency band according to the first bark frequency band, and the core suppression frequency band and the adaptive suppression frequency band both include the first bark frequency band.
[0289] Since the first bark band is the band corresponding to the maximum sharpness, the first bark band has relatively high sibilance energy. The core suppression band and the adaptive suppression band determined based on the first bark band can better determine the band with relatively high sibilance energy, thereby improving the sibilance suppression effect.
[0290] The specific description of determining the first bark band by the electronic device can refer to the related description of step S410, and will not be described herein again.
[0291] The following describes the process of determining the core suppression band and the adaptive suppression band by the electronic device according to the first bark band.
[0292] 1. The electronic device determines the core suppression band according to the first bark band
[0293] In the step of determining the core suppression band by the electronic device according to the first bark band, the electronic device determines the core suppression interval according to the first bark band, converts the core suppression interval from the bark domain to the linear frequency domain, and obtains the core suppression band corresponding to the core suppression interval.
[0294] In some embodiments, the electronic device determines the core suppression interval according to the first bark band and the bandwidth parameter, to obtain the core suppression band corresponding to the core suppression interval, and the core suppression band includes the first bark band and the band extended based on the bandwidth parameter.
[0295] It can be understood that the bandwidth parameter is used to expand the band.
[0296] For example, the bandwidth parameter is used to indicate the number M1 of bark bands. In implementation, the electronic device determines the core suppression interval according to M1 and the first bark band, to obtain the core suppression band.
[0297] In an example, the electronic device takes the first bark band as the center band, and expands M1 bark bands to both sides of the first bark band, to obtain the core suppression band.
[0298] In this embodiment, the electronic device takes the first bark band as the center band, and expands M1 bark bands to both sides of the first bark band, to obtain the core suppression interval, converts the core suppression interval from the bark domain to the linear frequency domain, and obtains the core suppression band corresponding to the core suppression interval.
[0299] Generally, the region with high sibilance energy is concentrated on both sides of the first bark band. Therefore, the core suppression band obtained by expanding M1 bark bands to both sides of the first bark band can well and comprehensively obtain the band with relatively high sibilance energy, thereby further improving the sibilance suppression effect.
[0300] The specific description of determining the core suppression interval by the electronic device with reference to the electronic device can refer to the related description of step S420 above, and will not be described again.
[0301] In other examples, the electronic device can also expand M1 bark frequency bands to either side of the first bark frequency band to obtain the core suppression interval, and convert the core suppression interval from the bark domain to the linear frequency domain to obtain the core suppression frequency band corresponding to the core suppression interval. For example, M1 = 35, and the first bark frequency band is the 165th bark. Then, the core suppression interval is [165 bark, 200 bark] or [130 bark, 165 bark].
[0302] It should be understood that the above manner of determining the core suppression frequency band according to the first bark frequency band and the bandwidth parameter is only illustrative. In other embodiments, the electronic device can also determine the core suppression frequency band according to the first bark frequency band and some rules.
[0303] For example, the core suppression interval can be determined in the manner of determining the adaptive suppression interval in step S430 above to obtain the core suppression frequency band. However, since the adaptive suppression frequency band itself is used to expand the core suppression frequency band, the first preset condition for limiting the expansion of the frequency band itself is relatively loose. When the core suppression frequency band is determined in a similar manner, the preset condition for limiting the expansion of the frequency band can be set to be relatively strict (for example, the first sharpness threshold is set to a larger value).
[0304] 2. The electronic device determines the adaptive suppression frequency band according to the first bark frequency band
[0305] In the step of determining the adaptive suppression frequency band according to the first bark frequency band by the electronic device, in some embodiments, the electronic device expands the frequency band to both sides of the first bark frequency band according to the sharpness distribution of the current audio frame to determine the adaptive suppression frequency band, wherein the frequency band sharpness of the adaptive suppression frequency band satisfies the first preset condition, and the first preset condition is used to limit the expansion of the adaptive suppression frequency band.
[0306] That is, in this embodiment, the electronic device expands the frequency band to both sides of the first bark frequency band according to the sharpness distribution of the current audio frame, and when the frequency band sharpness of the expanded adaptive suppression frequency band satisfies the first preset condition, the electronic device no longer expands the frequency band.
[0307] Since the adaptive suppression frequency band is extended based on the sharpness distribution of the current audio frame, and is related to the characteristics of the audio frame itself, the adaptive suppression frequency band obtained for each audio frame is more in line with the characteristics of each audio frame, so as to accurately determine the suppression frequency band of each audio frame, thereby better improving the effect of sibilance suppression, and improving the flexibility of sibilance processing.
[0308] In the above embodiment of extending the frequency band to both sides of the first bark frequency band to determine the adaptive suppression frequency band according to the sharpness distribution of the current audio frame, the adaptive suppression frequency band can be determined in the following two ways.
[0309] Way 1
[0310] The electronic device performs K comparison processes on the sharpness of each bark frequency band in the plurality of bark frequency bands on both sides of the first bark frequency band according to the sharpness distribution of the current audio frame, and determines the frequency band corresponding to the candidate suppression interval obtained in the Kth comparison process as the adaptive suppression frequency band in the case that the frequency band sharpness of the candidate suppression interval obtained in the Kth comparison process meets the first preset condition, K being an integer greater than 1.
[0311] That is, in way 1, the electronic device performs K comparison processes, and each comparison process obtains a candidate suppression interval. For the frequency band sharpness of the candidate suppression interval obtained in each comparison process, if the first preset condition is not met, the next comparison process is continued to expand the suppression interval, until the candidate suppression interval obtained in the Kth comparison process meets the first preset condition, and the suppression interval is no longer expanded, and the frequency band corresponding to the candidate suppression interval obtained in the Kth comparison process is determined as the final adaptive suppression frequency band.
[0312] In the 1st comparison process, the electronic device compares the sharpness corresponding to the two bark frequency bands on both sides of the first bark frequency band, and determines the bark frequency band corresponding to the maximum sharpness in the 1st comparison process as the initial candidate bark frequency band; the electronic device determines the interval between the first bark frequency band and the initial candidate bark frequency band as the candidate suppression interval in the 1st comparison process, and the frequency band sharpness of the candidate suppression interval in the 1st comparison process does not meet the first preset condition.
[0313] It should be understood that the two bark frequency bands on both sides of the first bark frequency band refer to two bark frequency bands located on both sides of the first bark frequency band with the same interval as the boundary frequency band of the first bark frequency band. Exemplarily, the two bark frequency bands on both sides of the first bark frequency band refer to the two bark frequency bands located on both sides of the first bark frequency band closest to the two boundary frequency bands of the candidate suppression interval.
[0314] The specific description of the first comparison process can refer to the related description of the first comparison process of step S430, and the initial candidate bark band and the candidate suppression interval here can correspond to the candidate bark band 1 and the candidate suppression interval 1 of step S430, respectively.
[0315] In the i-th comparison process, the electronic device compares the sharpness corresponding to the two bark bands on both sides of the candidate suppression interval obtained in the (i-1)-th comparison process, determines the bark band corresponding to the maximum sharpness in the i-th comparison process as the intermediate candidate bark band, i is greater than 1 and less than K; the electronic device determines the interval between the intermediate candidate bark band and the boundary bark band obtained in the (i-1)-th comparison process as the candidate suppression interval in the i-th comparison process, the boundary bark band is the frequency band far away from the intermediate candidate bark band in the candidate suppression interval obtained in the (i-1)-th comparison process, and the frequency band sharpness of the candidate suppression interval in the i-th comparison process does not satisfy the first preset condition.
[0316] It should be understood that the two bark bands on both sides of the candidate suppression interval represent two bark bands located on both sides of the candidate suppression interval with the same interval as the boundary frequency band of the candidate suppression interval. Exemplarily, the two bark bands on both sides of the candidate suppression interval represent the two bark bands closest to the two boundary frequency bands of the candidate suppression interval.
[0317] For the boundary bark band, in a general understanding, it is the frequency band far away from the intermediate candidate bark band obtained in the current comparison process in the candidate suppression interval obtained in the last comparison process.
[0318] The i-th comparison process here can be any one of the K comparison processes except the first comparison process and the K-th comparison process. For each comparison process except the first comparison process and the K-th comparison process, the candidate suppression interval of the i-th comparison process is obtained through the i-th comparison process. In addition, when i is equal to 2, the (i-1)-th comparison process is the first comparison process.
[0319] The specific description of the ith comparison process can refer to the description of the second comparison process or the third comparison process of step S430. When the ith comparison process herein corresponds to the third comparison process above, the intermediate candidate bark band herein can correspond to the candidate bark band 2 above, and the boundary bark band herein can correspond to the boundary bark band 1 above. When the ith comparison process herein corresponds to the third comparison process above, the intermediate candidate bark band herein can correspond to the candidate bark band 3 above, and the boundary bark band herein can correspond to the boundary bark band 2 above.
[0320] For the K-i th comparison process after the ith comparison process, the electronic device continues the K-i th comparison process until the adaptive suppression band is determined in the K th comparison process.
[0321] In an implementation, in the K th comparison process, the candidate suppression interval of the K th comparison process is obtained in the same manner as the ith comparison process, the frequency band sharpness of the candidate suppression interval obtained in the K th comparison process is judged, and it is determined that the frequency band sharpness of the candidate suppression interval obtained in the K th comparison process satisfies the first preset condition. Then, the candidate suppression interval obtained in the K th comparison process is determined as the adaptive suppression interval, the adaptive suppression interval is converted from the bark domain to the linear frequency domain to obtain the adaptive suppression band corresponding to the adaptive interval.
[0322] In the above embodiments, the electronic device performs K comparison processes on the sharpness corresponding to each of the plurality of bark bands on both sides of the first bark band according to the sharpness distribution of the current audio frame, determines the maximum sharpness in the sharpness corresponding to the two bark bands on both sides of the first bark band in each comparison process, determines the candidate suppression interval of each comparison process based on the maximum sharpness obtained in each comparison process, and realizes the purpose of expanding the suppression interval by a part of the interval in each comparison process. Ultimately, the expansion of the suppression interval is completed through multiple comparison processes to obtain the adaptive suppression band. Based on this way of cyclically traversing the sharpness of the current audio frame, the adaptive suppression band can be obtained as much as possible based on the characteristics of the current audio frame to comprehensively search for the bark band with a larger numerical value. The sharpness can be more accurately obtained to conform to the adaptive suppression band of the current audio frame.
[0323] Regarding the first preset condition, in some embodiments, the first preset condition is that the frequency band sharpness of the candidate suppression interval obtained in each comparison process is greater than a first sharpness threshold.
[0324] For example, the first sharpness threshold is the product of the frame sharpness of the current audio frame and a coefficient, and the coefficient is less than 1.
[0325] For a detailed description of the first preset condition, please refer to the relevant description in step S430 above, which will not be repeated here.
[0326] Since the sharpness of the current audio frame is related to the current audio frame, the first sharpness threshold is designed as the product of the sharpness of the current audio frame and the coefficient. This makes the first preset condition related to each audio frame, and the resulting adaptive suppression frequency band is more consistent with each audio frame. This helps to more accurately determine the suppression frequency band of each audio frame, thereby improving the sibilance suppression effect.
[0327] Furthermore, for a detailed description of Method 1, please refer to the relevant description of step S430 above, which will not be repeated here.
[0328] Method 2
[0329] In other embodiments, using bark intervals as the granularity, the electronic device can also gradually expand the suppression interval by adjusting the sharpness of each bark interval in multiple bark intervals on both sides of the first bark frequency band, and finally obtain an adaptive suppression interval to obtain an adaptive suppression frequency band.
[0330] For example, the electronic device performs multiple comparison processes. In each comparison process, the maximum sharpness of the sharpness of the two bark intervals on both sides of the first bark frequency band is determined. Based on the maximum sharpness, the candidate suppression interval for each comparison process is determined. Each comparison process extends the suppression interval by a part of the frequency band. Finally, by extending the suppression interval through multiple comparison processes, an adaptive suppression interval is obtained to obtain the adaptive suppression frequency band.
[0331] It should be understood that the two bark intervals on both sides of the first bark band are not exactly the same in any two comparison processes. In the current comparison process, the two bark intervals on both sides of the first bark band are the two bark intervals on both sides of the candidate suppression interval obtained in the previous comparison process.
[0332] It should also be understood that a bark interval includes at least one bark band, and, for example, the band widths of the two bark intervals are the same in each comparison process. Furthermore, the band widths of the two bark intervals in any two comparison processes can also be the same.
[0333] In the first comparison process, the electronic device compares the sharpness corresponding to the two bark intervals on both sides of the first bark frequency band, determines the maximum sharpness in the sharpness corresponding to the two bark intervals, retains the bark interval corresponding to the maximum sharpness (denoted as candidate bark interval 1), records the interval between the first bark frequency band and the candidate bark interval 1 as the candidate suppression interval 1, accumulates and sums the sharpness corresponding to each bark frequency of the candidate suppression interval 1, and obtains the frequency band sharpness of the candidate suppression interval 1 from the sum result. If the frequency band sharpness of the candidate suppression interval 1 does not satisfy the first preset condition, the expansion of the suppression interval is continued.
[0334] Taking the first preset condition that the frequency band sharpness of the candidate suppression interval obtained in each comparison process is greater than the first sharpness threshold as an example. For example, taking the first bark frequency band as the 165th bark as an example, the two bark intervals on both sides of the 165th bark are [166 bark-167 bark] and [163 bark-164 bark], the sharpness of the two bark intervals is compared, the sharpness corresponding to [166 bark-167 bark] is greater than the sharpness corresponding to [163 bark-164 bark], therefore, the interval between the 165th bark and [166 bark-167 bark] including the 165th bark and [166 bark-167 bark] is recorded as the candidate suppression interval 1, and the sharpness corresponding to each bark frequency in the candidate suppression interval 1 is accumulated and summed, the frequency band sharpness of the candidate suppression interval 1 obtained is less than the first sharpness threshold, and the expansion of the suppression interval is continued.
[0335] In the second comparison process, taking the candidate suppression interval 1 obtained in the first comparison process as a reference, the electronic device compares the sharpness corresponding to the two bark intervals on both sides of the candidate suppression interval 1, determines the maximum sharpness in the sharpness corresponding to the two bark intervals, retains the bark interval corresponding to the maximum sharpness (denoted as candidate bark interval 2), records the interval between the candidate suppression interval 1 and the candidate bark interval 2 including the candidate suppression interval 1 and the candidate bark interval 2 as the candidate suppression interval 2, accumulates and sums the sharpness corresponding to each bark frequency of the candidate suppression interval 2, and the frequency band sharpness of the candidate suppression interval 2 obtained from the sum result does not satisfy the first preset condition, and the expansion of the suppression interval is continued.
[0336] Subsequently, in the same way, the comparison process is continued for multiple times until the sum of the sharpness of each bark band corresponding to the extended candidate compression interval K (i.e., the band sharpness) satisfies a first preset condition, the expansion of the interval is stopped, and the candidate compression interval K obtained in the last comparison process is determined as the adaptive compression interval to obtain the adaptive compression band.
[0337] It should be understood that the above manner of expanding the band to both sides of the first bark band according to the sharpness distribution of the current audio frame to determine the adaptive compression band is only illustrative and should not be construed as limiting the embodiments of the present application.
[0338] In other embodiments, the electronic device can determine the adaptive compression band according to the first bark band and a band parameter.
[0339] For example, the band parameter is used to indicate the number M2 of bark bands, and the electronic device expands M2 bark bands to both sides of the first bark band to obtain the adaptive compression interval, taking the first bark band as the center frequency band, to obtain the adaptive compression band.
[0340] For another example, the electronic device expands M2 bark bands to either side of the first bark band to obtain the adaptive compression interval, taking the first bark band as the center frequency band, to obtain the adaptive compression band.
[0341] It should be understood that in this embodiment, the number M2 of bark bands indicated by the band parameter used to determine the adaptive compression band is greater than the number M1 of bark bands indicated by the bandwidth parameter used to determine the core compression band, and therefore, the frequency band range of the adaptive compression band obtained based on this embodiment includes the frequency band range of the core compression band.
[0342] In step S630, the electronic device determines the first compression gain of the core compression band.
[0343] In some embodiments, the electronic device determines the first compression gain according to a gain coefficient and the initial compression gain, and the gain coefficient is related to the current audio frame.
[0344] That is, in this embodiment, the gain coefficient dynamically changes according to the current audio frame, and therefore, the gain coefficient can be different for different audio frames. In this way, the first compression gain can be more consistent and adaptive to each audio frame, which can improve the effect of sibilance suppression and improve the flexibility of sibilance processing.
[0345] The gain coefficient can be related to any audio parameter used to measure the audio. In an example, the gain coefficient is related to the frame sharpness of the current audio frame and the second sharpness threshold. Exemplarily, the relationship of the gain coefficient, the current audio frame and the second sharpness threshold can be represented by the first formula in step S450 above, and the specific description of the first formula can refer to the related description above, and will not be repeated here.
[0346] In this way, by associating the gain coefficient with the frame sharpness of the audio frame and the second sharpness threshold, the dynamic change of the gain coefficient with different audio frames can be conveniently realized.
[0347] In step S640, in the case that the frequency range of the adaptive suppression frequency band includes the frequency range of the core suppression frequency band, or the frequency range of the adaptive suppression frequency band has a partial intersection with the frequency range of the core suppression frequency band, at least one extended suppression frequency band is determined.
[0348] The specific description of this step can refer to the related description of step S470 above, and will not be repeated here.
[0349] In step S650, the electronic device determines the second suppression gain of each extended suppression frequency band.
[0350] In some embodiments, the electronic device determines the second suppression gain of each extended suppression frequency band in the frequency change direction from the first boundary frequency to the second boundary frequency of each extended suppression frequency band, based on the first suppression gain, and in the manner of increasing one gain step for each incremental frequency band width. Exemplarily, the gain step can be obtained according to the second formula in step S480 above, and the specific description of the second formula can refer to the related description above, and will not be repeated here.
[0351] The specific description of this step can refer to the related description of step S480 above, and will not be repeated here.
[0352] In step S670, the electronic device applies the first suppression gain to the energy of the current audio frame in the core suppression frequency band, and applies the second suppression gain of each extended suppression frequency band to the energy of the current audio frame in each extended suppression frequency band, to perform the sibilance suppression.
[0353] The specific description of step S670 can refer to the related description of step S360 above, and will not be repeated here.
[0354] The sibilance processing method provided in this application defines an audio frame with a defined suppression frequency band including a core suppression frequency band and an adaptive suppression frequency band for extending the core suppression frequency band. When the frequency range of the adaptive suppression frequency band includes the frequency range of the core suppression frequency band, or when there is partial overlap between the frequency ranges of the adaptive suppression frequency band and the core suppression frequency band, the sibilance suppression frequency band is expanded using the adaptive suppression frequency band. This results in an extended suppression frequency band in addition to the core suppression frequency band, increasing the frequency range of the suppression frequency band. Performing sibilance processing on a larger suppression frequency band yields better sibilance suppression. Furthermore, since the adaptive suppression frequency band extends the core suppression frequency band, and the gain of the core suppression frequency band is the first suppression gain, the second suppression gain of the extended suppression frequency band increases from the first suppression gain to 1. This avoids the sound quality degradation caused by using a smaller gain to suppress the energy of the extended suppression frequency band. Therefore, the gradually increasing second suppression gain can mitigate sound quality degradation, improving sibilance suppression while minimizing sound quality loss. In summary, the sibilance processing method provided in this application embodiment achieves a good balance between sound quality and sibilance suppression, and can effectively improve the sibilance suppression effect while minimizing sound quality damage.
[0355] In some embodiments, method 600 further includes: applying a first suppression gain to the energy of the current audio frame in the core suppression band when the band range of the core suppression band includes the band range of the adaptive suppression band.
[0356] Although the frequency range of the core suppression band includes that of the adaptive suppression band, the frequency range of the core suppression band itself is relatively small. Using the core suppression band as the final suppression band and performing sibilance processing on it can achieve sibilance suppression with minimal or acceptable impact on sound quality. Moreover, since the first suppression gain is itself the gain of the core suppression band, directly using the first suppression gain to perform sibilance processing on the core suppression band avoids increasing the algorithmic complexity due to additional steps, making the algorithm easier to implement.
[0357] It should be understood that the sequence numbers of the processes in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. For example, step S630 may be executed after step S640 or step S650.
[0358] The above, combined with Figures 1 to 12 The method for processing dental sounds provided in the embodiments of this application is described in detail below. Figures 13 to 14 This application provides a detailed description of the apparatus and electronic device for sibilance processing according to embodiments thereof.
[0359] Figure 13 FIG. 7 is an exemplary block diagram of an apparatus 700 for sibilance processing provided by an embodiment of the present application. The apparatus 700 is an electronic device, and can also be a chip in an electronic device.
[0360] The apparatus 700 is configured to perform each process and step corresponding to the terminal device in the method 600 described above. The apparatus 700 includes a processing unit 710 configured to perform the following steps:
[0361] determine that a current audio frame is a sibilance frame;
[0362] determine a core suppression frequency band of the current audio frame;
[0363] determine an adaptive suppression frequency band of the current audio frame;
[0364] determine a first suppression gain of the core suppression frequency band, the first suppression gain being less than 1;
[0365] in a case where a frequency band range of the adaptive suppression frequency band includes a frequency band range of the core suppression frequency band, or in a case where there is a partial intersection between the frequency band range of the adaptive suppression frequency band and the frequency band range of the core suppression frequency band, determine at least one extended suppression frequency band, each of the extended suppression frequency bands being a frequency band other than the core suppression frequency band in the adaptive suppression frequency band;
[0366] determine a second suppression gain of each of the extended suppression frequency bands, the second suppression gain of each of the extended suppression frequency bands increasing from the first suppression gain to 1 in a frequency variation direction from a first boundary frequency closest to the core suppression frequency band to a second boundary frequency farthest from the core suppression frequency band in each of the extended suppression frequency bands;
[0367] apply the first suppression gain to energy of the current audio frame in the core suppression frequency band, and apply the second suppression gain of each of the extended suppression frequency bands to energy of the current audio frame in each of the extended suppression frequency bands.
[0368] Optionally, the processing unit 710 is specifically configured to: determine a first bark frequency band in a bark domain corresponding to a maximum sharpness according to a sharpness distribution of the current audio frame; and determine the core suppression frequency band according to the first bark frequency band, the core suppression frequency band including the first bark frequency band.
[0369] Optionally, the processing unit 710 is specifically configured to: take the first bark frequency band as a center frequency band, and extend M1 bark frequency bands to both sides of the first bark frequency band respectively to obtain the core suppression frequency band.
[0370] Optionally, the processing unit 710 is specifically configured to: determine a first bark band in a bark domain corresponding to a maximum sharpness according to the sharpness distribution of the current audio frame; and determine the adaptive suppression band according to the first bark band, the adaptive suppression band including the first bark band.
[0371] Optionally, the processing unit 710 is specifically configured to: expand the first bark band to both sides according to the sharpness distribution of the current audio frame to determine the adaptive suppression band, wherein a band sharpness of the adaptive suppression band satisfies a first preset condition, and the first preset condition is used to limit expansion of the adaptive suppression band.
[0372] Optionally, the processing unit 710 is specifically configured to: perform K times of comparison processes on sharpnesses corresponding to each bark band of a plurality of bark bands on both sides of the first bark band according to the sharpness distribution of the current audio frame, and determine a band corresponding to a candidate suppression interval obtained in a Kth comparison process as the adaptive suppression band in a case where a band sharpness of the candidate suppression interval obtained in the Kth comparison process satisfies the first preset condition, K being an integer greater than 1; wherein,
[0373] in a first comparison process, comparing sharpnesses corresponding to two bark bands on both sides of the first bark band, and determining a bark band corresponding to a maximum sharpness in the first comparison process as an initial candidate bark band;
[0374] determining an interval between the first bark band and the initial candidate bark band including the first bark band and the initial candidate bark band as a candidate suppression interval in the first comparison process, a band sharpness of the candidate suppression interval in the first comparison process not satisfying the first preset condition;
[0375] in an ith comparison process, comparing sharpnesses corresponding to two bark bands on both sides of a candidate suppression interval obtained in an (i-1)th comparison process, and determining a bark band corresponding to a maximum sharpness in the ith comparison process as an intermediate candidate bark band, i being greater than 1 and less than K;
[0376] determining an interval between the intermediate candidate bark band and a boundary bark band obtained in the i-1th comparison process as a candidate suppression interval in the ith comparison process, the boundary bark band being a bark band far away from the intermediate candidate bark band in the candidate suppression interval obtained in the i-1th comparison process, and the frequency band sharpness of the candidate suppression interval in the ith comparison process not satisfying the first preset condition;
[0377] performing K-i comparison processes after the ith comparison process until the adaptive suppression frequency band is determined in the Kth comparison process.
[0378] Optionally, the first preset condition is that the frequency band sharpness of the candidate suppression interval obtained in each comparison process is greater than a first sharpness threshold.
[0379] Optionally, the first sharpness threshold is a product of a frame sharpness of the current audio frame and a coefficient, and the coefficient is less than 1.
[0380] Optionally, the processing unit 710 is specifically configured to determine the first suppression gain according to a gain coefficient and an initial suppression gain, the gain coefficient being related to the current audio frame.
[0381] Optionally, the gain coefficient is related to a frame sharpness of the current audio frame and a second sharpness threshold.
[0382] Optionally, the processing unit 710 is specifically configured to determine, according to the first suppression gain, a second suppression gain of each of the extended suppression frequency bands in a manner that each gain step is increased by one when each frequency band width is increased by one in a frequency variation direction from the first boundary frequency to the second boundary frequency of each of the extended suppression frequency bands.
[0383] Optionally, the gain step is obtained according to the following formula:
[0384]
[0385] wherein step is the gain step, P is the number of frequency points in each of the extended suppression frequency bands, and gain is the first suppression gain.
[0386] Optionally, the processing unit 710 is further configured to, in a case where a frequency band range of the core suppression frequency band includes a frequency band range of the adaptive suppression frequency band, apply the first suppression gain to energy of the current audio frame in the core suppression frequency band.
[0387] Optionally, the processing unit 710 is specifically configured to determine a target feature value of the current audio frame.
[0388] In a case where the target feature value satisfies a second preset condition, the current audio frame is determined as a sibilance frame.
[0389] Optionally, the target feature value is a frame sharpness of the current audio frame, and the second preset condition is that the frame sharpness of the current audio frame is greater than a second sharpness threshold.
[0390] It should be understood that the processing unit 710 can be used to execute each step of the electronic device in the method 600, and the specific description can refer to the related description above, which will not be repeated here.
[0391] In an embodiment of the present application, Figure 13 The device can also be a chip or a chip system, such as a system on chip (SoC).
[0392] Figure 14 is a schematic structural diagram of an electronic device 800 provided by an embodiment of the present application. The electronic device 800 is used to execute the corresponding steps and / or processes in the above method embodiments.
[0393] The electronic device 800 includes a processor 810, a transceiver 820, and a memory 830. The processor 810, the transceiver 820, and the memory 830 communicate with each other through an internal connection path. The processor 810 can realize the functions of the processor 810 in various possible implementations in the electronic device 800. The memory 830 is used to store instructions, and the processor 800 is used to execute the instructions stored in the memory 830, or in other words, the processor 810 can call these stored instructions to realize the functions of the processor 810 in the electronic device 800.
[0394] Optionally, the memory 830 can include read-only memory and random access memory, and provide instructions and data to the processor. Part of the memory can also include non-volatile random access memory. For example, the memory can also store device type information. The processor 810 can be used to execute the instructions stored in the memory, and when the processor 810 executes the instructions stored in the memory, the processor 810 is used to execute the steps and / or processes of the above method embodiments corresponding to the network device or the terminal device.
[0395] In addition, Figure 14 The electronic device 800 shown can correspond to Figure 1 The electronic device 100 shown, the processor 810 in the electronic device 800 can correspond to the processor 110 in the electronic device 100, and the transceiver 820 in the electronic device 800 can correspond to the antenna 1 and the antenna 2 in the electronic device 100.
[0396] The electronic device 800 is configured to perform each process and step corresponding to the electronic device in the method 600 described above, and the steps performed by the processor 810 can correspond to each step performed by the processing unit 710 in the apparatus 700. For details, refer to the foregoing description, which will not be repeated here.
[0397] It should be understood that the processor of the apparatus described above in the embodiments of the present application can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0398] In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor or the instruction in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware processor execution completion, or executed by the combination of hardware and software units in the processor. The software unit can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory, and the processor executes the instructions in the memory to complete the steps of the above method in combination with the hardware thereof. To avoid repetition, it will not be described in detail here.
[0399] The embodiments of the present application provide a computer program product, when the computer program product is run in a terminal device, so that the terminal device executes the technical solutions in the above embodiments. The implementation principle and technical effects are similar to those of the above method-related embodiments, which will not be repeated here.
[0400] The embodiments of the present application provide a readable storage medium, which contains instructions, when the instructions are run in a terminal device, so that the terminal device executes the technical solutions of the above embodiments. The implementation principle and technical effects are similar, which will not be repeated here.
[0401] The embodiments of the present application provide a chip, which is used to execute instructions, when the chip is run, executes the technical solutions in the above embodiments. The implementation principle and technical effects are similar, which will not be repeated here.
[0402] In the embodiments described above, all or some of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or some of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded into and executed by a computer, all or some of the processes or functions according to the embodiments described in the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a high-density digital video disc (digital video disc, DVD)), or a semiconductor medium (such as a solid state disk (solid state disk, SSD)), etc.
[0403] It should be understood that the "embodiments" mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0404] It should also be understood that in the present application, "when", "if" and "if" all refer to the corresponding processing of the UE or the base station under certain objective conditions, not the time limit, and do not require the UE or the base station to have a judgment action when implementing, nor does it mean that there are other limitations.
[0405] Those of ordinary skill in the art can understand that the various numbers such as first, second, etc. involved in the present application are only for the convenience of description and do not limit the scope of the embodiments of the present application, nor do they represent the order of precedence.
[0406] In the present application, the element using the singular is intended to represent "one or more" rather than "one and only one", unless otherwise specifically indicated. In the present application, "at least one" is intended to represent "one or more", and "multiple" is intended to represent "two or more", unless otherwise specifically indicated.
[0407] The term "and / or" in the present application is merely used to describe an associated relationship, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, B exists alone, where A can be singular or plural, and B can be singular or plural.
[0408] The term "at least one of" or "at least one" in the present application means all or any combination of the listed items, for example, "at least one of A, B and C" can mean that A exists alone, B exists alone, C exists alone, A and B exist together, B and C exist together, A, B and C exist together, where A can be singular or plural, B can be singular or plural, and C can be singular or plural.
[0409] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0410] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0411] In the several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other ways. For example, the device embodiments described above are only schematic, and the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0412] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0413] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0414] The functions, if realized in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the present application that essentially contributes to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0415] The same or similar parts in each embodiment of the present application can be mutually referred to. In each embodiment of the present application, and each implementation / implementation method / realization method in each embodiment, if there is no special description and no logical conflict, the terms and / or descriptions of different embodiments, and each implementation / implementation method / realization method in each embodiment are consistent and can be mutually referred to. The technical features of different embodiments, and each implementation / implementation method / realization method in each embodiment can be combined to form new embodiments, implementations, implementation methods, or realization methods according to their inherent logical relationship. The above described embodiments of the present application do not constitute a limitation on the protection scope of the present application.
[0416] The above merely describes a preferred embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims. In summary, the above merely describes a preferred embodiment of the technical solution of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be covered in the protection scope of the present application.
Claims
1. A method for processing dental sounds, characterized in that, include: The current audio frame is determined to be an sibilant frame; Determine the core suppression frequency band and adaptive suppression frequency band of the current audio frame; Determine a first suppression gain for the core suppression frequency band, wherein the first suppression gain is less than 1; The frequency range of the adaptive suppression band includes the frequency range of the core suppression band, or, if there is a partial overlap between the frequency range of the adaptive suppression band and the frequency range of the core suppression band, at least one extended suppression band is determined, and each extended suppression band is a frequency band in the adaptive suppression band other than the core suppression band. The second suppression gain of each extended suppression band is determined, and the second suppression gain of each extended suppression band is increased from the first suppression gain to 1 according to the frequency variation direction from the first boundary frequency closest to the core suppression band to the second boundary frequency furthest from the core suppression band. The first suppression gain is applied to the energy of the current audio frame in the core suppression band, and the second suppression gain of each extended suppression band is applied to the energy of the current audio frame in each of the extended suppression bands.
2. The method according to claim 1, characterized in that, Before determining the core suppression band and adaptive suppression band of the current audio frame, the method further includes: Based on the sharpness distribution of the current audio frame, determine the first bark frequency band in the bark domain corresponding to the maximum sharpness; and, determining the core suppression frequency band and adaptive suppression frequency band of the current audio frame includes: Based on the first bark band, the core suppression band and the adaptive suppression band are determined, and both the core suppression band and the adaptive suppression band include the first bark band.
3. The method according to claim 2, characterized in that, Based on the first bark band, the core suppression band is determined, including: Taking the first bark frequency band as the center frequency band, M1 bark frequency bands are extended to both sides of the first bark frequency band to obtain the core suppression frequency band, where M1 is an integer greater than or equal to 1.
4. The method according to claim 2 or 3, characterized in that, Determining the adaptive suppression frequency band based on the first bark frequency band includes: Based on the sharpness distribution of the current audio frame, frequency bands are extended to both sides of the first bark frequency band to determine the adaptive suppression frequency band, wherein the frequency band sharpness of the adaptive suppression frequency band satisfies a first preset condition, and the first preset condition is used to limit the extension of the adaptive suppression frequency band.
5. The method according to claim 4, characterized in that, The step of expanding the frequency band to both sides of the first bark band based on the sharpness distribution of the current audio frame to determine the adaptive suppression frequency band includes: Based on the sharpness distribution of the current audio frame, the sharpness of each bark frequency band in multiple bark frequency bands on both sides of the first bark frequency band is compared K times. If the frequency band sharpness of the candidate suppression interval obtained in the Kth comparison process meets the first preset condition, the frequency band corresponding to the candidate suppression interval obtained in the Kth comparison process is determined as the adaptive suppression frequency band, where K is an integer greater than 1; where, In the first comparison process, the sharpness of the two bark frequency bands on both sides of the first bark frequency band is compared, and the bark segment corresponding to the maximum sharpness in the first comparison process is determined as the initial candidate bark frequency band. The interval between the first bark frequency band and the initial candidate bark frequency band, including the first bark frequency band and the initial candidate bark frequency band, is determined as the candidate suppression interval in the first comparison process. The frequency band sharpness of the candidate suppression interval in the first comparison process does not meet the first preset condition. In the i-th comparison process, the sharpness of the two bark frequency bands on both sides of the candidate suppression interval obtained in the (i-1)-th comparison process is compared, and the bark frequency band corresponding to the maximum sharpness in the i-th comparison process is determined as the intermediate candidate bark frequency band, where i is greater than 1 and less than K. The interval between the intermediate candidate bark frequency band and the boundary bark frequency band, including the intermediate candidate bark frequency band and the boundary bark frequency band obtained in the (i-1)th comparison process, is determined as the candidate suppression interval in the i-th comparison process. The boundary bark frequency band is the frequency band in the candidate suppression interval obtained in the (i-1)th comparison process that is far away from the intermediate candidate bark frequency band. The frequency band sharpness of the candidate suppression interval in the i-th comparison process does not meet the first preset condition. The process involves performing Ki comparisons after the i-th comparison until the adaptive suppression frequency band is determined during the K-th comparison.
6. The method according to claim 5, characterized in that, The first preset condition is that the frequency band sharpness of the candidate suppression interval obtained in each comparison process is greater than the first sharpness threshold.
7. The method according to claim 6, characterized in that, The first sharpness threshold is the product of the frame sharpness of the current audio frame and a coefficient, wherein the coefficient is less than 1.
8. The method according to any one of claims 1 to 7, characterized in that, Determining the first suppression gain of the core suppression frequency band includes: The first suppression gain is determined based on the gain coefficient and the initial suppression gain, wherein the gain coefficient is related to the current audio frame.
9. The method according to claim 8, characterized in that, The gain coefficient is related to the frame sharpness and the second sharpness threshold of the current audio frame.
10. The method according to any one of claims 1 to 9, characterized in that, Determining the second suppression gain for each extended suppression band includes: Based on the frequency change direction from the first boundary frequency to the second boundary frequency of each extended suppression band, and taking the first suppression gain as a reference, the second suppression gain of each extended suppression band is determined by increasing the gain step by one for each increase in band width.
11. The method according to claim 10, characterized in that, The gain step size is obtained according to the following formula: Where step is the gain step size, P is the number of frequency points in each extended suppression band, and gain is the first suppression gain.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: When the frequency range of the core suppression band includes the frequency range of the adaptive suppression band, the first suppression gain is applied to the energy of the current audio frame in the core suppression band.
13. The method according to any one of claims 1 to 12, characterized in that, The step of determining that the current audio frame is an sibilant frame includes: Determine the target feature value of the current audio frame; If the target feature value satisfies the second preset condition, the current audio frame is determined to be an sibilant frame.
14. The method according to claim 13, characterized in that, The target feature value is the frame sharpness of the current audio frame; and the second preset condition is that the frame sharpness of the current audio frame is greater than the second sharpness threshold.
15. An electronic device, characterized in that, include: Memory, used to store computer instructions; A processor for invoking computer instructions stored in the memory to perform the method as described in any one of claims 1 to 14.
16. A computer-readable storage medium, characterized in that, Used to store computer instructions for implementing the method as described in any one of claims 1 to 14.
17. A chip, characterized in that, The chip includes: Memory: Used to store instructions; A processor for retrieving and executing the instructions from the memory, causing an electronic device on which the chip system is mounted to perform the method as described in any one of claims 1 to 14.
Citation Information
Patent Citations
Tooth tone adjustment method and device, electronic equipment and computer readable storage medium
CN112951266A
Tooth sound processing method and device, electronic equipment and storage medium
CN116189696A