Sibilance processing method and apparatus, electronic device, and storage medium
By detecting the feature values of audio frames frame by frame and dynamically adjusting the sibilance filter parameters, the problems of poor real-time performance and unreasonable parameter settings in existing technologies are solved, achieving efficient sibilance removal in real-time and recording scenarios and improving audio quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2026-03-27
AI Technical Summary
Existing sibilance processing methods have poor real-time performance, making it impossible to effectively remove sibilance in scenarios with high real-time requirements. Furthermore, manually setting parameters can easily lead to over- or incomplete filtering.
By detecting the feature values of audio frames one by one and dynamically adjusting the sibilance filter parameters, real-time sibilance detection and adjustment processing of each audio frame is achieved. Dynamic EQ or multi-band compression processing is used to avoid fixed parameter settings.
It improves the real-time performance and accuracy of sibilance processing, enabling efficient removal of sibilance in both real-time and recording scenarios, enhancing audio quality, and avoiding problems of over- or incomplete filtering.
Smart Images

Figure CN116189696B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, in particular to a sibilance processing method and device, electronic equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] Sibilance refers to all the fricative sounds emitted by a person, corresponding to a higher sharpness, which is generally not suitable for human ears to listen to. For audio acquisition software (such as singing software), the sibilance in the acquired audio data is usually processed after the audio data is acquired, so that each frame of data in the audio data is within a suitable sharpness range, avoiding damage to human hearing caused by high sharpness sibilance.
[0003] However, the real-time performance of the sibilance processing method in the related art is poor. SUMMARY
[0004] Embodiments of the present application provide a sibilance processing method and device, electronic equipment and storage medium to improve the real-time performance of sibilance adjustment processing.
[0005] In one aspect, a sibilance processing method is provided, including: obtaining a current audio frame; determining a target feature value of the current audio frame; in response to the target feature value of the current audio frame meeting a preset condition, determining that the current audio frame is a sibilance frame of sibilance, and determining an adjustment parameter of the sibilance, and performing sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance.
[0006] In some embodiments, the sibilance adjustment processing based on the adjustment parameter of the sibilance in response to the target feature value of the current audio frame meeting the preset condition, including: in response to a previous frame of the current audio frame being a non-sibilance frame and the target feature value of the current audio frame being greater than or equal to a first threshold value, determining that the current audio frame is a first sibilance frame of the sibilance, and setting the adjustment parameter of the current audio frame as the adjustment parameter of the sibilance; setting and starting a sibilance filter based on the adjustment parameter of the sibilance to perform sibilance adjustment processing on the current audio frame.
[0007] In some embodiments, the method further includes: in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being greater than or equal to the first threshold value, determining the current audio frame as a sibilance frame other than the first sibilance frame of the sibilance, keeping the adjustment parameter of the sibilance as the adjustment parameter of the first sibilance frame of the sibilance, and keeping the sibilance filter in an open state to perform sibilance adjustment processing on the current audio frame.
[0008] In some embodiments, the method further includes: in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being less than the first threshold value, determining the current audio frame as a non-sibilance frame, and stopping the sibilance filter.
[0009] In some embodiments, the method further includes: in response to the previous frame of the current audio frame being the non-sibilance frame and the target feature value of the current audio frame being less than the first threshold value, determining the current audio frame as a non-sibilance frame, and keeping the sibilance filter in a stopped state.
[0010] In some embodiments, the method further includes: in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being greater than or equal to the first threshold value, determining the current audio frame as a sibilance frame other than the first sibilance frame of the sibilance, keeping the adjustment parameter of the sibilance as the adjustment parameter of the first sibilance frame of the sibilance, and keeping the sibilance filter in an open state to perform sibilance adjustment processing on the current audio frame.
[0011] In some embodiments, the method further includes: in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being less than the first threshold value, determining the current audio frame as a non-sibilance frame, and stopping the sibilance filter.
[0012] In some embodiments, the method further includes: based on the target feature value of the current audio frame, displaying a first image of the target feature value changing over time on a display interface.
[0013] In some embodiments, the adjustment parameter of the sibilance includes a center frequency at which the sibilance energy is concentrated and an effective frequency band of the sibilance that needs to be processed by the sibilance adjustment processing, and the method further includes: determining a loudness of the current audio frame; determining, based on the loudness, the center frequency at which the sibilance energy of the current audio frame is concentrated and an effective bandwidth at which the energy of the current audio frame is attenuated to the preset attenuation ratio; and determining, based on the effective bandwidth and the center frequency at which the sibilance energy is concentrated, the effective frequency band of the current audio frame.
[0014] In some embodiments, the method further includes: displaying, based on the effective bandwidth of the current audio frame and the center frequency at which the sibilance energy of the current audio frame is concentrated, the effective bandwidth of the sibilance and the center frequency at which the sibilance energy of the sibilance is concentrated on a display interface.
[0015] In some embodiments, the adjustment parameter of the sibilance further includes a response time of the sibilance adjustment processing and a maximum attenuation amount of the sibilance adjustment processing, and the method further includes: obtaining the response time of the sibilance adjustment processing and the filtering ratio of the sibilance adjustment processing input by a display interface; determining, based on the filtering ratio and a preset trigger threshold of the sibilance adjustment processing, the maximum attenuation amount of the sibilance adjustment processing.
[0016] In some embodiments, the determining the target feature value of the current audio frame includes: in response to the frame energy of the current audio frame being greater than a second threshold and the zero-crossing rate of the current audio frame being greater than a third threshold, determining the target feature value of the current audio frame.
[0017] In some embodiments, the target feature value includes a sharpness used to measure the degree of sound sharpness.
[0018] In some embodiments, the obtaining the current audio frame includes: collecting the current audio frame by a real-time audio collection device.
[0019] Another aspect of the embodiments of the present application provides a sibilance processing apparatus, including: an obtaining unit configured to obtain a current audio frame; a determining unit configured to determine a target feature value of the current audio frame; a processing unit configured to, in response to the target feature value of the current audio frame satisfying a preset condition, determine that the current audio frame belongs to a sibilance frame of sibilance, determine an adjustment parameter of the sibilance, and perform sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance.
[0020] In some embodiments, the processing unit, when used for determining, in response to the target feature value of the current audio frame satisfying a preset condition, that the current audio frame belongs to a sibilance frame of a sibilance, and determining an adjustment parameter of the sibilance, performing sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance, is further used for: determining, in response to a previous frame of the current audio frame being a non-sibilance frame and the target feature value of the current audio frame being greater than or equal to a first threshold value, that the current audio frame is a first sibilance frame of the sibilance, and setting the adjustment parameter of the current audio frame as the adjustment parameter of the sibilance; and setting and starting a sibilance filter based on the adjustment parameter of the sibilance, to perform sibilance adjustment processing on the current audio frame.
[0021] In some embodiments, the processing unit, when used for determining, in response to the target feature value of the current audio frame satisfying a preset condition, that the current audio frame belongs to a sibilance frame of a sibilance, and determining an adjustment parameter of the sibilance, performing sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance, is further used for: determining, in response to a previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being greater than or equal to the first threshold value, that the current audio frame is a sibilance frame other than the first sibilance frame of the sibilance, maintaining the adjustment parameter of the sibilance as the adjustment parameter of the first sibilance frame of the sibilance; and maintaining an open state of the sibilance filter, to perform sibilance adjustment processing on the current audio frame.
[0022] In some embodiments, the processing unit is further used for determining, in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being less than the first threshold value, that the current audio frame is a non-sibilance frame, and stopping the sibilance filter.
[0023] In some embodiments, the processing unit is further used for determining, in response to the previous frame of the current audio frame being the non-sibilance frame and the target feature value of the current audio frame being less than the first threshold value, that the current audio frame is a non-sibilance frame, and maintaining a stop state of the sibilance filter.
[0024] In some embodiments, the determining unit, when used for determining the target feature value of the current audio frame, is further used for: obtaining an initial feature value of the current audio frame; and performing smoothing processing on the initial feature value to obtain the target feature value of the current audio frame.
[0025] In some embodiments, the determining unit is further configured to: determine a number M of audio frames for smoothing based on a frame length of the current audio frame; obtain target feature values of M-1 audio frames before the current audio frame; and smooth the initial feature value of the current audio frame based on the target feature values of the M-1 audio frames to obtain the target feature value of the current audio frame.
[0026] In some embodiments, the apparatus further includes a first display unit configured to display, on a display interface, a first image of the target feature value over time based on the target feature value of the current audio frame.
[0027] In some embodiments, the adjustment parameter of the sibilance includes a center frequency at a sibilance energy concentration and an effective frequency band of the sibilance that needs to be subjected to the sibilance adjustment processing, and the apparatus further includes a first parameter determining unit configured to determine a loudness of the current audio frame; determine the center frequency at the sibilance energy concentration of the current audio frame and an effective bandwidth at which the energy of the current audio frame decays to the preset decay ratio based on the loudness; and determine the effective frequency band of the current audio frame based on the effective bandwidth and the center frequency at the sibilance energy concentration.
[0028] In some embodiments, the apparatus further includes a second display unit configured to display, on a display interface, the effective bandwidth of the sibilance and the center frequency at the sibilance energy concentration of the sibilance based on the effective bandwidth of the current audio frame and the center frequency at the sibilance energy concentration of the current audio frame.
[0029] In some embodiments, the adjustment parameter of the sibilance further includes a response time of the sibilance adjustment processing and a maximum decay amount of the sibilance adjustment processing, and the apparatus further includes a second parameter determining unit configured to obtain the response time of the sibilance adjustment processing and a filter-out ratio of the sibilance adjustment processing input by a display interface; and determine the maximum decay amount of the sibilance adjustment processing based on the filter-out ratio and a preset trigger threshold of the sibilance adjustment processing.
[0030] In some embodiments, the determining unit is further configured to: determine the target feature value of the current audio frame in response to the frame energy of the current audio frame being greater than a second threshold and the zero-crossing rate of the current audio frame being greater than a third threshold.
[0031] In some embodiments, the target feature value includes a sharpness for measuring a degree of sound sharpness.
[0032] In some embodiments, the acquisition unit, when used for acquiring the current audio frame, is further configured to: acquire the current audio frame by using a real-time audio acquisition device.
[0033] Another aspect of the embodiments of the present application provides an electronic device, comprising: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the method according to any one of the preceding embodiments via execution of the executable instructions.
[0034] Another aspect of the embodiments of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method according to any one of the preceding embodiments.
[0035] Another aspect of the embodiments of the present application provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the method according to any one of the preceding embodiments.
[0036] The sibilance processing method and device, electronic device and storage medium provided by the embodiments of the present application can acquire a current audio frame, determine a target feature value of the current audio frame, determine that the current audio frame is a sibilance frame of sibilance in response to the target feature value of the current audio frame meeting a preset condition, and determine an adjustment parameter of the sibilance, and then perform sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance, so that it can be determined whether each audio frame is sibilance, and if so, the sibilance adjustment processing is performed on the audio frame, thereby eliminating sibilance frame by frame and improving the quality of the audio. Since the sibilance processing method provided by the embodiments of the present application can perform sibilance detection and adjustment processing frame by frame, it can be applied to real-time scenes with high real-time requirements, that is, sibilance detection and adjustment processing can be performed on each audio frame acquired in real time, thereby improving the real-time performance of sibilance adjustment processing. At the same time, the sibilance processing method can also be applied to recording and broadcasting scenes with low real-time requirements, so as to perform sibilance detection and adjustment processing on each audio frame in a recording file, thereby improving the scene diversity of the sibilance processing method. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0038] Figure 1 Structure diagram of a system applying the sibilance processing method of the embodiments of the present application;
[0039] Figure 2A flowchart of a sibilance processing method provided by an embodiment of the present application is shown in FIG. 1.
[0040] Figure 3 A comparison diagram of initial characteristic values and target characteristic values provided by an embodiment of the present application is shown in FIG. 3.
[0041] Figure 4 A diagram of loudness provided by an embodiment of the present application is shown in FIG. 4.
[0042] Figure 5 A flowchart of a sibilance processing method provided by another embodiment of the present application is shown in FIG. 5.
[0043] Figure 6 A structural diagram of a sibilance processing apparatus provided by an embodiment of the present application is shown in FIG. 6.
[0044] Figure 7 A structural diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 7. DETAILED DESCRIPTION
[0045] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0046] The embodiments of the present application provide a sibilance processing method, apparatus, electronic device and storage medium. Specifically, the sibilance processing method of the embodiments of the present application can be executed by an electronic device, which can be a terminal or a server, etc. The terminal can be a smart phone, a tablet computer, a notebook computer, a smart voice interaction device, a smart home appliance, a wearable smart device, a flying vehicle, a smart vehicle terminal, etc. The terminal can also include a client, which can be an audio client, a video client, a browser client, an instant messaging client or an applet, etc. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms, etc.
[0047] First, the proper nouns related to the embodiments of the present application are explained.
[0048] Level: the level of an audio signal, which can represent the volume.
[0049] Equalizer (EQ): Each frequency band of human voice has its own sound characteristics. In music post-production, a certain frequency band needs to be attenuated by a certain value to obtain better sound effects. For sibilance, the frequency band of sibilance is relatively high, and usually only attenuation is considered.
[0050] Classification of EQ: ① Static EQ refers to the attenuation of a certain frequency band of audio, which does not change dynamically with signal level; ② Dynamic EQ refers to the gain or attenuation of a certain frequency band of audio, which changes dynamically with signal level. The greater the level, the greater the attenuation, and vice versa. Dynamic EQ can make the volume of a certain segment of audio not fluctuate greatly.
[0051] Parameters of dynamic EQ: ① Q value, defined as the frequency band of EQ, i.e. the frequency band that needs gain or attenuation; ② Frequency, defined as the center frequency of EQ, i.e. the center frequency of the sibilance energy concentration; ③ Threshold, the signal exceeding the threshold will trigger EQ; ④ Range, defined as the maximum attenuation of dynamic EQ in dB; ⑤ Attack, defined as the response time of dynamic EQ from triggering to full operation; ⑥ Release, defined as the release time of dynamic EQ from full operation to cancellation. The actual attenuation Gain of each frame of audio signal can be dynamically calculated by the above 6 parameters. If the level of the current audio frame is x dB, then the actual maximum attenuation Gain = Range*(1-x / Threshold).
[0052] Dry sound: refers to the sound after recording without any processing.
[0053] Sibilance (ess / sibilance): Sibilance in a broad sense refers to all fricative sounds emitted by a person. In this embodiment, it refers to the fricative sound emitted by a person when singing or speaking, which usually appears at the beginning of consonants such as "s", "c", and "q" and has a frequency of 2k-10kHz.
[0054] Sharpness (unit: acum): the degree of harshness of sound. According to the principle of psychoacoustics, the more concentrated the energy of sibilance in the frequency spectrum, the more harsh the sibilance, the higher the center frequency value of the sibilance, and the greater the overall energy of the sibilance.
[0055] Loudness: loudness is the loudness of sound perceived by the human ear. Frequency ranges with the same acoustic characteristics are usually divided into a Bark frequency band, and the human auditory frequency band can usually be divided into 25 Bark frequency bands.
[0056] DeEsser: a human voice sibilance removal plug-in. Audio engineers use De-esser to remove sibilance from human voice. Such plug-ins generally have adjustable parameters, real-time processing, and real-time monitoring functions.
[0057] bypass: bypass, for audio processor, when bypass parameter is true, enter non-working state, when false, enter working state.
[0058] external UI: show the state variable inside DeEsser to external user.
[0059] In order to improve the quality of audio, it is usually necessary to remove the sibilance in the dry sound. In the related art, there is a sibilance processing method. For recording and broadcasting files, the sibilance is discriminated based on the sharpness value, so that the problem of false detection of normal human voice can be solved. However, the detection process and the processing process of the method are separated, that is, the sibilance of the entire audio file is detected first, and then the sibilance of the entire audio file is processed, so it cannot be used in some scenes with high real-time requirements.
[0060] In addition, in the related art, the sibilance processing also uses DeEsser plug-in, including Waves RDeEsser, Izotope Rx9 De-ess, Fabfilter Pro-DS and the like. The implementation of DeEsser plug-in generally uses multi-section compressor or dynamic EQ. DeEsser plug-in needs to input manual adjustment parameters in the display interface. The adjustment parameters generally include: ①Frequency, center frequency or cutoff frequency of sibilance energy concentration; ②Type: filter type used for sibilance filtering (generally three kinds of LowPass / Notch / Peak Filter) ③Threshold: level threshold triggering work, similar to the Threshold of dynamic EQ; ④Range: similar to the Range of dynamic EQ; ⑤Speed: similar to the Attack of dynamic EQ; ⑥Monitor: sibilance monitoring, after being selected, only the filtered sibilance can be heard.
[0061] However, the DeEsser plug-in in the related art requires manual setting of multiple parameters. For example, for a certain audio file, a fixed Frequency needs to be manually set. However, the Frequency of different sibilants is usually different, and the Frequency of sibilants in different pronunciations of different people is also different. For example, in Chinese songs, the initial sound "z" of "走" is often sung as "c", and the frequency of "z" is lower than that of "c"; in Chinese songs, the initial sound "t" of "他" is often sung as "ch" (the frequency of "t" is lower than that of "ch"); similarly, the initial sound "j" of "就" is sung as "q" (the frequency of "j" is lower than that of "q"). This is the common phenomenon of "voiced consonants being sung as voiceless consonants" in Chinese songs, and there are differences in the singing methods of different people, which will also cause differences in the sibilants of the same Chinese character for different people. The fixed Frequency will cause the DeEsser plug-in to filter out some sibilants too much, while some sibilants are not filtered cleanly, resulting in a poor sibilant filtering effect.
[0062] Moreover, it takes a lot of time to manually adjust to the appropriate adjustment parameters. Once the parameters are set unreasonably (such as the Frequency being set too low or the Threshold being set too low), it is easy to misdetect normal human voices, and the filtered sound will become dull, resulting in damage to the human voice timbre.
[0063] To solve at least one of the above problems, the embodiments of the present application provide a sibilant processing method, device, electronic device and storage medium. By obtaining the current audio frame; determining the target feature value of the current audio frame; in response to the target feature value of the current audio frame satisfying a preset condition, determining that the current audio frame is a sibilant frame belonging to a sibilant, and determining the adjustment parameter of the sibilant, and performing sibilant adjustment processing on the current audio frame based on the adjustment parameter of the sibilant, so as to determine whether each audio frame is a sibilant. If so, perform sibilant adjustment processing on the audio frame, thereby eliminating sibilants frame by frame and improving the quality of the audio. Since the sibilant processing method provided by the embodiments of the present application can perform sibilant detection and adjustment processing frame by frame, it can be applied to real-time scenarios with high requirements for real-time performance, that is, sibilant detection and adjustment processing can be performed according to each audio frame obtained in real time, improving the real-time performance of sibilant adjustment processing. At the same time, the sibilant processing method can also be applied to recording and broadcasting scenarios with low requirements for real-time performance to perform sibilant detection and adjustment processing on each audio frame in the recording file, improving the scene diversity of the sibilant processing method. Figure 1 FIG. is a schematic structural diagram of a system applying the sibilant processing method of the embodiments of the present application. Please refer to Figure 1 This system includes a terminal 10 and a server 20, etc.; the terminal 10 and the server 20 are connected through a network, for example, through a wired or wireless network connection, etc.
[0064] The terminal 10 can be configured to display a graphical user interface. The terminal can be configured to interact with a user through the graphical user interface, for example, by downloading and installing a corresponding client through the terminal, by invoking and running a corresponding applet, by logging into a website to present a corresponding graphical user interface, and the like. In the embodiments of the present application, the terminal 10 can be installed with an audio processing application, and display an audio waveform image and the like through the audio processing application. The server 20 can be configured to obtain a current audio frame, determine a target feature value of the current audio frame, in response to the target feature value of the current audio frame satisfying a preset condition, determine that the current audio frame is a sibilance frame of sibilance, determine an adjustment parameter of the sibilance, and perform sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance.
[0065] It should be noted that, although the application program is taken as an audio processing application for illustration, those skilled in the art should understand that the application program can be other suitable programs, such as an audio playing application, a video playing application, a video processing application, and the like. In addition, the application program can be an application installed on a desktop computer, an application installed on a mobile terminal, an applet embedded in an application, and the like, which are not specially limited in the present disclosure.
[0066] It should be noted that the above application scenarios are only for facilitating understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0067] The following will be described in detail. It should be noted that the description order of the following embodiments is not a limitation on the priority order of the embodiments.
[0068] Figure 2 A flowchart of a sibilance processing method provided by an embodiment of the present application is shown in FIG. 1. Figure 2 The embodiment of the present application provides a sibilance processing method 100, including the following steps S110 to S130.
[0069] In step S110, a current audio frame is obtained.
[0070] In step S120, a target feature value of the current audio frame is determined.
[0071] In step S130, in response to the target feature value of the current audio frame satisfying a preset condition, it is determined that the current audio frame is a sibilance frame of sibilance, and an adjustment parameter of the sibilance is determined. The current audio frame is subjected to sibilance adjustment processing based on the adjustment parameter of the sibilance.
[0072] In step S110, the audio data can be subjected to frame processing, and the current audio frame can be obtained. It can be understood that the audio data can be real-time data or recorded data, which can be set according to requirements. The length of each audio frame can be N, where N is a positive integer less than or equal to 2048 if the audio data is real-time data. If N is greater than 2048, the real-time performance will be lost.
[0073] In some embodiments, in step S110, the current audio frame can be obtained by using a real-time audio acquisition device. The real-time audio acquisition device can be used in a real-time scenario with human voice, such as singing recording, live anchor recording, live singing, and speech, so as to perform sibilance adjustment processing on real-time audio data and improve the audio quality.
[0074] In step S120, the target feature value of the current audio frame can be determined. The target feature value can be used to distinguish normal sound and sibilance, so as to determine the working range of sibilance adjustment processing. In some embodiments, the target feature value includes sharpness, which is used to measure the sharpness of sound, so as to accurately distinguish normal sound and sibilance. Of course, in other embodiments, the target feature value can also include loudness and other audio parameters.
[0075] It can be understood that a sibilance usually lasts for 40-400 ms, and an audio frame with 1024 sampling points only lasts for 23 ms. Therefore, a sibilance usually includes multiple audio frames. For ease of description, the audio frames in a sibilance are referred to as sibilance frames, and the audio frames of normal sound that are not sibilance are referred to as non-sibilance frames.
[0076] In step S130, if the target feature value of the current audio frame meets a preset condition, it can be determined that the current audio frame is a sibilance frame of sibilance. The adjustment parameter of the sibilance can be determined, and the current audio frame can be subjected to sibilance adjustment processing based on the adjustment parameter of the sibilance. In addition, if the target feature value of the current audio frame does not meet the preset condition, it can be determined that the current audio frame is a non-sibilance frame, and the audio frame can not be processed and can be directly copied.
[0077] It can be understood that in the present embodiment, if the current audio frame is a sibilance frame, the adjustment parameter of the sibilance can be calculated by using the related parameters of the current audio frame, and then the current audio frame can be subjected to sibilance adjustment processing by using the adjustment parameter of the sibilance. The sibilance adjustment processing can be dynamic EQ processing or multi-section compression processing.
[0078] In this embodiment, the target feature value can be used to distinguish sibilance frames from normal sound frames frame by frame, so that sibilance frames can be processed in a targeted manner, and the real-time performance of sibilance processing is improved. The sibilance processing can be used in real-time scenes as well as in recording and broadcasting scenes with low real-time requirements.
[0079] In addition, compared with the DeEsser plug-in in the related art, the target feature value can be used to determine whether the sibilance is sibilance. If so, the adjustment parameter of the sibilance can be determined in a targeted manner, and the sibilance can be adjusted. Thus, the problem of excessive sibilance filtering or poor sibilance filtering effect caused by setting a fixed parameter can be avoided, and the accuracy of sibilance processing is improved.
[0080] In some embodiments, in step S130, in response to the target feature value of the current audio frame meeting a preset condition, it is determined that the current audio frame is a sibilance frame of sibilance, and an adjustment parameter of the sibilance is determined. The sibilance of the current audio frame is adjusted based on the adjustment parameter of the sibilance. The sibilance adjustment processing can include the following substep one and substep two.
[0081] Substep one, in response to the previous frame of the current audio frame being a non-sibilance frame and the target feature value of the current audio frame being greater than or equal to a first threshold value, the current audio frame is determined to be the first sibilance frame of sibilance, and the adjustment parameter of the current audio frame is set as the adjustment parameter of the sibilance.
[0082] Substep two, based on the adjustment parameter of the sibilance, a sibilance filter is set and started to adjust the sibilance of the current audio frame.
[0083] The first threshold value can be a critical value for determining sibilance. Taking the target feature value as an example, the first threshold value can be a critical sharpness value of sibilance and normal sound. If the target feature value is greater than or equal to the critical sharpness value, the current audio frame can be considered as a sibilance frame. If the target feature value is less than the critical sharpness value, the current audio frame can be considered as a non-sibilance frame. The non-sibilance frame can be an audio frame of normal sound other than sibilance. The previous frame of the current audio frame being a non-sibilance frame can be understood as the target feature value of the previous audio frame of the current audio frame being less than the first threshold value.
[0084] If the previous frame of the current audio frame is a non-sibilance frame, and the target feature value of the current audio frame is greater than or equal to the first threshold value, the current audio frame can be determined to be the first sibilance frame of sibilance, that is, the sibilance starts from the current audio frame. At this time, the adjustment parameter of the current audio frame can be calculated based on the current audio frame, and the adjustment parameter can be set as the adjustment parameter of the sibilance.
[0085] Since the current sibilance frame is the first sibilance frame of sibilance, the parameter of the sibilance filter can be set based on the adjustment parameter of the sibilance, and the sibilance filter can be started, so that the sibilance of the current audio frame can be adjusted.
[0086] It can be understood that the sibilance adjustment processing can be implemented by means of a sibilance filter, which can be a dynamic EQ filter (hereinafter referred to as dynamic EQ). The sub-step two can set the parameters of the dynamic EQ and start to enable the dynamic EQ to perform the sibilance adjustment processing on the current audio frame. Compared with a static EQ, the dynamic EQ takes into account the different influences of different sibilances on the filtering degree, filters more sibilances with high energy and filters less sibilances with low energy, thereby improving the accuracy of sibilance processing.
[0087] Of course, in other embodiments, the sibilance filter can also be a multi-section compression filter or the like, which can be set according to actual conditions.
[0088] In this embodiment, by judging the type of the previous frame of the current audio frame and whether the current audio frame is greater than or equal to the first threshold value, the opening time of the sibilance filter and the adjustment parameter for the sibilance adjustment processing can be determined.
[0089] In some embodiments, in response to the target feature value of the current audio frame satisfying the preset condition in step S130, it is determined that the current audio frame belongs to the sibilance frame of the sibilance, and the adjustment parameter of the sibilance is determined. The sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance can also include sub-step three and sub-step four.
[0090] In response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being greater than or equal to the first threshold value, the current audio frame is determined to be the remaining sibilance frame of the sibilance except the first sibilance frame, and the adjustment parameter of the sibilance is kept as the adjustment parameter of the first sibilance frame of the sibilance.
[0091] The sub-step four keeps the opening state of the sibilance filter to perform the sibilance adjustment processing on the current audio frame.
[0092] In this embodiment, if the previous frame of the current audio frame is the sibilance frame and the target feature value of the current audio frame is greater than or equal to the first threshold value, it can be determined that the current audio frame is the remaining sibilance frame of the sibilance. The remaining sibilance frame is the remaining sibilance frame of the sibilance except the first sibilance frame. If the sibilance includes four sibilance frames, the remaining sibilance frame can be the second, third or fourth sibilance frame of the sibilance.
[0093] Since the adjustment parameter of the sibilance has been calculated when the first sibilance frame of the sibilance is detected, and the sibilance filter has been set and opened, when the current audio frame is the remaining sibilance frame, the parameters of the sibilance filter can be used, and the opening state of the sibilance filter can be kept, so as to continue the sibilance adjustment processing on the remaining sibilance frame.
[0094] It can be understood that in the embodiment, the first sibilance frame of the sibilance can be used to determine the adjustment parameter of the sibilance, and the parameter is used to process the plurality of sibilance frames of the entire sibilance, so that the calculation amount of the method 100 can be reduced, and one adjustment parameter can be used for adjustment processing for each sibilance, to realize the targeted dynamic adjustment of the sibilance.
[0095] In some embodiments, the method 100 can further include: in response to the previous frame of the current audio frame being a sibilance frame and the target feature value of the current audio frame being less than the first threshold, determining that the current audio frame is a non-sibilance frame, and stopping the sibilance filter.
[0096] The target feature value of the current audio frame being less than the first threshold indicates that the current audio frame is a non-sibilance frame, and if the previous frame of the current audio frame is a sibilance frame and the current sibilance frame is a non-sibilance frame, it indicates that the current audio frame is normal sound, at which time the sibilance filter can be stopped, i.e., the current audio frame is not subjected to sibilance adjustment processing, so as to clearly determine the stopping time of the sibilance adjustment processing.
[0097] In some embodiments, the method 100 can further include: in response to the previous frame of the current audio frame being a non-sibilance frame and the target feature value of the current audio frame being less than the first threshold, determining that the current audio frame is a non-sibilance frame, and maintaining the stopped state of the sibilance filter.
[0098] If the current audio frame and the previous frame of the current audio frame are both non-sibilance frames, it can be determined that the current audio frame is a non-sibilance frame, and the current audio frame is a non-first frame of the normal sound, at which time the stopped state of the sibilance filter can be maintained, i.e., the current audio frame is not subjected to sibilance adjustment processing.
[0099] By the above method, the starting frame and the ending frame of each sibilance are determined, the influence of the sharpness of each audio frame in a sibilance on the harshness of the sibilance is considered, and by accurately determining the audio frame at which the sibilance starts and the audio frame at which the sibilance ends, the starting time and the stopping time of the sibilance filter can be accurately determined, the adjustment processing of the sibilance of the audio data can be realized, the sibilance can be accurately detected in real time and accurately filtered, and the sibilance processing effect is improved.
[0100] In some embodiments, the determination of the target feature value of the current audio frame in step S120 can include: obtaining an initial feature value of the current audio frame; and performing smoothing processing on the initial feature value to obtain the target feature value of the current audio frame.
[0101] It can be understood that in the embodiment, the target feature value can be a target feature value after smoothing. Taking sharpness as an example, the initial feature value can be an initial sharpness calculated according to the current audio frame. For example, the initial sharpness of the current audio frame can be calculated by using the standard document DIN 45692, that is, the initial sharpness can be calculated by using the loudness of the current audio frame.
[0102] Figure 3 An initial feature value and a target feature value provided by the embodiment of the application are compared in a graph. Curve A1 is a sibilance waveform graph of "xi", curve A2 is an initial sharpness before smoothing, and the trend of the curve is jagged, and curve A3 is a sharpness after smoothing, and the trend is relatively flat. The curve after smoothing can be used as an explicit UI, so that the sharpness of the sibilance can be scientifically represented, and at the same time, the sharpness after smoothing can also be used to judge the working interval of the sibilance filter.
[0103] In some embodiments, the initial feature value is smoothed to obtain the target feature value of the current audio frame, which can include the following steps one to three.
[0104] Step one, based on the frame length of the current audio frame, determine the number M of audio frames used for smoothing.
[0105] Step two, obtain the target feature values of the M-1 audio frames before the current audio frame.
[0106] Step three, based on the target feature values of the M-1 audio frames, smooth the initial feature value of the current audio frame to obtain the target feature value of the current audio frame.
[0107] In the embodiment, the initial sharpness value can be smoothed by 2048 sampling points. Taking the frame length N as an example, M=2048 / N initial sharps of consecutive frames need to be smoothed, so that the real-time requirement can be met. For example, if the frame length N=2048, M=1 does not need to be smoothed; if N=1024, M=2, the initial sharps of the consecutive 2 need to be smoothed; if N=512, M=4, the initial sharps of the consecutive 4 need to be smoothed, and so on.
[0108] It can be understood that the larger the frame length N is, the better the frequency resolution is, and the more accurate the sibilance adjustment parameters (such as Frequency and Q value) are, but the higher the calculation complexity is. On the contrary, the smaller N is, the lower the calculation complexity is. In combination with the above factors, in some embodiments, N can be set to 512, and the smoothing number M can be set to 4, so that the calculation complexity and the accuracy of the sibilance adjustment parameters can be balanced.
[0109] In the smoothing process, the initial feature value of the current audio frame and the target feature values of the previous M-1 audio frames can be considered comprehensively. For example, taking N=512 and M=4 as an example, assuming that the current audio frame is the 20th audio frame of the audio data, the sharpness of the 17th-19th audio frames of the audio data can be obtained, and the initial sharpness of the current audio frame is obtained, and then the four data are comprehensively smoothed, and the result of the processing is taken as the sharpness of the 20th audio frame.
[0110] In some embodiments, if the current audio frame is the 1st audio frame of the audio data, the sharpness thereof can be the initial sharpness. If the current audio frame is the 2nd audio frame of the audio data, the sharpness thereof can be obtained by smoothing the sharpness of the 1st audio frame and the initial sharpness of the current audio frame, and so on. If there are less than M-1 audio frames before the current audio frame, the sharpness of the current audio frame can be obtained by smoothing the sharpness of all the audio frames before the current audio frame and the initial sharpness of the current audio frame.
[0111] The smoothing method can be various, for example, a real-time calculation "sliding smoothing method" can be used.
[0112] The embodiment comprehensively considers the influence of audio frames with different frame lengths on the sharpness calculation, and can combine the initial feature value of the current audio frame and the target feature values of the previous M-1 audio frames to perform smoothing processing, so as to obtain the target feature value of the current audio frame.
[0113] In some embodiments, the method 100 can further include: based on the target feature value of the current audio frame, displaying a first image of the target feature value changing over time on a display interface, that is, the sharpness curve after the smoothing processing can be used for external UI, so that the sharpness of the sibilance can be represented scientifically, and the user can intuitively observe the sharpness of the sibilance of the audio data, and thus the sounder can have a clearer cognition of his own voice, so that better recording quality can be obtained in the next recording, and the sounder can also understand how to process the voice.
[0114] In some embodiments, the adjustment parameter of the sibilance includes: a center frequency of a sibilance energy concentration place and an action frequency band of the sibilance that needs to be adjusted, and the method 100 can further include: determining the loudness of the current audio frame; based on the loudness, determining the center frequency of the sibilance energy concentration place and the action bandwidth of the current audio frame when the energy of the current audio frame decays to a preset decay ratio; and based on the action bandwidth and the center frequency of the sibilance energy concentration place, determining the action frequency band of the current audio frame.
[0115] The center frequency of the sibilance energy concentration place corresponds to the Frequency of the dynamic EQ, and the action frequency band of the sibilance that needs to be adjusted corresponds to the Q value of the dynamic EQ.
[0116] Figure 4 The diagram of loudness provided by the embodiment of the present application; please refer to Figure 4 , the abscissa is loudness, and the ordinate is energy ratio k%, after obtaining the current audio frame, the loudness of the current audio frame can be calculated first, thereby obtaining Figure 4 the curve shown in the figure, and the point B with the maximum loudness is the sibilance energy concentration point, thus the center frequency of the sibilance energy concentration point can be obtained by calculating the frequency of the point B.
[0117] In addition, the ordinate of the point B can correspond to the energy ratio k% being 100%, according to the diagram, the action bandwidth of the current audio frame when the energy is attenuated to the preset attenuation ratio k% can be calculated, wherein, generally, 20<k<40. If k is 40, the straight line with the ordinate being 40 in the figure intersects with the loudness curve, and the action bandwidth of the range C of the two intersection points is the action bandwidth of the current audio frame when the energy is attenuated to the preset attenuation ratio, and based on Q=Frequency / action bandwidth, the Q value of the dynamic EQ can be calculated.
[0118] The calculation of the adjustment parameters of the sibilance in the embodiment can be performed when the target characteristic value is determined in step S120, or the parameter calculation can be performed when the current audio frame is determined to be the first sibilance frame of the sibilance in step S130, and the specific setting can be made according to the actual situation.
[0119] The embodiment can adaptively calculate the adjustment parameters such as the frequency and the Q value of the sibilance according to the sibilance, without manual setting, thereby improving the user experience, and compared with setting the fixed adjustment parameters, the adjustment parameters of the sibilance are more accurate, and the filtering effect of the sibilance can be improved.
[0120] In some embodiments, the method 100 can further include: displaying the action bandwidth of the sibilance and the center frequency of the sibilance energy concentration point of the sibilance on the display interface based on the action bandwidth of the current audio frame and the center frequency of the sibilance energy concentration point of the sibilance of the current audio frame.
[0121] It can be understood that the action bandwidth and the center frequency of the sibilance energy concentration point of each audio frame can be calculated, thereby the action bandwidth curve of the entire audio data and the center frequency curve of the sibilance energy concentration point of the sibilance can be displayed on the display interface, thereby the sound cognition of the sound recorder can be more clear, and better recording quality can be obtained in the next recording, and the sound processor can be more familiar with how to process the sound.
[0122] In addition, it can be understood that the action bandwidth and the frequency of each audio frame can be used for an explicit UI, and the action bandwidth and the frequency of the first sibilance frame can be used to set the sibilance filter for sibilance adjustment processing.
[0123] In some embodiments, the adjustment parameters of the sibilance can further include a response time of the sibilance adjustment processing and a maximum attenuation amount of the sibilance adjustment processing, and the method 100 can further include: obtaining the response time of the sibilance adjustment processing and a filtering ratio of the sibilance adjustment processing input by the display interface; determining the maximum attenuation amount of the sibilance adjustment processing based on the filtering ratio and a preset trigger threshold of the sibilance adjustment processing.
[0124] The response time of the sibilance adjustment processing corresponds to the Attack of the dynamic EQ, and the maximum attenuation amount of the sibilance adjustment processing corresponds to the Range of the dynamic EQ. The preset trigger threshold of the sibilance adjustment processing corresponds to the Threshold of the dynamic EQ, and can be a preset parameter.
[0125] The user can input the Attack and the filtering ratio Ratio of the sibilance adjustment processing through the display interface. The Attack can be adjusted by the user through a plurality of input parameters Speed set on the display interface (for example, three grades 1ms / 5ms / 10ms), and the Ratio can be valued from 0.1 to 1.
[0126] After obtaining the Ratio set by the user, the Range can be calculated by Range=Ratio*Threshold, and the maximum attenuation amount Gain of the sibilance adjustment processing can be calculated by Gain=Range*(1-x / Threshold)=Ratio*(Threshold-x).
[0127] Since the Attack and the Ratio two parameters can affect the strength of the sibilance, different people have different feelings about the strength of the sibilance, and therefore, it is more reasonable for the user to adjust these two parameters through the display interface, which improves the user's experience. Compared with the DeEsser plug-in in the related art, the user needs to set fewer parameters, and does not need to spend much time adjusting the parameters, thereby improving the accuracy of sibilance detection and improving the effect of sibilance processing. For the user (audio engineer or other user), the productivity can be liberated, and there is no need to manually set the frequency range of the sibilance, and there is no need to worry about being difficult to adjust a fixed frequency that meets all sibilance ranges, and there is no need to worry about that the sibilance filtering is too much to affect the normal human voice part.
[0128] Of course, in some embodiments, the sibilance adjustment parameter can also include a working time Release of the sibilance adjustment processing and a trigger threshold Threshold of the sibilance adjustment processing, which can be preset by the system.
[0129] In some embodiments, the step S120 of determining the target feature value of the current audio frame can include: in response to the frame energy of the current audio frame being greater than a second threshold and the zero-crossing rate of the current audio frame being greater than a third threshold, determining the target feature value of the current audio frame.
[0130] Wherein, the zero-crossing rate (ZCR) refers to the ratio of the sign change of a signal, for example, the signal changes from positive to negative or vice versa. The second threshold can be a frame energy threshold, and the third threshold can be a zero-crossing rate threshold, which can be set according to actual conditions.
[0131] It can be understood that after obtaining the current audio frame and before determining the target feature value of the current audio frame, the method 100 can further include pre-processing the current audio frame, which can determine whether the frame energy of the current audio frame is less than the second threshold and the zero-crossing rate of the current audio frame is greater than the third threshold. If the frame energy is greater than the second threshold and the zero-crossing rate is greater than the third threshold, the target feature value of the current audio frame can be determined. If the frame energy and the zero-crossing rate threshold conditions are not met, the sibilance detection and processing of the audio frame can be directly skipped.
[0132] Since the audio frame with too small frame energy or too small zero-crossing rate is not considered as a sibilance frame, the audio frame with too small frame energy or too small zero-crossing rate can be filtered out in advance, thereby saving computing resources.
[0133] Figure 5 The flowchart of the sibilance processing method provided for another embodiment of the present application is shown in the figure. Please refer to Figure 5 In one embodiment, taking frame length N=512 and M=4 as an example, after obtaining the current audio frame, it can be first determined whether the frame energy (rms) is greater than the second threshold (rms_thres) and the zero-crossing rate (zcr) is greater than the third threshold (zcr_thres). If rms>rms_thres&zcr>zcr_thres is not met, the sibilance detection and processing steps can be skipped, and the audio data of the current audio frame, i.e., the audio data with a length of 512, can be directly copied.
[0134] If rms>rms_thres&zcr>zcr_thres is met, the audio features are detected, the sharpness of the current audio frame, the center frequency (Frequency) at the sibilance energy concentration, and the action bandwidth when the energy decays to a preset decay ratio k% are calculated.
[0135] If the length of the frames continuously collected is not less than 2048, i.e. the total length of the frames collected is greater than or equal to 2048, it means that 4 audio frames have been exceeded, and then one smoothing process is performed for 4 audio frames with a total length of 2048 frames (including the current audio frame and 3 audio frames before the current audio frame).
[0136] In addition, after detecting the audio features, 4 cases are included as follows:
[0137] ① If the last audio frame (prev_ess) is not a sibilance frame and the sharpness of the current audio frame is greater than or equal to the first threshold, i.e. prev_ess = false and sharpness >= sharpness_thres, it is determined that the current audio frame is the first sibilance frame of the sibilance, at this time the parameter crnt_ess of the current audio frame can be set to true (indicating that the current audio frame is a sibilance frame), bypass is set to false, dynamic EQ is reset and started to work, and the parameter setting of dynamic EQ is as follows:
[0138] (1) The value of Frequency is set to the Frequency calculated in the step of detecting audio features.
[0139] (2) The value of Q is set to the parameter calculated based on the Frequency and the action bandwidth of energy attenuation to k%.
[0140] (3) Attack can be determined according to the gear selection of the input parameter Speed for the display interface.
[0141] (4) Release can be set to 150 ms according to experience.
[0142] (5) Threshold can be set to a lower value of -60 dB (acting on signals greater than -60 dB).
[0143] (6) Range can be calculated from Threshold and the parameter Ratio input by the user through the display interface, the formula is Range = Ratio * Threshold; therefore, the maximum attenuation amount Gain = Range * (1-x / Threshold) = Ratio * (Threshold-x).
[0144] ② If the last audio frame is a sibilance frame and the sharpness of the current audio frame is not less than the first threshold, i.e. prev_ess = true and sharpness >= sharpness_thres, the current audio frame belongs to the remaining sibilance frames of the current sibilance, and dynamic EQ can maintain the working state.
[0145] If the previous audio frame is a zither frame and the sharpness of the current audio frame is lower than the first threshold, i.e. prev_ess=true and sharpness<sharpness_thres, the current audio frame is the first non-zither frame entering the non-zither, crnt_ess is set to false (indicating that the current audio frame is a non-zither frame), bypass is set to true, and the dynamic EQ stops working.
[0146] If the previous audio frame is a non-zither frame and the sharpness of the current audio frame is lower than the first threshold, i.e. prev_ess=false and sharpness<sharpness_thres, the current audio frame belongs to the non-zither frame, and the dynamic EQ remains in the stop state.
[0147] Whether the dynamic EQ works can be determined by its internal bypass state, and when the bypass is switched, a smoothing processing mechanism can be set to ensure that the audio frame signals before and after the dynamic EQ are continuous.
[0148] In addition, in the above four cases, the dynamic EQ of each audio frame only processes 512 sample point data, i.e. one frame length. After processing, it can be offset by 512 to enter the next audio frame cycle. The processed audio frame and the copied audio frame are output in time sequence to obtain real-time processed audio data.
[0149] To better implement the zither processing method of the embodiments of the present application, the embodiments of the present application further provide a zither processing device. Please refer to Figure 6 , Figure 6 The structure diagram of the zither processing device provided by the embodiments of the present application. Wherein, the zither processing device 700 can include the following units.
[0150] The acquisition unit 701 is configured to acquire a current audio frame.
[0151] The determination unit 702 is configured to determine a target feature value of the current audio frame.
[0152] The processing unit 703 is configured to, in response to the target feature value of the current audio frame satisfying a preset condition, determine that the current audio frame belongs to a zither frame of a zither, and determine an adjustment parameter of the zither, and perform zither adjustment processing on the current audio frame based on the adjustment parameter of the zither.
[0153] In some embodiments, the processing unit, when used for determining, in response to the target feature value of the current audio frame satisfying a preset condition, that the current audio frame belongs to a sibilance frame of a sibilance, and determining an adjustment parameter of the sibilance, performing sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance, is further used for: determining, in response to a previous frame of the current audio frame being a non-sibilance frame and the target feature value of the current audio frame being greater than or equal to a first threshold, that the current audio frame is a first sibilance frame of the sibilance, and setting the adjustment parameter of the current audio frame as the adjustment parameter of the sibilance; and setting and starting a sibilance filter based on the adjustment parameter of the sibilance, to perform sibilance adjustment processing on the current audio frame.
[0154] In some embodiments, the processing unit, when used for determining, in response to the target feature value of the current audio frame satisfying a preset condition, that the current audio frame belongs to a sibilance frame of a sibilance, and determining an adjustment parameter of the sibilance, performing sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance, is further used for: determining, in response to a previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being greater than or equal to the first threshold, that the current audio frame is a sibilance frame other than the first sibilance frame of the sibilance, maintaining the adjustment parameter of the sibilance as the adjustment parameter of the first sibilance frame of the sibilance; and maintaining an open state of the sibilance filter, to perform sibilance adjustment processing on the current audio frame.
[0155] In some embodiments, the processing unit is further used for determining, in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being less than the first threshold, that the current audio frame is a non-sibilance frame, and stopping the sibilance filter.
[0156] In some embodiments, the processing unit is further used for determining, in response to the previous frame of the current audio frame being the non-sibilance frame and the target feature value of the current audio frame being less than the first threshold, that the current audio frame is a non-sibilance frame, and maintaining a stop state of the sibilance filter.
[0157] In some embodiments, the determining unit, when used for determining the target feature value of the current audio frame, is further used for: obtaining an initial feature value of the current audio frame; and performing smoothing processing on the initial feature value to obtain the target feature value of the current audio frame.
[0158] In some embodiments, the determining unit is further configured to: determine a number M of audio frames for smoothing based on a frame length of the current audio frame; obtain target feature values of M-1 audio frames before the current audio frame; and smooth the initial feature value of the current audio frame based on the target feature values of the M-1 audio frames to obtain the target feature value of the current audio frame.
[0159] In some embodiments, the apparatus further includes a first display unit configured to display, on a display interface, a first image of the target feature value over time based on the target feature value of the current audio frame.
[0160] In some embodiments, the adjustment parameter of the sibilance includes a center frequency at a sibilance energy concentration and an effective frequency band of the sibilance that needs to be subjected to the sibilance adjustment processing, and the apparatus further includes a first parameter determining unit configured to determine a loudness of the current audio frame; determine the center frequency at the sibilance energy concentration of the current audio frame and an effective bandwidth at which the energy of the current audio frame decays to the preset decay ratio based on the loudness; and determine the effective frequency band of the current audio frame based on the effective bandwidth and the center frequency at the sibilance energy concentration.
[0161] In some embodiments, the apparatus further includes a second display unit configured to display, on a display interface, the effective bandwidth of the sibilance and the center frequency at the sibilance energy concentration of the sibilance based on the effective bandwidth of the current audio frame and the center frequency at the sibilance energy concentration of the current audio frame.
[0162] In some embodiments, the adjustment parameter of the sibilance further includes a response time of the sibilance adjustment processing and a maximum decay amount of the sibilance adjustment processing, and the apparatus further includes a second parameter determining unit configured to obtain the response time of the sibilance adjustment processing and a filter-out ratio of the sibilance adjustment processing input by a display interface; and determine the maximum decay amount of the sibilance adjustment processing based on the filter-out ratio and a preset trigger threshold of the sibilance adjustment processing.
[0163] In some embodiments, the determining unit is further configured to: in response to the frame energy of the current audio frame being greater than a second threshold and the zero-crossing rate of the current audio frame being greater than a third threshold, determine the target feature value of the current audio frame.
[0164] In some embodiments, the target feature value includes a sharpness for measuring a degree of sound sharpness.
[0165] In some embodiments, the acquisition unit, when acquiring the current audio frame, is further configured to: acquire the current audio frame via a real-time audio acquisition device. In some embodiments, this application also provides an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the methods implemented in any of the above embodiments via executing the executable instructions.
[0166] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 800 may be... Figure 1 The terminal or server shown. For example... Figure 7 As shown, the electronic device 800 may include: a communication interface 801, a memory 802, a processor 803, and a communication bus 804. The communication interface 801, memory 802, and processor 803 communicate with each other via the communication bus 804. The communication interface 801 is used for data communication between the electronic device 800 and external devices. The memory 802 can be used to store software programs and modules, and the processor 803 runs the software programs and modules stored in the memory 802, such as the software programs for the corresponding operations in the above method embodiments.
[0167] In some embodiments, the processor 803 may invoke software programs and modules stored in the memory 802 to perform the following operations: acquire the current audio frame; determine the target feature value of the current audio frame; in response to the target feature value of the current audio frame satisfying a preset condition, determine that the current audio frame belongs to a sibilant frame, determine the sibilant adjustment parameter, and perform sibilant adjustment processing on the current audio frame based on the sibilant adjustment parameter.
[0168] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods of any of the above embodiments. For brevity, further details are omitted here.
[0169] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods of any of the above embodiments. For brevity, further details are omitted here.
[0170] This application also provides a computer program that includes computer instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the corresponding processes in the methods described in the embodiments of this application. For the sake of brevity, these details will not be elaborated further here.
[0171] It should be understood that the processor of the embodiments of the present application can be an integrated circuit chip with a processing capability of signals. In the implementation process, each step of the method embodiments described above can be completed by the integrated logic circuit of hardware in the processor or the instructions in the form of software. The processor described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware coding processor for execution, or a combination of hardware and software modules in the coding processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the storage, and the processor reads the information in the storage, and combines the hardware to complete the steps of the above method.
[0172] It can be appreciated that the memory in the embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM can be used, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory of the system and method described herein is intended to include, but not be limited to, these and any other suitable types of memory.
[0173] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0175] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the division of the units is only a logical function division, and there can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, or the among different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0176] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0177] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit.
[0178] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, and various program codes that can be stored in the medium.
[0179] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A sibilance processing method, characterized by, The method comprises: obtaining a current audio frame; determining a target feature value of the current audio frame; in response to the target feature value of the current audio frame meeting a preset condition, determining that the current audio frame is a sibilance frame of sibilance and determining an adjustment parameter of the sibilance, the adjustment parameter comprising a center frequency at a sibilance energy concentration of the sibilance, an effective frequency band of the sibilance requiring the sibilance adjustment processing, a response time of the sibilance adjustment processing, and a maximum attenuation amount of the sibilance adjustment processing; performing sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance, comprising: in response to a previous frame of the current audio frame being a non-sibilance frame and the target feature value of the current audio frame being greater than or equal to a first threshold value, determining that the current audio frame is a first sibilance frame of the sibilance, setting the adjustment parameter of the current audio frame as the adjustment parameter of the sibilance, and setting and starting a sibilance filter based on the adjustment parameter of the sibilance to perform sibilance adjustment processing on the current audio frame; in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being greater than or equal to the first threshold value, determining that the current audio frame is a remaining sibilance frame of the sibilance except for the first sibilance frame, keeping the adjustment parameter of the sibilance as the adjustment parameter of the first sibilance frame of the sibilance, and keeping the sibilance filter in a started state to perform sibilance adjustment processing on the current audio frame; in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being less than the first threshold value, determining that the current audio frame is a non-sibilance frame, and stopping the sibilance filter; based on the target feature value of the current audio frame, displaying a first image of the target feature value changing over time on a display interface; based on the effective bandwidth of the current audio frame and the center frequency at the sibilance energy concentration of the current audio frame, displaying the effective bandwidth of the sibilance and the center frequency at the sibilance energy concentration of the sibilance on the display interface.
2. The method of claim 1, wherein, The method further comprises: in response to the previous frame of the current audio frame being the non-sibilance frame and the target feature value of the current audio frame being less than the first threshold value, determining that the current audio frame is a non-sibilance frame, and keeping the sibilance filter in a stopped state.
3. The method of claim 1, wherein, The method further comprises: obtaining an initial feature value of the current audio frame; performing smoothing processing on the initial feature value to obtain the target feature value of the current audio frame.
4. The method of claim 3, wherein, The method further comprises: based on the frame length of the current audio frame, determining a number M of audio frames used for smoothing processing; obtaining target feature values of M-1 audio frames before the current audio frame; performing smoothing processing on the initial feature value of the current audio frame based on the target feature values of the M-1 audio frames to obtain the target feature value of the current audio frame.
5. The method of claim 1, wherein, The method further comprises: determining a loudness of the current audio frame; determine a center frequency at which energy of the current audio frame is concentrated and a bandwidth at which energy of the current audio frame attenuates to a preset attenuation ratio based on the loudness; determine an effective frequency band of the current audio frame based on the center frequency and the bandwidth.
6. The method of claim 1, wherein, The method further comprises: obtain a response time of the sibilance adjustment processing and a filtering ratio of the sibilance adjustment processing input by the display interface; determine a maximum attenuation amount of the sibilance adjustment processing based on the filtering ratio and a preset triggering threshold of the sibilance adjustment processing.
7. The method of claim 1, wherein, The determining the target feature value of the current audio frame comprises: determine the target feature value of the current audio frame in response to a frame energy of the current audio frame being greater than a second threshold and a zero-crossing rate of the current audio frame being greater than a third threshold.
8. The method of claim 1, wherein the target feature value comprises a sharpness used to measure a degree of sound sharpness.
9. The method of claim 1, wherein, The obtaining the current audio frame comprises: collecting the current audio frame by a real-time audio collection device.
10. An apparatus for sibilance processing, the apparatus comprising: comprises: an obtaining unit, configured to obtain a current audio frame; a determining unit, configured to determine a target feature value of the current audio frame; a processing unit, configured to, in response to the target feature value of the current audio frame satisfying a preset condition, determine that the current audio frame is a sibilance frame of sibilance and determine an adjustment parameter of the sibilance, the adjustment parameter comprising a center frequency at which energy of the sibilance is concentrated, an effective frequency band of the sibilance that needs to be processed by the sibilance adjustment processing, a response time of the sibilance adjustment processing, and a maximum attenuation amount of the sibilance adjustment processing; the processing unit is configured to perform sibilance adjustment processing on the current audio frame based on the adjustment parameter of the sibilance, comprising: in response to a previous frame of the current audio frame being a non-sibilance frame and the target feature value of the current audio frame being greater than or equal to a first threshold, determining that the current audio frame is a first sibilance frame of the sibilance, taking the adjustment parameter of the current audio frame as the adjustment parameter of the sibilance, and setting and starting a sibilance filter based on the adjustment parameter of the sibilance to perform sibilance adjustment processing on the current audio frame; in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being greater than or equal to the first threshold, determining that the current audio frame is a remaining sibilance frame of the sibilance except the first sibilance frame, keeping the adjustment parameter of the sibilance as the adjustment parameter of the first sibilance frame of the sibilance, and keeping a starting state of the sibilance filter to perform sibilance adjustment processing on the current audio frame; in response to the previous frame of the current audio frame being the sibilance frame and the target feature value of the current audio frame being less than the first threshold, determining that the current audio frame is a non-sibilance frame, and stopping the sibilance filter; a first display unit, configured to display, based on the target feature value of the current audio frame, a first image of the target feature value changing over time on a display interface; The second display unit is configured to display, on a display interface, the active bandwidth of the sibilance and the center frequency at the sibilance energy concentration of the current audio frame based on the active bandwidth of the current audio frame and the center frequency at the sibilance energy concentration of the current audio frame.
11. The apparatus of claim 10, wherein, The processing unit is further configured to determine that the current audio frame is a non-sibilance frame and maintain a stop state of the sibilance filter in response to a previous frame of the current audio frame being the non-sibilance frame and a target feature value of the current audio frame being less than the first threshold value.
12. The apparatus of claim 10, wherein, The determining unit is further configured to: obtain an initial feature value of the current audio frame; smooth the initial feature value to obtain the target feature value of the current audio frame.
13. The apparatus of claim 12, wherein, The determining unit is further configured to: determine a number M of audio frames for smoothing based on a frame length of the current audio frame; obtain target feature values of M-1 audio frames before the current audio frame; smooth the initial feature value of the current audio frame based on the target feature values of the M-1 audio frames to obtain the target feature value of the current audio frame.
14. The apparatus of claim 10, wherein, The apparatus further includes: a first parameter determining unit configured to determine a loudness of the current audio frame, determine a center frequency at a sibilance energy concentration of the current audio frame and an active bandwidth of the current audio frame when an energy of the current audio frame attenuates to a preset attenuation ratio based on the loudness, and determine an active frequency band of the current audio frame based on the active bandwidth and the center frequency at the sibilance energy concentration.
15. The apparatus of claim 10, wherein, The apparatus further includes: a second parameter determining unit configured to obtain a response time of the sibilance adjustment processing and a filtering ratio of the sibilance adjustment processing input by a display interface, and determine a maximum attenuation amount of the sibilance adjustment processing based on the filtering ratio and a trigger threshold value of the preset sibilance adjustment processing.
16. The apparatus of claim 10, wherein, The determining unit is further configured to determine the target feature value of the current audio frame in response to a frame energy of the current audio frame being greater than a second threshold value and a zero-crossing rate of the current audio frame being greater than a third threshold value.
17. The apparatus of claim 10, wherein, The target feature value includes a sharpness for measuring a degree of sound sharpness.
18. The apparatus of claim 10, wherein, The obtaining unit is further configured to: obtain the current audio frame by using a real-time audio acquisition device.
19. An electronic device, comprising: include: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the method of any one of claims 1-9 by executing the executable instructions.
20. A computer-readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium and is executed by the processor to implement the method of any one of claims 1-9.
21. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-9.
Citation Information
Patent Citations
Tooth tone adjustment method and device, electronic equipment and computer readable storage medium
CN112951266A