Audio signal processing method and device, medium and electronic equipment
By converting the audio signal from the time domain to the frequency domain and performing segmented processing, the gain value at each frequency point is determined, thus solving the problem of frequency perception imbalance in the time domain dynamic range control method and improving the audio signal processing effect and listening consistency.
Patent Information
- Application Number
- CN202410500802.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-10-24
AI Technical Summary
In existing technologies, time-domain dynamic range control methods cannot personalize the processing of audio signals according to the perceived intensity of different frequencies, resulting in some frequency signals being perceived as too low or too high, leading to poor audio signal processing performance.
The audio signal is converted from the time domain to the frequency domain by short-time Fourier transform, divided into multiple signal segments, and further divided into sub-spectral information according to frequency thresholds. The gain value of each frequency point is determined, and finally the signal is restored to the time domain by inverse short-time Fourier transform for precise volume adjustment.
It enables appropriate volume control at each frequency point, improves audio signal processing, ensures consistent loudness across different frequency components, and enhances the listening experience.
Smart Images

Figure CN120835247A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, and in particular, to an audio signal processing method and device, a medium and an electronic device. BACKGROUND
[0002] Dynamic Range Control (DRC) is a signal amplitude adjustment method that can map the dynamic range of an input audio signal to a specified dynamic range. This helps to adjust the dynamic range of the audio signal, making the audio more balanced during playback and avoiding excessively high or low volumes. With the increasing number of audio signal distribution channels, such as radio, television, Internet, etc., different playback devices and environments have certain limitations on the dynamic range of the audio. Dynamic range control can process these audio signals to adapt to different playback environments and device requirements.
[0003] In related technologies, a time-domain dynamic range control method is used to process audio signals, which compresses when the signal exceeds a threshold value and amplifies when the signal is below another threshold value. Correspondingly, the frequency components are processed with equal amplitude. However, the human ear perceives different frequencies of signals with different intensities, which can cause some frequency signals to be perceived as too small or some frequency signals to be perceived as too large, resulting in poor audio signal processing results. SUMMARY
[0004] To overcome the problems in the related art, the present disclosure provides an audio signal processing method, device, medium and electronic device.
[0005] According to a first aspect of an embodiment of the present disclosure, an audio signal processing method is provided, which comprises: performing short-time Fourier transform on a to-be-processed audio signal to obtain frequency spectrum information of a plurality of continuous signal segments; dividing the frequency spectrum information of each signal segment into a plurality of sub-frequency spectrum information according to a preset frequency threshold; determining a gain value corresponding to each frequency point according to the amplitude value corresponding to all frequency points in the sub-frequency spectrum information; obtaining an amplitude value after gain of each frequency point according to the gain value and the amplitude value corresponding to each frequency point; performing short-time inverse Fourier transform on the amplitude value after gain of each frequency point in the plurality of continuous signal segments to obtain an audio output signal.
[0006] Optionally, before the amplitude value after gain of each frequency point is obtained according to the gain value and the amplitude value corresponding to each frequency point, the audio signal processing method further comprises: performing smoothing processing on the gain value corresponding to each frequency point.
[0007] Optionally, the smoothing processing on the gain value corresponding to each frequency point comprises: determining a target smoothing coefficient; obtaining the smoothed gain value corresponding to the frequency point according to the target smoothing coefficient, a historical smoothed gain value and the gain value corresponding to the frequency point, wherein the historical smoothed gain value represents the smoothed gain value corresponding to the frequency point with the same order as the frequency point in a previous signal segment of a signal segment to which the frequency point belongs.
[0008] Optionally, the determining of the target smoothing coefficient comprises: determining the target smoothing coefficient from a first smoothing coefficient and a second smoothing coefficient according to the historical smoothed gain value and the gain value corresponding to the frequency point, wherein the first smoothing coefficient represents a smoothing coefficient corresponding to a starting time, the second smoothing coefficient represents a smoothing coefficient corresponding to a releasing time, and the first smoothing coefficient is smaller than the second smoothing coefficient.
[0009] Optionally, the determining of the target smoothing coefficient from the first smoothing coefficient and the second smoothing coefficient according to the historical smoothed gain value and the gain value corresponding to the frequency point comprises: in a case where the historical smoothed gain value is greater than the gain value corresponding to the frequency point, determining the first smoothing coefficient as the target smoothing coefficient; in a case where the historical smoothed gain value is less than or equal to the gain value corresponding to the frequency point, determining the second smoothing coefficient as the target smoothing coefficient.
[0010] Optionally, the determining of the gain value corresponding to each frequency point according to the amplitude value corresponding to each frequency point in the sub-spectrum information comprises: determining an overall amplitude value according to the amplitude value corresponding to each frequency point in the sub-spectrum information; determining an overall gain value according to the overall amplitude value and a preset frequency domain compression threshold value; determining the gain value corresponding to each frequency point according to the overall gain value and a preset frequency point gain adjustment coefficient.
[0011] Optionally, the determining of the overall gain value according to the overall amplitude value and the preset frequency domain compression threshold value comprises: calculating a ratio of the preset frequency domain compression threshold value to the overall amplitude value; in a case where the ratio of the preset frequency domain compression threshold value to the overall amplitude value is less than 1, determining the ratio of the preset frequency domain compression threshold value to the overall amplitude value as the overall gain value; in a case where the ratio of the preset frequency domain compression threshold value to the overall amplitude value is greater than or equal to 1, determining 1 as the overall gain value.
[0012] Optionally, the determining the gain value corresponding to each frequency point according to the overall gain value and a preset frequency point gain adjustment coefficient comprises: converting the overall gain value into a gain decibel value; determining the gain value corresponding to each frequency point according to the gain decibel value and the preset frequency point gain adjustment coefficient.
[0013] According to a second aspect of the embodiments of the present disclosure, an audio signal processing apparatus is provided, which comprises: a first processing module configured to perform short-time Fourier transform on a to-be-processed audio signal to obtain frequency spectrum information of a plurality of continuous signal segments; a second processing module configured to divide the frequency spectrum information of each signal segment into a plurality of sub-frequency spectrum information according to a preset frequency threshold; a third processing module configured to determine a gain value corresponding to each frequency point according to an amplitude value corresponding to each frequency point in the sub-frequency spectrum information; a fourth processing module configured to obtain an amplitude value after gain processing of each frequency point according to the gain value corresponding to each frequency point and the amplitude value; a fifth processing module configured to perform short-time inverse Fourier transform on the amplitude value after gain processing of each frequency point in the plurality of continuous signal segments to obtain an audio output signal.
[0014] According to a third aspect of the embodiments of the present disclosure, a non-transitory computer readable storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the steps of the audio signal processing method provided in any one of the first aspect of the present disclosure.
[0015] According to a fourth aspect of the embodiments of the present disclosure, an electronic device is provided, which comprises: a memory having a computer program stored thereon; a processor configured to execute the computer program in the memory to implement the steps of the audio signal processing method provided in any one of the first aspect of the present disclosure.
[0016] According to the technical scheme, the to-be-processed audio signal is converted from time domain to frequency domain by short-time Fourier transform, and is divided into multiple signal segments, then the signal segments are divided into multiple sub-spectrum information according to the frequency threshold, the gain value is determined based on each sub-spectrum information for adjustment, and finally the short-time inverse Fourier transform is used to restore to time domain, complete the processing of the audio signal. The short-time Fourier transform can highlight the local characteristics of the audio signal, which is beneficial to the frequency domain signal analysis. The signal segments are further segmented to determine the gain value, so that the determined gain value is more reasonable, and then the volume of each frequency point is properly processed, thereby improving the audio signal processing effect.
[0017] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, and are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation on the present disclosure. In the drawings: Figure 1 is a flowchart of an audio signal processing method according to an exemplary embodiment.
[0019] Figure 2 is a flowchart of another audio signal processing method according to an exemplary embodiment.
[0020] Figure 3 is a flowchart of sub-step S204 according to an exemplary embodiment.
[0021] Figure 4 is a flowchart of sub-step S3 according to an exemplary embodiment.
[0022] Figure 5 is a schematic diagram of an audio signal processing method according to an exemplary embodiment.
[0023] Figure 6 is a flowchart of sub-step S33 according to an exemplary embodiment.
[0024] Figure 7 is a block diagram of an audio signal processing device according to an exemplary embodiment.
[0025] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0026] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely intended for illustration and explanation of the present disclosure and are not intended to limit the present disclosure.
[0027] In the following description, the words "first", "second", etc. are used only to distinguish the described purposes, and cannot be understood as indicating or implying relative importance, nor as indicating or implying order.
[0028] DRC is a signal amplitude adjustment method that can map the dynamic range of an input audio signal to a specified dynamic range. This helps to adjust the dynamic range of the audio signal, making the audio more balanced during playback, avoiding excessively high or low volume. With the increase of audio signal circulation channels such as radio, television, Internet, etc., different playback devices and environments have certain limitations on the dynamic range of audio. Dynamic range control can process these audio signals to adapt to different playback environments and device requirements.
[0029] In the related art, a time-domain dynamic range control method is used to process audio signals, which is compressed when a threshold signal is exceeded and amplified when another threshold signal is below. Correspondingly, the frequency components are processed with equal amplitude. However, the human ear perceives different frequencies of signals differently, which may result in some frequency signals being perceived too small or some frequency signals being perceived too large, leading to poor audio signal processing effect.
[0030] To solve the above technical problems, the short-time Fourier transform is used to convert the to-be-processed audio signal from time domain to frequency domain and divide it into multiple signal segments, then the signal segments are divided into multiple sub-spectrum information according to the frequency threshold, and the gain value is determined based on each sub-spectrum information for adjustment, and finally the short-time inverse Fourier transform is used to restore to time domain, completing the processing of the audio signal. Short-time Fourier transform can highlight the local characteristics of the audio signal, which is beneficial to frequency domain signal analysis. The signal segment is further segmented to determine the gain value, so that the determined gain value is more reasonable, and then the volume of each frequency point is appropriately processed, improving the audio signal processing effect.
[0031] Please refer to Figure 1 , Figure 1 is a flowchart of an audio signal processing method according to an exemplary embodiment. The audio signal processing method can be applied to an electronic device, and can include steps S1-S5.
[0032] Step S1, performing short-time Fourier transform on the to-be-processed audio signal to obtain frequency spectrum information of multiple continuous signal segments.
[0033] The audio signal to be processed is divided into a series of equal-length short-time frames, each of which is a signal segment, and a plurality of consecutive short-time frames constitute a plurality of consecutive signal segments. Then, a window function is applied to the signal segment for windowing processing, and Fourier transform is performed on the windowed signal segment to obtain the frequency spectrum information corresponding to the signal segment. The above processing is performed on each signal segment to obtain the frequency spectrum information of the plurality of consecutive signal segments. The window function can be, but is not limited to, a Hamming window, a Hanning window, a rectangular window, etc. The frequency spectrum information can include a plurality of frequency points and an amplitude value representing the volume size information corresponding to each frequency point.
[0034] In step S2, the frequency spectrum information of each signal segment is divided into a plurality of sub-frequency spectrum information according to a preset frequency threshold.
[0035] In order to achieve accurate processing, the signal can be segmented according to the frequency range characteristics of the main components of the music signal. The preset frequency threshold can be one frequency threshold or a plurality of frequency thresholds, which are not limited in the embodiment.
[0036] When the preset frequency threshold is one frequency threshold, which is a first frequency threshold, the frequency spectrum information of the signal segment is divided into a plurality of sub-frequency spectrum information. It can be understood that the frequency spectrum information of one signal segment is divided into two sub-frequency spectrum information. One sub-frequency spectrum information includes frequency points less than the first frequency threshold and the amplitude values corresponding to the frequency points. The other sub-frequency spectrum information includes frequency points greater than or equal to the first frequency threshold and the amplitude values corresponding to the frequency points.
[0037] When the preset frequency threshold is two frequency thresholds, which are a second frequency threshold and a third frequency threshold, the second frequency threshold is less than the third frequency threshold. The frequency spectrum information of the signal segment is divided into a plurality of sub-frequency spectrum information. It can be understood that the frequency spectrum information of one signal segment is divided into three sub-frequency spectrum information. The first sub-frequency spectrum information includes frequency points less than the second frequency threshold and the amplitude values corresponding to the frequency points. The second sub-frequency spectrum information includes frequency points greater than or equal to the second frequency threshold and less than the third frequency threshold and the amplitude values corresponding to the frequency points. The third sub-frequency spectrum information includes frequency points greater than or equal to the third frequency threshold and the amplitude values corresponding to the frequency points.
[0038] For example, it is generally considered that 300Hz or below is low frequency, 300Hz~6500Hz is medium frequency, and 6500Hz or above is high frequency, so the second frequency threshold can be set as 300Hz, and the third frequency threshold can be set as 6500Hz, and the frequency spectrum information of a signal segment is divided into three sub-frequency spectrum information, the first sub-frequency spectrum information includes the frequency points of low frequency below 300Hz and the corresponding amplitude values, the second sub-frequency spectrum information includes the frequency points of medium frequency in the interval of 300Hz~6500Hz and the corresponding amplitude values, and the third sub-frequency spectrum information includes the frequency points of high frequency above 6500Hz and the corresponding amplitude values.
[0039] The frequency spectrum information of each signal segment is processed in the above manner, and multiple sub-frequency spectrum information of the frequency spectrum information of each signal segment can be obtained.
[0040] According to the amplitude values corresponding to all the frequency points in the sub-frequency spectrum information, the gain value corresponding to each frequency point can be determined.
[0041] According to the amplitude values corresponding to all the frequency points in the sub-frequency spectrum information, the gain value corresponding to each frequency point in the sub-frequency spectrum information can be obtained.
[0042] According to the amplitude values corresponding to all the frequency points in the sub-frequency spectrum information, the gain value corresponding to each frequency point in the sub-frequency spectrum information can be obtained.
[0043] According to the amplitude values corresponding to all the frequency points in the sub-frequency spectrum information, the gain value corresponding to each frequency point in the sub-frequency spectrum information can be obtained.
[0044] According to the amplitude values corresponding to all the frequency points in the sub-frequency spectrum information, the gain value corresponding to each frequency point in the sub-frequency spectrum information can be obtained.
[0045] According to the amplitude values corresponding to all the frequency points in the sub-frequency spectrum information, the gain value corresponding to each frequency point in the sub-frequency spectrum information can be obtained.
[0046] According to the amplitude values corresponding to all the frequency points in the sub-frequency spectrum information, the gain value corresponding to each frequency point in the sub-frequency spectrum information can be obtained.
[0047] The audio signal to be processed is converted from time domain to frequency domain by short-time Fourier transform, and divided into multiple signal segments. Then, the signal segments are divided into multiple sub-spectrum information according to a frequency threshold. The gain value is determined based on each sub-spectrum information for adjustment. Finally, the short-time inverse Fourier transform is used to restore to time domain, and the processing of the audio signal is completed. The short-time Fourier transform can highlight the local characteristics of the audio signal, which is beneficial to the analysis of the frequency domain signal. The gain value is determined by further segmenting the signal segments, so that the determined gain value is more reasonable. Then, the volume of each frequency point is properly processed, and the effect of the audio signal processing is improved.
[0048] In a possible implementation, before step S4, the audio signal processing method further includes: The gain value corresponding to each frequency point is smoothed.
[0049] Referring to Figure 2 FIG. 1 is a flowchart of an audio signal processing method according to an example embodiment. The audio signal processing method can be applied to an electronic device. The audio signal processing method can include steps S201-S206.
[0050] In step S201, short-time Fourier transform is performed on the audio signal to be processed to obtain the spectrum information of multiple continuous signal segments.
[0051] In step S202, the spectrum information of each signal segment is divided into multiple sub-spectrum information according to a preset frequency threshold.
[0052] In step S203, the gain value corresponding to each frequency point is determined according to the amplitude of all frequency points in the sub-spectrum information.
[0053] In step S204, the gain value corresponding to each frequency point is smoothed.
[0054] In step S205, the amplitude of each frequency point after gain is obtained according to the gain value and the amplitude corresponding to each frequency point.
[0055] In step S206, the amplitude of each frequency point after gain in the multiple continuous signal segments is subjected to short-time inverse Fourier transform to obtain an audio output signal.
[0056] By smoothing the gain value corresponding to each frequency point, the transition between the frequency points with the same order in adjacent signal segments can be made natural, so that the final audio output signal is mixed, and the quality of the audio output signal is improved.
[0057] It should be noted that the detailed description of steps S201, S202, S203, S205 and S206 can be referred to steps S1, S2, S3, S4 and S5 respectively, and the present embodiment will not be described here.
[0058] In a possible implementation, referring to Figure 3 , step S204 can include steps S2041 and S2042.
[0059] Step S2041, determining a target smoothing coefficient. The target smoothing coefficient can be customized by the user according to the actual situation, and the target smoothing coefficient can also be determined from a plurality of smoothing coefficients according to the gain value.
[0060] Step S2042, obtaining the smoothed gain value corresponding to the frequency point according to the target smoothing coefficient, the historical smoothing gain value and the gain value corresponding to the frequency point.
[0061] Wherein, the historical smoothing gain value represents the smoothed gain value corresponding to the frequency point with the same order as the frequency point in the previous signal segment of the signal segment to which the frequency point belongs.
[0062] For example, the tenth frequency point in the fifth signal segment is currently processed, and the historical smoothing gain value represents the smoothed gain value corresponding to the tenth frequency point in the fourth signal segment.
[0063] The smoothed gain value corresponding to the frequency point = target smoothing coefficient x historical smoothing gain value + (1-target smoothing coefficient) x gain value corresponding to the frequency point.
[0064] For example, a·Y06+(1-a)·G16=Y16 Wherein, a is the target smoothing coefficient, Y06 is the smoothed gain value corresponding to the sixth frequency point in the 0th signal segment, Y06 is the historical smoothing gain value, G16 represents the gain value corresponding to the sixth frequency point in the first signal segment, and Y16 represents the smoothed gain value corresponding to the sixth frequency point in the first signal segment.
[0065] It should be understood that in the process of processing the smoothed gain value corresponding to the frequency point in the first signal segment, the first signal segment does not have historical smoothing gain value, so the historical smoothing gain value corresponding to the first signal segment can be set to a fixed value, for example, 1.
[0066] a·Y16+(1-a)·G26=Y26 Wherein, a is a target smoothing coefficient, Y16 is a smoothed gain value corresponding to the sixth frequency point in the first signal segment, Y16 is a historical smoothed gain value, G26 represents a gain value corresponding to the sixth frequency point in the second signal segment, and Y26 represents a smoothed gain value corresponding to the sixth frequency point in the second signal segment.
[0067] a·Y26+ (1-a)·G36=Y36 Wherein, a is a target smoothing coefficient, Y26 is a smoothed gain value corresponding to the sixth frequency point in the second signal segment, Y26 is a historical smoothed gain value, G36 represents a gain value corresponding to the sixth frequency point in the third signal segment, and Y36 represents a smoothed gain value corresponding to the sixth frequency point in the third signal segment.
[0068] In a possible implementation, determining the target smoothing coefficient can include: determining the target smoothing coefficient from the first smoothing coefficient and the second smoothing coefficient according to the historical smoothed gain value and the gain value corresponding to the frequency point; in a case where the historical smoothed gain value is greater than the gain value corresponding to the frequency point, determining the first smoothing coefficient as the target smoothing coefficient; in a case where the historical smoothed gain value is less than or equal to the gain value corresponding to the frequency point, determining the second smoothing coefficient as the target smoothing coefficient.
[0069] Wherein, the first smoothing coefficient represents a smoothing coefficient corresponding to a start-up time, and the second smoothing coefficient represents a smoothing coefficient corresponding to a release time, and the first smoothing coefficient is less than the second smoothing coefficient.
[0070] The calculation formula of the first smoothing coefficient is: at1 = exp(-1·fs / atk1) Wherein, at1 represents the first smoothing coefficient, fs represents a sampling rate of the audio signal, and atk1 represents the start-up time, which can be set according to actual requirements, and the set range can be, for example, 10-40 ms.
[0071] The calculation formula of the second smoothing coefficient is: at2 = exp(-1·fs / atk2) Wherein, at2 represents the second smoothing coefficient, fs represents a sampling rate of the audio signal, and atk2 represents the release time, which can be set according to actual requirements, and the set range can be, for example, 0.5-1 s.
[0072] By defining independent start-up time and release time for each frequency point, more precise control can be achieved, and then a smoothed gain value is obtained through smoothing transition in units of frames.
[0073] According to the equal loudness curve theory, the human ear has different perceptions of different frequencies, for example, the lower the frequency, the louder the sound needs to be to be as loud as a higher frequency sound. If the same gain compression is used for all frequency points of the low frequency, then after compression, the frequency points that are lower in frequency will have a smaller perceived loudness than the frequency points that are higher in frequency. In order to solve this problem, the gain of the frequency points that are lower in frequency needs to be increased a little, and vice versa, the gain of the frequency points that are higher in frequency needs to be decreased a little.
[0074] In one possible implementation, referring to Figure 4 and Figure 5 , step S3 can include steps S31-S33.
[0075] In step S31, the overall amplitude is determined according to the amplitudes corresponding to all frequency points in the sub-spectrum information.
[0076] The square sum of the amplitudes corresponding to all frequency points in the sub-spectrum information is calculated, and the overall amplitude of the sub-spectrum information is obtained by taking the square root.
[0077] For example, the amplitudes of all frequency points (for example, M) in the sub-spectrum information are S0, S1, …, SM, respectively. M-1 The calculation formula of the overall amplitude is: E=
[0078] Wherein, E represents the overall amplitude.
[0079] The above processing is performed on each sub-spectrum information, and the overall amplitude corresponding to each sub-spectrum information is obtained. For example, three sub-spectrum information corresponding to one signal end can obtain three overall amplitudes, which are the overall amplitude corresponding to the low frequency, the overall amplitude corresponding to the medium frequency, and the overall amplitude corresponding to the high frequency.
[0080] In step S32, the overall gain value is determined according to the overall amplitude and the preset frequency domain compression threshold.
[0081] Determining the overall gain value according to the overall amplitude and the preset frequency domain compression threshold can be understood as: calculating the ratio of the preset frequency domain compression threshold to the overall amplitude; in the case that the ratio of the preset frequency domain compression threshold to the overall amplitude is less than 1, the ratio of the preset frequency domain compression threshold to the overall amplitude is determined as the overall gain value; in the case that the ratio of the preset frequency domain compression threshold to the overall amplitude is greater than or equal to 1, 1 is determined as the overall gain value.
[0082] For example, the calculation formula of the overall gain value is: G = min(1,T / E) Wherein, G represents the overall gain value, T represents the preset frequency domain compression threshold, and E represents the overall amplitude.
[0083] The above processing is performed on each sub-spectrum information, and the overall gain value corresponding to each sub-spectrum information is obtained. For example, three sub-spectrum information corresponding to one signal end can obtain three overall gain values, which are the overall gain value corresponding to the low frequency, the overall gain value corresponding to the medium frequency, and the overall gain value corresponding to the high frequency.
[0084] In step S33, the gain value corresponding to each frequency point is determined according to the overall gain value and the preset frequency point gain adjustment coefficient.
[0085] In a possible implementation, the preset frequency point gain adjustment coefficient can be a decreasing array, for example, C0, C1, …, C M-1 .
[0086] The calculation formula of the gain value corresponding to the frequency point is: G i = G·C i Wherein, G i characterizes the gain value corresponding to the i-th frequency point, G characterizes the overall gain value, and C i characterizes the i-th value in the preset frequency point gain adjustment coefficient.
[0087] In a possible implementation, referring to Figure 6 , step S33 can also include step S331 and step S332.
[0088] In step S331, the overall gain value is converted into a gain decibel value.
[0089] The calculation formula of the gain decibel value is: G db = 20·log 10 (G) Wherein, G db characterizes the gain decibel value, and G characterizes the overall gain value.
[0090] In step S332, the gain value corresponding to each frequency point is determined according to the gain decibel value and the preset frequency point gain adjustment coefficient.
[0091] The preset frequency point gain adjustment coefficient composed of the increasing array is determined through the self-defined db domain attenuation coefficient curve, and the lower the frequency, the smaller the attenuation coefficient, so that the sound of lower frequency is compressed a little. For example, the preset frequency point gain adjustment coefficient is C db0 , C db1 , …, C dbM-1 , such as 0.8, 0.81, …, 0.8+(M-1)×0.01, The calculation formula of the gain value corresponding to the frequency point is: G i = pow 10 (G db ·C dbi / 20) Among them, G i Represents the gain value corresponding to the i-th frequency point, G db Characterize the gain decibel value, C dbi Represents the i-th value in the preset frequency point gain adjustment coefficient.
[0092] By performing the above processing on each sub-spectrum information, the gain value corresponding to each frequency point in each sub-spectrum information can be obtained, that is, the frequency point gain value. For example, the three sub-spectrum information corresponding to a signal end can obtain the frequency point gain value corresponding to the low frequency, the frequency point gain value corresponding to the mid-frequency, and the frequency point gain value corresponding to the high frequency.
[0093] This method precisely controls the amplitude of each frequency point and even aligns the loudness of the processed results, so that the perceived loudness of sounds with different frequency components is consistent, achieving a better listening experience.
[0094] To implement the above method embodiment, this embodiment provides an audio signal processing device, such as Figure 7 As shown, Figure 7 FIG1 is a block diagram of an audio signal processing apparatus according to an exemplary embodiment. The audio signal processing apparatus 500 can be applied to an electronic device, and the audio signal processing apparatus 500 may include: The first processing module 501 is configured to perform a short-time Fourier transform on the audio signal to be processed to obtain spectrum information of multiple continuous signal segments; The second processing module 502 is configured to divide the spectrum information of each signal segment into a plurality of sub-spectrum information according to a preset frequency threshold; The third processing module 503 is configured to determine a gain value corresponding to each frequency point according to the amplitude values corresponding to all frequency points in the sub-spectrum information; The fourth processing module 504 is configured to obtain the amplitude of each frequency point after gain according to the gain value and amplitude corresponding to each frequency point; The fifth processing module 505 is configured to perform an inverse short-time Fourier transform on the amplitude of each frequency point after gain in the plurality of continuous signal segments to obtain an audio output signal.
[0095] Optionally, the audio signal processing device 500 further includes: The sixth processing module is configured to perform smoothing processing on the gain value corresponding to each frequency point.
[0096] Optionally, the sixth processing module comprises: a first sub-processing module configured to determine a target smoothing coefficient; a second sub-processing module configured to obtain a smoothed gain value corresponding to a frequency point according to the target smoothing coefficient, a historical smoothed gain value, and a gain value corresponding to the frequency point, wherein the historical smoothed gain value represents a smoothed gain value corresponding to a frequency point of the same order as the frequency point in a previous signal segment of a signal segment to which the frequency point belongs.
[0097] Optionally, the first sub-processing module is specifically configured to: determine the target smoothing coefficient from a first smoothing coefficient and a second smoothing coefficient according to the historical smoothed gain value and the gain value corresponding to the frequency point, wherein the first smoothing coefficient represents a smoothing coefficient corresponding to a start-up time, the second smoothing coefficient represents a smoothing coefficient corresponding to a release time, and the first smoothing coefficient is smaller than the second smoothing coefficient.
[0098] Optionally, the first sub-processing module is specifically configured to: determine the first smoothing coefficient as the target smoothing coefficient in a case where the historical smoothed gain value is greater than the gain value corresponding to the frequency point; determine the second smoothing coefficient as the target smoothing coefficient in a case where the historical smoothed gain value is less than or equal to the gain value corresponding to the frequency point.
[0099] Optionally, the third processing module comprises: a third sub-processing module configured to determine an overall amplitude value according to amplitudes corresponding to all frequency points in the sub-spectrum information; a fourth sub-processing module configured to determine an overall gain value according to the overall amplitude value and a preset frequency domain compression threshold value; a fifth sub-processing module configured to determine a gain value corresponding to each frequency point according to the overall gain value and a preset frequency point gain adjustment coefficient.
[0100] Optionally, the fourth sub-processing module is specifically configured to: calculate a ratio of the preset frequency domain compression threshold value to the overall amplitude value; determine the ratio of the preset frequency domain compression threshold value to the overall amplitude value as the overall gain value in a case where the ratio of the preset frequency domain compression threshold value to the overall amplitude value is less than 1; determine 1 as the overall gain value in a case where the ratio of the preset frequency domain compression threshold value to the overall amplitude value is greater than or equal to 1.
[0101] Optionally, the fifth sub-processing module is specifically configured to: perform decibel value conversion on the overall gain value to obtain a gain decibel value; According to the gain decibel value and a preset frequency point gain adjustment coefficient, a gain value corresponding to each frequency point is determined.
[0102] As to the audio signal processing apparatus in the above-mentioned embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the audio signal processing method, and thus will not be described in detail here.
[0103] Figure 8 is a block diagram of an electronic device 700 according to an example embodiment. As shown, the electronic device 700 can include a processor 701, a memory 702. The electronic device 700 can also include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705. Figure 8
[0104] The processor 701 is configured to control overall operations of the electronic device 700 to complete all or part of the steps of the above-described audio signal processing method. The memory 702 is configured to store various types of data to support operations of the electronic device 700, which can include, for example, instructions for operating any application or method on the electronic device 700, and application-related data, such as contact data, transmitted and received messages, pictures, audio, video, and the like. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk, or an optical disk. The multimedia component 703 can include a screen and an audio component. The screen can be, for example, a touch screen, and the audio component is configured to output and / or input audio signals. For example, the audio component can include a microphone configured to receive external audio signals. The received audio signals can be further stored in the memory 702 or transmitted through the communication component 705. The audio component further includes at least one speaker configured to output audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules, which can be a keyboard, a mouse, a button, and the like. The buttons can be virtual buttons or physical buttons. The communication component 705 is configured to perform wired or wireless communication between the electronic device 700 and other devices. The wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, and the like, or a combination of one or more of them, is not limited herein. Therefore, the communication component 705 can include, for example, a Wi-Fi module, a Bluetooth module, an NFC module, and the like.
[0105] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements for performing the above-described audio signal processing method.
[0106] In another exemplary embodiment, a non-transitory computer-readable storage medium including program instructions that, when executed by a processor, implement the steps of the above-described audio signal processing method is also provided. For example, the computer-readable storage medium can be the above-described memory 702 including program instructions that are executable by the processor 701 of the electronic device 700 to complete the above-described audio signal processing method.
[0107] In another exemplary embodiment, a computer program product is also provided, which contains a computer program executable by a programmable apparatus, the computer program having code portions for performing the above-described audio signal processing method when executed by the programmable apparatus.
[0108] The preferred embodiments of the present disclosure are described in detail above with reference to the accompanying drawings, but the present disclosure is not limited to the specific details of the above-described embodiments. Various simple modifications can be made to the technical solutions of the present disclosure within the scope of the technical concept of the present disclosure, and these simple modifications all belong to the protection scope of the present disclosure.
[0109] In addition, it should be noted that each of the specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, various possible combinations are not described again by the present disclosure.
[0110] Furthermore, any combination of the various different embodiments of the present disclosure can also be made, as long as it does not deviate from the idea of the present disclosure, it should also be considered as disclosed by the present disclosure.
Claims
1. A method of audio signal processing, characterized by, The audio signal processing method comprises: performing short-time Fourier transform on a to-be-processed audio signal to obtain spectral information of a plurality of continuous signal segments; dividing the spectral information of each signal segment into a plurality of sub-spectral information according to a preset frequency threshold; determining a gain value corresponding to each frequency point according to an amplitude value corresponding to all frequency points in the sub-spectral information; obtaining an amplitude value after gain of each frequency point according to the gain value corresponding to each frequency point and the amplitude value; performing short-time inverse Fourier transform on the amplitude value after gain of each frequency point in the plurality of continuous signal segments to obtain an audio output signal.
2. The audio signal processing method of claim 1, wherein, Before the step of obtaining the amplitude value after gain of each frequency point according to the gain value corresponding to each frequency point and the amplitude value, the audio signal processing method further comprises: performing smoothing processing on the gain value corresponding to each frequency point.
3. The audio signal processing method of claim 2, wherein, The step of performing smoothing processing on the gain value corresponding to each frequency point comprises: determining a target smoothing coefficient; obtaining a smoothed gain value corresponding to a frequency point according to the target smoothing coefficient, a historical smoothed gain value and the gain value corresponding to the frequency point, wherein the historical smoothed gain value represents a smoothed gain value corresponding to a frequency point of the same order as the frequency point in a previous signal segment of a signal segment to which the frequency point belongs.
4. The audio signal processing method of claim 3, wherein, The step of determining the target smoothing coefficient comprises: determining the target smoothing coefficient from a first smoothing coefficient and a second smoothing coefficient according to the historical smoothed gain value and the gain value corresponding to the frequency point, wherein the first smoothing coefficient represents a smoothing coefficient corresponding to a start time, the second smoothing coefficient represents a smoothing coefficient corresponding to a release time, and the first smoothing coefficient is smaller than the second smoothing coefficient.
5. The audio signal processing method of claim 4, wherein, The step of determining the target smoothing coefficient from the first smoothing coefficient and the second smoothing coefficient according to the historical smoothed gain value and the gain value corresponding to the frequency point comprises: in a case where the historical smoothed gain value is greater than the gain value corresponding to the frequency point, determining the first smoothing coefficient as the target smoothing coefficient; in a case where the historical smoothed gain value is less than or equal to the gain value corresponding to the frequency point, determining the second smoothing coefficient as the target smoothing coefficient.
6. The audio signal processing method of claim 1, wherein, The step of determining the gain value corresponding to each frequency point according to the amplitude value corresponding to all frequency points in the sub-spectral information comprises: determining an overall amplitude value according to the amplitude value corresponding to all frequency points in the sub-spectral information; determining an overall gain value according to the overall amplitude value and a preset frequency domain compression threshold value; determining the gain value corresponding to each frequency point according to the overall gain value and a preset frequency point gain adjustment coefficient.
7. The audio signal processing method of claim 6, wherein, The step of determining the overall gain value according to the overall amplitude value and the preset frequency domain compression threshold value comprises: calculating a ratio of the preset frequency domain compression threshold value to the overall amplitude value; in a case where the ratio of the preset frequency domain compression threshold value to the overall amplitude value is less than 1, determining the ratio of the preset frequency domain compression threshold value to the overall amplitude value as the overall gain value; in a case where the ratio of the preset frequency domain compression threshold value to the overall amplitude value is greater than or equal to 1, determining 1 as the overall gain value.
8. The audio signal processing method of claim 6, wherein, The method comprises: The overall gain value is converted into a gain decibel value; The gain decibel value and a preset frequency point gain adjustment coefficient are used to determine a gain value corresponding to each frequency point.
9. An audio signal processing apparatus, characterized by comprising: The audio signal processing device comprises: A first processing module configured to perform short-time Fourier transform on a to-be-processed audio signal to obtain frequency spectrum information of a plurality of continuous signal segments; A second processing module configured to divide the frequency spectrum information of each signal segment into a plurality of sub-frequency spectrum information according to a preset frequency threshold; A third processing module configured to determine a gain value corresponding to each frequency point according to an amplitude value corresponding to each frequency point in the sub-frequency spectrum information; A fourth processing module configured to obtain an amplitude value after gain processing of each frequency point according to the gain value corresponding to each frequency point and the amplitude value; A fifth processing module configured to perform short-time inverse Fourier transform on the amplitude value after gain processing of each frequency point in the plurality of continuous signal segments to obtain an audio output signal.
10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the audio signal processing method in any one of claims 1-8.
11. An electronic device, comprising: The program is executed by the processor to implement the steps of the audio signal processing method in any one of claims 1-8. The program is executed by the processor to implement the steps of the audio signal processing method in any one of claims 1-8.