Audio resampling method and device based on frequency domain processing
Through the method based on frequency domain processing, the sampling rate conversion of the audio signal and windowing it in the frequency domain is solved, and the spectral aliasing and noise problems in the prior art are achieved, and high-quality resampling is achieved.
Patent Information
- Application Number
- CN202510103714.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
Existing resampling techniques have challenges in preventing the introduction of noise from spectral aliasing and interpolation, resulting in a degradation in signal quality.
The frequency domain processing method is adopted to convert the sampling rate through the time domain to the frequency domain conversion, and windowing is performed in the frequency domain to achieve a smooth transition, thereby eliminating the time domain signal ringing phenomenon caused by spectrum truncation.
High-quality resampling is achieved, spectrum aliasing and noise introduction are avoided, and signal quality is improved.
Smart Images

Figure CN120048296A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio signal processing, and in particular, to an audio resampling method and apparatus based on frequency domain processing. Background Art
[0002] Audio resample is a key concept in the field of digital signal processing, which involves converting a digital audio signal from one sampling rate to another. Its purposes are as follows:
[0003] Format conversion: When a media file is converted from one format to another, resampling is usually involved to ensure that it meets the standards and requirements of the new format.
[0004] Device compatibility: Different playback devices or media platforms may have specific requirements for the sampling rate of audio. Resampling ensures that media files are compatible with different devices.
[0005] Network transmission: In order to adapt to network bandwidth limitations, the sampling rate is usually reduced to reduce the file size for faster transmission.
[0006] Mixing: When multiple audio signals are mixed, if the sampling rates of the input audio signals are inconsistent, it is necessary to resample the input signals to a unified sampling rate for mixing.
[0007] Figure 1 An existing frequency domain-based signal decomposition and reconstruction technology is given, including the following steps: By overlapping and framing an audio signal, then windowing, and converting the time-domain signal to the frequency domain through a fast Fourier transform (FFT) to obtain a spectrum, and then performing an inverse fast Fourier transform with normalization (normalized IFFT) on the spectrum to obtain the time-domain signal of each frame. Then, through windowing (optional), and finally overlapping and adding each frame, the time-domain signal can be perfectly reconstructed. This framework only involves the decomposition from the time domain to the frequency domain and the signal reconstruction from the frequency domain to the time domain, and the input and output have the same sampling rate, without involving sampling rate conversion.
[0008] Existing resampling technologies are usually carried out in the time domain, and the problems involved are as follows:
[0009] (1) Anti-aliasing in downsampling: When reducing the sampling rate, a low-pass filter is usually required to eliminate components higher than the Nyquist frequency corresponding to the new sampling rate to prevent aliasing. However, there are two problems with anti-aliasing using a low-pass filter: A. When the transition band is narrow, it is difficult to filter out components higher than the Nyquist frequency cleanly, resulting in partial aliasing. B. To prevent aliasing, increasing the transition band of the low-pass filter sacrifices some bandwidth, resulting in partial distortion.
[0010] (2) Upsampling interpolation: When increasing the sampling rate, new sample points need to be inserted between samples, usually accomplished through various interpolation techniques such as linear interpolation, polynomial interpolation, or more complex interpolation methods. This process introduces additional noise or distortion to the original signal. Summary of the Invention
[0011] Aiming at the problem that current resampling techniques are difficult to prevent spectral aliasing and interpolation introduces noise, resulting in a decline in signal quality, the present invention proposes a new method and device, adopting a new frequency-domain processing-based technology to perform high-quality and anti-aliasing resampling on audio signals with common sampling rates. The technical solution is as follows:
[0012] An audio resampling method based on frequency-domain processing, comprising the following steps:
[0013] Time-domain to frequency-domain step: Overlap and frame the input signal with the first sampling rate, apply a time-domain window, and convert it from the time domain to the frequency domain through FFT to obtain the first spectrum at the first sampling rate;
[0014] Sampling rate conversion step: Normalize the first spectrum with the number of sample points of the input signal for FFT as a parameter to obtain the normalized first spectrum; Under the condition that the spectral resolutions of the second spectrum and the first spectrum are the same, construct the second spectrum at the second sampling rate according to the first spectrum at the first sampling rate; Select a window function with a smooth transition function to window the second spectrum;
[0015] Frequency-domain to time-domain step: Convert the second spectrum at the second sampling rate from the frequency domain to the time domain through IFFT, and then overlap and add to obtain the time-domain value corresponding to the second sampling rate.
[0016] Further, the expression of the spectral resolution is: SR / N, where SR is the sampling rate in Hz, and N represents the number of input sample points for FFT or IFFT.
[0017] Further, the method for constructing the second spectrum at the second sampling rate according to the first spectrum at the first sampling rate is:
[0018] When the second sampling rate is greater than the first sampling rate and the spectral range of the second spectrum is more than that of the first spectrum, upsampling is adopted: For the spectral overlapping region of the second spectrum and the first spectrum, directly copy the spectrum at the first sampling rate to the spectrum at the second sampling rate, and assign 0 to the extra spectrum at the second sampling rate;
[0019] When the second sampling rate is less than the first sampling rate and the spectral range of the second spectrum is less than that of the first spectrum, downsampling is adopted: For the spectral overlapping region of the second spectrum and the first spectrum, directly copy the spectrum at the first sampling rate to the spectrum at the second sampling rate.
[0020] Further, for the smoothing transition window, at several points at the left and right ends of the rectangular window at the mutation position, a smoothing transition function is used for the smoothing transition from 0 to 1 and from 1 to 0, and the other positions remain 1 unchanged.
[0021] Further, the smoothing transition function includes a sigmoid function.
[0022] Further, the expression of the sigmoid function is:
[0023]
[0024] The value-taking method is: x takes several points in the range of -10 to 10, and S(x) of each point is obtained.
[0025] Further, the frame length of the overlapping frame segmentation is 10 ms; the applicable first sampling rates include 8 kHz, 16 kHz, 24 kHz, 32 kHz, 44.1 kHz, 48 kHz, 96 kHz, 192 kHz; the applicable second sampling rates include 8 kHz, 16 kHz, 24 kHz, 32 kHz, 44.1 kHz, 48 kHz, 96 kHz, 192 kHz.
[0026] Further, the frame length of the overlapping frame segmentation is 8 ms; the applicable first sampling rates include 8 kHz, 16 kHz, 24 kHz, 32 kHz, 48 kHz, 96 kHz, 192 kHz; the applicable second sampling rates include 8 kHz, 16 kHz, 24 kHz, 32 kHz, 48 kHz, 96 kHz, 192 kHz.
[0027] An audio resampling device based on frequency domain processing, comprising:
[0028] A time domain to frequency domain unit, configured to perform overlapping frame segmentation on an audio signal with a first sampling rate, then window it, and convert it from the time domain to the frequency domain through FFT to obtain a first spectrum at the first sampling rate;
[0029] A sampling rate conversion unit, configured to normalize the first spectrum with the number of samples of the input signal for FFT as a parameter to obtain a normalized first spectrum; under the condition that the spectral resolutions of the second spectrum and the first spectrum are the same, construct a second spectrum at a second sampling rate according to the first spectrum at the first sampling rate; select a window function with a smoothing transition function to window the second spectrum;
[0030] A frequency domain to time domain unit, configured to convert the second spectrum at the second sampling rate from the frequency domain to the time domain through IFFT, and then perform overlapping addition to obtain the time domain value corresponding to the second sampling rate.
[0031] Further, it further includes a frequency-domain signal processing unit, which is configured to perform frequency-domain processing on the second spectrum at the second sampling rate output by at least one of the sampling rate conversion units, and then output it to the frequency-domain to time-domain unit at the second sampling rate.
[0032] The present invention has the following beneficial effects:
[0033] 1. High-quality resampling can be achieved, which is suitable for scenarios where the sampling rate conversion requirements of audio signals are needed.
[0034] 2. This resampling method can be combined with other algorithms that need to be processed in the frequency domain to achieve the purpose of one-step resampling and other functions, thereby saving computing power. Description of the Drawings
[0035] Figure 1 is a flowchart of the existing signal decomposition and reconstruction processing based on the frequency domain;
[0036] Figure 2 is a basic flowchart of the resampling process of the present invention;
[0037] Figure 3 is a schematic diagram of upsampling;
[0038] Figure 4 is a schematic diagram of downsampling;
[0039] Figure 5 is a schematic diagram of windowing under upsampling;
[0040] Figure 6 is a schematic diagram of windowing under downsampling;
[0041] Figure 7 is a curve diagram of the sigmoid function;
[0042] Figure 8 is a flowchart of the combination of the resampling process and the echo cancellation process of the present invention;
[0043] Figure 9 is a flowchart of the combination of the resampling process and the noise reduction process of the present invention. Detailed Embodiments
[0044] To further illustrate the embodiments, the present invention provides drawings. These drawings are a part of the disclosure of the present invention, which are mainly used to illustrate the embodiments and can be combined with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible embodiments and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are usually used to represent similar components.
[0045] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0046] As Figure 2 shown, within the framework of the existing signal decomposition and reconstruction method based on frequency-domain processing, the present invention proposes a resampling method based on frequency-domain processing involving sampling rate conversion, including the following steps:
[0047] First, assume that the input sampling rate of the input signal is SR1, and the output sampling rate of the output signal is SR2, where SR1 is not equal to SR2.
[0048] (1) Convert the input signal from the time domain to the frequency domain: Through the existing resampling method based on frequency-domain processing, obtain the spectrum spec_1 at the SR1 sampling rate by overlapping framing, windowing in the time domain, and FFT.
[0049] (2) Normalize the spectrum spec_1: spec_1 = spec_1 / N (N is the number of samples for FFT of the input time-domain data)
[0050] (3) Construct the spectrum at the target sampling rate SR2: Construct the spectrum spec_2 at the SR2 sampling rate based on the spectrum spec_1 at the SR1 sampling rate.
[0051] It should be noted in step (3) that:
[0052] ① It is necessary to ensure that the spectral resolution of the spectrum spec_1 at the SR1 sampling rate is the same as that of the spectrum spec_2 at the SR2 sampling rate;
[0053] ② Find the region where the spectral positions in the spectrum spec_2 overlap with those in the spectrum spec_1, and assign the values of the spectrum spec_1 to the corresponding positions in the spectrum spec_2; if the spectrum of spec_2 is more than that of spec_1, assign 0 to the extra spectrum.
[0054] More specifically, the audio processing also includes two cases of upsampling and downsampling. They are described as follows.
[0055] Upsampling:
[0056] Refer to Figure 3 , SR2 > SR1, the spectrum at the SR2 sampling rate is more than that at the SR1 sampling rate. For the overlapping spectral region, directly copy the spectrum at the SR1 sampling rate to the spectrum at the SR2 sampling rate, and assign 0 to the extra spectrum of SR2.
[0057] Downsampling:
[0058] Refer to Figure 4, SR2 < SR1. The spectrum at the SR2 sampling rate has fewer components than that at the SR1 sampling rate. For the overlapping region of the spectra, directly copy the spectrum at the SR1 sampling rate to the spectrum at the SR2 sampling rate.
[0059] (4) Spectrum windowing. The mapping from spec_1 to spec_2, for the overlapping part of the frequencies, is equivalent to truncating with a rectangular window. After performing the inverse Fourier transform to restore to the time-domain signal, the spectrum truncation will exhibit ringing in the time domain. Therefore, window the spectrum of spec_2 to achieve a smooth transition of the spectrum and eliminate the ringing phenomenon of the time-domain signal caused by spectrum truncation. Preferably, a sigmoid function can be selected for the smooth transition (from 0 to 1 and from 1 to 0), but it is not limited to the sigmoid function, and other windows for smooth transition can also be selected.
[0060] spec_2 = spec_2 * window
[0061] At both ends of the window function window, at several points in the mutation position, use the sigmod function for smooth transition from 0 to 1 and from 1 to 0. Keep the other positions as 1.
[0062] Through the smooth transition, the ringing effect generated when transforming back from the frequency domain to the time domain by the rectangular window can be eliminated.
[0063] It should be noted that the window functions in time-domain windowing include but are not limited to the Hamming window function; in frequency-domain windowing, the window functions that can be used for smooth transition include but are not limited to the sigmod function.
[0064] Specifically, step (4) also needs to be processed separately for upsampling and downsampling:
[0065] Upsampling:
[0066] Fill the part of spec_2 that is higher than the frequency of spec_1 with 0, and directly copy the value of spec_1 for the overlapping part with spec_1, which is equivalent to Figure 4 the rectangular window shown by the dashed line.
[0067] In the rectangular window from 0 to 1 and from 1 to 0, there is no smooth transition, and ringing will occur when transforming back from the frequency domain to the time domain. For this, change the places from 0 to 1 and from 1 to 0 to smooth transition, as Figure 5 the window function shown by the blue curve in.
[0068] Downsampling:
[0069] Directly copy the value of spec_1 for the overlapping part of spec_2 with spec_1, which is equivalent to as Figure 6 the rectangular window shown by the red dashed line in.
[0070] Similarly, when the rectangular window transitions from 0 to 1 and from 1 to 0, it is not a smooth transition, and a ringing effect will occur when transforming from the frequency domain to the time domain. Therefore, the transitions from 0 to 1 and from 1 to 0 are changed to smooth transitions, such as Figure 6 the window function represented by the blue curve in
[0071] In this embodiment, the window function adopts the sigmoid function, as shown in Figure 7 . The expression of the sigmoid function is:
[0072]
[0073] The specific value-taking method: x takes several points in the range of -10 to 10, and the value of S(x) is obtained as the weight coefficient to achieve smooth transitions from 0 to 1 and from 1 to 0.
[0074] (5) Then, perform frequency-domain to time-domain processing through IFFT. Input spec_2, and the corresponding time-domain values at the SR2 sampling rate are obtained through processing. Then, through windowing (optional) and overlap-and-add, the output signal is obtained, thereby achieving resampling from SR1 to SR2.
[0075] It should be noted that:
[0076] ① The core process of this method is: based on the spectrum spec_1 of the data with a sampling rate of SR1, reconstruct the spectrum spec_2 with a sampling rate of SR2, and then perform the transformation from the frequency domain to the time domain through spec_2 to obtain the time-domain data with a sampling rate of SR2, achieving resampling from the input sampling rate SR1 to the output sampling rate SR2.
[0077] ② The spectrum spec_1 with a sampling rate of SR1 and the spectrum spec_2 with a sampling rate of SR2 must have the same spectral resolution.
[0078] ③ After constructing spec_2 through spec_1, it is necessary to perform frequency-domain windowing on spec_2 to achieve smooth transitions of the spectrum, thereby eliminating the ringing phenomenon in the time domain.
[0079] In this embodiment, when data is transformed to the frequency domain at different sampling rates, it has the same spectral resolution. The following is an example:
[0080] In the case of a given frame duration t (unit: ms) and 50% overlapping frame division, that is, the frame shift time is equal to the frame duration, both are t, and the sampling rate is SR (unit: Hz), the number of sample points of the frame shift t and the number of sample points corresponding to the frame length t are both equal to t * SR / 1000. This number of sample points must be an integer. At this time, the corresponding spectral resolutions of data at different sampling rates are equal. Using this feature, we can mutually convert the sampling rates that meet this condition.
[0081] In terms of the sampling rate of audio, common ones are 8 kHz, 16 kHz, 24 kHz, 32 kHz, 44.1 kHz, 48 kHz, 96 kHz, and 192 kHz.
[0082] Calculated with a frame shift time t = 10 ms and 50% overlapping frame division, it is compatible with various sampling rates in the following table. Using other frame shift times (in ms) is not compatible with 44.1 kHz. Except for 44.1 kHz, other frequencies can be compatible and mutually converted.
[0083] It should be noted that the spectral resolution (in Hz) = SR / N
[0084] Among them, SR: sampling rate; N: represents the number of sample points input when performing the fast Fourier transform FFT (or inverse fast Fourier transform IFFT).
[0085] Ensure that the spectral resolutions are equal, that is: SR1 / N1 = SR2 / N2
[0086] Among them, SR1 and SR2 represent the input sampling rate and the output sampling rate respectively
[0087] N1 and N2 represent the number of sample points input for the fast Fourier transform of the input signal and the number of sample points for the inverse fast Fourier transform of the output signal respectively.
[0088] It should be noted that the following method can be used to ensure that data transformed into the frequency domain at different sampling rates has the same spectral resolution:
[0089] When the given frame length time t (in ms) and 50% overlapping frame division, that is, the frame shift time and the frame length time are equal, both are t, and the sampling rate is SR (in Hz), the number of sample points of the frame shift t and the number of sample points corresponding to the frame length t are both equal to t*SR / 1000. This number of sample points must be an integer. At this time, the spectral resolutions of data with different sampling rates are equal. Using this feature, we can mutually convert the sampling rates that meet this condition.
[0090] In terms of the sampling rate of audio, common ones are 8 kHz, 16 kHz, 24 kHz, 32 kHz, 44.1 kHz, 48 kHz, 96 kHz, and 192 kHz.
[0091] Calculated with a frame shift time t = 10 ms and 50% overlapping frame division, it is compatible with various sampling rates in the following table. Using other frame shift times (in ms) is not compatible with 44.1 kHz. Except for 44.1 kHz, other frequencies can be compatible and mutually converted.
[0092] Explanation is given with a frame length \(t = 10\) ms and 50% overlapping framing:
[0093]
[0094] Explanation is given with a frame length \(t = 8\) ms and 50% overlapping framing:
[0095]
[0096]
[0097] It can be seen that when the frame length \(t = 8\) ms and 50% overlapping framing is used, 44.1 kHz cannot satisfy that the number of sample points is an integer. In this case, except for 44.1 kHz, other sampling rates can be converted to each other.
[0098] In summary, the following conclusions can be obtained:
[0099] ① When the time \(t\) corresponding to the frame length is equal under different sampling rates and 50% overlapping framing is used, the resolution of the spectrum corresponding to the fast Fourier transform and its inverse transform is equal.
[0100] ② Under the constraint that the number of sample points can only be an integer, find a frame time \(t\) that satisfies that the corresponding sample points under each sampling rate are integers. Different sampling rates that meet this condition can be converted to each other by this method.
[0101] Example 2
[0102] This method can be used for echo cancellation of microphones.
[0103] When the sampling rates of the microphone signal and the reference signal are inconsistent, the previous processing method must first perform resampling so that the microphone input signal and the reference signal have the same sampling rate SR3 as the target output signal, and then perform echo cancellation in the time domain or frequency domain.
[0104] As Figure 8 shown, echo cancellation and the resampling method of the present invention can be combined to perform echo cancellation in the frequency domain, achieving the effect of completing the two algorithms of resampling and echo cancellation in one step.
[0105] Example 3
[0106] This resampling method can be used for audio noise reduction.
[0107] As Figure 9As shown, when the sampling rate requirement of the output audio signal is inconsistent with that of the input signal and noise reduction processing needs to be performed on the signal, the noise reduction processing can be combined with the resampling method of the present invention to perform resampling and perform noise reduction processing in the frequency domain, achieving the effect of completing the two algorithms of resampling and noise reduction processing in one step.
[0108] That is to say, when there is a need to perform signal processing in the frequency domain and at the same time the sampling rates of the input signal and the output signal need to be converted, this method is suitable for combining resampling and frequency domain signal processing organically to achieve resampling and frequency domain processing in one step. There is no need for a separate resampling step.
[0109] In summary, the present resampling method has the following beneficial effects:
[0110] 1. Using the present resampling method can achieve high-quality resampling, which is suitable for scenarios with sampling rate conversion requirements of audio signals.
[0111] 2. The present resampling method can be combined with other algorithms that need to be processed in the frequency domain to achieve the purpose of completing resampling and other functions in one step, thereby saving computing power.
[0112] The present resampling method involves concepts such as FFT, IFFT, and normalized IFFT, which are supplemented and explained as follows:
[0113] FFT: It refers to a fast implementation method of the Discrete Fourier Transform (DFT).
[0114] IFFT: It refers to a fast implementation method of the Inverse Discrete Fourier Transform (IDFT).
[0115] Normalized IFFT: It refers to a fast implementation method of the normalized Inverse Discrete Fourier Transform (IDFT). The difference between the normalized IFFT and the IFFT is that the normalized IFFT normalizes the result of the IFFT (multiplies by 1 / N).
[0116] Formula expression:
[0117] Let x(n) be a finite sequence of length M (time-domain signal), and the N-point discrete Fourier transform of x(n) is defined as: (transformation from time domain to frequency domain)
[0118]
[0119] The inverse discrete Fourier transform of X(k) is: (transformation from frequency domain to time domain)
[0120]
[0121] The inverse discrete Fourier transform of X(k) (with a coefficient of 1 / N in front) is the transformation from the frequency domain to the time domain:
[0122]
[0123] In the formula:
[0124] As Figure 2 shown, the present invention also provides an audio resampling device based on frequency domain processing, including:
[0125] A time domain to frequency domain unit 10, configured to perform overlapping framing and time domain windowing on an audio signal at a sampling rate of SR1, and convert it from the time domain to the frequency domain through FFT to obtain a first spectrum at a sampling rate of SR1.
[0126] A sampling rate conversion unit 20, configured to normalize the spectrum at a sampling rate of SR1 with the number of samples of the input signal for FFT as a parameter to obtain a normalized first spectrum; under the condition that the spectral resolutions of the second spectrum and the first spectrum are the same, construct a second spectrum at a sampling rate of SR2 based on the first spectrum at a sampling rate of SR1, and remove high-frequency components through downsampling or fill zeros for high-frequency components through upsampling during this process; at a sampling rate of SR2, window the second spectrum with a window function having a smooth transition function.
[0127] A frequency domain to time domain unit 30, configured to convert the second spectrum at a sampling rate of SR2 from the frequency domain to the time domain through IFFT, perform time domain windowing (optional), and then perform overlapping addition to obtain the corresponding time domain value at a sampling rate of SR2.
[0128] Among them, the time domain to frequency domain unit 10 and the sampling rate conversion unit 20 are arranged in pairs to convert an audio signal at a sampling rate of SR1 to a spectrum at a sampling rate of SR2.
[0129] Combined with Figure 8 shown, the audio resampling device further includes a frequency domain signal processing unit 40, configured to perform processing such as echo cancellation. The inputs of the audio resampling device are respectively a microphone signal at a sampling rate of SR1 and a reference signal at a sampling rate of SR2. These two signals are respectively output as spectra at a sampling rate of SR3 after passing through the time domain to frequency domain unit and the sampling rate conversion unit. The frequency domain signal processing unit 40 is configured to perform frequency domain processing on the two spectra at a sampling rate of SR3, and then output them to the frequency domain to time domain unit 30 at a sampling rate of SR3, and finally output an output signal (time domain signal) at a sampling rate of SR3.
[0130] Combined with Figure 9As shown, the audio resampling device further includes a frequency-domain signal processing unit 50, which is used to perform functions such as noise reduction and voice conversion. The frequency-domain signal processing unit 50 is used to perform frequency-domain processing on the spectrum at the SR2 sampling rate output by one of the sampling rate conversion units 20, and then output it to the frequency-domain to time-domain unit 30 at the SR2 sampling rate, and finally output an output signal (time-domain signal) with a sampling rate of SR2.
[0131] It should be noted that the time-domain to frequency-domain unit 10, the sampling rate conversion unit 20, the frequency-domain to time-domain unit 30, and the frequency-domain signal processing units 40 and 50 can all be implemented by means of software (computer programs), firmware, and hardware combination.
[0132] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all of them are within the protection scope of the present invention.
Claims
1. An audio resampling method based on frequency domain processing, characterized in that: The following steps are involved: The step of converting the time domain to the frequency domain is as follows: overlapping and framing the input signal of the first sampling rate, windowing the time domain, and converting the signal from the time domain to the frequency domain by FFT to obtain a first spectrum at the first sampling rate; Sampling rate conversion step: normalizing the first spectrum with the number of sample points of the FFT of the input signal as a parameter to obtain a normalized first spectrum; constructing a second spectrum at a second sampling rate according to the first spectrum at the first sampling rate under the condition that the spectrum resolution of the second spectrum is consistent with that of the first spectrum; and windowing the second spectrum using a window function with a smooth transition function; Frequency domain to time domain conversion step: convert the second spectrum at the second sampling rate from the frequency domain to the time domain through IFFT, and then overlap and add to obtain the time domain value corresponding to the second sampling rate.
2. The audio resampling method based on frequency domain processing according to claim 1, characterized in that: The spectrum resolution is expressed as: SR / N, where SR is the sampling rate in Hz, and N represents the number of sample points input when performing FFT or IFFT.
3. The audio resampling method based on frequency domain processing according to claim 1, characterized in that: The method of constructing the second spectrum at the second sampling rate according to the first spectrum at the first sampling rate is: When the second sampling rate is greater than the first sampling rate, and the spectrum range of the second spectrum is greater than the spectrum range of the first spectrum, upsampling is adopted: for the spectrum overlap area between the second spectrum and the first spectrum, the spectrum at the first sampling rate is directly copied to the spectrum at the second sampling rate, and the extra spectrum at the second sampling rate is assigned 0; When the second sampling rate is lower than the first sampling rate, and the spectral range of the second spectrum is lower than the spectral range of the first spectrum, downsampling is adopted: for the spectral overlapping area between the second spectrum and the first spectrum, the spectrum at the first sampling rate is directly copied to the spectrum at the second sampling rate.
4. The audio resampling method based on frequency domain processing according to claim 1, characterized in that: The smooth transition window is a rectangular window where a smooth transition function is used at several points at the mutation positions on the left and right ends to perform smooth transition from 0 to 1 and from 1 to 0, and other positions remain unchanged at 1.
5. The audio resampling method based on frequency domain processing as claimed in claim 4, characterized in that: The smooth transition function includes a sigmoid function.
6. The audio resampling method based on frequency domain processing as claimed in claim 5, characterized in that: The expression of the sigmoid function is: The value selection method is: x takes several points in the range of -10 to 10, and calculates S(x) at each point.
7. The audio resampling method based on frequency domain processing according to claim 1, characterized in that: The frame duration of the overlapping frames is 10ms; the applicable first sampling rates include 8kHz, 16kHz, 24kHz, 32kHz, 44.1kHz, 48kHz, 96kHz, and 192kHz; the applicable second sampling rates include 8kHz, 16kHz, 24kHz, 32kHz, 44.1kHz, 48kHz, 96kHz, and 192kHz.
8. The audio resampling method based on frequency domain processing according to claim 1, characterized in that: The frame duration of the overlapping frames is 8ms; the applicable first sampling rates include 8kHz, 16kHz, 24kHz, 32kHz, 48kHz, 96kHz, and 192kHz; the applicable second sampling rates include 8kHz, 16kHz, 24kHz, 32kHz, 48kHz, 96kHz, and 192kHz.
9. An audio resampling device based on frequency domain processing, characterized in that: include: A time domain to frequency domain conversion unit, used to perform overlapping framing and time domain windowing on the audio signal of the first sampling rate, and convert the time domain to the frequency domain through FFT to obtain a first spectrum at the first sampling rate; The sampling rate conversion unit is used to normalize the first spectrum with the number of sample points of the FFT of the input signal as a parameter to obtain a normalized first spectrum; under the condition that the spectrum resolution of the second spectrum is consistent with that of the first spectrum, construct a second spectrum at a second sampling rate according to the first spectrum at the first sampling rate; and select a window function with a smooth transition function to window the second spectrum; The frequency domain to time domain unit converts the second spectrum at the second sampling rate from the frequency domain to the time domain through IFFT, and then overlaps and adds them to obtain the time domain value corresponding to the second sampling rate.
10. The audio resampling device based on frequency domain processing according to claim 9, characterized in that: It also includes a frequency domain signal processing unit for performing frequency domain processing on a second spectrum at a second sampling rate output by at least one of the sampling rate conversion units, and then outputting the second spectrum at the second sampling rate to the frequency domain to time domain conversion unit.
Citation Information
Cited By
Signal processing method and device, electronic equipment and storage medium
CN121742742A