Audio coding and decoding method based on dynamic sampling
By employing multi-feature fusion decision-making and adaptive filtering interpolation, the problems of switching jitter and aliasing distortion in dynamic sampling technology are solved, achieving efficient compression and high-quality audio encoding and decoding, suitable for different platforms and audio signals.
Patent Information
- Application Number
- CN202511097430.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing dynamic sampling technology suffers from switching jitter, aliasing distortion, and uneven switching points, resulting in low compression efficiency and degraded sound quality.
A multi-feature fusion decision mechanism is adopted, which comprehensively analyzes the signal by instantaneous amplitude, short-time energy, short-time average amplitude and short-time zero-crossing rate, dynamically adjusts the sampling rate, and performs adaptive filtering and interpolation processing at the decoding end to smooth the sampling rate switching point.
It achieves more stable sampling rate switching, improves compression efficiency, ensures the continuity of sound quality and a natural listening experience, reduces the computational burden on encoding equipment, and enhances the robustness and applicability of the system.
Smart Images

Figure CN120913573A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio technology, and in particular to an audio coding method based on dynamic sampling. BACKGROUND
[0002] Current mainstream digital audio processing technologies, such as CD audio or MP3, AAC, etc., usually use a fixed sampling rate to encode and store audio signals. Although this method is simple to implement, it has inherent defects: in order to adapt to the most complex part of the signal (such as the transient or high-frequency overtone of music), a higher sampling rate (such as 44.1 kHz or 48 kHz) must always be used. However, audio signals often contain a large amount of relatively simple content, such as silence, steady vowels, or low-frequency instrument sounds. Using a high sampling rate to process these simple contents will result in a large amount of data redundancy, increasing the cost of storage and transmission.
[0003] To solve this problem, the industry has proposed dynamic sampling rate (or variable sampling rate) technology. The core idea is to dynamically adjust the sampling rate according to the real-time characteristics of the signal: use a high sampling rate for complex parts of the signal and a low sampling rate for simple parts. However, existing dynamic sampling technology still has one or more of the following technical problems: Switching jitter and audio quality degradation: Early dynamic sampling technology only relies on a single signal feature (such as an amplitude threshold) to determine the sampling rate switching. This simple criterion is easily affected by signal energy fluctuations, resulting in frequent and unnecessary switching (i.e., switching jitter) near the critical point, which not only affects compression efficiency but also may introduce audible switching noise, reducing the auditory experience.
[0004] Aliasing distortion risk: During the process of reducing the sampling rate from high to low, if effective and high-quality anti-aliasing filtering is not performed, the high-frequency components in the original signal will be incorrectly folded into the low-frequency region, resulting in severe aliasing distortion. This distortion is irreversible and will permanently damage the audio quality.
[0005] Switching point is not smooth: When decoding and playing, the splicing of data blocks with different sampling rates can easily produce waveform discontinuity, which is manifested as clicks / pops in the auditory sense. How to smoothly handle these "joints", especially in the case of complex and variable signal characteristics, is a major challenge for existing technologies. Existing solutions usually use fixed transition methods, which lack adaptability to signal content and are difficult to achieve the best balance between preserving audio quality and achieving smoothness.
[0006] Therefore, there is an urgent need for a more intelligent and fine dynamic sampling codec method, which can not only accurately judge the signal complexity to select the appropriate sampling rate, but also can be fine in each link of encoding and decoding, especially in the moment of sampling rate switching, so as to significantly improve the compression efficiency while maximizing the subjective and objective quality of the audio. SUMMARY
[0007] The present application mainly solves the technical problems of the prior art, such as switching jitter, aliasing distortion risk, and non-smooth switching point, and provides an audio codec method based on dynamic sampling.
[0008] The present application mainly solves the above technical problems through the following technical scheme: an audio codec method based on dynamic sampling, comprising an encoding process and a decoding process, wherein the encoding process comprises: S1: discretizing the analog audio signal to obtain a discretized audio signal, and performing time domain analysis on the discretized audio signal to obtain instantaneous amplitude A n , short-time energy E n , short-time average amplitude M n , and short-time zero-crossing rate Z n ; the discretization is essentially a sampling, but the sampling frequency is much higher than that of ordinary audio, for example, 192KHz; S2: updating the short-time energy threshold and the short-time zero-crossing rate threshold according to the following formula: E th (n) =θE th (n-1) +(1-θ)E n ; Z th (n) =θZ th (n-1) +(1-θ)Z n ; In the formula, E th (n) is the current short-time energy threshold, E th (n-1) is the short-time energy threshold of the previous moment, Z th (n) is the current short-time zero-crossing rate threshold, Z th (n-1) is the short-time zero-crossing rate threshold of the previous moment, and θ is a forgetting factor; S3: determining the current sampling rate by the following formula: ; In the formula, f s (n) is the sampling rate selected at the nth moment, f max , fmid and f base are high, medium and base sampling rates respectively, E th (n) is a current short-time energy threshold, Z th (n) is a current short-time zero-crossing rate threshold, A th is an instantaneous amplitude threshold, M th is a short-time average amplitude threshold, a is a first adjustment coefficient, β is a second adjustment coefficient, and γ is a third adjustment coefficient; S4: resample the discretized audio signal according to the sampling rate determined in step S3 to generate a sampling data stream; S5: continuously monitor the sampling rate determined in step S3, and when the sampling rate changes, cut the sampling data stream before the change into an audio segment; S6: for each audio segment, encode it according to its corresponding sampling rate and package it into a data frame containing a segment marker, sampling rate information, data length, encoded data and a check bit; the format of the encoded audio data is: 4 bytes of segment marker + 4 bytes of sampling rate + 4 bytes of encoded data length + several bytes of encoded data + 2 bytes of check bit; S7: sequentially connect all data frames and add a file header and a file tail to obtain the final encoded audio file; The decoding process is to take out each segment of data of the audio file according to the segment marker, and then decode and play it according to the corresponding sampling rate.
[0009] As a preferred embodiment, when the sampling rate is switched from a higher sampling rate f_high to a lower sampling rate f_new in the decoding process, the following processing is performed: input the last h milliseconds of audio data at the higher sampling rate and the first h milliseconds of audio data at the lower sampling rate into an anti-aliasing low-pass filter with a cutoff frequency lower than f_new / 2 to obtain filtered data, and the low-pass filter is: ; where A_stop is 60 dB, A_pass is 1 dB, w_stop is 0.5f_new, w_pass is 0.45f_new, and f_new is the lower sampling rate.
[0010] The length of the data (2h milliseconds) input into the anti-aliasing low-pass filter is determined as follows: If the audio data before and after the switching point is a strong periodic signal and the fundamental period is T_pitch, then 2h milliseconds is 2T_pitch, i.e. input one fundamental period length of data before and after the switching point into the low-pass filter; If the audio data before and after the switching point is transient signal, h = 1, that is, 1 ms of data before and after the switching point is input into the low-pass filter; If the audio data before and after the switching point is neither strong periodic signal nor transient signal, h = 10, that is, 10 ms of data before and after the switching point is input into the low-pass filter.
[0011] By controlling different input lengths for different signal types, the sound quality can be preserved to the greatest extent and smoothing can be achieved, and strong adaptability is achieved.
[0012] The determination method of strong periodic signal is as follows: (1) Framing and windowing: the audio signal x(n) is divided into overlapping frames (for example, frame length 30 ms, overlap 50%); a window function (for example, Hamming window) is applied to each frame to obtain x_w(n).
[0013] (2) Calculate the autocorrelation function: the autocorrelation function R(τ) of the processed signal frame x_w(n) is calculated; (3) Normalize the autocorrelation function: in order to eliminate the influence of energy size on the judgment, the autocorrelation function needs to be normalized, the method is: R'(τ) = R(τ) / R(0); Wherein R(0) is the total energy of the signal frame, which is also the maximum value of the autocorrelation function, so that the value range of the normalized R'(τ) is between [-1, 1]; (4) Find the best candidate period: in the range of [τ_min, τ_max], find the maximum peak value of the normalized autocorrelation function R'(τ). Let this maximum peak value be R_peak, and the corresponding delay (position) be τ_pitch; for the audio signal with a sampling rate of 48 kHz, in order to effectively detect the fundamental frequency of the human voice, the lowest frequency of the search range can be set to 60 Hz, and the highest frequency can be set to 500 Hz, and the corresponding period search range [τ_min, τ_max] (in sampling points) is [96, 800]; (5) Joint determination and output: compare the found maximum peak value R_peak with the preset periodicity intensity threshold Th_tonal (for example, Th_tonal = 0.85); If R_peak > Th_tonal, the signal frame is a strong periodic signal, and the fundamental frequency period is τ_pitch.
[0014] The determination method of transient signal is as follows: (1) The audio signal is divided into a plurality of very short (not more than 5 ms) sub-frames, for example, a frame of 20 ms of data is further divided into 8 sub-frames of 2.5 ms; (2) Calculate the energy E_sub(i) of each sub-frame: (3) Find the ratio of the maximum sub-frame energy and the minimum sub-frame energy in a frame, E_ratio = max(E_sub) / min(E_sub); (4) If E_ratio is greater than a threshold Th_energy_ratio (e.g. Th_energy_ratio = 100, i.e. the energy changes more than 100 times in a short time), then the signal is determined to be transient.
[0015] As a preference, when switching from a lower sampling rate f_new to a higher sampling rate f_high during decoding and playing, the last g milliseconds of the audio data at the lower sampling rate are interpolated to recover the high frequency components: ; where R = f_high / f_new is the interpolation factor; S sampled [n] is the interpolated signal, S original [m] is the signal before interpolation, and sinc is a normalized function.
[0016] The length of the audio data (g milliseconds) that needs to be interpolated is 10-20 milliseconds.
[0017] The signal is filtered and interpolated during decoding, but not during encoding, mainly to avoid sudden jumps when the sampling rate of the signal changes. Filtering and interpolation actually replace some data, making the transition smoother when the sampling rate of the signal changes.
[0018] As a preference, the high sampling rate is 48KHz, the medium sampling rate is 44.1KHz, and the basic sampling rate is 24KHz.
[0019] As a preference, in step S1, the instantaneous amplitude is calculated as follows: the continuous signal s(t) is discretized as s[m], and the signal envelope is extracted by Hilbert transform: ; where H{s(t)} is defined as: ; is a convolution operation; P.V. represents the Cauchy principal value, which is used to handle the singularity of the integral at τ = t, and the kernel function 1 / πt is the impulse response of the Hilbert transform.
[0020] As a preference, the instantaneous amplitude threshold A th and the short-time average amplitude threshold M th are determined by global analysis, the first adjustment coefficient α is 1.5, the second adjustment coefficient β is 0.8, and the third adjustment coefficient γ is 2.0.
[0021] As a preference, the forgetting factor θ is 0.9.
[0022] The substantial effects brought by the present application are: (1) accurate decision, stable switching, and higher compression efficiency: the present application adopts a multi-feature fusion decision mechanism, and analyzes signals in four dimensions of instantaneous amplitude, short-time energy, short-time average amplitude, and short-time zero-crossing rate. Compared with a scheme relying on a single feature, the present application can more accurately identify different types of audio content such as transient, strong periodicity (stationary strong signal), and silence. In combination with a dynamically updated threshold, the present application effectively avoids sampling rate jitter caused by temporary fluctuations in signal energy, so that the switching of the sampling rate is more stable and reasonable, thereby realizing higher compression efficiency under the premise of ensuring sound quality.
[0023] (2) adaptive smoothing processing at the decoding end, and higher sound quality fidelity: the present application performs adaptive length smoothing processing on the switching point at the decoding end according to the signal type. By determining whether the signal near the switching point is strong periodicity, transient, or other types in real time, the present application dynamically adjusts the window length of filtering or interpolation (for example, an integer multiple of the length of the fundamental frequency period of the periodic signal, and an extremely short length of the transient signal). Compared with a fixed-length transition scheme, this processing method can not only ensure the continuity of the phase when processing stationary signals, but also avoid blurring the sharpness of transient signals when processing them, thereby preserving the original sound quality to the greatest extent and realizing seamless and natural listening experience.
[0024] (3) simplified encoding, flexible decoding, and consideration of efficiency and quality: the present application places the complex anti-aliasing filtering and smoothing transition task at the decoding end, so that the encoding process is simplified, and the computational pressure on the encoding device (especially real-time encoding device) is reduced. At the decoding end, different complexity smoothing algorithms can be selected according to the processing capability of the playback device and the requirement of the user on sound quality. This design philosophy of "simplified encoding and optimized decoding" provides high flexibility and scalability for implementing dynamic sampling technology on platforms with different performance.
[0025] (4) improved robustness and applicability of the system: by introducing adaptive threshold updating and adaptive processing based on signal content, the present application has good adaptability to audio signals of different sources and different dynamic ranges (such as speech, music, and environmental sound). The system can automatically track the long-term statistical characteristics of the signal, and does not need to manually adjust parameters for specific audio, thereby greatly improving the robustness and universality of the method. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 is a flow chart of an audio encoding process based on dynamic sampling of the present application. DETAILED DESCRIPTION
[0027] The technical solutions of the present application are further specifically described below by examples in combination with the drawings.
[0028] Embodiment: An audio coding method based on dynamic sampling, including an encoding process and a decoding process, as shown in the following figure. Figure 1 S1: Discretize the analog audio signal to obtain a discretized audio signal, and perform time domain analysis on the discretized audio signal to obtain the instantaneous amplitude A n , the short-time energy E n , the short-time average amplitude M n , and the short-time zero-crossing rate Z n ; the discretization is essentially a sampling, but the sampling frequency is much higher than that of ordinary audio; S2: Update the short-time energy threshold and the short-time zero-crossing rate threshold according to the following formula: E th (n) = θE th (n-1) + (1- θ)E n ; Z th (n) = θZ th (n-1) + (1- θ)Z n ; In the formula, E th (n) is the current short-time energy threshold, E th (n-1) is the short-time energy threshold at the last moment, Z th (n) is the current short-time zero-crossing rate threshold, Z th (n-1) is the short-time zero-crossing rate threshold at the last moment, and θ is a forgetting factor (θ = 0.9); this step can avoid the fixed threshold being insensitive to the dynamic range of the signal; S3: Determine the current sampling rate by the following formula: ; In the formula, f s (n) is the sampling rate selected at the nth moment, f max , f mid , and f base are respectively the high (48KHz), medium (44.1KHz), and basic (24KHz) sampling rates, E th (n) is the current short-time energy threshold, Z th (n) is the current short-time zero-crossing rate threshold, and A th is the instantaneous amplitude threshold, and M th is a short-time average amplitude threshold, a is a first adjustment coefficient, β is a second adjustment coefficient, and γ is a third adjustment coefficient; high sampling rate: energy and zero-crossing rate are both high (transient signal), or amplitude exceeds threshold (to avoid clipping); medium sampling rate: amplitude is high but zero-crossing rate is low (stationary strong signal); basic sampling rate: low energy, low amplitude, or mute section; S4: resampling the discretized audio signal according to the sampling rate determined in step S3 to generate a sampling data stream; S5: continuously monitoring the sampling rate determined in step S3, and when the sampling rate changes, cutting the sampling data stream before the change into an audio segment; S6: for each audio segment, encoding according to the corresponding sampling rate and packaging into a data frame containing a segment marker, sampling rate information, data length, encoded data, and a check bit; the format of the encoded audio data is: 4 bytes of segment marker + 4 bytes of sampling rate + 4 bytes of encoded data length + several bytes of encoded data + 2 bytes of check bit; S7: sequentially connecting all data frames and adding a file header and a file tail to obtain the final encoded audio file; The dynamic sampling and encoding process needs to use segmented encoding, and the encoding format can be selected from the existing encoding formats (such as MP3\AAC, etc.). The audio data after dynamic sampling can be regarded as the framing of audio, and for each frame, there can be different sampling rates. In the encoding process, the TAG marker needs to be identified first, the data of the current data frame is read, and then encoded in the traditional way, and finally the segmented encoded audio data is assembled into a file, which can realize a dynamic sampling and encoding file of audio.
[0029] The decoding process is to take out each segment of data of the audio file according to the segment marker, and then decode and play according to the corresponding sampling rate.
[0030] For the analog signal of audio, taking a recording of human voice as an example, the audio signal produced by a person speaking is fluctuating and relatively dense and frequent, while the audio signal produced by a person not speaking is relatively flat. According to these two different situations, high-frequency sampling (44100Hz\48000Hz or higher sampling) can be used for the audio segment with large signal fluctuation and high density, and low-frequency sampling (22050Hz\24000Hz or lower sampling) can be used for the audio segment with small signal fluctuation and flatness.
[0031] Here, time domain analysis of the audio analog signal is needed. The characteristics of the signal in the time domain can more directly analyze the changes of the waveform signal of the audio, mainly for signal transient amplitude, short-time energy, short-time average amplitude, short-time zero-crossing rate, etc. to capture the changes of the audio signal.
[0032] The instantaneous amplitude directly reflects the instantaneous value of the signal amplitude, which is calculated by: ; Where H{s(t)} is defined as: ; is a convolution operation; P.V. represents the Cauchy principal value, which is used to deal with the singularity of the integral at τ = t, and the kernel function 1 / πt is the impulse response of the Hilbert transform.
[0033] The instantaneous amplitude threshold A th and the short-time average amplitude threshold M th The first adjustment coefficient α is 1.5, the second adjustment coefficient β is 0.8, and the third adjustment coefficient γ is 2.0, which are determined by global analysis.
[0034] In the area where the amplitude changes rapidly (such as music transient), the sampling rate is increased to reduce distortion, and in the stable area, the sampling rate is reduced to save resources.
[0035] The short-time energy can be used to represent the size and super-segment information of the speech signal, and the calculation method is as follows: ; Where h(n) = ω(n) 2 ; E n represents the short-time energy when the window function is added at the n-th point of the signal, and the Hamming window is selected here: ; If X w (n) is used to represent the signal after windowing processing of x(n), the length of the window function is N, and the short-time energy can be represented as: ; The size information of the current speech signal can be determined by the calculation and analysis of the short-time energy. Due to the limitations of the short-time energy, the short-time average amplitude value is still needed to measure the current signal change: ; The above indicators can reflect the change amplitude of the continuous speech signal. In order to more accurately analyze the input speech signal, a new analysis method is added here, the short-time zero-crossing rate is used to count whether a segment of speech signal meets the standard of increasing or decreasing the sampling rate. The analysis method of short-time zero-crossing is as follows: ; Here, the window amplitude is 1 / 2N, which means taking the average of the zero-crossing rate in the window range. , considering that the non-zero value range of w(n-m) is n-m≥0, n-m≤N-1, the function can also be rewritten as: ; The above analyzes the results of the four indexes of signal instantaneous amplitude, short-time energy, short-time average amplitude value, and short-time zero-crossing rate to measure the sampling rate currently required.
[0036] During the decoding process, when the sampling rate is switched from a higher sampling rate f_high to a lower sampling rate f_new, the following processing is performed: The last h milliseconds of the audio data at the higher sampling rate and the first h milliseconds of the audio data at the lower sampling rate are input into an anti-aliasing low-pass filter with a cutoff frequency lower than f_new / 2 to obtain filtered data, and the low-pass filter is: ; In the formula, A_stop is 60 dB, A_pass is 1 dB, w_stop is 0.5f_new, w_pass is 0.45f_new, and f_new is the lower sampling rate.
[0037] The data length (2h milliseconds) input into the anti-aliasing low-pass filter is determined in the following manner: If the audio data before and after the switching point is a strong periodic signal, and the fundamental frequency period is T_pitch, then 2h milliseconds is 2 times T_pitch, that is, 1 fundamental frequency period length of data before and after the switching point is input into the low-pass filter; If the audio data before and after the switching point is a transient signal, then h=1, that is, 1 millisecond of data before and after the switching point is input into the low-pass filter; If the audio data before and after the switching point is neither a strong periodic signal nor a transient signal, then h=10, that is, 10 milliseconds of data before and after the switching point is input into the low-pass filter.
[0038] By controlling different input lengths for different signal types, the sound quality can be preserved to the greatest extent and smoothing can be achieved, with strong adaptability.
[0039] The determination method of the strong periodic signal is: (1) Framing and windowing: the audio signal x(n) is divided into overlapping frames (for example, frame length 30 ms, overlap 50%); a window function (for example, Hamming window) is applied to each frame to obtain x_w(n).
[0040] (2) Calculate the autocorrelation function: calculate the autocorrelation function R(τ) of the processed signal frame x_w(n); (3) Normalized autocorrelation function: in order to eliminate the influence of energy size on the judgment, the autocorrelation function needs to be normalized, the method is: R'(τ)=R(τ) / R(0); Wherein R(0) is the total energy of the signal frame, which is also the maximum value of the autocorrelation function, so that the value range of the normalized R'(τ) is between [-1, 1]; (4) Finding the best candidate period: in the range of [τ_min,τ_max], find the maximum peak value of the normalized autocorrelation function R'(τ). Let this maximum peak value be R_peak, and its corresponding delay (position) be τ_pitch; For the audio signal with a sampling rate of 48kHz, in order to effectively detect the fundamental frequency of the human voice, the lowest frequency of the search range can be set to 60Hz, and the highest frequency can be set to 500Hz. The corresponding period search range [τ_min,τ_max] (in sampling points) is [96, 800]; (5) Joint determination and output: compare the found maximum peak value R_peak with the preset periodicity intensity threshold Th_tonal (for example, Th_tonal=0.85); If R_peak>Th_tonal, the signal frame is a strong periodic signal, and the fundamental frequency period is τ_pitch.
[0041] The determination method of transient signal is: (1) Divide the audio signal into several very short (not more than 5ms) sub-frames (sub-frames), for example, a frame of 20ms data is subdivided into 8 sub-frames of 2.5ms; (2) Calculate the energy E_sub(i) of each sub-frame: (3) Find the ratio E_ratio=max(E_sub) / min(E_sub) of the maximum sub-frame and the minimum sub-frame in a frame; (4) If E_ratio is greater than the threshold Th_energy_ratio (for example, Th_energy_ratio=100, that is, the energy instantaneously changes more than 100 times), the signal is determined to be a transient signal.
[0042] When decoding and playing, when switching from a lower sampling rate to a higher sampling rate, the last several milliseconds (g milliseconds) of audio data of the lower sampling rate are recovered by interpolation: ; Wherein R=f_high / f_new is the interpolation multiple; S sampled [n] is the interpolated signal, S original [m] is the signal before interpolation, and sinc is a normalized function. Polynomial interpolation or FFT interpolation can also be used.
[0043] The length of audio data that needs to be interpolated (g milliseconds) is 10-20 milliseconds.
[0044] The signal is filtered and interpolated at the time of decoding, but not at the time of encoding. This is mainly to avoid sudden jumps when the signal sampling rate changes. Filtering and interpolation actually replace some data, making the signal sampling rate change more smoothly.
[0045] The specific embodiments described herein are merely illustrative of the spirit of the present application. Those skilled in the art can make various modifications or additions to the specific embodiments described or adopt similar ways to replace them without departing from the spirit of the present application or exceeding the scope defined by the appended claims.
[0046] Although the terms short-time energy, short-time zero-crossing rate, and adjustment coefficient are used frequently herein, the possibility of using other terms is not excluded. These terms are used only to facilitate the description and explanation of the essence of the present application; any interpretation of them as additional limitations is contrary to the spirit of the present application.
Claims
1. A method of audio coding based on dynamic sampling, characterized by, The encoding process comprises the following steps: S1: discretize the analog audio signal to obtain a discretized audio signal, perform time domain analysis on the discretized audio signal to obtain an instantaneous amplitude A n , a short-time energy E n , a short-time average amplitude M n , and a short-time zero-crossing rate Z n ; S2: updating the short-time energy threshold and the short-time zero-crossing rate threshold according to the following formula: E th (n) = θE th (n-1) + (1 - θ)E n ; Z th (n) = θZ th (n-1) + (1 - θ)Z n ; where E th (n) is the current short-time energy threshold, E th (n-1) is the previous short-time energy threshold, Z th (n) is the current short-time zero-crossing rate threshold, Z th (n-1) is the previous short-time zero-crossing rate threshold, θ is a forgetting factor. S3: determining the current sampling rate by the following formula: ; wherein f s (n) is the sampling rate used at the nth moment, f max , f mid and f base are high, medium and basic sampling rates, respectively, E th (n) is the current short-time energy threshold, Z th (n) is the current short-time zero-crossing rate threshold, A th is the instantaneous amplitude threshold, M th is the short-time average amplitude threshold, and α, β and γ are first, second and third adjustment coefficients, respectively. S4: resampling the discretized audio signal according to the sampling rate determined in step S3 to generate a sampling data stream; S5: continuously monitoring the sampling rate determined in step S3, and when the sampling rate changes, cutting the sampling data stream before the change into an audio segment; S6: for each audio segment, encoding according to the corresponding sampling rate and packaging into a data frame containing a segment marker, sampling rate information, data length, encoded data and a check bit; S7: sequentially connecting all data frames and adding a file header and a file tail to obtain the final encoded audio file; The decoding process is to take out each segment of data of the audio file according to the segment marker, and then decode and play according to the corresponding sampling rate.
2. The audio codec method based on dynamic sampling according to claim 1, characterized in that, In the decoding process, when the sampling rate is switched from a higher sampling rate f_high to a lower sampling rate f_new, the following processing is performed: The last several seconds of audio data at the higher sampling rate and the first several seconds of audio data at the lower sampling rate are input into an anti-aliasing low-pass filter with a cutoff frequency lower than f_new / 2 to obtain filtered data, and the low-pass filter is: ; where A_stop is 60 dB, A_pass is 1 dB, w_stop is 0.5f_new, w_pass is 0.45f_new, and f_new is the lower sampling rate.
3. The audio codec method based on dynamic sampling according to claim 1 or 2, characterized in that, When switching from the lower sampling rate f_new to the higher sampling rate f_high during decoding and playing, the last several seconds of audio data at the lower sampling rate are restored to high-frequency components by interpolation: ; where R = f_high / f_new is the interpolation factor; S sampled [n] is the interpolated signal, S original [m] is the non-interpolated signal, and sinc is the normalization function.
4. The audio codec method based on dynamic sampling according to claim 1, characterized in that, The high sampling rate is 48 KHz, the medium sampling rate is 44.1 KHz, and the basic sampling rate is 24 KHz.
5. The audio coding method based on dynamic sampling according to claim 1, characterized in that, In step S1, the instantaneous amplitude is calculated by the following method: the continuous signal s(t) is discretized as s[m], and the signal envelope is extracted by Hilbert transform: ; where H{s(t)} is defined as: ; is the convolution operation; P.V. denotes the Cauchy principal value, used to deal with the singularity of the integral at τ = t, and the kernel function 1 / πt is the impulse response of the Hilbert transform.
6. The audio coding method based on dynamic sampling according to claim 1, characterized in that, instantaneous amplitude threshold A th and short-time average amplitude threshold M th By global analysis, the first adjustment coefficient a is 1.5, the second adjustment coefficient β is 0.8, and the third adjustment coefficient γ is 2.
0.
7. The audio codec method based on dynamic sampling according to claim 1, characterized in that, The forgetting factor θ is 0.9.
Citation Information
Patent Citations
Audio frequency decoder and audio frequency decoding method
CN101079296A
Voice processing method and device, electronic equipment and storage medium
CN111402908A
Speech feature processing method and device, equipment and medium
CN120340477A
Method for monitoring and for compression of digitized signals
US20010050953A1
A method of PCM code stream speech detection and the apparatus
WO2007121648A1