Frequency shift processing method for audio dynamic frequency band mapping

The audio frequency shifting processing method based on dynamic frequency band division and harmonic correction solves the problems of rigid frequency band division and harmonic relationship destruction in existing technologies, realizes dynamic adaptation of audio signals and optimization of human hearing perception, and improves the clarity and naturalness of audio.

CN121963754APending Publication Date: 2026-05-01PIONEER TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PIONEER TECH (SHANGHAI) CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing frequency shifting technology cannot dynamically adapt to the time-varying characteristics of audio signals and ignores the laws of human auditory perception, resulting in inflexible frequency band division, destruction of harmonic relationships, and affecting the clarity and naturalness of audio signals.

Method used

The frequency shifting processing method, which involves dynamic frequency band division, harmonic correction, and auditory optimization, includes preprocessing, feature quantization, dynamic frequency band division, dynamic mapping and precise frequency shift calculation, audio signal reconstruction and auditory optimization, and employs techniques such as short-time Fourier transform, low-pass filtering, pre-emphasis processing and harmonic correction.

Benefits of technology

It significantly improves the naturalness, clarity, and recognizability of audio processing, dynamically adapts to the time-varying characteristics of audio, deeply integrates the laws of human auditory perception, and strictly maintains harmonic conservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963754A_ABST
    Figure CN121963754A_ABST
Patent Text Reader

Abstract

The invention discloses a frequency shift processing method for audio dynamic frequency band mapping, and the method is characterized in that the method comprises the following steps: S1, preprocessing and feature quantification: preprocessing an original audio, converting the preprocessed original audio from a time domain signal to a frequency domain signal through short-time Fourier transform, and extracting the core feature parameters of the frequency domain signal; s2, dynamic frequency band division: constructing a decision function, and dividing the frequency domain signal into a plurality of sub-frequency bands based on a dynamic frequency band division rule; s3, dynamic mapping and accurate frequency shift calculation: constructing a frequency band scoring function, performing quantitative scoring on the sub-frequency bands, screening out a target frequency band, and shifting the rest of the sub-frequency bands to the target frequency band to obtain frequency domain signals after frequency shift; and S4, audio signal reconstruction and auditory optimization: filtering the frequency domain signal after frequency shift, converting the filtered frequency domain signal back to a time domain signal through inverse short-time Fourier transform, and dynamically optimizing the frequency band gain of the time domain signal through an equal-loudness curve.
Need to check novelty before this filing date? Find Prior Art

Description

A frequency shifting processing method for audio dynamic frequency band mapping Technical Field

[0001] This invention relates to the field of audio signal processing technology, and more specifically, to a frequency shifting processing method for dynamic audio frequency band mapping. Background Technology

[0002] Frequency shifting, a key technology in audio signal processing, achieves functions such as hearing adaptation, noise avoidance, and sound effect optimization through selective spectrum shifting. It has wide applications in scenarios such as in-vehicle communication, digital hearing aids, and instrument sound effect processing. Audio signals exhibit unique, strong time-varying characteristics; their fundamental frequency dynamically changes with the speaker type and content. For example, the fundamental frequency of male speech is typically in the range of 85-180Hz, female speech in the range of 165-255Hz, and children's speech can reach 200-255Hz. Furthermore, the harmonic structure is closely related to the fundamental frequency. Simultaneously, the human auditory system exhibits high sensitivity to the 1-4kHz frequency band and displays masking effects and isoloudness curve characteristics. This means that strong noise can mask weak signals, and different frequencies produce different loudness perceptions at the same sound pressure level. Moreover, the naturalness of audio signals such as speech and musical instruments is highly dependent on the conservation of the frequency interval between the fundamental wave and its harmonics; any disruption of harmonic relationships will lead to sound quality degradation.

[0003] Current frequency shifting technologies face multiple challenges. Frequency band allocation mechanisms lack flexibility, generally employing equal-width band divisions or preset fixed segmentation strategies, failing to respond to dynamic changes in audio signal characteristics. For example, when processing low-frequency male speech, fixed wide-band divisions force the merging of low-frequency harmonics, resulting in a broken harmonic structure after frequency shifting; while when processing high-frequency female speech, fixed narrow-band divisions not only increase computational complexity but also cause insufficient adaptation capabilities in the mid-to-high frequency bands. The mapping decision process fails to fully integrate the laws of human auditory perception, over-relying on a single signal-to-noise ratio parameter and ignoring key factors such as masking effects and sensitive frequency band characteristics. In digital hearing aid applications, high-frequency signals are indiscriminately shifted to the mid-frequency band. If environmental noise exists in the target frequency band, the shifted signal is easily masked by noise, or produces a harsh or distorted sound because the target frequency band exceeds the comfortable hearing range of the human ear. Frequency shifting calculations generally neglect harmonic correlation, achieving frequency shifting only through simple spectral stretching or compression, leading to a mismatch between the fundamental and harmonic frequency intervals. A typical example is when the 3-8kHz high-frequency band is compressed to 1-3kHz, the fixed scaling ratio destroys the harmonic relationship, the output speech exhibits robotic pitch distortion characteristics, the total harmonic distortion is significantly increased, and the speech clarity and naturalness are seriously damaged.

[0004] The aforementioned defects cause serious problems in core application scenarios. In the field of digital hearing aids, fixed-band frequency shifting cannot synchronously adapt to the personalized hearing curves and real-time speech dynamics of different hearing-impaired patients, resulting in decreased speech recognition accuracy. In in-vehicle audio communication scenarios, static frequency shifting mechanisms are unable to dynamically avoid sudden noise frequency bands, causing speech signals to be masked by engine noise or wind noise, reducing communication reliability. In the process of musical instrument sound effect processing, harmonic distortion causes the timbre characteristics to deviate from the original performance, affecting the realism of the sound effects and the user experience. Therefore, there is an urgent need to develop a frequency shifting processing scheme that can dynamically adapt to the time-varying characteristics of audio, deeply integrate the laws of human ear perception, and strictly maintain harmonic conservation.

[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0006] The purpose of this invention is to provide a frequency shifting processing method for dynamic frequency band mapping of audio, which can dynamically adapt to the time-varying characteristics of audio, deeply integrate the laws of human ear perception and strictly maintain harmonic conservation, thereby significantly improving the naturalness, clarity and recognizability of audio processing.

[0007] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0008] A frequency shifting processing method for audio dynamic frequency band mapping includes the following steps:

[0009] S1. Preprocessing and Feature Quantization

[0010] The original audio is preprocessed, and the preprocessed original audio is converted from a time domain signal to a frequency domain signal through short-time Fourier transform. The fundamental frequency, harmonic frequency interval, auditory masking threshold, frequency band signal-to-noise ratio, and power ratio of sensitive frequency bands are extracted from the frequency domain signal.

[0011] S2. Dynamic frequency band allocation

[0012] A decision function is constructed based on the fundamental frequency, frequency band signal-to-noise ratio, and power ratio of sensitive frequency bands, and the frequency domain signal is divided into several sub-frequency bands based on the dynamic frequency band division rules.

[0013] S3. Dynamic Mapping and Precise Frequency Shift Calculation

[0014] Based on the auditory masking threshold, a frequency band scoring function is constructed to quantify and score the signal quality and perceptual adaptability of the sub-frequency bands. A target frequency band is selected from several sub-frequency bands, and the remaining sub-frequency bands are shifted to the target frequency band to obtain the frequency domain signal after frequency shifting. Based on the harmonic frequency interval, the harmonic relationship of the frequency domain signal after frequency shifting is maintained by the harmonic correction formula.

[0015] S4. Audio Signal Reconstruction and Auditory Optimization

[0016] The frequency-domain signal after frequency shifting is filtered, and the filtered frequency-domain signal is converted back to the time-domain signal through inverse short-time Fourier transform. The frequency band gain of the time-domain signal is then dynamically optimized using equal loudness curves.

[0017] Furthermore, the preprocessing in S1 includes:

[0018] Analog-to-digital conversion: Performing analog-to-digital conversion on the input audio signal so that the converted audio signal satisfies the Nyquist sampling criterion;

[0019] Low-pass filtering: A low-pass filter is used to remove high-frequency noise and clutter while retaining the effective frequency band;

[0020] Pre-emphasis processing: The signal-to-noise ratio of high-frequency signals is improved by using the following pre-emphasis formula to match the sensitivity of the human ear to high frequencies;

[0021]

[0022] in, The pre-emphasized output signal at time t The original audio sample value at time t, t is the original audio sample value at time t-1, and 0.97 is the pre-emphasis coefficient.

[0023] Furthermore, the preprocessed original audio in S1 is framed using a Hanning window, and the original audio is converted from a time-domain signal to a frequency-domain signal using a short-time Fourier transform, the expression of which is:

[0024]

[0025]

[0026] in, This represents the complex amplitude value of the m-th frame at frequency domain index k. This represents summing over all sampled points within a frame. This represents the time-domain signal after pre-emphasis processing. This represents the Hanning window function. The Fourier transform kernel function is represented by N; N represents the number of sampling points in each frame of signal; and n is the index of the sampling point position in the current frame.

[0027] Furthermore, in S1, the extraction formulas for the fundamental frequency, harmonic frequency interval, auditory masking threshold, frequency band signal-to-noise ratio, and sensitive frequency band power ratio of the frequency domain signal are as follows:

[0028] The fundamental frequency is calculated using an autocorrelation function:

[0029]

[0030]

[0031] in, For the fundamental frequency, for The main peak position, Here, T is the autocorrelation function value, and T is the frame length. The pre-emphasized output signal at time t The signal after time delay. For time delay, For time integral infinitesimal elements;

[0032] The harmonic frequency interval is equal to the fundamental frequency;

[0033] Auditory masking threshold:

[0034] in, For auditory masking threshold, As the reference sound pressure level, This represents the critical frequency band offset. Noise level;

[0035] Frequency band signal-to-noise ratio:

[0036] in, For frequency band signal-to-noise ratio, For the signal power spectral density, The noise power spectral density;

[0037] Power percentage in sensitive frequency bands:

[0038] in, The numerator represents the signal power within the frequency range of 1000Hz to 4000Hz, while the denominator represents the total signal power across the entire frequency band.

[0039] Furthermore, in S2, the expression for the decision function is:

[0040]

[0041]

[0042] in, Let be the decision function. For the fundamental frequency, To preset the threshold for fundamental frequency variation, For the power ratio of sensitive frequency bands, This represents the global average signal-to-noise ratio. The maximum signal-to-noise ratio threshold is defined by α, β, and γ, which are weighting coefficients, K is the total number of frequency points, and k is the frequency index. This represents the signal-to-noise ratio at a single frequency point.

[0043] Furthermore, the dynamic frequency band allocation rules in S2 are as follows:

[0044] (a) Adaptive adjustment rule for the number of sub-bands:

[0045] The number of sub-bands M∈{3,4,5,6} is determined by the value of the decision function D:

[0046] When D < 0.6, M = 3; when 0.6 ≤ D < 1, M = 4; when D ≥ 1, M = 5 or 6.

[0047] (b) The sub-band boundaries adopt an exponential boundary formula:

[0048]

[0049] in, Let be the upper limit frequency of the (i+1)th sub-band. Let M be the upper limit frequency of the i-th sub-band, and M be the number of sub-bands;

[0050] (c) Optimization rules for interference scenarios:

[0051] If a certain sub-frequency band satisfies If so, the sub-band will be divided into two narrowband sub-bands;

[0052] in, For interference power, The signal power spectral density;

[0053] (d) High-frequency enhancement rules:

[0054] When the power proportion of the high-frequency band 8–20kHz is greater than 0.3, the bandwidth of that high-frequency band is extended.

[0055] Furthermore, the S3. dynamic mapping and precise frequency shift calculation step includes:

[0056] S31. Target Frequency Band Scoring and Screening

[0057] Construct a frequency band scoring function, and select several sub-frequency bands with higher scores as target frequency bands. Its expression is as follows:

[0058]

[0059] in, For frequency band scoring functions, For signal quality weights, For frequency point signal-to-noise ratio, The maximum signal-to-noise ratio threshold. To mask the adaptation weights, For auditory masking threshold, For the signal power spectral density, To suppress interference weights, Interference power;

[0060] S32. Sub-band to target band dynamic mapping

[0061] Quality first principle: will meet The low-quality sub-bands are mapped to the high-quality target bands. ;

[0062] in, These are the quality assessment parameters for sub-bands. These are the quality assessment parameters for the target frequency band; harmonic matching principle: the ratio of the sub-band bandwidth to the target frequency band bandwidth must satisfy:

[0063]

[0064] in, For the bandwidth of the sub-band, For the bandwidth of the target frequency band, For signal harmonic intervals, For adaptation coefficients;

[0065] S33. Harmonic Conservation Frequency Shift Calculation

[0066] Frequency shift compression ratio calculation: Based on the frequency range of the sub-band and the target band, the formula is:

[0067]

[0068] in, This is the frequency shift compression ratio. and These are the lowest and highest frequencies of the sub-band, respectively. and These are the lowest and highest frequencies of the target frequency band, respectively.

[0069] Frequency domain shift: A linear interpolation algorithm is used to achieve precise frequency spectrum shifting. The formula is as follows:

[0070]

[0071] in, For the mapped signal at the frequency point The amplitude and phase, For the original signal at frequency points Amplitude and phase information, The starting frequency domain index for the target frequency band. This is the starting frequency domain index for the sub-band;

[0072] Harmonic correction: To maintain the harmonic relationships of the audio after frequency shifting, the harmonic frequencies are corrected using a harmonic correction formula, which is as follows:

[0073]

[0074] in, The frequency of the nth harmonic in the target frequency band. Let n be the fundamental frequency of the target frequency band, and n be the harmonic order. This represents the harmonic frequency interval.

[0075] Furthermore, the S4. audio signal reconstruction and auditory optimization step includes:

[0076] S41. Filter the frequency-domain signal after frequency shifting.

[0077] The filtering employs an improved quadrature mirror filter bank, including an analysis filter and a synthesis filter;

[0078] The analysis filter adopts a Kaiser window design FIR structure, and its cutoff frequency is calculated using the following formula:

[0079]

[0080] in, The cutoff frequency, This is the lower limit frequency of the i-th sub-band after dynamic partitioning. This represents the upper limit frequency of the (i+1)th sub-band after dynamic partitioning. Sampling rate;

[0081] The synthesized filter is generated based on the reconstruction conditions of the analysis filter using the following formula:

[0082]

[0083] in, This is the time-domain impulse response of the synthesized filter, used to reassemble the split sub-band signals into a complete time-domain signal. To analyze the time-domain impulse response of the filter, the input signal is decomposed into different sub-frequency bands; N is the filter order, and n is the index of the time-domain sampling point;

[0084] S42. Time-domain signal reconstruction

[0085] The filtered frequency domain signal is converted back to the time domain signal using the inverse short-time Fourier transform, and the reconstruction formula is as follows:

[0086]

[0087] in, The reconstructed time-domain signal is given, where t is the time-domain sampling point position, m is the frame index, and N is the number of FFT points. For frequency points, For the target frequency domain signal, The inverse Fourier transform kernel function. For the Hanning window function;

[0088] S43. Auditory Perception Gain Optimization

[0089] Based on the equal loudness curve, the gain value is dynamically adjusted according to frequency, and the gain function is:

[0090]

[0091] in, For frequency The gain value applied at that point, For the target loudness level, The reconstructed time-domain signal in frequency The measured loudness level at the location.

[0092] In summary, the present invention has the following beneficial effects:

[0093] By using dynamic frequency band division, target frequency band selection, and frequency shift calculation, this technology solves the problems of lack of flexibility in frequency band division, lack of integration of human hearing perception in mapping decisions, and neglect of harmonic correlation in frequency shift calculation in existing technologies. It can dynamically adapt to the time-varying characteristics of audio, deeply integrate the laws of human hearing perception, and strictly maintain harmonic conservation, thereby significantly improving the naturalness, clarity, and recognizability of audio processing. Attached Figure Description

[0094] Figure 1 is a flowchart of the frequency shifting processing method for audio dynamic frequency band mapping according to the present invention. Detailed Implementation

[0095] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below with reference to the figures and specific embodiments.

[0096] As shown in Figure 1, this invention proposes a frequency shifting processing method for audio dynamic frequency band mapping, comprising the following steps:

[0097] S1. Preprocessing and Feature Quantization

[0098] The original audio is preprocessed, and the preprocessed original audio is converted from a time domain signal to a frequency domain signal through short-time Fourier transform. The fundamental frequency, harmonic frequency interval, auditory masking threshold, frequency band signal-to-noise ratio, and power ratio of sensitive frequency bands are extracted from the frequency domain signal.

[0099] S2. Dynamic frequency band allocation

[0100] A decision function is constructed based on the fundamental frequency, the signal-to-noise ratio of the frequency band, and the power ratio of the sensitive frequency band. The frequency domain signal is then divided into several sub-bands based on the dynamic frequency band division rules.

[0101] S3. Dynamic Mapping and Precise Frequency Shift Calculation

[0102] A frequency band scoring function is constructed to quantify and score the signal quality and perceptual adaptability of sub-frequency bands. The target frequency band is selected from several sub-frequency bands, and the remaining sub-frequency bands are shifted to the target frequency band to obtain the frequency domain signal after frequency shifting. Based on the harmonic frequency interval, the harmonic relationship of the frequency domain signal after frequency shifting is maintained by the harmonic correction formula.

[0103] S4. Audio Signal Reconstruction and Auditory Optimization

[0104] The frequency-domain signal after frequency shifting is filtered, and the filtered frequency-domain signal is converted back to the time-domain signal through inverse short-time Fourier transform. The frequency band gain of the time-domain signal is then dynamically optimized using equal loudness curves.

[0105] To facilitate understanding of this embodiment, some key terms are explained below:

[0106] The fundamental frequency is the lowest periodic vibration frequency in an audio signal. It determines pitch perception and is closely related to the harmonic structure of the audio signal.

[0107] The harmonic frequency interval refers to the frequency difference between the fundamental frequency and each harmonic in an audio signal. Its conservation plays an important role in maintaining the naturalness and timbre consistency of the audio.

[0108] The auditory masking threshold refers to the sound pressure level of the weakest sound (the masked sound) that the human ear can perceive in the presence of one or more loud sounds (masking sounds). This threshold reflects the degree to which the human ear's ability to perceive weak signals is affected by strong signals.

[0109] The signal-to-noise ratio (SNR) of a frequency band is the ratio of signal power to noise power within a specific frequency band, used to quantify the signal quality of that band.

[0110] Sensitive frequency band power ratio refers to the proportion of audio signal power in the frequency band sensitive to human ears (e.g., 1kHz to 4kHz) to the total frequency band power. It is used to assess the energy distribution of audio signals in key areas of human ear perception.

[0111] The decision function is a mathematical model built on parameters such as fundamental frequency, frequency band signal-to-noise ratio, and power ratio of sensitive frequency bands. It is used to comprehensively evaluate the dynamic characteristics and perceived importance of audio signals, thereby guiding subsequent dynamic frequency band allocation.

[0112] Dynamic frequency band allocation rules refer to a strategy that adaptively adjusts the number and boundaries of sub-bands based on the output value of the decision function to ensure that the frequency band allocation can match the time-varying characteristics of the audio signal in real time.

[0113] The frequency band scoring function is a mathematical function used to quantify the signal quality and perceptual adaptability of a sub-frequency band. It comprehensively considers factors such as signal-to-noise ratio, auditory masking threshold, and signal power spectral density to select the most suitable target frequency band as the frequency shifting target.

[0114] The harmonic correction formula is a mathematical expression used to correct harmonic frequencies during frequency shifting. It aims to maintain the frequency interval between the fundamental frequency and each harmonic of the audio signal after frequency shifting, thereby avoiding timbre distortion.

[0115] The equal loudness curve refers to the sound pressure level curve required for the human ear to perceive the same loudness at different frequencies. By dynamically adjusting the frequency band gain with reference to this curve, the auditory perception effect of the reconstructed time-domain signal can be optimized.

[0116] This embodiment provides a frequency shifting processing method for dynamic audio frequency band mapping, and its specific implementation process is as follows:

[0117] In step S1, Preprocessing and Feature Quantization, the original audio is preprocessed. This preprocessing may include basic noise suppression and gain adjustment. The preprocessed audio is then converted from a time-domain signal to a frequency-domain signal using a short-time Fourier transform (SFT). This conversion can employ a standard SFT algorithm, dividing the time-domain audio signal into frames and performing the SFT. Based on this, the fundamental frequency, harmonic frequency intervals, auditory masking threshold, frequency band signal-to-noise ratio (SNR), and sensitive frequency band power percentage are extracted from the frequency-domain signal. These features can be extracted using various signal processing techniques. For example, the fundamental frequency can be obtained through peak detection or cepstral analysis; the harmonic frequency intervals can be calculated based on integer multiples of the fundamental frequency; the auditory masking threshold can be estimated based on a psychoacoustic model; the frequency band SNR can be calculated by comparing the energy of the signal and noise within a specific frequency band; and the sensitive frequency band power percentage can be obtained by integrating the power within a specific frequency range and comparing it with the total power.

[0118] In step S2, dynamic frequency band partitioning, a decision function is constructed based on the extracted fundamental frequency, frequency band signal-to-noise ratio, and sensitive frequency band power ratio. This decision function can fuse these features through linear weighting or nonlinear combination to form a comprehensive decision index. For example, these feature values ​​can be simply weighted and summed to obtain a value reflecting the dynamic characteristics of the audio. The frequency domain signal is then divided into several sub-bands based on the dynamic frequency band partitioning rules. These rules can dynamically adjust the number and boundaries of sub-bands according to the output value of the decision function. For example, when the decision function value is high, the number of sub-bands can be increased for finer partitioning; when the decision function value is low, the number of sub-bands can be reduced to simplify processing. The boundaries of the sub-bands can be set using a linearly increasing or fixed-proportion method.

[0119] In step S3, Dynamic Mapping and Precise Frequency Shift Calculation, a frequency band scoring function is constructed to quantitatively score the signal quality and perceptual adaptability of sub-frequency bands. This scoring function can be designed by combining factors such as the signal-to-noise ratio, energy distribution, and matching degree with human hearing characteristics of the sub-frequency bands. A target frequency band is selected from several sub-frequency bands, and the remaining sub-frequency bands are shifted to the target frequency band to obtain the frequency-shifted frequency domain signal. The target frequency band can be selected from the sub-frequency band with the highest score. The frequency shifting of the remaining sub-frequency bands can be achieved through simple frequency translation or linear scaling. Based on the harmonic frequency interval, the harmonic relationship of the frequency-shifted frequency domain signal is maintained through a harmonic correction formula. This correction can be performed after the frequency shifting operation is completed, analyzing the frequency-shifted frequency domain signal to identify its harmonic structure and fine-tuning the frequencies of each harmonic according to the original harmonic frequency interval to ensure that they maintain the preset frequency relationship.

[0120] In step S4, audio signal reconstruction and auditory optimization, the frequency-shifted frequency domain signal is filtered. This filtering can be performed using a standard digital filter to process the frequency-shifted signal, removing unnecessary frequency components or smoothing the spectrum. Subsequently, the filtered frequency domain signal is converted back to a time domain signal using an inverse short-time Fourier transform (ISFT). This conversion can employ a standard IFT algorithm to reassemble the filtered frequency domain signal into a time domain signal. Finally, the frequency band gain of the time domain signal is dynamically optimized using equal-loudness curves. This optimization can dynamically adjust the gain of the time domain signal in different frequency bands based on the differences in perceived loudness of sounds at different frequencies. For example, appropriate gain compensation can be applied to low-frequency and high-frequency signals based on a preset equal-loudness curve model, so that sounds of different frequencies appear to have similar loudness to the human ear.

[0121] This method effectively solves the problem of rigid frequency band division in traditional frequency shifting algorithms through dynamic frequency band partitioning, achieving real-time adaptation of the fundamental frequency and harmonic structure of the audio signal. By integrating mapping decisions based on human auditory characteristics with frequency band scoring, it improves the perceived naturalness of the audio and avoids problems such as signal masking or "harshness." Simultaneously, it maintains harmonic conservation during frequency shifting, significantly reducing total harmonic distortion and ensuring the clarity of the output audio is consistent with the original timbre. Therefore, this method can provide superior audio processing results in scenarios such as digital hearing aids, in-vehicle audio communication, and musical instrument sound effect processing.

[0122] In some of the above-mentioned solutions of the present invention, preprocessing is proposed to convert the original audio into a frequency domain signal and extract features. However, in this process, the audio signal may not meet the sampling criteria, resulting in aliasing distortion. High-frequency noise and clutter residues affect the purity of the signal, and the signal-to-noise ratio of the high-frequency signal is insufficient to match the sensitivity of the human ear to high frequencies.

[0123] In response, this invention further proposes that the preprocessing in S1 includes: analog-to-digital conversion, low-pass filtering, and pre-emphasis processing.

[0124] Analog-to-digital conversion (ADC) is the process of converting an input audio signal from analog to digital, ensuring that the converted signal meets the Nyquist sampling criterion. ADC transforms a continuously changing analog audio signal into a discrete digital signal. Its purpose is to digitize the analog signal for subsequent digital signal processing and to ensure that the converted digital signal meets the Nyquist sampling criterion, which requires the sampling frequency to be at least twice the highest frequency of the signal. This effectively avoids aliasing distortion during digitization, guaranteeing the integrity and accuracy of the signal.

[0125] Low-pass filtering removes high-frequency noise and clutter while preserving the effective frequency band. In essence, it processes signals using a low-pass filter to remove high-frequency noise and clutter above a certain cutoff frequency, while retaining the signal components within the effective frequency band. This helps improve signal purity, reduces noise interference with subsequent feature extraction and frequency shifting, and thus enhances the overall audio processing quality.

[0126] Pre-emphasis processing enhances the signal-to-noise ratio of high-frequency signals using the following pre-emphasis formula, adapting to the human ear's sensitivity to high frequencies:

[0127]

[0128] in, The pre-emphasized output signal at time t The original audio sample value at time t, t is the original audio sample value at time t-1, and 0.97 is the pre-emphasis coefficient.

[0129] Pre-emphasis processing is a technique that improves the signal-to-noise ratio by boosting the energy of the high-frequency components of an audio signal. Its function is to compensate for potential attenuation of high-frequency signals during transmission or processing, and to adapt to the human ear's sensitivity to high-frequency sounds, allowing for more efficient utilization of high-frequency information in subsequent processing, thereby improving audio clarity and perceived quality.

[0130] Through the above technical solutions, during the preprocessing of the original audio, analog audio signals are converted into digital signals that meet the Nyquist sampling criterion via analog-to-digital conversion, effectively avoiding signal aliasing distortion and providing high-quality digital input for subsequent processing. Low-pass filtering removes high-frequency noise and clutter, ensuring the purity of the effective frequency band and reducing noise interference with feature extraction. Finally, pre-emphasis processing improves the signal-to-noise ratio of high-frequency signals, making high-frequency information more prominent and better suited to the human ear's sensitivity to high frequencies. These preprocessing steps work synergistically to lay a solid foundation for the accurate extraction of subsequent features such as fundamental frequency, harmonic frequency spacing, auditory masking threshold, frequency band signal-to-noise ratio, and power proportion of sensitive frequency bands. This significantly improves the quality and perceptual adaptability of the audio signal, thereby optimizing the accuracy and effectiveness of dynamic frequency band division and precise frequency shift calculation, ultimately resulting in a reconstructed audio signal with higher clarity and naturalness.

[0131] In some embodiments of the present invention described above, a method is proposed to convert the original audio signal from the time domain to the frequency domain using short-time Fourier transform in order to extract frequency domain features. However, in its implementation, frame-segmentation may lead to spectral leakage, affecting the accuracy of feature extraction.

[0132] To address this, the present invention further proposes that the preprocessed original audio in S1 be framed using a Hanning window, and that the original audio be converted from a time-domain signal to a frequency-domain signal using a short-time Fourier transform, the expression of which is:

[0133]

[0134]

[0135] in, This represents the complex amplitude value of the m-th frame at frequency domain index k. This represents summing over all sampled points within a frame. This represents the time-domain signal after pre-emphasis processing. This represents the Hanning window function. The Fourier transform kernel function is represented by N; N represents the number of sampling points in each frame of signal; and n is the index of the sampling point position in the current frame.

[0136] The Hanning window is a commonly used window function that smoothly decays to zero at the beginning and end of a signal, effectively reducing spectral leakage caused by signal truncation. Compared to other window functions such as the rectangular window, the Hanning window provides better frequency resolution and lower sidelobe levels, resulting in more accurate frequency domain analysis results. Its function is to avoid introducing spurious frequency components due to abrupt changes in signal framing, ensuring the accuracy of subsequent feature extraction such as fundamental frequency and harmonic frequency intervals. In practical applications, the Hanning window can be implemented by pre-calculating the window function coefficients and multiplying them point-by-point with the time-domain signal of each frame. Another implementation method is to utilize the hardware acceleration capabilities of a digital signal processor (DSP) or field-programmable gate array (FPGA) to quickly apply the window function through lookup tables or parallel multipliers to meet real-time processing requirements.

[0137] Short-Time Fourier Transform (STFT) is a method for converting a time-domain signal into a time-frequency domain representation. It achieves this by dividing the signal into frames and performing a Fourier transform on each frame. Its core function is to reveal the instantaneous frequency components of an audio signal and their changes over time, providing a foundation for subsequent feature extraction and frequency domain processing.

[0138] The pre-emphasized time-domain signal serves as the input to the short-time Fourier transform. Pre-emphasis processing aims to enhance the high-frequency components of the audio signal, making them more prominent in spectral analysis. The human ear is highly sensitive to mid-to-high frequency ranges, and high-frequency signals are prone to attenuation during transmission. Pre-emphasis processing compensates for this attenuation and makes high-frequency features easier to identify and analyze in the frequency domain. This signal is typically acquired by passing the original audio signal through analog-to-digital conversion and low-pass filtering, followed by a first-order high-pass filter. Besides the formula mentioned above, other forms of pre-emphasis filters can be used, such as filters with different coefficients or higher orders, to adapt to different audio characteristics or application scenarios; however, their core purpose is to enhance the relative strength of the high-frequency signal.

[0139] By employing the Hanning window for framing, this invention effectively solves the problem of spectral leakage in traditional framing methods. The smooth attenuation characteristic of the Hanning window reduces spectral energy diffusion caused by signal truncation, making the representation of features such as the fundamental frequency and harmonic frequency intervals of the frequency domain signal more concentrated and accurate, thereby avoiding feature extraction errors caused by spectral leakage. Based on this, the pre-emphasized time-domain signal is converted into a frequency-domain signal using a short-time Fourier transform, strictly adhering to the provided mathematical expression to ensure the accuracy and consistency of the conversion process. The pre-emphasis processing effectively enhances the high-frequency signal in the frequency domain, which matches the sensitivity of the human ear to high frequencies, providing more auditory-perceptual-compliant frequency domain data for subsequent dynamic frequency band division and frequency shifting. This accurate and low-leakage frequency domain conversion provides reliable input for the dynamic frequency band division in subsequent step S2, enabling the decision function to make judgments based on more accurate fundamental frequency, frequency band signal-to-noise ratio, and the power proportion of sensitive frequency bands. Simultaneously, it lays the foundation for dynamic mapping and precise frequency shift calculation in S3, ensuring the accuracy of frequency shift compression ratio and frequency domain shift, and helping to maintain the harmonic relationships of the frequency domain signal after frequency shift. Finally, in the audio signal reconstruction and auditory optimization stage of S4, filtering and gain adjustment can be performed based on more accurate frequency domain information, thereby significantly improving the clarity, naturalness, and auditory comfort of the reconstructed audio, effectively avoiding problems such as "robotic" pitch distortion caused by spectral distortion in traditional methods.

[0140] In some of the solutions described above in this invention, features such as fundamental frequency, harmonic frequency spacing, auditory masking threshold, frequency band signal-to-noise ratio, and sensitive frequency band power ratio are extracted in the preprocessing and feature quantization steps to support subsequent dynamic frequency band division and frequency shifting. However, in this process, the lack of standardized extraction formulas may lead to inaccurate or inconsistent feature calculations. For example, misjudgment of the fundamental frequency may destroy the harmonic structure, deviation in the estimation of the auditory masking threshold may ignore the noise masking effect, and quantization errors in the frequency band signal-to-noise ratio and sensitive frequency band power ratio may reduce the adaptability of frequency band division, thereby affecting the dynamic audio characteristic matching and human ear perception optimization effect of the entire method.

[0141] To address this, the present invention further proposes the following formula for extracting the fundamental frequency, harmonic frequency interval, auditory masking threshold, frequency band signal-to-noise ratio, and power proportion of sensitive frequency bands during the feature quantization process of frequency domain signals:

[0142] The fundamental frequency is calculated using an autocorrelation function:

[0143]

[0144]

[0145] in, For the fundamental frequency, for The main peak position, Here, T is the autocorrelation function value, and T is the frame length. The pre-emphasized output signal at time t The signal after time delay. For time delay, For time integral infinitesimal elements;

[0146] The harmonic frequency interval is equal to the fundamental frequency;

[0147] Auditory masking threshold:

[0148] in, For auditory masking threshold, As the reference sound pressure level, This represents the critical frequency band offset. Noise level;

[0149] Frequency band signal-to-noise ratio:

[0150] in, For frequency band signal-to-noise ratio, For the signal power spectral density, The noise power spectral density;

[0151] Power percentage in sensitive frequency bands:

[0152] in, The numerator represents the signal power within the frequency range of 1000Hz to 4000Hz, while the denominator represents the total signal power across the entire frequency band.

[0153] Specifically, the extraction of the fundamental frequency is fundamental to audio signal analysis, and its accuracy directly affects the subsequent identification and processing of harmonic structures. This invention uses the autocorrelation function to calculate the fundamental frequency. The autocorrelation function identifies periodicity by measuring the similarity between the signal and its delayed version.

[0154] The harmonic frequency interval is the difference between the harmonic frequencies in an audio signal. For naturally produced audio signals, this interval is usually equal to the fundamental frequency. This invention sets the harmonic frequency interval to be equal to the fundamental frequency to simplify calculations and maintain the naturalness of the harmonic relationships.

[0155] Auditory masking threshold quantifies the human ear's ability to perceive weak signals in the presence of strong signals, and is key to optimizing audio perception quality. This invention constructs an auditory masking threshold based on a reference sound pressure level, critical band frequency offset, and noise level.

[0156] The signal-to-noise ratio (SNR) of a frequency band is used to evaluate the signal quality within a specific frequency range. This invention calculates the SNR by the ratio of the signal power spectral density to the noise power spectral density. The signal power spectral density and noise power spectral density can be estimated using various methods. For example, the noise power spectral density can be measured during periods without signal, or estimated from the noisy signal using minimum statistical methods. Alternatively, noise suppression techniques based on Wiener filtering or spectral subtraction can be employed to calculate a more accurate signal power spectral density while estimating the noise, thus obtaining a more accurate SNR.

[0157] The power ratio of sensitive frequency bands measures the proportion of signal energy in the total energy within the frequency range (1000Hz to 4000Hz) that the human ear is most sensitive to.

[0158] Through the above technical solutions, this invention solves the problem of inaccurate feature quantization by providing a standardized feature extraction formula, ensuring the accuracy and consistency of subsequent audio processing steps. The fundamental frequency is calculated using an autocorrelation function. Leveraging the sensitivity of the autocorrelation function to periodic signals, the fundamental frequency is determined by the location of the main peak, avoiding misjudgments caused by dynamic changes in the fundamental frequency in traditional methods. This lays a solid foundation for maintaining the harmonic structure in subsequent dynamic frequency band division and frequency shifting. The harmonic frequency interval is directly equal to the fundamental frequency, simplifying the calculation process and maintaining the naturalness of the harmonic relationship. This effectively prevents audio distortion caused by incorrect harmonic interval estimation, ensuring the naturalness of the audio after frequency shifting. The auditory masking threshold is constructed based on the reference sound pressure level, critical band frequency offset, and noise level, accurately quantifying the masking effect of the human ear. This allows subsequent dynamic frequency band division and dynamic mapping decisions to effectively avoid noise interference areas, improving the perceived quality of the audio. The signal-to-noise ratio (SNR) of a frequency band, calculated as the ratio of signal power spectral density to noise power spectral density, accurately assesses signal quality and provides a reliable basis for prioritizing high-quality regions during dynamic frequency band division and for quality-priority principles during dynamic mapping. The power proportion of sensitive frequency bands focuses on the 1000Hz to 4000Hz range, which is sensitive to human hearing. By quantifying the power proportion, the energy distribution of audio in key perceptual frequency bands is optimized, thereby improving overall auditory comfort, especially in the audio signal reconstruction and auditory optimization steps, enabling better gain adjustment. These formulas collectively achieve the standardization and accurate extraction of audio features, providing precise and reliable input for the dynamic adaptation capability of the entire method, the integration of human auditory perception patterns, and the maintenance of harmonic conservation. This effectively overcomes the problems of rigid frequency band division, poor perceptual effect of mapping decisions, and severe frequency shift distortion caused by inaccurate feature quantization in existing technologies, significantly improving the clarity, naturalness, and auditory comfort of the output audio.

[0159] In some of the above-mentioned solutions of the present invention, a decision function is proposed for dynamically dividing frequency bands. However, in its implementation process, the specific implementation of the decision function is not clearly defined, which may cause the frequency band division to fail to accurately adapt to the fundamental frequency changes, power ratio of sensitive frequency bands and signal-to-noise ratio characteristics of audio signals, thereby affecting the accuracy and efficiency of dynamic frequency band division.

[0160] In response, this invention further proposes that the expression for the decision function in S2 is:

[0161]

[0162]

[0163] in, Let be the decision function. For the fundamental frequency, To preset the threshold for fundamental frequency variation, For the power ratio of sensitive frequency bands, This represents the global average signal-to-noise ratio. The maximum signal-to-noise ratio threshold is defined by α, β, and γ, which are weighting coefficients, K is the total number of frequency points, and k is the frequency index. This represents the signal-to-noise ratio at a single frequency point.

[0164] The decision function D is a key indicator used to evaluate the characteristics of the current audio frame and guide subsequent dynamic frequency band allocation. Its value comprehensively reflects the pitch stability of the audio, the energy distribution of the frequency bands sensitive to the human ear, and the overall and local signal-to-noise ratio. The calculation result of the decision function D directly affects the adaptive adjustment of the number and boundaries of sub-bands, ensuring that the frequency band allocation can respond in real time to the dynamic changes of the audio signal and the environmental noise conditions.

[0165] The fundamental frequency is the lowest periodic vibration frequency in an audio signal, reflecting the pitch of the sound. In speech signals, the fundamental frequency represents the frequency of vocal cord vibration and is crucial for distinguishing different speakers and recognizing speech content.

[0166] The preset fundamental frequency change threshold is a reference value used to measure the degree of change of the current fundamental frequency relative to a certain benchmark or historical value. The purpose of setting this threshold is to identify significant fluctuations in the fundamental frequency, thereby triggering or adjusting the frequency band allocation strategy to accommodate rapid changes in pitch in speech or music.

[0167] The power percentage of the sensitive frequency band quantifies the sensitivity of the human ear to a specific frequency range (e.g., 1000Hz to 4000Hz). This parameter reflects the energy concentration of the frequency band in the audio signal that is most important for human hearing perception.

[0168] The global average signal-to-noise ratio (SNR) is the average ratio of signal power to noise power across the entire audio frequency band, used to assess the overall noise level of the current audio environment. This parameter provides a macroscopic assessment of audio quality and is an important basis for dynamic frequency band allocation decisions.

[0169] The maximum signal-to-noise ratio (SNR) threshold is a preset reference value that indicates the upper limit of the SNR that should be achieved under ideal or acceptable audio quality. This threshold is used in the decision function to evaluate the difference between the current global average SNR and the ideal state, thereby guiding frequency band allocation to optimize the SNR. The maximum SNR threshold can be set according to the needs of the application scenario.

[0170] Weighting coefficients α, β, and γ are used to adjust the degree of influence of each feature in the decision function D on the final decision result. The settings of these coefficients α, β, and γ reflect the importance of each audio characteristic in frequency band allocation decisions under different application scenarios. These weighting coefficients α, β, and γ can be manually adjusted through expert experience to adapt to specific application needs; or they can be trained on large amounts of labeled data using machine learning algorithms, such as Support Vector Machines (SVMs) or neural networks, to automatically learn and optimize these weighting coefficients α, β, and γ to achieve the best frequency band allocation results.

[0171] The total number of frequency points K represents the number of discrete frequency points in the frequency domain signal, usually related to the number of FFT points N in the Short-Time Fourier Transform (STFT). The frequency index k is the sequence number of each discrete frequency point in the frequency domain signal, from 0 to K-1. The signal-to-noise ratio (SNR) at a single frequency point represents the ratio of signal power to noise power at a specific frequency point k. It provides more granular noise distribution information than the global average SNR, helping to identify frequency bands with localized noise concentrations.

[0172] Through the aforementioned technical solution, the decision function D comprehensively evaluates the characteristics of audio signals by considering multiple key features, including fundamental frequency, power proportion of sensitive frequency bands, global average signal-to-noise ratio, and signal-to-noise ratio at a single frequency point. Specifically, by comparing the fundamental frequency with a preset fundamental frequency change threshold, the decision function D can sensitively capture dynamic changes in the pitch of the audio signal, such as fluctuations in the speaker's fundamental frequency or changes in the pitch of an instrument, thereby ensuring that subsequent frequency band allocation can adapt to these time-varying characteristics in real time. Simultaneously, the introduction of power proportion of sensitive frequency bands allows the decision-making process to prioritize the energy distribution of frequency bands most sensitive to the human ear. This helps allocate more resources or more refined allocation strategies to frequency bands crucial to auditory perception during frequency band allocation, thereby improving the perceived naturalness and clarity of the final audio.

[0173] Furthermore, the decision function D integrates the global average signal-to-noise ratio (SNR) and the single-frequency SNR, combined with the maximum SNR threshold, enabling the system to comprehensively assess the noise level of the current audio environment. By considering both global and local SNR, the decision function D can identify frequency bands with concentrated noise and guide the system to adopt avoidance or optimization strategies in subsequent frequency band allocation, effectively reducing the masking effect of noise on the target signal. The introduction of weighting coefficients α, β, and γ gives the decision function D great flexibility, allowing for dynamic adjustment of the importance of each feature in the decision based on different application scenarios (such as hearing aids, in-vehicle communication, and musical instrument processing) and user needs, thereby achieving highly customized frequency band allocation strategies.

[0174] Through the above technical solution, this invention solves the problems of rigidity and insufficient adaptability in traditional frequency band division. The decision function D can reflect the fundamental tone changes of the audio signal, human auditory perception characteristics, and noise environment in real time and accurately, providing a scientific and quantitative basis for subsequent dynamic frequency band division. This allows the number and boundaries of sub-bands to be adaptively adjusted, avoiding the problems of harmonic structure breakage or excessive computational overhead caused by fixed frequency band division, significantly improving the accuracy and efficiency of audio processing, and ultimately optimizing the user's listening experience.

[0175] In some embodiments of the present invention described above, a dynamic frequency band partitioning rule is proposed to divide a frequency domain signal into several sub-frequency bands. However, in its implementation, the number of sub-frequency bands may not be able to adapt to the real-time changing characteristics of the audio signal, resulting in rigid partitioning; the boundary partitioning may not be precise enough to adapt to the differences in human hearing sensitivity; sub-frequency bands are easily affected by noise in interference scenarios, reducing signal quality; and high-frequency signals may not be sufficiently enhanced, affecting the perception effect. To address these issues, the present invention further proposes a dynamic frequency band partitioning rule in S2, which specifically includes:

[0176] This invention proposes an adaptive adjustment rule for the number of sub-bands. This rule aims to dynamically adjust the number M of sub-bands based on the real-time characteristics of the audio signal. The decision function D comprehensively considers multiple audio features such as fundamental frequency, band signal-to-noise ratio, and the power proportion of sensitive bands, reflecting the complexity and importance of the current audio signal. When the value of the decision function D is small (e.g., D < 0.6), it indicates that the audio signal is relatively simple or has low information density. In this case, using fewer sub-bands (M=3) is sufficient to effectively capture its main features and reduce computational overhead. When the value of the decision function D is moderate (e.g., 0.6 ≤ D < 1), the audio signal complexity increases, requiring an appropriate number of sub-bands (M=4) to provide a finer division. When the value of the decision function D is large (e.g., D ≥ 1), the audio signal may contain rich details or have high perceptual importance. In this case, increasing the number of sub-bands (M=5 or 6) can provide higher frequency resolution, thereby better preserving audio details. This adaptive adjustment avoids the rigidity problem of fixed-number sub-band division, allowing the band division to better match the dynamic changes of the audio signal.

[0177] This invention proposes a rule for determining the frequency boundaries of sub-bands using an exponential boundary formula. This rule is used to determine the frequency boundaries of each sub-band, and its core lies in employing an exponential boundary formula.

[0178]

[0179] in, Let be the upper limit frequency of the (i+1)th sub-band. Let M be the upper limit frequency of the i-th sub-band, and M be the number of sub-bands;

[0180] This exponential division method aligns with the human ear's frequency perception characteristics; that is, the human ear is more sensitive to frequency changes in the low-frequency region and relatively less sensitive to frequency changes in the high-frequency region. By using an exponential formula, the bandwidth of the low-frequency sub-bands can be relatively narrow, while the bandwidth of the high-frequency sub-bands can be relatively wide. This results in a more precise division across the entire frequency domain that better matches human auditory perception, optimizing the perceptual adaptability of frequency band division.

[0181] This invention proposes an optimization rule for interference scenarios. This rule aims to improve the signal quality of sub-frequency bands in interference-prone environments. When the interference power of a sub-frequency band exceeds the signal power spectral density, it indicates that the sub-frequency band is significantly affected by noise or interference. To effectively isolate and handle such interference, this rule splits the sub-frequency band into two narrow-band sub-frequency bands. By refining the interfered broadband sub-frequency band into narrower sub-frequency bands, the interference source can be located more accurately, and subsequent processing (such as filtering and noise reduction) can be targeted to optimize these narrow-band sub-frequency bands, thereby reducing the impact of interference on signal quality and improving signal purity.

[0182] This invention proposes a high-frequency enhancement rule. This rule aims to ensure that the high-frequency region most sensitive to the human ear is sufficiently enhanced to optimize overall perceived quality. When the signal power proportion in the 8–20 kHz high-frequency band is greater than 0.3, it indicates that this high-frequency band contains important audio information and has a significant impact on the clarity and naturalness of the audio. In this case, the rule expands the bandwidth of this high-frequency band. This bandwidth expansion can be achieved by adjusting the upper limit frequency of the highest sub-band or redistributing the sub-band boundaries of the high-frequency region, thereby providing more processing space and higher resolution for the high-frequency signal. This ensures that these key high-frequency components are better preserved and enhanced during frequency shifting, thus improving the perceived quality of the final output audio.

[0183] Through the aforementioned dynamic frequency band division rules, this invention effectively solves the problems of rigid frequency band division, inaccurate boundaries, insufficient interference suppression, and insufficient high-frequency enhancement in traditional methods. Specifically, the adaptive adjustment of the number of sub-bands allows the frequency band division to respond in real time to the dynamic changes in the audio signal, avoiding resource waste or information loss caused by fixed division. The application of the exponential boundary formula makes the division of frequency band boundaries more in line with the auditory perception characteristics of the human ear, especially providing higher resolution in the low-frequency region, while balancing efficiency and perception in the high-frequency region. The interference scenario optimization rule, through the refined splitting of the interfered sub-bands, can effectively isolate local noise, improve signal purity, and provide higher quality input for subsequent frequency shifting processing. The high-frequency enhancement rule ensures sufficient attention and processing of high-frequency signals that are sensitive to the human ear. When high-frequency information is rich, more details are retained by expanding the bandwidth, thereby significantly improving the clarity, naturalness, and overall perceived quality of the audio after frequency shifting. The synergistic effect of these rules makes the entire frequency shifting process highly adaptive, accurate, and robust during the frequency band division stage, laying a solid foundation for subsequent dynamic mapping and precise frequency shifting calculations, and ultimately achieving better audio processing results.

[0184] In some of the embodiments of the present invention, dynamic mapping and precise frequency shift calculation are proposed to realize frequency shifting of sub-bands and maintain harmonic relationships. However, in the process of implementation, traditional methods may lead to inaccurate spectrum shifting and destruction of harmonic relationships, thereby affecting audio clarity and naturalness. Specifically, after frequency shifting, the audio will have harmonic distortion, signal quality degradation, and insufficient perceptual adaptability.

[0185] In response, this invention further proposes S3. The dynamic mapping and precise frequency shift calculation steps include:

[0186] S31. Target Frequency Band Scoring and Screening

[0187] Construct a frequency band scoring function, and select several sub-frequency bands with higher scores as target frequency bands. Its expression is as follows:

[0188]

[0189] in, For frequency band scoring functions, For signal quality weights, For frequency point signal-to-noise ratio, The maximum signal-to-noise ratio threshold. To mask the adaptation weights, For auditory masking threshold, For the signal power spectral density, To suppress interference weights, Interference power;

[0190] S32. Sub-band to target band dynamic mapping

[0191] Quality first principle: will meet The low-quality sub-bands are mapped to the high-quality target bands. ;

[0192] in, These are the quality assessment parameters for sub-bands. These are the quality assessment parameters for the target frequency band; harmonic matching principle: the ratio of the sub-band bandwidth to the target frequency band bandwidth must satisfy:

[0193]

[0194] in, For the bandwidth of the sub-band, For the bandwidth of the target frequency band, For signal harmonic intervals, For adaptation coefficients;

[0195] S33. Harmonic Conservation Frequency Shift Calculation

[0196] Frequency shift compression ratio calculation: Based on the frequency range of the sub-band and the target band, the formula is:

[0197]

[0198] in, This is the frequency shift compression ratio. and These are the lowest and highest frequencies of the sub-band, respectively. and These are the lowest and highest frequencies of the target frequency band, respectively.

[0199] Frequency domain shift: A linear interpolation algorithm is used to achieve precise frequency spectrum shifting. The formula is as follows:

[0200]

[0201] in, For the mapped signal at the frequency point The amplitude and phase, For the original signal at frequency points Amplitude and phase information, The starting frequency domain index for the target frequency band. This is the starting frequency domain index for the sub-band;

[0202] Harmonic correction: To maintain the harmonic relationships of the audio after frequency shifting, the harmonic frequencies are corrected using a harmonic correction formula, which is as follows:

[0203]

[0204] in, The frequency of the nth harmonic in the target frequency band. Let n be the fundamental frequency of the target frequency band, and n be the harmonic order. This represents the harmonic frequency interval.

[0205] S31. Target Frequency Band Scoring and Selection aims to quantify the signal quality and perceptual adaptability of the sub-frequency bands by constructing a frequency band scoring function, and to select target frequency bands from several sub-frequency bands. Constructing the frequency band scoring function is to comprehensively evaluate the signal quality and perceptual adaptability of the sub-frequency bands, providing a quantitative basis for subsequent target frequency band selection. The signal quality weight in the frequency band scoring function measures the importance of signal quality in the overall score. This weight can be preset according to the application scenario; for example, a higher weight can be assigned in voice communication. Alternatively, it can be dynamically adjusted according to the real-time environmental noise level; for example, the higher the noise, the higher the signal quality weight. The signal-to-noise ratio (SNR) at frequency point k directly reflects the signal strength relative to noise at a specific frequency point, and its calculation can be obtained by comparing the signal power spectral density and the noise power spectral density. The maximum SNR threshold is used to normalize the SNR or set an upper limit to prevent extremely high SNR from having an excessive impact on the scoring function, ensuring a reasonable range for the score. Masking adaptation weight measures the importance of the human ear masking effect in the overall score. This weight can be set based on human auditory models (such as equal-loudness curves and critical band theory). The auditory masking threshold represents the lowest sound pressure level that the human ear can perceive at a specific frequency. Signals below this threshold will be masked. By considering this threshold, it is possible to avoid shifting signals to frequency bands imperceptible to the human ear. Signal power spectral density represents the energy distribution of a signal at different frequencies and is an important parameter for evaluating signal strength and characteristics. Interference suppression weight measures the importance of interference suppression in the overall score. This weight can be preset based on the type and intensity of environmental interference; or dynamically adjusted based on the characteristics of real-time interference signals (such as the spectral characteristics of engine noise). Interference power represents the intensity of the interference signal at a specific frequency. By considering interference power, it is possible to avoid shifting signals to frequency bands with severe interference. The signal quality and perceptual adaptability of the sub-frequency bands are quantitatively scored by integrating the aforementioned indicators through a frequency band scoring function to generate a comprehensive score for each sub-frequency band. This allows the system to objectively compare the merits of different sub-frequency bands, thus providing data support for subsequent selection. The target frequency band is then selected from several sub-frequency bands. The selection process can be based on a preset scoring threshold, selecting all sub-frequency bands with scores higher than that threshold; or it can select several sub-frequency bands with the highest scores as candidate target frequency bands for further optimization.

[0206] S32. Sub-band to target band dynamic mapping employs a quality-first principle and a harmonic matching principle. The quality-first principle aims to ensure better auditory quality in the frequency-shifted audio signal. When the quality assessment parameter of a sub-band is lower than that of the target band, the lower-quality sub-band is preferentially mapped to the higher-quality target band. The quality assessment parameter can be calculated based on the signal-to-noise ratio, total harmonic distortion, or a simplified form of the band scoring function. The harmonic matching principle aims to maintain the harmonic structure of the frequency-shifted audio signal, avoiding "robotic" or distortion. The adaptation coefficient can be a fixed value or an adjustable parameter used to balance harmonic matching and spectral efficiency.

[0207] S33. Harmonic conservation-type frequency shift calculation, which includes frequency shift compression ratio calculation, frequency domain shift, and harmonic correction. The frequency shift compression ratio calculation determines the frequency scaling ratio required when shifting a sub-band to the target band. This is achieved by using the ratio of the lowest to the highest frequency of both the sub-band and the target band, ensuring that the relative width of the entire band is maintained. Frequency domain shifting uses a linear interpolation algorithm to achieve precise spectrum shifting. The linear interpolation algorithm generates new spectrum data points by performing linear estimation between the original spectrum data points, thus achieving a smooth and continuous spectrum shift and avoiding spectrum breaks or distortions that may occur in traditional methods.

[0208] Through the above technical solutions, this invention quantifies the signal quality, human auditory adaptability, and interference suppression capability of sub-frequency bands by constructing a comprehensive frequency band scoring function. This allows for intelligent selection of the most suitable region as the target frequency band, effectively avoiding signal shifting to frequency bands with severe noise masking or low human auditory sensitivity, and significantly improving the perceptual adaptability of the shifted audio. Simultaneously, by introducing a quality-priority principle and a harmonic matching principle for dynamic mapping, it ensures that low-quality sub-frequency bands can be mapped to high-quality target frequency bands, maintaining the bandwidth ratio and harmonic spacing matching between the sub-frequency band and the target frequency band during the mapping process, effectively preventing the breakage of the harmonic structure after frequency shifting. Furthermore, through precise frequency shift compression ratio calculation, frequency domain shifting based on a linear interpolation algorithm, and key harmonic correction steps, this invention achieves precise spectrum shifting and actively corrects harmonic frequency deviations that may occur during the frequency shifting process. This fundamentally solves the problems of inaccurate spectrum shifting and harmonic relationship disruption in traditional frequency shifting methods, significantly reducing harmonic distortion in the shifted audio and ensuring audio clarity and naturalness. These synergistic steps work together to ensure that frequency shifting improves audio quality while preserving the auditory characteristics and naturalness of the original audio to the greatest extent possible.

[0209] In some of the embodiments of the present invention described above, audio signal reconstruction and auditory optimization are proposed to reconstruct the frequency-shifted signal and optimize auditory perception. However, in the process of implementation, inaccurate filtering may lead to signal distortion and sub-band merging errors, and improper gain adjustment may fail to adapt to the loudness curve characteristics of the human ear, affecting the naturalness and clarity of the audio.

[0210] In response, this invention further proposes audio signal reconstruction and auditory optimization steps, specifically including the following:

[0211] S41. Filter the frequency-domain signal after frequency shifting.

[0212] The filtering employs an improved quadrature mirror filter bank, including an analysis filter and a synthesis filter;

[0213] The analysis filter adopts a Kaiser window design FIR structure, and its cutoff frequency is calculated using the following formula:

[0214]

[0215] in, The cutoff frequency, This is the lower limit frequency of the i-th sub-band after dynamic partitioning. This represents the upper limit frequency of the (i+1)th sub-band after dynamic partitioning. Sampling rate;

[0216] The synthesized filter is generated based on the reconstruction conditions of the analysis filter using the following formula:

[0217]

[0218] in, This is the time-domain impulse response of the synthesized filter, used to reassemble the split sub-band signals into a complete time-domain signal. To analyze the time-domain impulse response of the filter, the input signal is decomposed into different sub-frequency bands; N is the filter order, and n is the index of the time-domain sampling point;

[0219] S42. Time-domain signal reconstruction

[0220] The filtered frequency domain signal is converted back to the time domain signal using the inverse short-time Fourier transform, and the reconstruction formula is as follows:

[0221]

[0222] in, The reconstructed time-domain signal is given, where t is the time-domain sampling point position, m is the frame index, and N is the number of FFT points. For frequency points, For the target frequency domain signal, The inverse Fourier transform kernel function. For the Hanning window function;

[0223] S43. Auditory Perception Gain Optimization

[0224] Based on the equal loudness curve, the gain value is dynamically adjusted according to frequency, and the gain function is:

[0225]

[0226] in, For frequency The gain value applied at that point, For the target loudness level, The reconstructed time-domain signal in frequency The measured loudness level at the location.

[0227] In step S41, the frequency-shifted frequency domain signal is filtered. This filtering process employs an improved quadrature mirror filter bank, which includes analysis filters and synthesis filters. The improved quadrature mirror filter bank is a filter bank specifically designed for subband decomposition and reconstruction, capable of achieving perfect or near-perfect signal reconstruction, effectively avoiding aliasing distortion and phase distortion that may be introduced by traditional filtering methods.

[0228] Analysis filters are used to decompose input signals into different sub-bands. The Kaiser window is a window function with adjustable parameters. By adjusting its shape parameters, the transition band width and stopband attenuation of the filter can be flexibly controlled, thereby achieving precise matching of the boundaries of dynamically divided sub-bands.

[0229] The synthesis filter is responsible for recombining the processed sub-band signals into a complete time-domain signal. Its design is closely related to the analysis filter to ensure distortion-free reconstructed signal. Besides the formula mentioned above, the synthesis filter can also be designed directly based on the polyphase components of the analysis filter, or generated by minimizing reconstruction errors through optimization algorithms. This precise filtering ensures accurate frequency band separation and merging of the frequency-domain signal after frequency shifting before reconstruction, significantly reducing the possibility of signal distortion and sub-band merging errors.

[0230] In step S42, the inverse short-time Fourier transform (ISFT) is a technique for converting a frequency domain representation (such as a spectrogram) back to a time domain waveform. It performs an IFT on each frequency domain frame and smoothly concatenates these short-time signals using the overlap-add (OLA) method to generate a continuous time-domain audio signal. Besides the overlap-add method, the overlap-save method can also be used for time-domain signal reconstruction. This step ensures that the frequency-shifted audio signal can be smoothly and continuously converted back to the time domain, maintaining the coherence of the original audio.

[0231] In step S43, auditory perception gain optimization is performed based on the equal-loudness curve. The equal-loudness curve describes the differences in perceived loudness of sounds at different frequencies; that is, the sound pressure level required to achieve the same perceived loudness varies at different frequencies. By dynamically adjusting the gain, this invention can compensate for the differences in auditory sensitivity of the human ear at different frequencies, making the reconstructed audio more balanced and natural to the ear. In addition to directly adjusting the gain based on the equal-loudness curve, multi-band compressors / expanders or perception-weighted filters can also be used to achieve gain optimization. In this way, the problem of improper gain adjustment leading to an inability to adapt to the characteristics of the human ear's equal-loudness curve can be effectively solved, significantly improving the naturalness and clarity of the audio.

[0232] In summary, by employing an improved orthogonal mirror filter bank for precise filtering, combined with inverse short-time Fourier transform for smooth time-domain signal reconstruction, and based on equal-loudness curves for dynamic auditory perception gain optimization, the audio signal reconstruction and auditory optimization scheme of this invention effectively solves the problems of signal distortion and sub-band merging errors caused by inaccurate filtering, as well as the impact of improper gain adjustment on audio naturalness and clarity. These techniques work synergistically to ensure that the frequency-shifted audio signal maintains high fidelity during reconstruction and ultimately presents an excellent perceptual effect that conforms to the characteristics of human hearing, thereby significantly improving the overall audio quality and user experience.

[0233] In this document, the terms "upper," "lower," "front," "back," "left," "right," "top," "bottom," "inner," "outer," "vertical," and "horizontal," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only used for the clarity of expressing the technical solution and for the convenience of description, and therefore should not be construed as limiting the present invention.

[0234] In this document, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.

[0235] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A frequency shifting processing method for audio dynamic frequency band mapping, characterized in that, Includes the following steps: S1. Preprocessing and Feature Quantization: The original audio is preprocessed, and the preprocessed original audio is converted from a time-domain signal to a frequency-domain signal through short-time Fourier transform. The fundamental frequency, harmonic frequency interval, auditory masking threshold, frequency band signal-to-noise ratio, and sensitive frequency band power ratio of the frequency-domain signal are extracted. S2. Dynamic Frequency Band Division: A decision function is constructed based on the fundamental frequency, frequency band signal-to-noise ratio, and sensitive frequency band power ratio. The frequency-domain signal is divided into several sub-frequency bands based on dynamic frequency band division rules. S3. Dynamic Mapping and Precise Frequency Shift Calculation: A frequency band scoring function is constructed based on the auditory masking threshold. The signal quality and perceptual adaptability of the sub-frequency bands are quantitatively scored. The target frequency band is selected from several sub-frequency bands, and the remaining sub-frequency bands are shifted to the target frequency band to obtain the frequency-shifted frequency-domain signal. Based on the harmonic frequency interval, the harmonic relationship of the frequency-shifted frequency-domain signal is maintained through a harmonic correction formula. S4. Audio signal reconstruction and auditory optimization: The frequency domain signal after frequency shifting is filtered, and the filtered frequency domain signal is converted back to the time domain signal through inverse short-time Fourier transform. The frequency band gain of the time domain signal is dynamically optimized through equal loudness curves.

2. The frequency shifting processing method for audio dynamic frequency band mapping according to claim 1, characterized in that, The preprocessing in S1 includes: analog-to-digital conversion: performing analog-to-digital conversion on the input audio signal so that the converted audio signal meets the Nyquist sampling criterion; low-pass filtering: using a low-pass filter to remove high-frequency noise and clutter, retaining the effective frequency band; pre-emphasis processing: improving the signal-to-noise ratio of the high-frequency signal through the following pre-emphasis formula to match the sensitivity of the human ear to high frequencies; in, The pre-emphasized output signal at time t The original audio sample value at time t, t is the original audio sample value at time t-1, and 0.97 is the pre-emphasis coefficient.

3. The frequency shifting processing method for audio dynamic frequency band mapping according to claim 2, characterized in that, The preprocessed original audio in S1 is framed using a Hanning window, and then converted from a time-domain signal to a frequency-domain signal using a short-time Fourier transform. The expression for this conversion is: in, This represents the complex amplitude value of the m-th frame at frequency domain index k. This represents summing over all sampled points within a frame. This represents the time-domain signal after pre-emphasis processing. This represents the Hanning window function. The Fourier transform kernel function is represented by N; N represents the number of sampling points in each frame of signal; and n is the index of the sampling point position in the current frame.

4. The frequency shifting processing method for audio dynamic frequency band mapping according to claim 1, characterized in that, In S1, the extraction formulas for the fundamental frequency, harmonic frequency interval, auditory masking threshold, frequency band signal-to-noise ratio, and sensitive frequency band power ratio of the frequency domain signal are as follows: The fundamental frequency is calculated using the autocorrelation function. in, For the fundamental frequency, for The main peak position, Here, T is the autocorrelation function value, and T is the frame length. The pre-emphasized output signal at time t The signal after time delay. For time delay, For time integral infinitesimals; harmonic frequency intervals equal to the fundamental frequency; auditory masking threshold: in, For auditory masking threshold, As the reference sound pressure level, This represents the critical frequency band offset. Noise level; Signal-to-noise ratio in the frequency band: in, For frequency band signal-to-noise ratio, For the signal power spectral density, Noise power spectral density; power percentage in sensitive frequency bands: in, The numerator represents the signal power within the frequency range of 1000Hz to 4000Hz, while the denominator represents the total signal power across the entire frequency band.

5. The frequency shifting processing method for audio dynamic frequency band mapping according to claim 1, characterized in that, In S2, the expression for the decision function is: in, Let be the decision function. For the fundamental frequency, To preset the threshold for fundamental frequency variation, For the power ratio of sensitive frequency bands, This represents the global average signal-to-noise ratio. The maximum signal-to-noise ratio threshold is defined by α, β, and γ, which are weighting coefficients, K is the total number of frequency points, and k is the frequency index. This represents the signal-to-noise ratio at a single frequency point.

6. The frequency shifting processing method for audio dynamic frequency band mapping according to claim 1, characterized in that, The dynamic frequency band division rules in S2 are as follows: (a) Adaptive adjustment rule for the number of sub-bands: The number of sub-bands M∈{3,4,5,6} is determined by the value of the decision function D: when D<0.6, M=3; when 0.6≤D<1, M=4; when D≥1, M=5 or 6; (b) The boundaries of the sub-bands adopt the exponential boundary formula: in, Let be the upper limit frequency of the (i+1)th sub-band. (c) Interference scenario optimization rules: If a certain sub-band satisfies Then, the sub-band will be divided into two narrowband sub-bands; among which, For interference power, (d) High-frequency enhancement rule: When the signal power ratio of the 8–20 kHz high-frequency band is >0.3, the bandwidth of the high-frequency band is extended.

7. The frequency shifting processing method for audio dynamic frequency band mapping according to claim 1, characterized in that, The S3. Dynamic Mapping and Precise Frequency Shift Calculation step includes: S31. Target Frequency Band Scoring and Screening. Constructing a frequency band scoring function, selecting several sub-frequency bands with higher scores as target frequency bands, the expression of which is as follows: in, For frequency band scoring functions, For signal quality weights, For frequency point signal-to-noise ratio, The maximum signal-to-noise ratio threshold. To mask the adaptation weights, For auditory masking threshold, For the signal power spectral density, To suppress interference weights, For interference power; S32. Sub-band-target band dynamic mapping quality priority principle: will satisfy The low-quality sub-bands are mapped to the high-quality target bands. ;in, These are the quality assessment parameters for sub-bands. These are the quality assessment parameters for the target frequency band; Harmonic matching principle: The ratio of the bandwidth of the sub-band to the bandwidth of the target frequency band must satisfy: in, For the bandwidth of the sub-band, For the bandwidth of the target frequency band, For signal harmonic intervals, For the adaptation coefficient; S33. Harmonic conservation type frequency shift calculation: frequency shift compression ratio calculation: based on the frequency range of the sub-band and the target frequency band, the formula is: in, This is the frequency shift compression ratio. and These are the lowest and highest frequencies of the sub-band, respectively. and These are the lowest and highest frequencies of the target frequency band, respectively; Frequency domain shift: A linear interpolation algorithm is used to achieve precise spectrum shifting, the formula is: in, For the mapped signal at the frequency point The amplitude and phase, For the original signal at frequency points Amplitude and phase information, The starting frequency domain index for the target frequency band. This is the starting frequency domain index for the sub-band; Harmonic correction: To maintain the harmonic relationships of the audio after frequency shifting, the harmonic frequencies are corrected using the harmonic correction formula, which is: in, The frequency of the nth harmonic in the target frequency band. Let n be the fundamental frequency of the target frequency band, and n be the harmonic order. This represents the harmonic frequency interval.

8. The frequency shifting processing method for audio dynamic frequency band mapping according to claim 1, characterized in that, The S4. audio signal reconstruction and auditory optimization step includes: S41. Filtering the frequency-shifted frequency domain signal. The filtering adopts an improved orthogonal mirror filter bank, including an analysis filter and a synthesis filter; the analysis filter adopts a Kaiser window design FIR structure, and its cutoff frequency is calculated using the following formula: in, The cutoff frequency, This is the lower limit frequency of the i-th sub-band after dynamic partitioning. This represents the upper limit frequency of the (i+1)th sub-band after dynamic partitioning. The sampling rate is used; the synthesized filter is generated based on the reconstruction conditions of the analysis filter using the following formula: in, This is the time-domain impulse response of the synthesized filter, used to reassemble the split sub-band signals into a complete time-domain signal. To analyze the time-domain impulse response of the filter, the input signal is decomposed into different sub-frequency bands; N is the filter order, and n is the index of the time-domain sampling point; S42. Time-domain signal reconstruction converts the filtered frequency-domain signal back to the time-domain signal using the inverse short-time Fourier transform. The reconstruction formula is: in, The reconstructed time-domain signal is given, where t is the time-domain sampling point position, m is the frame index, and N is the number of FFT points. For frequency points, For the target frequency domain signal, The inverse Fourier transform kernel function. The Hanning window function is used; S43. Auditory perception gain optimization: Based on the aforementioned equal loudness curves, the gain value is dynamically adjusted according to frequency. The gain function is: in, For frequency The gain value applied at that point, For the target loudness level, The reconstructed time-domain signal in frequency The measured loudness level at the location.