High-frequency compensation and sound field broadening method and system based on audio content

By using time-frequency analysis and dynamic filter generation, combined with sound image localization stability verification, the problems of sound quality degradation and unstable sound image localization in audio processing are solved. This achieves synergistic effects of high-frequency compensation and sound field widening, thereby enhancing the auditory experience of audio content.

CN121789699APending Publication Date: 2026-04-03JINXUAN ELECTRONICS (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing audio processing technologies are prone to sound quality degradation and unstable sound image positioning in high-frequency compensation and sound field widening, making it difficult to balance enhancement effects and sound quality fidelity.

Method used

By extracting the high-frequency energy distribution, spectral sparsity, and spatial cue features of the audio signal through time-frequency analysis, a high-frequency compensation filter is dynamically generated. Combined with a sound image localization stability feedback verification mechanism, multi-channel signal processing and time-frequency domain weighted fusion are performed.

Benefits of technology

It achieves synergistic effects of high-frequency compensation and sound field widening, ensuring sound quality fidelity and stable sound image positioning, and enhancing the naturalness and immersion of the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789699A_ABST
    Figure CN121789699A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital signal processing, and discloses a high-frequency compensation and sound field broadening method and system based on audio content, and the method comprises the steps: carrying out the time-frequency analysis of an input audio signal, so as to extract the high-frequency energy distribution feature, the spectrum sparsity feature and the sound field space clue feature of the input audio signal, dynamically generating a high-frequency compensation filter matched with the content of the input audio signal according to the high-frequency energy distribution characteristic and the spectrum sparsity characteristic so as to perform content perception compensation on the high-frequency component in the input audio signal to obtain a compensation signal; generating a multi-channel output signal in which the input audio signal has a spatial expansion sense; calculating the sound image positioning stability of the multi-channel output signal; and when the sound image positioning stability accords with a preset sound image positioning stability threshold value, performing time-frequency domain weighted fusion on the multi-channel output signal and the compensation signal to output a target audio signal. According to the invention, the balance degree between the enhancement effect and the tone quality fidelity of the audio content can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for high-frequency compensation and sound field widening based on audio content, belonging to the field of digital signal processing technology. Background Technology

[0002] High-frequency compensation and sound field widening methods for audio content refer to the use of intelligent signal processing technology to adaptively restore or enhance high-frequency details and spatial information lost due to recording, compression, and other reasons, based on the characteristics of different audio materials. The aim is to solve the defects of unnatural sound quality, flat sound field, or blurred positioning in traditional processing. Its core significance lies in breaking through the technical contradiction that enhancement equals distortion. While maintaining the original mixing intention and stable sound image positioning, it significantly improves the brightness, clarity, and immersion of the listening experience. It is especially suitable for scenarios such as streaming media audio optimization, restoration of old recordings, and enhancement of sound effects in consumer electronics. It is a key technical path to achieve the unity of high fidelity and presence.

[0003] High-frequency compensation often uses preset equalization curves, which can easily over-amplify sound sources containing noise or distortion, resulting in harsh listening or degraded sound quality. Sound field widening often relies on simple phase inversion, delay, or static head correlation function processing, which can easily destroy the original sound image localization when expanding the sound field, causing the sound field to become hollow or the sound localization to drift, affecting the naturalness of the hearing, thus resulting in a poor balance between enhancement effect and sound quality fidelity. Summary of the Invention

[0004] This invention provides a method and system for high-frequency compensation and sound field widening based on audio content, the main purpose of which is to improve the balance between the enhancement effect and sound quality fidelity of audio content.

[0005] To achieve the above objectives, the present invention provides a high-frequency compensation and sound field broadening method based on audio content, comprising: Time-frequency analysis is performed on the input audio signal to extract its high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics. Based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, a high-frequency compensation filter that matches the content of the input audio signal is dynamically generated to perform content-aware compensation on the high-frequency components in the input audio signal, thereby obtaining a compensation signal. Based on the aforementioned sound field spatial cue features, a multi-channel output signal with a sense of spatial expansion is generated from the input audio signal; Calculate the acoustic image localization stability of the multi-channel output signal; When the sound image localization stability meets the preset sound image localization stability threshold, the multi-channel output signal and the compensation signal are fused in the time and frequency domain to output the target audio signal.

[0006] Optionally, based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, a high-frequency compensation filter matching the content of the input audio signal is dynamically generated, including: Using the high-frequency energy distribution characteristics and spectral sparsity characteristics as inputs, the target frequency response curve for high-frequency compensation is determined through a predefined feature-parameter mapping function; Define a gain constraint factor that is negatively correlated with the aforementioned spectral sparsity characteristics; Based on the target frequency response curve and the gain constraint factor, the coefficient vector of the digital filter is calculated to generate a high-frequency compensation filter that matches the content of the input audio signal, wherein the digital filter is selected from peak filters.

[0007] Optionally, calculating the coefficient vector of the digital filter based on the target frequency response curve and the gain constraint factor includes: The target frequency response curve under the gain constraint factor is calculated using the following formula:

[0008] in, This represents the constrained frequency response curve under the gain constraint factor. This represents the gain constraint factor. Represents the target frequency response curve; Based on the constrained frequency response curve, the quality factor of the digital filter is calculated to calculate the transfer function of the digital filter. The coefficient vector of the digital filter is calculated based on the transfer function.

[0009] Optionally, the quality factor of the digital filter is calculated based on the constrained frequency response curve, including: Extract the key parameters used to define the digital filter from the constrained frequency response curve, wherein the key parameters include the center frequency, the peak gain at the center frequency, and the target bandwidth; Based on the center frequency, the peak gain at the center frequency, and the target bandwidth, the quality factor of the digital filter is calculated using the following formula:

[0010] in, The quality factor represents the digital filter. Indicates the center frequency. Center frequency Peak gain at that point Indicates the target bandwidth.

[0011] Optionally, time-frequency analysis is performed on the input audio signal to extract high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics of the input audio signal, including: By performing time-frequency transformation on the input audio signal, the complex spectrum of the input audio signal is obtained, and the amplitude spectrum of each frequency point in the complex spectrum is further calculated. From the amplitude spectrum, extract the high-frequency energy distribution features and spectral sparsity features of the input audio signal; From the complex spectrum, the acoustic field spatial cue features of the input audio signal are extracted.

[0012] Optionally, extracting the high-frequency energy distribution features and spectral sparsity features of the input audio signal from the amplitude spectrum includes: Based on the amplitude spectrum, the sum of squares of the amplitudes at each frequency point within the predefined high-frequency sub-band range is calculated as the signal energy of the high-frequency sub-band. The signal energy is then normalized and compared with the energy of the reference frequency band to obtain the high-frequency energy distribution characteristics. Based on the amplitude spectrum, the amplitude data sequence within the high-frequency sub-band is extracted, and the kurtosis statistics of the amplitude data sequence are calculated as a spectral sparsity feature of the input audio signal.

[0013] Optionally, extracting the sound field spatial cue features of the input audio signal from the complex spectrum includes: For the complex spectrum of the multi-channel signal, the phase difference and amplitude ratio between channels corresponding to the same frequency point are calculated, and then aggregated within the psychoacoustic critical band to obtain the sound field spatial cue features.

[0014] Optionally, based on the spatial cue features of the sound field, a multi-channel output signal with a sense of spatial expansion of the input audio signal is generated, including: Define the binaural cue preservation constraints for the input audio signal; The spatial cue features of the sound field are analyzed to obtain the target sound field width parameter and the target lateral sound energy ratio parameter, so as to perform adaptive spatial rendering on the input audio signal under the constraint of maintaining the binaural cue, so as to generate the spatial extension signal of the input audio signal; The spatially extended signal is synthesized into a multi-channel output signal in stereo format with left and right channels.

[0015] Optionally, calculating the acoustic image localization stability of the multi-channel output signal includes: Extract binaural cue features of the multi-channel output signal within the psychoacoustic critical frequency band, wherein the binaural cue features include the processed inter-channel phase difference and the processed inter-channel amplitude ratio; The binaural cue offset of the psychoacoustic critical band is calculated based on the binaural cue features to calculate the acoustic image localization stability of the multi-channel output signal. The binaural cue offset includes phase difference offset and amplitude ratio offset.

[0016] To address the aforementioned problems, the present invention also provides a high-frequency compensation and sound field broadening system based on audio content, the system comprising: The signal time-frequency analysis module is used to perform time-frequency analysis on the input audio signal to extract the high-frequency energy distribution characteristics, spectral sparsity characteristics and sound field spatial cue characteristics of the input audio signal; The high-frequency signal compensation module is used to dynamically generate a high-frequency compensation filter that matches the content of the input audio signal based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, so as to perform content-aware compensation on the high-frequency components in the input audio signal and obtain a compensation signal. A multi-channel signal output module is used to generate a multi-channel output signal with a sense of spatial expansion of the input audio signal based on the spatial cue features of the sound field; The acoustic image stability analysis module is used to calculate the acoustic image localization stability of the multi-channel output signal; The target audio signal output module is used to perform time-frequency domain weighted fusion of the multi-channel output signal and the compensation signal to output the target audio signal when the sound image positioning stability meets the preset sound image positioning stability threshold.

[0017] Firstly, regarding audio fidelity, this invention, through comprehensive analysis of high-frequency energy distribution and spectral sparsity characteristics, can accurately distinguish harmonic structures and noise components in the signal. This ensures that the generated high-frequency compensation filter is no longer blindly boosted, while effectively suppressing the amplification of noise and compression artifacts. Thus, while restoring the high-frequency airiness and detail, it completely overcomes the metallic and plastic-like distortions caused by traditional methods, making the sound more natural and realistic. Secondly, regarding sound field construction, this invention expands the space based on sound field spatial cues and creatively introduces a feedback verification mechanism for sound image localization stability. This design precisely solves the problems of existing sound field widening technologies. The core contradiction of "a wide center leads to emptiness, and emptiness leads to dispersion" is addressed by this invention. It ensures that while pursuing a grand and immersive experience, the sound image positions of core elements such as vocals and bass drums do not shift or become blurred, avoiding issues like a floating or hollow sound. More importantly, this inherent guarantee of stability naturally ensures the compatibility of the output signal in a mono system, preventing sound loss due to phase cancellation and greatly expanding the application scenarios of this invention. Finally, through time-frequency domain weighted fusion, this invention organically combines high-frequency compensation enhanced for fidelity with sound field widening verified for stability, dynamically adjusting the weights of compensation and widening to achieve synergistic effects rather than mutual interference. Therefore, this invention can improve the balance between enhancement effects and sound quality fidelity in audio content. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a method for high-frequency compensation and sound field broadening based on audio content, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the generation of a multi-channel output signal that achieves high-frequency compensation and sound field widening based on audio content, according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the modules for implementing the high-frequency compensation and sound field widening method based on audio content according to an embodiment of the present invention; Figure 4 A schematic diagram of a computer device for a high-frequency compensation and sound field broadening method based on audio content provided in an embodiment of the present invention; The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0020] This application provides a method for high-frequency compensation and sound field widening based on audio content. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for high-frequency compensation and sound field widening based on audio content can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.

[0021] Reference Figure 1 The diagram shown is a flowchart illustrating a method for high-frequency compensation and sound field broadening based on audio content according to an embodiment of the present invention. In this embodiment, the method for high-frequency compensation and sound field broadening based on audio content includes: S1. Perform time-frequency analysis on the input audio signal to extract the high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics of the input audio signal.

[0022] This invention performs time-frequency analysis on the input audio signal to extract the high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial clue characteristics of the input audio signal, providing a data basis for subsequent signal compensation and sound field widening.

[0023] In detail, the step of performing time-frequency analysis on the input audio signal to extract the high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics of the input audio signal includes: By performing time-frequency transformation on the input audio signal, the complex spectrum of the input audio signal is obtained, and the amplitude spectrum of each frequency point in the complex spectrum is further calculated. From the amplitude spectrum, extract the high-frequency energy distribution features and spectral sparsity features of the input audio signal; From the complex spectrum, the acoustic field spatial cue features of the input audio signal are extracted.

[0024] Wherein, the complex spectrum refers to the frequency domain representation obtained by performing a time-frequency transformation on a time-domain audio frame; each frequency point refers to the discretized index of the complex spectrum on the frequency axis; the amplitude spectrum refers to the spectrum derived from the complex spectrum that contains only the amplitude information of each frequency point; the high-frequency energy distribution characteristics refer to the characteristics characterizing the relative intensity of the high-frequency components of the signal; the spectral sparsity characteristics refer to the characteristic values ​​characterizing whether the energy of the high-frequency components of the signal is concentrated in a few discrete frequency points or widely distributed; and the sound field spatial cue characteristics refer to the characteristics characterizing the apparent direction, width, and spatial diffusion of the sound source. Further, extracting the high-frequency energy distribution features and spectral sparsity features of the input audio signal from the amplitude spectrum includes: Based on the amplitude spectrum, the sum of squares of the amplitudes at each frequency point within the predefined high-frequency sub-band range is calculated as the signal energy of the high-frequency sub-band. The signal energy is then normalized and compared with the energy of the reference frequency band to obtain the high-frequency energy distribution characteristics. Based on the amplitude spectrum, the amplitude data sequence within the high-frequency sub-band is extracted, and the kurtosis statistics of the amplitude data sequence are calculated as a spectral sparsity feature of the input audio signal.

[0025] Wherein, the high-frequency sub-band refers to a predefined continuous or discontinuous frequency range on the frequency axis of the amplitude spectrum; the sum of squares refers to squaring the amplitude of each frequency point within the high-frequency sub-band and then summing all these squared values; the signal energy refers to a scalar value obtained by calculating the sum of squares of the amplitudes of each frequency point within the high-frequency sub-band, used to quantify the strength of the signal within the frequency band; the reference frequency band energy refers to a benchmark energy value used for comparison with the signal energy of the high-frequency sub-band; the amplitude data sequence refers to a one-dimensional data sequence composed of the amplitudes of all frequency points within the high-frequency sub-band arranged in frequency order; and the kurtosis statistic refers to a statistical moment used to describe the shape of the probability distribution.

[0026] Further, extracting the sound field spatial cue features of the input audio signal from the complex spectrum includes: For the complex spectrum of the multi-channel signal, the phase difference and amplitude ratio between channels corresponding to the same frequency point are calculated, and then aggregated within the psychoacoustic critical band to obtain the sound field spatial cue features.

[0027] Wherein, "channel-to-channel" refers to the different channels of a multi-channel audio signal, such as the left channel and the right channel; "same frequency point" refers to the frequency point with the same frequency index in the complex spectrum of different channels when comparing channels; "phase difference" refers to the difference between the phase angle of the complex spectrum of the left channel and the phase angle of the complex spectrum of the right channel at the same frequency point; "amplitude ratio" refers to the ratio between the amplitude of the amplitude spectrum of the left channel and the amplitude spectrum of the right channel at the same frequency point; and "psychoacoustic critical band" refers to the frequency band on a non-uniform frequency scale divided based on the frequency resolution characteristics of the human auditory system.

[0028] S2. Based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, a high-frequency compensation filter that matches the content of the input audio signal is dynamically generated to perform content-aware compensation on the high-frequency components in the input audio signal, thereby obtaining a compensation signal.

[0029] Based on the aforementioned high-frequency energy distribution characteristics and spectral sparsity characteristics, this invention dynamically generates a high-frequency compensation filter that matches the content of the input audio signal, enabling precise differentiation between harmonic structures and noise components in the signal. This ensures that the generated high-frequency compensation filter no longer blindly boosts the signal, but rather performs content-aware, refined compensation. It intelligently injects energy into missing harmonic details in the music while effectively suppressing the amplification of noise and compression artifacts, thereby restoring high-frequency airiness and detail.

[0030] Specifically, the step of dynamically generating a high-frequency compensation filter that matches the content of the input audio signal based on the high-frequency energy distribution characteristics and spectral sparsity characteristics includes: Using the high-frequency energy distribution characteristics and spectral sparsity characteristics as inputs, the target frequency response curve for high-frequency compensation is determined through a predefined feature-parameter mapping function; Define a gain constraint factor that is negatively correlated with the aforementioned spectral sparsity characteristics; Based on the target frequency response curve and the gain constraint factor, the coefficient vector of the digital filter is calculated to generate a high-frequency compensation filter that matches the content of the input audio signal, wherein the digital filter is selected from peak filters.

[0031] The feature-parameter mapping function is a predefined mathematical function that maps the high-frequency energy distribution characteristics and spectral sparsity characteristics of the input signal to a set of filter target parameters. The target frequency response curve is a desired amplitude response curve defined in the frequency domain, indicating the ideal gain value that the filter should provide at different frequencies to compensate for the high-frequency loss of the input signal. The gain constraint factor is a scalar coefficient between 0 and 1, whose value is dynamically calculated based on the spectral sparsity characteristics. When the high-frequency components of the signal are detected to be sparse, the value is close to 1, allowing sufficient compensation; when the spectrum is detected to be dense, the value becomes smaller, used to proportionally attenuate the target gain to suppress distortion and noise amplification. The coefficient vector is a set of parameters used to define the difference equation of a digital filter. The high-frequency compensation filter is a digital filter dynamically generated according to the aforementioned steps, used to selectively boost the gain of a specific high-frequency band of the input audio signal. The peak filter is a special type of digital filter whose amplitude response presents a convex peak near a certain center frequency, used to accurately boost or attenuate this narrow frequency region, and is the preferred filter type for achieving selective high-frequency compensation.

[0032] Further, calculating the coefficient vector of the digital filter based on the target frequency response curve and the gain constraint factor includes: The target frequency response curve under the gain constraint factor is calculated using the following formula:

[0033] in, This represents the constrained frequency response curve under the gain constraint factor. This represents the gain constraint factor. Represents the target frequency response curve; Based on the constrained frequency response curve, the quality factor of the digital filter is calculated to calculate the transfer function of the digital filter. The coefficient vector of the digital filter is calculated based on the transfer function.

[0034] Here, the constrained frequency response curve refers to the corrected target frequency response function, and the quality factor is a dimensionless parameter used to characterize the frequency selectivity of the filter. For peak filters, a higher quality factor indicates that its frequency response curve is sharper and has a narrower bandwidth near the center frequency; a lower value indicates a flatter response and a wider bandwidth. The transfer function is a mathematical function in digital filters that describes the ratio of the output signal transformation to the input signal transformation.

[0035] Furthermore, calculating the quality factor of the digital filter based on the constrained frequency response curve includes: Extract the key parameters used to define the digital filter from the constrained frequency response curve, wherein the key parameters include the center frequency, the peak gain at the center frequency, and the target bandwidth; Based on the center frequency, the peak gain at the center frequency, and the target bandwidth, the quality factor of the digital filter is calculated using the following formula:

[0036] in, The quality factor represents the digital filter. Indicates the center frequency. Center frequency Peak gain at that point Indicates the target bandwidth.

[0037] Wherein, the center frequency refers to the frequency value corresponding to the maximum gain value obtained by the constraint frequency response curve, the peak gain refers to the maximum gain that can be achieved at the center frequency point, and the target bandwidth refers to the frequency range that reflects the main effect of the filter.

[0038] Optionally, the transfer function of the digital filter can be calculated using the bilinear transform method.

[0039] This invention performs content-aware compensation on the high-frequency components of the input audio signal, resulting in a compensated signal that enhances the realism of the sound. Specifically, the content-aware compensation of the high-frequency components in the input audio signal can be achieved by filtering the high-frequency components using a dynamically generated high-frequency compensation filter in the time domain. The compensated signal refers to the output signal obtained after processing the input audio signal through the dynamically generated high-frequency compensation filter.

[0040] S3. Based on the spatial cue features of the sound field, generate a multi-channel output signal with a sense of spatial expansion of the input audio signal.

[0041] Based on the spatial cue features of the sound field, this invention generates a multi-channel output signal with a sense of spatial expansion of the input audio signal, which effectively widens the sound field while maintaining the sound image positioning intent of the original mix to the greatest extent.

[0042] Specifically, generating a multi-channel output signal with a sense of spatial expansion of the input audio signal based on the spatial cue features of the sound field includes: Define the binaural cue preservation constraints for the input audio signal; The spatial cue features of the sound field are analyzed to obtain the target sound field width parameter and the target lateral sound energy ratio parameter, so as to perform adaptive spatial rendering on the input audio signal under the constraint of maintaining the binaural cue, so as to generate the spatial extension signal of the input audio signal; The spatially extended signal is synthesized into a multi-channel output signal in stereo format with left and right channels.

[0043] The binaural cue preservation constraint refers to the constraint on limiting the variation of the binaural time difference and binaural level difference of the processed audio signal relative to the corresponding features of the original input signal within the key psychoacoustic critical frequency band, ensuring that the variation is below the auditory threshold value for perceptible changes in sound image localization. The target sound field width parameter is used to quantify the desired subjective expansion of the sound field. The target lateral sound energy ratio parameter is the target ratio of the energy attributable to the lateral sound field to the total sound field energy in the desired output signal after spatial rendering processing, used to control the sense of enclosure and diffusion of the sound field. The spatial expansion signal refers to the intermediate audio signal obtained after the input audio signal has undergone the adaptive spatial rendering processing. The multi-channel output signal refers to the stereo format signal with independent left and right channels obtained after format conversion and allocation of the spatial expansion signal.

[0044] Optionally, in the step of analyzing the sound field spatial cue features to obtain the target sound field width parameter and the target lateral sound energy ratio parameter, the target sound field width parameter is obtained by dynamic mapping based on the part of the sound field spatial cue features that represents the original spatial width, and the target lateral sound energy ratio parameter is dynamically determined based on the part of the sound field spatial cue features that represents the sound source directionality and spatial diffusion.

[0045] Optionally, the adaptive spatial rendering of the input audio signal under the binaural cue preservation constraint can be achieved through HRTF convolution.

[0046] See Figure 2 The diagram illustrates the generation of a multi-channel output signal based on audio content-based high-frequency compensation and sound field widening, according to an embodiment of the present invention. First, the input audio signal is format-determined, and the appropriate processing path is selected based on the number of channels. Multi-channel signals are intelligently mixed into dual-channel signals via downmixing / format conversion. Dual-channel signals are directly transmitted. Single-channel signals are processed using virtual stereo to generate dual-channel signals, followed by level balancing and phase alignment to ensure loudness coordination and stable sound image positioning between the dual channels. Finally, a stereo format multi-channel output signal compatible with playback devices is output.

[0047] S4. Calculate the acoustic image localization stability of the multi-channel output signal.

[0048] The present invention introduces a feedback verification mechanism for the sound image positioning stability in calculating the sound image positioning stability of the multi-channel output signal, ensuring that while pursuing a grand and immersive experience, the sound image positions of core elements such as vocals and bass drum will not shift or become blurred.

[0049] In detail, the calculation of the acoustic image localization stability of the multi-channel output signal includes: Extract binaural cue features of the multi-channel output signal within the psychoacoustic critical frequency band, wherein the binaural cue features include the processed inter-channel phase difference and the processed inter-channel amplitude ratio; The binaural cue offset of the psychoacoustic critical band is calculated based on the binaural cue features to calculate the acoustic image localization stability of the multi-channel output signal. The binaural cue offset includes phase difference offset and amplitude ratio offset.

[0050] Wherein, the processed inter-channel phase difference refers to the difference between the average phase angle of the complex spectrum of the left channel and the average phase angle of the complex spectrum of the right channel within a certain psychoacoustic critical frequency band after time-frequency transformation of the multi-channel output signal; the processed inter-channel amplitude ratio refers to the square root of the ratio of the signal energy of the left channel to the signal energy of the right channel within a certain psychoacoustic critical frequency band after time-frequency transformation of the multi-channel output signal; the phase difference offset refers to the magnitude of the change in the processed inter-channel phase difference relative to the original phase difference of the corresponding critical frequency band extracted from the original input audio signal; the amplitude ratio offset refers to the magnitude of the change in the processed inter-channel amplitude ratio relative to the original amplitude ratio of the corresponding critical frequency band extracted from the original input audio signal; and the sound image localization stability refers to whether the spatial audio processing process maintains the stability of the original sound image localization in subjective hearing.

[0051] Optionally, the calculation of the acoustic image localization stability of the multi-channel output signal can be obtained by perceptually weighting the phase difference offset and amplitude ratio offset of all psychoacoustic critical frequency bands.

[0052] S5. When the sound image positioning stability meets the preset sound image positioning stability threshold, the multi-channel output signal and the compensation signal are fused in the time and frequency domain to output the target audio signal.

[0053] When the sound image localization stability meets a preset sound image localization stability threshold, this invention performs time-frequency domain weighted fusion of the multi-channel output signal and the compensation signal. This allows the output target audio signal to dynamically adjust the weights of compensation and widening according to the audio content's needs at different times and frequencies, achieving synergistic effects rather than mutual interference. The sound image localization stability threshold refers to a critical value used to determine whether the sound image localization stability is acceptable. The target audio signal refers to the final output audio signal after processing by the complete method described in this invention. The target audio signal is a discrete sequence in the time domain, and its perceptual characteristics simultaneously possess the following two enhancement effects: 1. The high-frequency components of the target audio signal receive compensation that matches the content and improves fidelity; 2. The sound field of the target audio signal gains a perceptible sense of spatial expansion while maintaining the original sound image localization stability.

[0054] Firstly, regarding audio fidelity, this invention, through comprehensive analysis of high-frequency energy distribution and spectral sparsity characteristics, can accurately distinguish harmonic structures and noise components in the signal. This ensures that the generated high-frequency compensation filter is no longer blindly boosted, while effectively suppressing the amplification of noise and compression artifacts. Thus, while restoring the high-frequency airiness and detail, it completely overcomes the metallic and plastic-like distortions caused by traditional methods, making the sound more natural and realistic. Secondly, regarding sound field construction, this invention expands the space based on sound field spatial cues and creatively introduces a feedback verification mechanism for sound image localization stability. This design precisely solves the problems of existing sound field widening technologies. The core contradiction of "a wide center leads to emptiness, and emptiness leads to dispersion" is addressed by this invention. It ensures that while pursuing a grand and immersive experience, the sound image positions of core elements such as vocals and bass drums do not shift or become blurred, avoiding issues like a floating or hollow sound. More importantly, this inherent guarantee of stability naturally ensures the compatibility of the output signal in a mono system, preventing sound loss due to phase cancellation and greatly expanding the application scenarios of this invention. Finally, through time-frequency domain weighted fusion, this invention organically combines high-frequency compensation enhanced for fidelity with sound field widening verified for stability, dynamically adjusting the weights of compensation and widening to achieve synergistic effects rather than mutual interference. Therefore, this invention can improve the balance between enhancement effects and sound quality fidelity in audio content.

[0055] like Figure 3 The diagram shown is a functional block diagram of the high-frequency compensation and sound field widening system based on audio content of the present invention.

[0056] The high-frequency compensation and sound field widening system 300 based on audio content described in this invention can be installed in an electronic device. Depending on the functions implemented, the high-frequency compensation and sound field widening system based on audio content may include a signal time-frequency analysis module 301, a high-frequency signal compensation module 302, a multi-channel signal output module 303, a sound image stability analysis module 304, and a target audio signal output module 305. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.

[0057] In this embodiment of the invention, the functions of each module / unit are as follows: The signal time-frequency analysis module 301 is used to perform time-frequency analysis on the input audio signal to extract the high-frequency energy distribution characteristics, spectral sparsity characteristics and sound field spatial cue characteristics of the input audio signal. The high-frequency signal compensation module 302 is used to dynamically generate a high-frequency compensation filter that matches the content of the input audio signal based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, so as to perform content-aware compensation on the high-frequency components in the input audio signal and obtain a compensation signal. The multi-channel signal output module 303 is used to generate a multi-channel output signal with a sense of spatial expansion of the input audio signal based on the spatial cue features of the sound field. The acoustic image stability analysis module 304 is used to calculate the acoustic image positioning stability of the multi-channel output signal; The target audio signal output module 305 is used to perform time-frequency domain weighted fusion of the multi-channel output signal and the compensation signal to output the target audio signal when the sound image positioning stability meets the preset sound image positioning stability threshold.

[0058] In detail, the modules in the high-frequency compensation and sound field widening system 300 based on audio content described in this embodiment of the invention employ the same methods as described above. Figure 1 The high-frequency compensation and sound field widening methods based on audio content described herein are the same techniques used and can produce the same technical effects, so they will not be elaborated here.

[0059] In one embodiment, a computer device is provided, which may be a server or a client, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements functions or steps on the server or client side of a high-frequency compensation and sound field broadening method based on audio content.

[0060] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Time-frequency analysis is performed on the input audio signal to extract its high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics. Based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, a high-frequency compensation filter that matches the content of the input audio signal is dynamically generated to perform content-aware compensation on the high-frequency components in the input audio signal, thereby obtaining a compensation signal. Based on the aforementioned sound field spatial cue features, a multi-channel output signal with a sense of spatial expansion is generated from the input audio signal; Calculate the acoustic image localization stability of the multi-channel output signal; When the sound image localization stability meets the preset sound image localization stability threshold, the multi-channel output signal and the compensation signal are fused in the time and frequency domain to output the target audio signal.

[0061] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Time-frequency analysis is performed on the input audio signal to extract its high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics. Based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, a high-frequency compensation filter that matches the content of the input audio signal is dynamically generated to perform content-aware compensation on the high-frequency components in the input audio signal, thereby obtaining a compensation signal. Based on the aforementioned sound field spatial cue features, a multi-channel output signal with a sense of spatial expansion is generated from the input audio signal; Calculate the acoustic image localization stability of the multi-channel output signal; When the sound image localization stability meets the preset sound image localization stability threshold, the multi-channel output signal and the compensation signal are fused in the time and frequency domain to output the target audio signal.

[0062] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0063] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0064] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0065] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0066] Finally, it should be noted that in the above embodiments, each embodiment can be combined with each other or independent. Deleting any one of them will not affect the technical implementation of other embodiments. The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for high-frequency compensation and sound field broadening based on audio content, characterized in that, The method includes: Time-frequency analysis is performed on the input audio signal to extract its high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics. Based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, a high-frequency compensation filter that matches the content of the input audio signal is dynamically generated to perform content-aware compensation on the high-frequency components in the input audio signal, thereby obtaining a compensation signal. Based on the aforementioned sound field spatial cue features, a multi-channel output signal with a sense of spatial expansion is generated from the input audio signal; Calculate the acoustic image localization stability of the multi-channel output signal; When the sound image localization stability meets the preset sound image localization stability threshold, the multi-channel output signal and the compensation signal are fused in the time and frequency domain to output the target audio signal.

2. The high-frequency compensation and sound field broadening method based on audio content as described in claim 1, characterized in that, Based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, a high-frequency compensation filter that matches the content of the input audio signal is dynamically generated, including: Using the high-frequency energy distribution characteristics and spectral sparsity characteristics as inputs, the target frequency response curve for high-frequency compensation is determined through a predefined feature-parameter mapping function; Define a gain constraint factor that is negatively correlated with the aforementioned spectral sparsity characteristics; Based on the target frequency response curve and the gain constraint factor, the coefficient vector of the digital filter is calculated to generate a high-frequency compensation filter that matches the content of the input audio signal, wherein the digital filter is selected from peak filters.

3. The high-frequency compensation and sound field broadening method based on audio content as described in claim 2, characterized in that, The step of calculating the coefficient vector of the digital filter based on the target frequency response curve and the gain constraint factor includes: The target frequency response curve under the gain constraint factor is calculated using the following formula: in, This represents the constrained frequency response curve under the gain constraint factor. This represents the gain constraint factor. Represents the target frequency response curve; Based on the constrained frequency response curve, the quality factor of the digital filter is calculated to calculate the transfer function of the digital filter. The coefficient vector of the digital filter is calculated based on the transfer function.

4. The high-frequency compensation and sound field broadening method based on audio content as described in claim 3, characterized in that, Based on the constrained frequency response curve, the quality factor of the digital filter is calculated, including: Extract the key parameters used to define the digital filter from the constrained frequency response curve, wherein the key parameters include the center frequency, the peak gain at the center frequency, and the target bandwidth; Based on the center frequency, the peak gain at the center frequency, and the target bandwidth, the quality factor of the digital filter is calculated using the following formula: in, The quality factor represents the digital filter. Indicates the center frequency. Center frequency Peak gain at that point Indicates the target bandwidth.

5. The high-frequency compensation and sound field broadening method based on audio content as described in claim 1, characterized in that, Time-frequency analysis is performed on the input audio signal to extract its high-frequency energy distribution characteristics, spectral sparsity characteristics, and sound field spatial cue characteristics, including: By performing time-frequency transformation on the input audio signal, the complex spectrum of the input audio signal is obtained, and the amplitude spectrum of each frequency point in the complex spectrum is further calculated. From the amplitude spectrum, extract the high-frequency energy distribution features and spectral sparsity features of the input audio signal; From the complex spectrum, the acoustic field spatial cue features of the input audio signal are extracted.

6. The high-frequency compensation and sound field broadening method based on audio content as described in claim 5, characterized in that, Extracting the high-frequency energy distribution features and spectral sparsity features of the input audio signal from the amplitude spectrum includes: Based on the amplitude spectrum, the sum of squares of the amplitudes at each frequency point within the predefined high-frequency sub-band range is calculated as the signal energy of the high-frequency sub-band. The signal energy is then normalized and compared with the energy of the reference frequency band to obtain the high-frequency energy distribution characteristics. Based on the amplitude spectrum, the amplitude data sequence within the high-frequency sub-band is extracted, and the kurtosis statistics of the amplitude data sequence are calculated as a spectral sparsity feature of the input audio signal.

7. The high-frequency compensation and sound field broadening method based on audio content as described in claim 5, characterized in that, Extracting the acoustic field spatial cue features of the input audio signal from the complex spectrum includes: For the complex spectrum of the multi-channel signal, the phase difference and amplitude ratio between channels corresponding to the same frequency point are calculated, and then aggregated within the psychoacoustic critical band to obtain the sound field spatial cue features.

8. The high-frequency compensation and sound field broadening method based on audio content as described in claim 1, characterized in that, Based on the aforementioned sound field spatial cue features, a multi-channel output signal with a sense of spatial expansion is generated from the input audio signal, including: Define the binaural cue preservation constraints for the input audio signal; The spatial cue features of the sound field are analyzed to obtain the target sound field width parameter and the target lateral sound energy ratio parameter, so as to perform adaptive spatial rendering on the input audio signal under the constraint of maintaining the binaural cue, so as to generate the spatial extension signal of the input audio signal; The spatially extended signal is synthesized into a multi-channel output signal in stereo format with left and right channels.

9. The high-frequency compensation and sound field broadening method based on audio content as described in claim 1, characterized in that, Calculating the acoustic image localization stability of the multi-channel output signal includes: Extract binaural cue features of the multi-channel output signal within the psychoacoustic critical frequency band, wherein the binaural cue features include the processed inter-channel phase difference and the processed inter-channel amplitude ratio; The binaural cue offset of the psychoacoustic critical band is calculated based on the binaural cue features to calculate the acoustic image localization stability of the multi-channel output signal. The binaural cue offset includes phase difference offset and amplitude ratio offset.

10. A high-frequency compensation and sound field widening system based on audio content, characterized in that, The system includes: The signal time-frequency analysis module is used to perform time-frequency analysis on the input audio signal to extract the high-frequency energy distribution characteristics, spectral sparsity characteristics and sound field spatial cue characteristics of the input audio signal; The high-frequency signal compensation module is used to dynamically generate a high-frequency compensation filter that matches the content of the input audio signal based on the high-frequency energy distribution characteristics and spectral sparsity characteristics, so as to perform content-aware compensation on the high-frequency components in the input audio signal and obtain a compensation signal. A multi-channel signal output module is used to generate a multi-channel output signal with a sense of spatial expansion of the input audio signal based on the spatial cue characteristics of the sound field; The acoustic image stability analysis module is used to calculate the acoustic image localization stability of the multi-channel output signal; The target audio signal output module is used to perform time-frequency domain weighted fusion of the multi-channel output signal and the compensation signal to output the target audio signal when the sound image positioning stability meets the preset sound image positioning stability threshold.

Citation Information

Patent Citations

  • Scene automatic identification-based business hotel guest room scene illumination system

    CN109874209A

  • Speech spectrum reconstruction method and system combining time domain half-wave rectification and weighted Gaussian mixture model decoder, terminal and medium

    CN121054016A

  • Method and apparatus to compensate for imperfections in sound field using peak and dip frequencies

    US20050119879A1

  • Spatial audio enhancement processing method and apparatus

    US20080031462A1

  • Apparatus and method for encoding an audio signal using a compensation value

    US20190189137A1