Audio signal processing method, electronic device, storage medium, and program product

CN122846012APending Publication Date: 2026-09-29上海勤宽科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610920336.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,传统听音室评估对场地、设备及人员依赖较强,测试成本高且结果易受环境声学条件与听音疲劳影响;现有虚拟重放方案多侧重单一处理环节,模块协同不足,对扬声器特性适配不灵活,且易受耳机自身频响差异干扰,导致重放保真度与评估一致性不足

Benefits of technology

[0055]本申请实施例提供的音频信号处理方法、电子设备、存储介质及程序产品,通过获取原始音频信号后,依次对其执行均衡补偿处理、根据扬声器特性进行虚拟化处理、基于与头相关脉冲响应的卷积处理以及根据耳机频响特性进行均衡校正处理,能够在信号链路中分别对原始信号频谱、目标扬声器重放特征、双耳空间听觉效果及耳机回放偏差进行针对性修正,进而降低对物理听音环境的依赖和测试成本,并提升虚拟扬声器重放的保真度、灵活适配能力以及音质评估结果的一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122846012A_ABST
    Figure CN122846012A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio signal processing method, an electronic device, a storage medium and a program product. The method comprises: obtaining an original audio signal; performing first-level processing on the original audio signal, that is, performing equalization compensation processing through a plurality of cascaded filters to generate a first intermediate signal; performing second-level processing on the first intermediate signal, that is, performing virtualization processing according to loudspeaker characteristics to generate a second intermediate signal; performing third-level processing on the second intermediate signal, that is, performing convolution processing with a head-related impulse response to generate a binaural stereo signal; and performing fourth-level processing on the binaural stereo signal, that is, performing equalization correction processing according to earphone frequency response characteristics to generate an output signal. The method is used to achieve audio playback closer to the target loudspeaker listening experience under non-physical listening room conditions, and improve the adaptability, fidelity and consistency of virtual evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speaker equalization technology, and in particular to an audio signal processing method, electronic device, storage medium, and program product. Background Technology

[0002] In the field of loudspeaker research and development and sound quality testing, existing technologies typically employ physical listening rooms for evaluation or use headphone-based virtual loudspeaker playback methods for auxiliary testing.

[0003] However, traditional listening room assessments are highly dependent on venues, equipment, and personnel, resulting in high testing costs and results that are easily affected by environmental acoustic conditions and listening fatigue. Existing virtual playback solutions often focus on a single processing stage, lacking modular coordination, and are not flexible in adapting to speaker characteristics. They are also susceptible to interference from differences in the frequency response of the headphones themselves, leading to insufficient playback fidelity and consistency in the assessment.

[0004] Therefore, how to reduce the environmental dependence and testing costs of loudspeaker sound quality evaluation, and improve the fidelity, adaptability and evaluation consistency of virtual playback, has become a technical problem that needs to be solved. Summary of the Invention

[0005] This application provides an audio signal processing method, electronic device, storage medium, and program product to achieve audio playback that more closely resembles the listening experience of the target loudspeaker under non-physical listening room conditions, thereby improving the adaptability, fidelity, and consistency of virtual evaluation.

[0006] In a first aspect, embodiments of this application provide an audio signal processing method, including:

[0007] Acquire the raw audio signal;

[0008] The first stage of processing is performed on the original audio signal, which is to perform equalization compensation processing through multiple cascaded filters to generate the first intermediate signal;

[0009] The first intermediate signal is processed in the second stage, which is to perform virtualization processing based on the speaker characteristics to generate the second intermediate signal;

[0010] The second intermediate signal undergoes a third level of processing, namely, convolution processing with head-related impulse responses to generate a binaural stereo signal;

[0011] The binaural stereo signal undergoes a fourth level of processing, namely, equalization correction based on the headphone frequency response characteristics, to generate the output signal.

[0012] In one possible embodiment, equalization compensation is performed using multiple cascaded filters, including:

[0013] Obtain the center frequency, quality factor, and gain of each filter;

[0014] The amplitude coefficient is calculated based on the gain, the angular frequency is calculated based on the center frequency and the sampling frequency, and the bandwidth coefficient is calculated based on the angular frequency and the quality factor.

[0015] The transfer function coefficients of each filter are calculated based on the amplitude coefficient and bandwidth coefficient. The transfer function coefficients include the numerator coefficient and the denominator coefficient.

[0016] Construct the transfer function based on the transfer function coefficients of each filter;

[0017] The original audio signal is filtered sequentially through multiple second-order infinite impulse response filters to obtain the first intermediate signal.

[0018] In one possible embodiment, the transfer function coefficients are calculated as follows:

[0019] The constant term of the numerator coefficient is the sum of the amplitude coefficient and the bandwidth coefficient. The first-order coefficient of the numerator coefficient is twice the negative of the cosine of the angular frequency. The second-order coefficient of the numerator coefficient is the constant term minus the product of the amplitude coefficient and the bandwidth coefficient.

[0020] The constant term of the denominator coefficients is the sum of the reciprocal of the amplitude coefficient and the bandwidth coefficient. The first-order coefficient of the denominator coefficients is twice the negative of the cosine of the angular frequency. The second-order coefficient of the denominator coefficients is the constant term minus the product of the reciprocal of the amplitude coefficient and the bandwidth coefficient.

[0021] In one possible embodiment, a second level of processing is performed on the first intermediate signal, namely, virtualization processing based on speaker characteristics to generate a second intermediate signal, including:

[0022] Obtain the impulse response or frequency response of the speaker;

[0023] The first intermediate signal is subjected to time-domain convolution processing based on the impulse response, or frequency-domain filtering processing based on the frequency response, to simulate the playback characteristics of a loudspeaker.

[0024] The output signal, after being processed by speaker virtualization, serves as the second intermediate signal.

[0025] In one possible embodiment, obtaining the frequency response of the loudspeaker includes:

[0026] Obtain the frequency domain characteristic file of the loudspeaker, which contains frequency, sound pressure level and phase information;

[0027] Interpolation is performed on the sound pressure level and phase at missing frequencies in the frequency domain characteristic file to obtain the loudspeaker frequency response over the complete frequency range.

[0028] In one possible embodiment, a third level of processing is performed on the second intermediate signal, namely, convolution processing with the head-related impulse response to generate a binaural stereo signal, including:

[0029] Acquire head-related impulse responses corresponding to the left and right ears of the listener;

[0030] The second intermediate signal is convolved with the left ear head correlation impulse response and the right ear head correlation impulse response respectively to generate the left channel signal and the right channel signal;

[0031] The left and right channel signals are output as binaural stereo signals.

[0032] In one possible embodiment, acquiring head-related impulse responses corresponding to the left and right ears of the listener includes:

[0033] Load a standard format head-related impulse response file, which supports multiple position selections for azimuth or elevation angles;

[0034] Based on the selected azimuth or pitch angle, extract the corresponding left and right ear impulse responses from the head-related impulse response file.

[0035] In one possible embodiment, a fourth level of processing is performed on the binaural stereo signal, namely, equalization correction processing based on the headphone frequency response characteristics, to generate an output signal, including:

[0036] Obtain the impulse response of the headphones;

[0037] Perform inverse operations on the impulse response of the headphones to generate equalization filter coefficients;

[0038] The binaural stereo signal is filtered according to the equalization filter coefficients to correct the frequency response characteristics of the headphones and obtain the output signal.

[0039] In one possible embodiment, the impulse response of the headphones is inversely processed to generate equalization filter coefficients, including:

[0040] Perform a Fourier transform on the impulse response of the headphones to obtain the frequency response of the headphones;

[0041] Calculate the power spectrum based on the frequency response of the headphones and its conjugate.

[0042] The frequency response of the equalization filter is calculated based on the power spectrum and the preset regularization factor.

[0043] The frequency response of the equalization filter is subjected to inverse Fourier transform to obtain the equalization filter coefficients in the time domain.

[0044] Secondly, embodiments of this application provide an audio signal processing apparatus, comprising:

[0045] The acquisition module is used for the raw audio signal;

[0046] The first processing module is used to perform the first-level processing on the original audio signal, that is, to perform equalization compensation processing through multiple cascaded filters to generate the first intermediate signal;

[0047] The second processing module is used to perform a second-level processing on the first intermediate signal, that is, to perform virtualization processing based on the speaker characteristics to generate a second intermediate signal.

[0048] The third processing module is used to perform third-level processing on the second intermediate signal, namely, to generate a binaural stereo signal by performing convolution processing with head-related impulse response.

[0049] The output module is used to perform the fourth level of processing on the binaural stereo signal, namely, equalization correction processing based on the frequency response characteristics of the headphones, and generate the output signal.

[0050] Thirdly, embodiments of this application provide an audio signal processing device, including: a memory and a processor;

[0051] The memory stores computer-executed instructions;

[0052] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0053] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0054] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0055] The audio signal processing method, electronic device, storage medium, and program product provided in this application, after acquiring the original audio signal, sequentially perform equalization compensation processing, virtualization processing based on speaker characteristics, convolution processing based on head-related impulse response, and equalization correction processing based on headphone frequency response characteristics. This enables targeted correction of the original signal spectrum, target speaker playback characteristics, binaural spatial hearing effect, and headphone playback deviation in the signal link, thereby reducing dependence on the physical listening environment and testing costs, and improving the fidelity, flexible adaptation capability, and consistency of sound quality evaluation results of virtual speaker playback. Attached Figure Description

[0056] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0057] Figure 1 Flowchart of the audio signal processing method provided in this application Figure 1 ;

[0058] Figure 2 Flowchart of the audio signal processing method provided in this application Figure 2 ;

[0059] Figure 3 A schematic diagram of the amplitude-frequency response curve of the loudspeaker provided in this application;

[0060] Figure 4 A schematic diagram of the phase frequency response curve of the loudspeaker provided in this application;

[0061] Figure 5 A schematic diagram of the impulse response waveform of the loudspeaker provided in this application;

[0062] Figure 6 A comparative diagram of the original audio signal and the signal waveforms after each stage of processing provided in this application;

[0063] Figure 7 A schematic diagram of the audio signal processing device provided in this application;

[0064] Figure 8 A schematic diagram of the audio signal processing device provided in this application.

[0065] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0066] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application.

[0067] In the fields of loudspeaker R&D, audio equipment tuning, and sound quality testing, it is often necessary to conduct auditory evaluations of the acoustic performance of target loudspeakers to determine their frequency response characteristics, spatial perception effects, and final sound consistency. Related applications typically involve processing steps such as audio signal acquisition, loudspeaker characteristic modeling, binaural playback, and headphone monitoring, and are suitable for scenarios such as product development verification, solution comparison testing, and routine quality assessment.

[0068] In the existing technology, one type of solution mainly relies on physical listening rooms and real speakers for listening tests, where testers play reference audio in a specific acoustic environment and make subjective judgments; another type of solution uses headphones to perform virtual speaker playback, which usually simulates the speaker playback effect by applying some acoustic processing to the audio signal. Its basic principle is to combine the input audio with preset speaker or space-related parameters to obtain a listening experience close to that of external speaker playback.

[0069] However, traditional listening room solutions are highly dependent on venue, equipment, and personnel, resulting in high construction and maintenance costs. Furthermore, test results are easily affected by room acoustics, environmental noise, and listening fatigue, leading to insufficient consistency between different batches and individuals. Existing headphone virtual playback solutions often optimize only a single processing stage, lacking a continuous processing chain from the original audio to the final headphone output. This results in incomplete speaker characteristic simulation and difficulty in balancing fidelity and compatibility.

[0070] Furthermore, when there is a lack of coordination between speaker characteristic simulation and binaural spatial processing, or when the frequency response differences of the headphones themselves are not corrected, additional distortion and coloration will be superimposed on the final output signal. This will cause the heard results to be affected by both speaker model errors and headphone response deviations, thereby weakening the comparability and credibility of the test results and making it difficult to meet the requirements of stability, repeatability and low cost implementation in R&D and evaluation scenarios.

[0071] In view of this, how to reduce the dependence of speaker sound quality evaluation on the external environment and improve the fidelity and evaluation consistency of virtual playback under headphone playback conditions has become an urgent technical problem to be solved. To solve the above problems, an audio signal processing method is provided. After acquiring the original audio signal, equalization compensation processing is first performed through multiple cascaded filters to generate a first intermediate signal. Then, virtualization processing is performed on the first intermediate signal according to the speaker characteristics to generate a second intermediate signal. Subsequently, convolution processing with head-related impulse response is performed to generate a binaural stereo signal. Finally, equalization correction processing is performed according to the headphone frequency response characteristics to generate an output signal.

[0072] This technical approach corresponds to a multi-level audio processing chain for headphone monitoring, which can be deployed on audio processing software platforms or related processing systems. The original audio is output after passing through equalization compensation, speaker virtualization, binaural convolution, and headphone correction in sequence. This allows for a more accurate reconstruction of the listening effect of the target speaker without relying on the conditions of a physical listening room, providing a more stable foundation for subsequent sound quality evaluation.

[0073] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0074] Figure 1 Flowchart of the audio signal processing method provided in this application Figure 1 ,like Figure 1 As shown, the method includes:

[0075] S101: Acquire the raw audio signal.

[0076] In this application, the raw audio signal serves as the input signal for the entire processing chain, providing the audio content to be processed and serving as the basic input for subsequent equalization compensation processing, speaker virtualization processing, head-related impulse response convolution processing, and headphone frequency response correction processing. Based on the above analysis, the raw audio signal can be an audio signal from a loaded WAV (Waveform Audio File Format) file, a FLAC (Free Lossless Audio Codec) file, a PC (Personal Computer) M (Pulse Code Modulation) stream, a encoded / decoded digital audio stream, or other forms of audio signal to be processed. In one possible embodiment, acquiring the raw audio signal includes the audio processing software reading the audio file from the local storage medium and parsing the sampled data in the file into a discrete digital sequence arranged in time order. In another possible embodiment, acquiring the raw audio signal may also include receiving a real-time audio stream from a sound card input interface, network audio interface, or external bus interface, and writing it into an input buffer for subsequent processing. It should be understood that the above examples are merely illustrative and not limiting.

[0077] In practice, the execution entity can be audio processing software, a digital signal processor, an audio algorithm module executed by a general-purpose processor, or a processing unit integrated into an audio evaluation system. This step establishes a data path corresponding to the input source and provides the acquired raw audio signal to subsequent processing links. After this step is completed, a digital audio sequence is generated that can be directly called by the first-level processing, triggering subsequent equalization compensation processing.

[0078] S102: Perform the first stage of processing on the original audio signal, that is, perform equalization compensation processing through multiple cascaded filters to generate the first intermediate signal.

[0079] In this application, the first-stage processing is used to perform preliminary equalization compensation on the original audio signal to correct the frequency response defects of the audio itself, making the processed signal closer to the predetermined target characteristics. Filters are used to perform frequency domain compensation on the original audio signal, adjusting the gain or attenuation of specific frequency bands. The equalization compensation processing applies multiple filters sequentially to the input signal, performing compensation on different frequency bands respectively. The first intermediate signal serves as the output signal after the first-stage processing, receiving the equalization compensation result and serving as the input for subsequent speaker virtualization processing. Based on the above analysis, it can be seen that the cascaded multiple filters can be combinations of filters capable of achieving equalization compensation, as long as equalization compensation processing of the original audio signal can be completed through cascading. It should be understood that the above example is merely illustrative and not limiting.

[0080] In practice, the target equalization curve for the input audio is first obtained, and this target curve is discretized into multiple frequency band compensation points. Then, multiple filters are configured according to preset filtering parameters and cascaded in a predetermined order. The original audio signal is processed sequentially through multiple filters, with the output of the previous filter serving as the input to the next filter, ultimately yielding the first intermediate signal. Optionally, stability and dynamic range control can be performed as needed. Based on the above analysis, by performing frequency response compensation before the original audio enters the speaker characteristic simulation, the subsequent speaker virtualization processing can be built on the corrected input, thereby reducing the deviation caused by the superposition of input frequency response defects and the speaker model. The first intermediate signal is closer to the preset reference state in the spectrum, facilitating consistency and repeatability in subsequent processing stages.

[0081] S103: Perform second-level processing on the first intermediate signal, that is, perform virtualization processing according to the speaker characteristics to generate a second intermediate signal.

[0082] In this application, the second-level processing is used to virtualize the first intermediate signal according to the playback attributes of the target speaker, so that the processed signal has acoustic characteristics corresponding to the target speaker. Speaker characteristics are used to characterize the playback attributes of the target speaker, and can be expressed as the speaker's acoustic response characteristics; speaker virtualization processing is used to simulate the effect of the input signal after being played through the target speaker; the second intermediate signal serves as the output signal after speaker virtualization processing, and continues to be used as the input for subsequent head-related impulse response convolution processing. Based on the above analysis, speaker virtualization processing according to speaker characteristics can include processing the first intermediate signal accordingly based on the speaker characteristics to obtain a signal corresponding to the playback effect of the target speaker. It should be understood that the above example is merely illustrative and not limiting.

[0083] In specific implementation, the speaker characteristics corresponding to the target speaker are obtained, and the first intermediate signal is virtualized based on these characteristics to obtain the second intermediate signal. Optionally, this processing can be accomplished using an appropriate implementation method. Based on the above analysis, this step introduces the acoustic transmission characteristics of the target speaker into the processing chain, so that the second intermediate signal no longer only retains the original audio content, but also includes the playback characteristics corresponding to the target speaker. The subsequent head-related impulse response convolution processing is thus established on the basis of the speaker model, enabling a more stable correspondence between the final headphone playback result and the listening experience of the target speaker.

[0084] S104: Perform third-level processing on the second intermediate signal, that is, perform convolution processing with the head-related impulse response to generate a binaural stereo signal.

[0085] In this application, the third-level processing is used to convert the signal after speaker virtualization into stereo audio suitable for binaural hearing reproduction. Head-related impulse response (HRI) is used to describe the response differences formed by the influence of the head, shoulders, and auricles when sound propagates to both ears. Convolution processing is used to mathematically synthesize the second intermediate signal and the propagation responses corresponding to the left and right ears, respectively. The binaural stereo signal consists of a left channel signal and a right channel signal. Based on the above analysis, the HRI may include a left-ear HRI and a right-ear HRI, so that the generated binaural stereo signal corresponds to the auditory perception result of the target speaker at a predetermined spatial location. In one possible embodiment, the system invokes the left-ear HRI and right-ear HRI corresponding to the predetermined spatial location. It should be understood that the above examples are merely illustrative and not limiting.

[0086] In practice, the relevant impulse response of the left ear head is read. Impulse response related to the right ear head Next, the second intermediate signal... respectively with , Perform convolution operation to obtain the left channel signal. and right channel signal After the two channel signals are synchronized in time, they are combined to form a binaural stereo signal. For continuous audio streams, this application uses frame-by-frame convolution to maintain output continuity and performs overlapping storage or overlapping addition processing between adjacent frames to eliminate block boundary distortion. Based on the above analysis, this step further maps the speaker virtualization processing result to the propagation difference reaching both ears, so that the binaural stereo signal simultaneously includes the time difference, level difference, and auricular spectrum features between the left and right ears. This allows the headphone to obtain auditory perception corresponding to the spatial position of the target speaker under playback conditions. The spatialization result is connected with the pre-amplifier speaker model, avoiding the situation where only spatial processing is performed without speaker characteristic constraints.

[0087] S105: Performs fourth-level processing on the binaural stereo signal, namely, equalization correction processing based on the frequency response characteristics of the headphones, and generates the output signal.

[0088] In this application, the fourth-level processing is used to compensate for the frequency response deviation introduced by the headphones themselves, so that the binaural stereo signal retains the target sound characteristics as much as possible after being played through the headphones. The headphone frequency response characteristics characterize the headphone's response at different frequencies. The equalization correction processing establishes an inverse compensation relationship based on these frequency response characteristics, performs correction filtering on the left and right channels respectively, and the output signal serves as the final audio signal output by the entire method, which can be used for subsequent headphone playback or sound quality evaluation. Based on the above analysis, it can be seen that the headphone frequency response characteristics can be obtained by measuring the target headphones, or by simulation or calibration databases; equalization correction processing based on the headphone frequency response characteristics can include performing inverse filtering on the binaural stereo signal, or it can include generating equalization filter coefficients based on the headphone frequency response characteristics with the opposite direction of its deviation, and using these equalization filter coefficients to process the binaural stereo signal. It should be understood that the above examples are merely illustrative and not limiting.

[0089] In practice, the frequency response or impulse response of the left and right channels of the headphones is first obtained. If the impulse response correction method is used, a frequency domain transformation is performed on the headphone impulse response to obtain the complex frequency response. and The inverse operation is performed within the defined frequency band to construct the correction transfer function. and To suppress abnormal gain caused by direct inversion at deep frequency points, this application sets a lower limit threshold for the amplitude response before inversion and sets a maximum value constraint for the correction gain at the high and low frequencies. The correction transfer function is then converted into time-domain filter coefficients, or directly multiplied in the frequency domain with the left and right channel spectra of the binaural stereo signal to obtain the corrected left and right channel outputs. If parametric equalization is used, multiple compensation frequency bands are extracted based on the peak and dip positions in the headphone frequency response curve. A corresponding center frequency, quality factor, and gain are configured for each frequency band, and cascaded equalization filter banks are constructed for each of the left and right channels. The binaural stereo signal is then filtered separately. After correction, unified gain management and limiting processing are performed on the left and right channels to generate the final output signal, which is then sent to the headphone playback interface or sound quality evaluation module. Based on the above analysis, it can be seen that, under the condition that the speaker characteristics and binaural spatial characteristics have been superimposed in the preamplifier, further independent compensation of the headphone frequency response characteristics can separate the reproduction deviation of the headphone device itself from the overall link, so that the final heard result is mainly controlled by the target speaker model and spatial model, and the listening result corresponding to the output signal has more stable consistency in different times and different environments.

[0090] Based on the above analysis, this application provides an audio signal processing method, including acquiring an original audio signal; performing a first-level processing on the original audio signal, namely, performing equalization compensation processing through multiple cascaded filters to generate a first intermediate signal; performing a second-level processing on the first intermediate signal, namely, performing speaker virtualization processing based on speaker characteristics to generate a second intermediate signal; performing a third-level processing on the second intermediate signal, namely, performing convolution processing with head-related impulse response to generate a binaural stereo signal; and performing a fourth-level processing on the binaural stereo signal, namely, performing equalization correction processing based on headphone frequency response characteristics to generate an output signal. In this application, the original audio signal sequentially undergoes four consecutive processing levels: equalization compensation, speaker virtualization processing, head-related impulse response convolution processing, and headphone correction. The output of the previous level serves as the input of the next level, forming a sequentially connected complete processing chain. The first-level processing unifies the input frequency response state, the second-level processing introduces the acoustic transmission characteristics of the target speaker, the third-level processing introduces the spatial propagation differences reaching both ears, and the fourth-level processing performs reverse compensation for headphone playback errors, thereby enabling the final output signal to more closely approximate the actual listening experience of the target speaker under headphone monitoring conditions. Furthermore, in one possible embodiment, each processing stage can be configured with a bypass state according to processing requirements, or the corresponding parameters can be dynamically adjusted based on the input audio, speaker characteristics, head-related impulse response, and headphone frequency response characteristics. It should be understood that the above examples are merely illustrative and not limiting.

[0091] In other words, this application achieves the above objectives through a four-stage cascaded processing: the first stage uses multiple cascaded filters for equalization compensation to compensate for frequency response defects in the audio source or acquisition link; the second stage performs virtualization processing based on speaker characteristics to simulate the speaker's playback effect in a free field; the third stage uses convolution processing with head-related impulse response (HRIR) to introduce spatial orientation information; and the fourth stage performs equalization correction based on the headphone's frequency response characteristics to eliminate sound quality degradation caused by uneven frequency response of the headphone itself. The entire link forms a complete signal processing chain of "acquisition / audio source compensation → speaker simulation → spatial positioning → headphone adaptation".

[0092] The audio signal processing method provided in this application, after acquiring the original audio signal, sequentially performs equalization compensation processing, virtualization processing based on speaker characteristics, convolution processing based on head-related impulse response, and equalization correction processing based on headphone frequency response characteristics. This enables targeted correction of the original signal spectrum, target speaker playback characteristics, binaural spatial hearing effect, and headphone playback deviation in the signal link, thereby reducing dependence on the physical listening environment and testing costs, and improving the fidelity, flexible adaptation capability, and consistency of sound quality evaluation results of virtual speaker playback.

[0093] Figure 2 Flowchart of the audio signal processing method provided in this application Figure 2 ,like Figure 2 As shown, in this embodiment... Figure 1 Based on the embodiments, the audio signal processing method is described in detail, which includes:

[0094] S301. Obtain the center frequency, quality factor, and gain of each filter; calculate the amplitude coefficient based on the gain, calculate the angular frequency based on the center frequency and sampling frequency, and calculate the bandwidth coefficient based on the angular frequency and quality factor; calculate the transfer function coefficients of each filter based on the amplitude coefficient and bandwidth coefficient, the transfer function coefficients include numerator and denominator coefficients; construct the transfer function based on the transfer function coefficients of each filter; filter the original audio signal sequentially through multiple second-order infinite impulse response filters to obtain the first intermediate signal.

[0095] Among them, the sampling frequency is the sampling rate of the original audio signal, the amplitude coefficient is used to convert the gain into the numerical form required for filter calculation, the angular frequency is used to characterize the discrete domain frequency parameters, and the bandwidth coefficient is used to reflect the bandwidth correlation characteristics of the filter.

[0096] In practical implementation, multiple pre-configured filter parameters can be read first. Each filter corresponds to a different center frequency, quality factor, and gain. Then, the amplitude coefficient is obtained based on the gain conversion, and the angular frequency is calculated by combining the center frequency and the sampling frequency. The bandwidth coefficient is further obtained from the angular frequency and the quality factor. Subsequently, the numerator and denominator coefficients of each filter are calculated based on the amplitude coefficient, bandwidth coefficient, and the cosine term of the angular frequency, and the coefficients are assembled into the corresponding transfer function. The second-order infinite impulse response filter adopts a feedback digital filtering structure, which can continuously compensate for different frequency bands in the original audio signal. After multiple filters are cascaded in series, the first intermediate signal after equalization compensation is output. In practical applications, other parameter configurations can also be selected for this filter bank, which is not limited in this application.

[0097] This equalization compensation process establishes filtering models for multiple frequency bands and executes filtering sequentially, so that the deviations of different frequency bands are corrected in the same processing link, and the first intermediate signal thus has a signal distribution that is closer to the target frequency response characteristics.

[0098] After the above processing, the frequency response distortion of the original audio signal can be corrected before entering the subsequent speaker virtualization processing, so that the subsequently generated binaural stereo signal has a more consistent frequency band basis and reduces the cumulative effect of the initial frequency response deviation in the cascaded processing.

[0099] In one optional embodiment, the constant term of the numerator coefficient is the sum of the amplitude coefficient and the bandwidth coefficient, the first-order coefficient of the numerator coefficient is twice the negative of the cosine of the angular frequency, and the second-order coefficient of the numerator coefficient is the constant term minus the product of the amplitude coefficient and the bandwidth coefficient; the constant term of the denominator coefficient is the sum of the reciprocal of the amplitude coefficient and the bandwidth coefficient, the first-order coefficient of the denominator coefficient is twice the negative of the cosine of the angular frequency, and the second-order coefficient of the denominator coefficient is the constant term minus the product of the reciprocal of the amplitude coefficient and the bandwidth coefficient.

[0100] In practical implementation, the amplitude coefficient is obtained by converting the filter gain, the bandwidth coefficient is calculated by combining the center frequency and quality factor with the sampling frequency, and the angular frequency is obtained by normalizing the center frequency and sampling frequency. The transfer function coefficients can be numerically written into the parameter table of the difference equation of a second-order infinite impulse response filter, and are called point-by-point by the audio processing module according to the set sampling period. If different digital signal processing platforms are used, the transfer function coefficients can still be solved using the same mathematical relationships and mapped to the corresponding platform's floating-point or fixed-point coefficient format. In practical applications, this calculation method can also be adapted to other filter implementations, which this application does not limit.

[0101] The method of constructing transfer function coefficients integrates amplitude, bandwidth, and angular frequency into the same filter parameter model, allowing the numerator and denominator to change collaboratively under the same frequency constraints. This results in the filter output achieving the expected equalization compensation. Based on the transfer function generated by these coefficients, the gain and attenuation relationships of the target frequency band of the original audio signal can be stably expressed after cascaded filtering, providing a more consistent input signal for subsequent speaker virtualization processing. Using this coefficient calculation method, the source of filter parameters is clear, facilitating repeated configuration based on different speaker and headphone characteristics, thereby improving the consistency and reproducibility of audio processing results.

[0102] S302. Obtain the impulse response or frequency response of the speaker; perform time-domain convolution processing on the first intermediate signal based on the impulse response, or perform frequency-domain filtering processing on the first intermediate signal based on the frequency response, to simulate the speaker playback characteristics; output the signal after speaker virtualization processing as the second intermediate signal.

[0103] Among them, temporal convolution processing is used to apply the loudspeaker impulse response to the first intermediate signal, and frequency domain filtering processing is used to apply the loudspeaker frequency response to the first intermediate signal to form acoustic output characteristics corresponding to the target loudspeaker.

[0104] In implementation, the impulse response data of the target loudspeaker under standard test conditions can be measured in advance and written into the parameter table of the processing module. After receiving the first intermediate signal, the processing module performs convolution operation on it according to the sampling points to obtain an output signal containing the loudspeaker cabinet resonance, unit transients, and frequency band attenuation characteristics. In another implementation, the complete frequency response curve of the loudspeaker can be obtained by frequency sweep measurement, and then a filter corresponding to the curve can be constructed in the frequency domain to perform amplitude and phase correction on the first intermediate signal, outputting a signal after loudspeaker virtualization processing. The processing module can be a digital signal processor, an audio algorithm chip, or a software module running on a terminal device. In practical applications, other models of this component can also be selected, and this application does not limit this.

[0105] This processing method utilizes the speaker's response data to directly perform feature mapping on the first intermediate signal. This allows the virtualized signal to carry the playback characteristics of the target speaker and serve as input for subsequent head-related impulse response convolution processing, resulting in a final output that more closely resembles the sound quality of a real speaker. Because the speaker's impulse response or frequency response is incorporated into the processing chain, the output signal provides a more complete simulation of the target speaker, thereby improving the consistency and comparability of sound quality evaluation under headphone monitoring conditions.

[0106] In an optional embodiment, a frequency domain characteristic file of the loudspeaker is obtained, which includes frequency, sound pressure level, and phase information;

[0107] Interpolation is performed on the sound pressure level and phase at missing frequencies in the frequency domain characteristic file to obtain the loudspeaker frequency response over the complete frequency range.

[0108] The frequency domain characteristic file of the loudspeaker can be output by the loudspeaker testing system or exported by acoustic measurement software in a readable data format for easy subsequent reading and calculation. In practical applications, the file can also adopt other data organization forms, which are not limited in this application.

[0109] After acquiring the frequency domain characteristic file, the system performs completion processing on the frequency points where no measured values ​​are recorded. Specifically, for the missing frequency intervals between adjacent known frequency points, a continuous variation relationship is established based on the sound pressure level and phase values ​​of adjacent points. Interpolation calculations are then performed on the sound pressure level and phase at the missing frequencies to generate estimated values ​​for the corresponding frequency points. Interpolation can be performed using linear interpolation, spline interpolation, or other interpolation methods that maintain the continuity of the frequency response, ensuring that the completed sound pressure level and phase are continuously distributed across the entire frequency range. After completion, the sound pressure level and phase of each frequency point are combined to obtain the loudspeaker frequency response over the complete frequency range.

[0110] During operation, the system first reads the speaker's frequency domain characteristic file, then extracts sound pressure level and phase information based on the existing frequency sampling points in the file, and performs interpolation to complete missing frequency points, ultimately forming complete frequency response data that can be used for speaker virtualization processing. This frequency response can cover all frequency information within the target frequency band, providing an input basis for subsequent frequency domain filtering or convolution modeling based on speaker characteristics.

[0111] Figure 3 The schematic diagram of the amplitude-frequency response curve of the loudspeaker provided in this application shows the gain characteristics of the loudspeaker at different frequencies, with the horizontal axis representing frequency (Hz) and the vertical axis representing amplitude (dB).

[0112] Figure 4 The schematic diagram of the phase response curve of the loudspeaker provided in this application shows that the horizontal axis is frequency (Hz) and the vertical axis is phase (°), which is used to show the phase characteristics of the loudspeaker at different frequencies.

[0113] Figure 5 The schematic diagram of the impulse response waveform of the loudspeaker provided in this application shows the impulse response characteristics of the loudspeaker in the time domain, with the horizontal axis representing time (s) and the vertical axis representing amplitude.

[0114] Figure 3 , Figure 4 and Figure 5 These represent three different domains for the same loudspeaker. Among them, Figure 3 (Amplitude-frequency response) and Figure 4(Phase response) together constitute the complete frequency domain description of the loudspeaker, and the two can be correlated through Fourier transform; Figure 5 (Impulse response) is the response to Figure 3 and Figure 4 The time-domain description obtained after performing the inverse Fourier transform.

[0115] In practical signal processing, both frequency domain methods and frequency domain methods can be used. Figure 3 , Figure 4 Speaker virtualization filtering can be performed, or a time-domain approach can be used. Figure 5 The two methods are mathematically equivalent and can be performed using equivalent convolution. After the above processing, the discrete data in the loudspeaker frequency domain characteristic file can be reconstructed into a continuous frequency response, and the correspondence between frequency, sound pressure level, and phase can be kept consistent. This provides data support for the accurate characterization of loudspeaker characteristics and gives the loudspeaker simulation results in subsequent audio processing a more complete frequency domain basis.

[0116] S303. Obtain head-related impulse responses corresponding to the left and right ears of the listener;

[0117] The second intermediate signal is convolved with the left ear head correlation impulse response and the right ear head correlation impulse response respectively to generate the left channel signal and the right channel signal;

[0118] The left and right channel signals are output as binaural stereo signals.

[0119] In this context, the left and right ears of the listener refer to the two receiving directions corresponding to the left and right auditory channels, respectively carrying the spatialized processing results of the left and right channel signals. The left and right ear head-related impulse responses correspond to the spatial auditory responses of the left and right ears, respectively, and serve as the impulse response kernels for convolution operations. The left and right channel signals represent the output components of the second intermediate signal after being synthesized from the left and right ear responses, and together they constitute the binaural stereo signal.

[0120] In practical implementation, head-related impulse response data corresponding to the listener can be acquired in advance. This data can come from a measurement database, standardized head model measurement results, or individualized acquisition results, and the data format includes discrete time-domain sampling sequences. After acquisition, the second intermediate signal is input into the convolution channels of the left and right ears respectively, and weighted summation is performed in the time domain according to sampling points to obtain the left and right channel signals. If time-domain convolution is used, the signal length can be aligned based on a fixed sampling rate, and zero-padding or equal-length extension methods can be used at the boundaries to ensure that the length of the convolution result meets the subsequent output requirements. After convolution, the left and right channel signals are encapsulated and output in a dual-channel audio format as input signals for subsequent headphone frequency response correction.

[0121] This processing method introduces head-related impulse responses corresponding to the left and right ears respectively, so that the second intermediate signal forms a differentiated propagation result in the binaural dimension, thereby enabling the output signal to carry more specific spatial positioning characteristics and providing intermediate audio with binaural sound field characteristics for subsequent headphone playback.

[0122] With this implementation method, the binaural stereo signal can retain spatial cues related to the head in the left and right channels, making the virtual playback result closer to the auditory performance formed by the target speaker at the listening position. Moreover, the binaural output structure is clear, which is convenient for connection with the subsequent headphone frequency response correction stage, and improves the consistency and reproduction stability of the overall processing link.

[0123] In an optional embodiment, a standard format head-related impulse response file is loaded, which supports multiple position selections for azimuth or elevation angles;

[0124] Based on the selected azimuth or pitch angle, extract the corresponding left and right ear impulse responses from the head-related impulse response file.

[0125] This file can pre-store multiple sets of left and right ear impulse responses corresponding to azimuth or elevation angles according to a unified data organization method, so that they can be directly retrieved after loading.

[0126] In its implementation, before entering the head-related impulse response (HIR) convolution processing, the system first calls the file parsing module to load a standard-format HIR file and reads its angle index, sampling rate, channel identifier, and corresponding impulse response sequence. The standard format can be organized using angle coordinates and left / right ear response matrices, allowing for rapid matching based on the current target listening position after the file is loaded. The current target listening position can be determined by speaker orientation settings, preset spatial monitoring parameters, or user input, and mapped to the corresponding azimuth or pitch angle values. The system then searches for matching entries in the index table based on this angle value. If data with a completely identical angle exists in the file, the corresponding left and right ear impulse responses are directly extracted. If discrete sampling positions exist, the adjacent data units are located using the index, and the corresponding responses are extracted, outputting the impulse response sequence for left and right channel convolution.

[0127] This processing method ensures consistency between the acquisition of head-related impulse responses and spatial location selection, accurately mapping binaural spatial information from different directions to the corresponding responses of the left and right ears, and providing a stable data source for subsequent convolution operations. Loading and indexing using standard format files allows for a unified data organization and retrieval method for response data at different azimuth or pitch angles, thus guaranteeing spatial positioning consistency and reproduction consistency during binaural stereo signal generation.

[0128] S304, Obtain the impulse response of the headphones;

[0129] Perform inverse operations on the impulse response of the headphones to generate equalization filter coefficients;

[0130] The binaural stereo signal is filtered according to the equalization filter coefficients to correct the frequency response characteristics of the headphones and obtain the output signal.

[0131] The inverse operation is used to solve the impulse response in reverse to obtain the filter parameters that can compensate for the headphone frequency response deviation; the binaural stereo signal is the left and right channel signal pair obtained after the third stage of processing; the output signal is the final playback signal after headphone correction.

[0132] In practical implementation, the headphone impulse response can be obtained by playing a short-time excitation signal to the target headphones and collecting the response results. The collected time-domain response can be truncated and normalized using a window function before being input into the inverse operation module. The inverse operation can calculate the conjugate spectrum based on the frequency domain representation of the impulse response, and combine it with a preset regularization factor to form a stable inverse filtering target. Then, the time-domain filter coefficients are obtained through inverse transformation. The equalization filter coefficients can be stored as a finite-length finite impulse response coefficient array, or constructed in a second-order cascaded form for execution in a digital signal processor, audio codec chip, or software audio engine. Subsequently, the binaural stereo signal is input into the equalization filters of the left and right channels for convolution or recursive filtering, so that the sound pressure response played to the headphones matches the target frequency response.

[0133] This equalization correction process maps the frequency response differences of the headphones themselves to compensable filter coefficients and applies corresponding corrections to the binaural stereo signals, ensuring that the final output signal maintains the spatial image and spectral characteristics formed by the pre-processing stage during headphone playback. Thus, the influence of the headphones' own coloration on the binaural signals is constrained within a predetermined range, and the output signal can be directly used for subsequent monitoring, comparison, or sound quality evaluation.

[0134] In one alternative embodiment, the impulse response of the headphones is subjected to a Fourier transform to obtain the frequency response of the headphones;

[0135] Calculate the power spectrum based on the frequency response of the headphones and its conjugate.

[0136] The frequency response of the equalization filter is calculated based on the power spectrum and the preset regularization factor.

[0137] The frequency response of the equalization filter is subjected to inverse Fourier transform to obtain the equalization filter coefficients in the time domain.

[0138] In practical implementation, the impulse response of the headphones can be represented by a measured discrete impulse response sequence, which is then converted into a complex spectrum via a Fast Fourier Transform (FFT). This spectrum is then multiplied by its conjugate to obtain the power spectrum, where the amplitude at each frequency point reflects the headphone's response intensity at the corresponding frequency. A regularization factor can be set as a frequency-dependent constant or frequency-dependent term and participates in the calculation along with the power spectrum to ensure the frequency response of the equalization filter remains stable at low-energy frequencies. After completing the frequency domain solution, an inverse Fourier transform is performed on the frequency response to obtain time-domain equalization filter coefficients of finite length. These coefficients can be directly used for subsequent filtering and correction of binaural stereo signals. In practical applications, the number of points, window functions, and zero-padding length used in the Fourier transform and inverse Fourier transform can be set according to the headphone response bandwidth and the target filtering length; this application does not impose any limitations on these settings.

[0139] The generation process of the equalization filter coefficients transforms the headphone's frequency response deviation into implementable time-domain correction coefficients, and avoids excessive amplification of local spectral valleys by the inverse operation under the constraint of the regularization factor. Therefore, when filtering the binaural stereo signal subsequently, the headphone's own coloration can be controlled within a predetermined range, making the output signal closer to the target acoustic result.

[0140] By adopting this method, the equalization filter coefficients can be constructed in a directional manner based on the actual headphone response, and its frequency domain solution and time domain recovery process have good numerical stability, thereby improving the consistency of headphone calibration and reducing the listening deviation caused by differences in headphone frequency response.

[0141] In an optional embodiment, the method can be executed by a hardware and software co-processing system, wherein the system input is the original audio signal (i.e., the original reference audio source signal_0), and the system output is the audio signal after multi-level processing and each intermediate level signal, for subsequent sound quality evaluation or playback.

[0142] The specific processing procedure is as follows:

[0143] Step 1: Obtain the raw audio signal:

[0144] First, the system loads the audio file to be processed from the storage medium or data interface. In this embodiment, the audio file is a standard format WAV file, which the system decodes and reads to obtain the digitized raw audio signal. ).

[0145] Step Two, First-Level Processing: Balanced Compensation Processing

[0146] For the original audio signal ( The first stage of processing is performed, which involves equalization compensation through multiple cascaded filters to generate the first intermediate signal. ,Right now ).

[0147] Specifically, this step uses N cascaded second-order infinite impulse response (IIR) filters to filter the original audio signal step by step. The frequency band characteristics of each filter are characterized by three parameters: center frequency f0 (Hz), quality factor Q, and gain G (dB). For each filter, its transfer function coefficients are calculated as follows:

[0148] First, calculate the amplitude coefficient based on the gain G. Calculate the angular frequency based on the center frequency f0 and the sampling frequency fs. Then based on angular frequency Calculate the bandwidth coefficient using the quality factor Q. .

[0149] Then, based on the above coefficients, the transfer function coefficients of each filter are calculated, including the numerator coefficients. and denominator coefficient The specific calculation method is as follows:

[0150] constant term of numerator coefficient The first-order coefficient of the numerator The second-order coefficient of the numerator coefficient ;

[0151] The constant term of the denominator coefficient The coefficient of the first term in the denominator The coefficient of the second term in the denominator .

[0152] The transfer functions of each filter are constructed based on the numerator and denominator coefficients mentioned above. Subsequently, the original audio signal is filtered sequentially through these N cascaded second-order IIR filters to obtain the first intermediate signal. .

[0153] Step 3, Second-level processing: Speaker virtualization processing;

[0154] For the first intermediate signal The second level of processing involves virtualization based on the speaker's characteristics to generate a second intermediate signal. .

[0155] In this embodiment, the system automatically detects and loads the speaker's characteristic file. This file supports two formats: one is a two-column time-domain impulse response file. Secondly, three frequency domain response files. .

[0156] If a frequency domain response file is loaded, the system performs linear interpolation on the sound pressure level and phase at the missing frequencies in the file to obtain the speaker frequency response over the complete frequency range. Then, the first intermediate signal is frequency-domain filtered based on this frequency response to simulate the playback characteristics of the speaker.

[0157] If a time-domain impulse response file is loaded Then, the first intermediate signal and the impulse response are directly convolved in the time domain. This achieves the same purpose of simulating the playback characteristics of a speaker.

[0158] After processing, the system outputs a signal that has undergone speaker virtualization processing, which serves as the second intermediate signal. .

[0159] Step 4, Third-level processing: Head-related impulse response convolution processing:

[0160] For the second intermediate signal The third level of processing involves convolution with the head-related impulse response (HRIR) to generate a binaural stereo signal. .

[0161] Specifically, the system loads a standard-format head-related impulse response file, which supports multiple position selections for azimuth or elevation angles. Based on the user-selected azimuth or elevation angle, the corresponding left ear impulse response is extracted from this file. and right ear impulse response .

[0162] Then, the second intermediate signal is convolved with the left ear head correlation impulse response and the right ear head correlation impulse response, respectively, to generate the left channel signal and the right channel signal, i.e.:

[0163] left channel signal Right channel signal .

[0164] The left and right channel signals are used as binaural stereo signals. Output.

[0165] Step 5, Fourth Level Processing: Headphone Equalization Correction Processing

[0166] For binaural stereo signals The fourth level of processing is performed, which involves equalization correction based on the headphone's frequency response characteristics to generate the output signal. .

[0167] In this embodiment, the system first acquires the impulse response of the currently playing headphone. Then, the impulse response is inversely processed to generate the equalization filter coefficients. The specific process is as follows:

[0168] Impulse response of headphones Perform a Fourier transform to obtain the frequency response H(f) of the headphones; based on this frequency response and its conjugate... Calculate the power spectrum Based on the power spectrum and the preset regularization factor β (within the range of 0 < β < 0.1), the frequency response of the equalization filter is calculated using the following formula:

[0169]

[0170] in, The maximum value of the power spectrum is represented by the regularization factor β, which is used to prevent excessive gain boosting at deep-spectrum notches, thus ensuring filter stability. Finally, the frequency response of the equalization filter is discussed. Performing an inverse Fourier transform yields the equalization filter coefficients in the time domain. .

[0171] Based on the equalization filter coefficients, temporal convolution processing is performed on the left and right channels of the binaural stereo signal respectively:

[0172]

[0173]

[0174] If the user chooses to bypass this level of processing, then directly... (Left and right channels remain unchanged). Through the above filtering process, the frequency response characteristics of the headphones can be corrected, the coloration of the sound by the headphones themselves can be compensated, the frequency response of the headphones will tend to be flat, and the final output signal will be obtained. ).

[0175] Based on the equalization filter coefficients, the binaural stereo signals (left and right channels respectively) are filtered to correct the frequency response characteristics of the headphones, compensate for the coloration of the sound by the headphones themselves, and finally obtain the output signal. ).

[0176] Step Six: Sound Quality Evaluation (Optional Step):

[0177] During the above processing stages, the system can also process the original audio signal ( ), first intermediate signal ( ), second intermediate signal ( ), binaural stereo signal ( ) and output signal ( The data are fed into the 5th level sound quality evaluation module (e.g., an audio analyzer or a subjective listening evaluation system) to compare and analyze the impact of each level of processing on sound quality, thereby assisting in debugging or verifying algorithm performance.

[0178] On the other hand, another audio signal processing method can also be provided. The main difference from the previous embodiment is that this embodiment combines the speaker virtualization processing and head-related impulse response convolution processing into one stage, and omits the separate simulation of speaker characteristics, instead directly using the head-related impulse response of the speaker and the acoustic environment for convolution.

[0179] The specific process is as follows:

[0180] Step 1: Obtain the raw audio signal:

[0181] Similar to Example 1, the system loads the WAV file to obtain the original audio signal. ).

[0182] Step Two, First-Level Processing, Balanced Compensation Processing:

[0183] For the original audio signal ( The first stage of processing is performed, which involves equalization compensation through multiple cascaded second-order IIR filters to generate the first intermediate signal. ,Right now The specific calculation method for this step is exactly the same as the first-level processing in Example 1, and will not be repeated here.

[0184] Step 3, Second-level processing: Head-related impulse response convolution processing of the loudspeaker and the acoustic environment:

[0185] For the first intermediate signal ( The second level of processing is performed, which involves virtualization based on the combined characteristics of the speaker and the acoustic environment to generate a second intermediate signal. ).

[0186] In this embodiment, the system directly loads a pre-measured or simulated head-related impulse response file that includes the combined effects of the speaker and the acoustic environment (such as room reflections and reverberation). This file is also in a standard format and supports multiple position selections for azimuth or pitch angles. Based on the user-selected azimuth or pitch angle, the corresponding left ear joint impulse response is extracted. Combined impulse response with the right ear .

[0187] Subsequently, the first intermediate signal is convolved with the left ear joint impulse response and the right ear joint impulse response respectively to generate the left channel signal and the right channel signal:

[0188] left channel signal Right channel signal .

[0189] The left and right channel signals are combined into a second intermediate signal. This signal is a binaural stereo signal that already includes speaker virtualization and binaural spatialization.

[0190] Step 4, Third-level processing: Headphone equalization calibration:

[0191] For the second intermediate signal ( The third level of processing is performed, which involves equalization correction based on the headphone's frequency response characteristics to generate an output signal. ).

[0192] The specific implementation of this step is exactly the same as the previous fourth-level processing: obtain the impulse response of the headphones, perform inverse operations to generate equalization filter coefficients (including Fourier transform, power spectrum calculation, regularization and inverse Fourier transform), then filter the left and right channels of the second intermediate signal respectively to correct the frequency response characteristics of the headphones, and finally obtain the output signal. ).

[0193] Step 5: Sound Quality Evaluation (Optional Step):

[0194] Similar to the previous embodiment, the system can convert the original audio signal ( ), first intermediate signal ( ), second intermediate signal ( ) and output signal ( The signal is sent to the Level 4 audio quality evaluation module to analyze the impact of each processing level on audio quality.

[0195] As can be seen from the above two embodiments, the audio signal processing method provided by the present invention, through four-level cascaded processing (equalization compensation, speaker virtualization, head-related impulse response convolution, and headphone equalization correction), can process ordinary stereo sound source signals into binaural output signals with realistic spatial sense and good headphone adaptability. Furthermore, each processing stage can be flexibly combined or adjusted according to actual application scenarios, exhibiting good adaptability and scalability.

[0196] Figure 6 This is a comparative diagram of the original audio signal and the signal waveforms after each stage of processing provided in this application. From top to bottom, they are: Original audio signal After the first level of EQ compensation After the second-level speaker virtualization Binaural stereo signal after third-level HRIR convolution (Left and right channels), output signal after fourth-level headphone equalization (Left and right channels). From Figure 6 The waveform comparison clearly shows that after EQ compensation, the original dry signal ( Low-frequency components are enhanced and mid-frequency standing waves are suppressed; after speaker virtualization ( The signal incorporates the resonance and attenuation characteristics of the loudspeaker; after HRIR convolution ( The left and right channels produce differential time delays and spectral coloration introduced by the head-related transfer function; after headphone equalization ( The inherent frequency response defects of the headphones are compensated, and the left and right channel signals restore spectral balance while maintaining spatial information. Figure 6 The waveform differences of each signal in the process verify the effectiveness of the four-level cascaded processing in this application.

[0197] The audio signal processing method provided in this application, after acquiring the original audio signal, sequentially performs equalization compensation processing, virtualization processing based on speaker characteristics, convolution processing based on head-related impulse response, and equalization correction processing based on headphone frequency response characteristics. This enables targeted correction of the original signal spectrum, target speaker playback characteristics, binaural spatial hearing effect, and headphone playback deviation in the signal link, thereby reducing dependence on the physical listening environment and testing costs, and improving the fidelity, flexible adaptation capability, and consistency of sound quality evaluation results of virtual speaker playback.

[0198] Figure 7 A schematic diagram of the audio signal processing device provided in this application is shown below. Figure 7 As shown, the audio signal processing device 40 provided in this embodiment includes:

[0199] Acquisition module 401 is used for the raw audio signal;

[0200] The first processing module 402 is used to perform the first-level processing on the original audio signal, that is, to perform equalization compensation processing through multiple cascaded filters to generate the first intermediate signal.

[0201] The second processing module 403 is used to perform a second-level processing on the first intermediate signal, that is, to perform virtualization processing according to the speaker characteristics to generate a second intermediate signal.

[0202] The third processing module 404 is used to perform third-level processing on the second intermediate signal, namely, to generate a binaural stereo signal by performing convolution processing with head-related impulse response.

[0203] The output module 405 is used to perform the fourth level of processing on the binaural stereo signal, that is, to perform equalization correction processing based on the frequency response characteristics of the headphones to generate the output signal.

[0204] In one possible embodiment, the first processing module 402 is used to obtain the center frequency, quality factor and gain of each filter;

[0205] The amplitude coefficient is calculated based on the gain, the angular frequency is calculated based on the center frequency and the sampling frequency, and the bandwidth coefficient is calculated based on the angular frequency and the quality factor.

[0206] The transfer function coefficients of each filter are calculated based on the amplitude coefficient and bandwidth coefficient. The transfer function coefficients include the numerator coefficient and the denominator coefficient.

[0207] Construct the transfer function based on the transfer function coefficients of each filter;

[0208] The original audio signal is filtered sequentially through multiple second-order infinite impulse response filters to obtain the first intermediate signal.

[0209] In one possible embodiment, the transfer function coefficients in the first processing module 402 are calculated as follows: the constant term of the numerator coefficient is the sum of the amplitude coefficient and the bandwidth coefficient; the first-order coefficient of the numerator coefficient is twice the negative of the cosine of the angular frequency; and the second-order coefficient of the numerator coefficient is the constant term minus the product of the amplitude coefficient and the bandwidth coefficient.

[0210] The constant term of the denominator coefficients is the sum of the reciprocal of the amplitude coefficient and the bandwidth coefficient. The first-order coefficient of the denominator coefficients is twice the negative of the cosine of the angular frequency. The second-order coefficient of the denominator coefficients is the constant term minus the product of the reciprocal of the amplitude coefficient and the bandwidth coefficient.

[0211] In one possible embodiment, the second processing module 403 is used to acquire the impulse response or frequency response of the speaker;

[0212] The first intermediate signal is subjected to time-domain convolution processing based on the impulse response, or frequency-domain filtering processing based on the frequency response, to simulate the playback characteristics of a loudspeaker.

[0213] The output signal, after being processed by speaker virtualization, serves as the second intermediate signal.

[0214] In one possible embodiment, the second processing module 403 is used to obtain the frequency domain characteristic file of the loudspeaker, which includes frequency, sound pressure level and phase information.

[0215] Interpolation is performed on the sound pressure level and phase at missing frequencies in the frequency domain characteristic file to obtain the loudspeaker frequency response over the complete frequency range.

[0216] In one possible embodiment, the third processing module 404 is used to acquire head-related impulse responses corresponding to the left and right ears of the listener;

[0217] The second intermediate signal is convolved with the left ear head correlation impulse response and the right ear head correlation impulse response respectively to generate the left channel signal and the right channel signal;

[0218] The left and right channel signals are output as binaural stereo signals.

[0219] In one possible embodiment, the third processing module 404 is used to load a standard format head-related impulse response file, which supports multiple position selections for azimuth or elevation angles.

[0220] Based on the selected azimuth or pitch angle, extract the corresponding left and right ear impulse responses from the head-related impulse response file.

[0221] In one possible embodiment, the output module 405 is used to acquire the impulse response of the headphones;

[0222] Perform inverse operations on the impulse response of the headphones to generate equalization filter coefficients;

[0223] The binaural stereo signal is filtered according to the equalization filter coefficients to correct the frequency response characteristics of the headphones and obtain the output signal.

[0224] In one possible embodiment, the output module 405 is used to perform a Fourier transform on the impulse response of the headphones to obtain the frequency response of the headphones;

[0225] Calculate the power spectrum based on the frequency response of the headphones and its conjugate.

[0226] The frequency response of the equalization filter is calculated based on the power spectrum and the preset regularization factor.

[0227] The frequency response of the equalization filter is subjected to inverse Fourier transform to obtain the equalization filter coefficients in the time domain.

[0228] The audio signal processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0229] Figure 8 This is a schematic diagram of the audio signal processing device provided in this application. Figure 8 As shown, the audio signal processing device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the audio signal processing device 50 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus 504.

[0230] In the specific implementation process, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above method.

[0231] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0232] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0233] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0234] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0235] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0236] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0237] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0238] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0239] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0240] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0241] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0242] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0243] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0244] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. An audio signal processing method, characterized in that, include: Acquire the raw audio signal; The original audio signal is processed in the first stage, which involves equalization compensation through multiple cascaded filters to generate a first intermediate signal. The first intermediate signal is processed in the second stage, that is, virtualization is performed according to the characteristics of the speaker to generate a second intermediate signal; The second intermediate signal is processed in the third stage, namely, by convolution with the head-related impulse response to generate a binaural stereo signal; The binaural stereo signal is processed in the fourth stage, which involves equalization correction based on the frequency response characteristics of the headphones to generate an output signal.

2. The method according to claim 1, characterized in that, Equalization compensation is performed through multiple cascaded filters, including: Obtain the center frequency, quality factor, and gain of each filter; The amplitude coefficient is calculated based on the gain, the angular frequency is calculated based on the center frequency and the sampling frequency, and the bandwidth coefficient is calculated based on the angular frequency and the quality factor. The transfer function coefficients of each filter are calculated based on the amplitude coefficient and the bandwidth coefficient, wherein the transfer function coefficients include numerator coefficients and denominator coefficients; The transfer function is constructed based on the transfer function coefficients of each filter; The original audio signal is sequentially filtered through the plurality of second-order infinite impulse response filters to obtain the first intermediate signal.

3. The method according to claim 2, characterized in that, The transfer function coefficients are calculated as follows: The constant term of the numerator coefficient is the sum of the amplitude coefficient and the bandwidth coefficient; the first-order term coefficient of the numerator coefficient is twice the negative of the cosine of the angular frequency; and the second-order term coefficient of the numerator coefficient is the constant term minus the product of the amplitude coefficient and the bandwidth coefficient. The constant term of the denominator coefficient is the sum of the reciprocal of the amplitude coefficient and the bandwidth coefficient. The first-order term coefficient of the denominator coefficient is twice the negative of the cosine of the angular frequency. The second-order term coefficient of the denominator coefficient is the constant term minus the product of the reciprocal of the amplitude coefficient and the bandwidth coefficient.

4. The method according to claim 1, characterized in that, The first intermediate signal undergoes a second level of processing, namely, virtualization processing based on the speaker characteristics, to generate a second intermediate signal, including: Obtain the impulse response or frequency response of the speaker; The first intermediate signal is subjected to time-domain convolution processing based on the impulse response, or to frequency-domain filtering processing based on the frequency response, to simulate the playback characteristics of a loudspeaker. The signal processed by speaker virtualization is output as the second intermediate signal.

5. The method according to claim 4, characterized in that, Obtain the frequency response of the loudspeaker, including: Obtain the frequency domain characteristic file of the loudspeaker, which includes frequency, sound pressure level and phase information; The sound pressure level and phase at the missing frequencies in the frequency domain characteristic file are interpolated to obtain the loudspeaker frequency response over the complete frequency range.

6. The method according to claim 1, characterized in that, The second intermediate signal undergoes a third level of processing, namely, convolution processing with the head-related impulse response to generate a binaural stereo signal, including: Acquire head-related impulse responses corresponding to the left and right ears of the listener; The second intermediate signal is convolved with the left ear head related impulse response and the right ear head related impulse response to generate the left channel signal and the right channel signal. The left channel signal and the right channel signal are output as binaural stereo signals.

7. The method according to claim 6, characterized in that, Acquire head-related impulse responses corresponding to the left and right ears of the listener, including: Load a standard format head-related impulse response file, which supports multiple position selections for azimuth or elevation angles; Based on the selected azimuth or pitch angle, the corresponding left ear impulse response and right ear impulse response are extracted from the head-related impulse response file.

8. The method according to claim 1, characterized in that, The binaural stereo signal undergoes a fourth level of processing, namely, equalization correction based on the headphone frequency response characteristics, to generate an output signal, including: Obtain the impulse response of the headphones; The impulse response of the headphones is inversely processed to generate equalization filter coefficients; The binaural stereo signal is filtered according to the equalization filter coefficients to correct the frequency response characteristics of the headphones and obtain the output signal.

9. The method according to claim 8, characterized in that, The impulse response of the headphones is inversely processed to generate equalization filter coefficients, including: The frequency response of the headphones is obtained by performing a Fourier transform on the impulse response of the headphones. The power spectrum is calculated based on the frequency response of the headphones and its conjugate. The frequency response of the equalization filter is calculated based on the power spectrum and the preset regularization factor. The frequency response of the equalization filter is subjected to inverse Fourier transform to obtain the equalization filter coefficients in the time domain.

10. An audio signal processing device, characterized in that, include: The acquisition module is used for the raw audio signal; The first processing module is used to perform first-level processing on the original audio signal, that is, to perform equalization compensation processing through multiple cascaded filters to generate a first intermediate signal. The second processing module performs a second-level processing on the first intermediate signal, namely, virtualization processing based on the speaker characteristics to generate a second intermediate signal; The third processing module is used to perform third-level processing on the second intermediate signal, namely, to generate a binaural stereo signal by performing convolution processing with head-related impulse response. The output module is used to perform fourth-level processing on the binaural stereo signal, namely, equalization correction processing based on the frequency response characteristics of the headphones, to generate an output signal.

11. An audio signal processing device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-9.

13. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-9.