Audio processing method, apparatus, device, and storage medium

By using a minimum-phase FIR low-pass filter and a linear-phase high-pass filter to process stereo signals, the problems of blurred sound image positioning and auditory imbalance caused by nonlinear phase distortion in stereo sound field adjustment are solved, achieving high fidelity and high-fidelity audio processing effects.

CN122093733APending Publication Date: 2026-05-26GUANGZHOU KUGOU COMP TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU KUGOU COMP TECH CO LTD
Filing Date
2026-02-13
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies, in stereo sound field adjustment, introduce nonlinear phase distortion due to IIR filters, resulting in blurred sound image positioning and auditory imbalance, which affects sound quality and fidelity.

Method used

The low-frequency signal is extracted using an FIR low-pass filter with minimum phase properties, and the mid-to-high frequency signal is processed by a linear phase high-pass filter to ensure the integrity of the energy distribution and phase characteristics of the low-frequency signal. The high-frequency signal is then mixed with the low-frequency signal after sound field adjustment.

Benefits of technology

It achieves high fidelity and high-fidelity stereo sound field processing, solving the problems of blurred sound image positioning and auditory imbalance, and ensuring the naturalness and accuracy of sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093733A_ABST
    Figure CN122093733A_ABST
Patent Text Reader

Abstract

This application discloses an audio processing method, device, and storage medium, belonging to the audio field. The method includes: filtering the left and right channel signals of the original audio using a first low-pass filter to obtain low-frequency left and right channel signals; the first low-pass filter is a finite impulse response low-pass filter with minimum phase property; filtering the left and right channel signals of the original audio using a first high-pass filter to obtain high-frequency left and right channel signals; and mixing the high-frequency left and right channel signals with the low-frequency left and right channel signals after sound field adjustment to obtain a stereo output signal. This solution achieves flexible sound field adjustment while completely solving the problem of blurred sound image positioning caused by nonlinear phase distortion in traditional solutions, as well as the problem of auditory imbalance caused by changes in low-frequency energy distribution, thereby achieving a high-fidelity and high-reproducibility stereo sound field processing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio signal processing, and in particular to an audio processing method, apparatus, device, and storage medium. Background Technology

[0002] When adjusting the sound field of stereo, the sound field can usually be expanded or narrowed by switching the left and right channels (LR) to the middle-side (MS) and scaling the M and S signals. However, this process often causes abnormal changes in the distribution of low-frequency energy. Since the low-frequency energy is mainly concentrated in the M signal after the LR to MS conversion, directly scaling the M and S signals will disrupt the balance of the original low-frequency response, thus affecting the naturalness and accuracy of the listening experience.

[0003] Currently, to solve the above problems, a frequency division processing scheme is usually adopted. This involves first extracting the low-frequency signal using a low-pass filter and keeping it unprocessed, while simultaneously using a high-pass filter to separate the mid-to-high-frequency signals. These mid-to-high-frequency signals are then converted into MS signals for sound field adjustment. Finally, the processed mid-to-high-frequency signals are mixed with the original low-frequency signal. Because infinite impulse response (IIR) filters have low order and high computational efficiency, this method commonly uses IIR filters for frequency division.

[0004] However, the above-mentioned solutions have obvious drawbacks. IIR filters introduce nonlinear phase distortion, which can easily lead to sound quality degradation in high-fidelity audio systems. In particular, it can cause blurred sound image localization in the auditory perception, thereby affecting the spatial reproduction and overall fidelity of the sound, resulting in poor audio processing effects. Summary of the Invention

[0005] This application provides an audio processing method, apparatus, device, and storage medium that, while achieving flexible sound field adjustment, completely solves the problems of blurred sound image positioning caused by nonlinear phase distortion in traditional solutions, as well as the auditory imbalance caused by changes in low-frequency energy distribution, thereby achieving high fidelity and high-fidelity stereo sound field processing effects. The technical solution is as follows: According to one aspect of this application, an audio processing method is provided, the method comprising: The left and right channel signals of the original audio are filtered by a first low-pass filter to obtain low-frequency left and right channel signals. The first low-pass filter is a finite impulse response low-pass filter with minimum phase property. The left and right channel signals of the original audio are filtered by the first high-pass filter to obtain high-frequency left and right channel signals; After the high-frequency left and right channel signals are adjusted for sound field, they are mixed with the low-frequency left and right channel signals to obtain a stereo output signal.

[0006] According to another aspect of this application, an audio processing apparatus is provided, the apparatus comprising: The low-pass filter module is used to filter the left and right channel signals of the original audio through the first low-pass filter to obtain low-frequency left and right channel signals. The first low-pass filter is a finite impulse response low-pass filter with minimum phase property. A high-pass filter module is used to filter the left and right channel signals of the original audio through the first high-pass filter to obtain high-frequency left and right channel signals; The mixing module is used to adjust the sound field of the high-frequency left and right channel signals and mix them with the low-frequency left and right channel signals to obtain a stereo output signal.

[0007] According to another aspect of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the audio processing method as described above.

[0008] According to another aspect of this application, a computer-readable storage medium is provided, wherein a computer program is stored therein, the computer program being loaded and executed by a processor to implement the audio processing method as described above.

[0009] According to another aspect of this application, a computer program product is provided, comprising a computer program executed by a processor to implement the audio processing methods provided in various alternative implementations of the foregoing aspects.

[0010] This application provides an audio processing solution that extracts low-frequency signals using an FIR low-pass filter with minimum phase properties. This ensures accurate frequency division while maintaining the transient response of the low-frequency signal, significantly reducing processing delay in the low-frequency path and minimizing signal distortion. By processing the mid-to-high frequency signals after high-pass filtering, while directly mixing the low-frequency signals after minimum phase filtering, the original low-frequency energy distribution and phase characteristics are fully preserved. This solution achieves flexible sound field adjustment while completely resolving the problems of blurred sound image positioning caused by nonlinear phase distortion and auditory imbalance caused by changes in low-frequency energy distribution in traditional solutions, thus achieving high fidelity and high-fidelity stereo sound field processing effects. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of the implementation environment of an audio processing method provided in an embodiment of this application; Figure 2 This is a flowchart of an audio processing method provided according to an embodiment of this application; Figure 3 This is a flowchart illustrating another audio processing method provided according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an audio processing device provided according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a computer device provided according to an embodiment of this application.

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0015] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0016] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0017] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the media resources involved in this application were obtained with full authorization.

[0018] Figure 1 This is a schematic diagram illustrating the implementation environment of an audio processing method according to an embodiment of this application. See also... Figure 1 The implementation environment specifically includes: terminal device 101 and server 102. Terminal device 101 can be connected to server 102 via wireless network or wired network.

[0019] Terminal device 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, tablet, smart TV, smart speaker, in-vehicle terminal, and laptop computer. An application can be installed and run on terminal device 101, adapted to scenarios such as code development, code review, operation and maintenance, and network security. This application is associated with server 102, which provides background services to terminal device 101.

[0020] Terminal device 101 can refer to one of a plurality of terminal devices. This embodiment uses terminal device 101 as an example. Those skilled in the art will know that the number of the above-mentioned terminal devices can be more or less. For example, there can be several, dozens or hundreds, or more terminal devices. This application embodiment does not limit the number or type of terminal devices.

[0021] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of servers can be more or less, and this embodiment does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services. In some embodiments, server 102 undertakes the main computing work, and terminal device 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal device 101 undertakes the main computing work; or, server 102 and terminal device 101 collaborate on computing using a distributed computing architecture. Server 102 can be connected to terminal device 101 and other terminal devices via a wireless network or a wired network. Optionally, the number of servers can be more or less, and this embodiment does not limit this.

[0022] Figure 2This is a flowchart of an audio processing method provided according to an embodiment of this application, such as... Figure 2 As shown, this method is executed by a computer device, which can be... Figure 1 The terminal device 101 shown includes the following steps.

[0023] Step 201: Filter the left and right channel signals of the original audio through the first low-pass filter to obtain low-frequency left and right channel signals.

[0024] In this embodiment, the left and right channel signals (L / R) of the original audio are the unprocessed raw stereo audio data. The first low-pass filter is a finite impulse response low-pass filter with minimum phase property. Low-pass means allowing low frequencies to pass. The low-frequency left and right channel signals (L_low / R_low) are the processed output. In other words, the low-frequency left and right channel signals are the components in the original audio below a specific frequency (i.e., the "cutoff frequency," such as 200Hz), mainly containing the energy of heavy sounds such as bass and kick drums in music. This step is to extract the low-frequency components from the input raw audio (the left and right channels of stereo) to form separate low-frequency signals.

[0025] It's important to note that sound signals are composed of sine waves of different frequencies. When a sound signal passes through a filter, not only may the amplitude of each frequency component be altered, but the position (phase) of the sound signal on the time axis will also shift. This shift is called phase distortion. Among all filters with the same amplitude-frequency response, the minimum-phase filter has the smallest group delay. That is, after a sound signal passes through a filter with the minimum phase property, the energy of the sound signal arrives earliest, and the waveform is "delayed" the least on the time axis.

[0026] Step 202: Filter the left and right channel signals of the original audio through the first high-pass filter to obtain high-frequency left and right channel signals.

[0027] In the embodiments of this application, this step is the second stage of the audio processing method, the purpose of which is to extract the mid-high frequency components from the input original audio (the left and right channels of stereo) to form a separate high frequency signal, and to ensure that this extraction process does not introduce any phase distortion that would destroy the sense of sound spatial positioning.

[0028] Correspondingly, the function of the first high-pass filter complements that of the first low-pass filter described above. The high-frequency left and right channel signals (L_high / R_high) are the outputs after processing by the first high-pass filter. The high-frequency left and right channel signals are the components of the original signal above a specific cutoff frequency (e.g., 200Hz, which is complementary to the low-pass filter), and contain information that determines sound detail, clarity, and spatiality, such as vocals, main instruments, and cymbals.

[0029] Optionally, the first high-pass filter is a finite impulse response high-pass filter with linear phase properties. Linear phase means that the time delay caused by the filter for all frequency components is constant. All frequencies are "shifted" by exactly the same distance on the time axis after passing through the filter, and their relative timing relationships are not distorted. The spatial sense and sound image localization of stereo sound are highly dependent on the tiny phase (time) difference between the left and right channels. Non-linear phase distortion randomly distorts the phase relationship between different frequencies, disrupting this delicate time difference and leading to blurred, diffused, and inaccurate sound image localization. The first high-pass filter does not change the shape of the composite waveform (assuming a flat amplitude response within the passband), only the overall delay. This is crucial for signals requiring subsequent complex processing (such as sound field processing), because the signal input to the processing module is itself clean and has accurate timing relationships.

[0030] Step 203: After adjusting the sound field of the high-frequency left and right channel signals, mix them with the low-frequency left and right channel signals to obtain a stereo output signal.

[0031] In this embodiment of the application, the low-frequency signal and high-frequency signal, which have been processed independently and optimized, are recombined according to preset rules to generate a new stereo signal that has been adjusted in sound field and whose overall sound quality has been enhanced with fidelity.

[0032] The sound field adjustment of the high-frequency left and right channel signals involves converting the linear-phase, pure, and phase-distortion-free mid-to-high-frequency left and right channel signals from the previous step into center-side signals for processing. This is the core algorithm for achieving the effect of "widening" or "narrowing" the sound field. Then, the adjusted M and S signals are converted back to the left and right channel format to obtain the sound field-adjusted high-frequency signals. Finally, the unprocessed, original low-frequency energy is seamlessly aligned and superimposed in the time domain with the precisely spatially adjusted mid-to-high-frequency details to obtain the stereo output signal.

[0033] This application provides an audio processing solution that extracts low-frequency signals using an FIR low-pass filter with minimum phase properties. This ensures accurate frequency division while maintaining the transient response of the low-frequency signal, significantly reducing processing delay in the low-frequency path and minimizing signal distortion. By processing the mid-to-high frequency signals after high-pass filtering, while directly mixing the low-frequency signals after minimum phase filtering, the original low-frequency energy distribution and phase characteristics are fully preserved. This solution achieves flexible sound field adjustment while completely resolving the problems of blurred sound image positioning caused by nonlinear phase distortion and auditory imbalance caused by changes in low-frequency energy distribution in traditional solutions, thus achieving high fidelity and high-fidelity stereo sound field processing effects.

[0034] The above Figure 2The diagram shows the main flow of an audio processing method provided in this application. The audio processing scheme will be further explained below. Figure 3 This is a schematic flowchart of another audio processing method provided according to an embodiment of this application. The method is executed by a computer device, which can be... Figure 1 The terminal device 101 shown is as follows: Figure 3 As shown, the method includes the following steps.

[0035] Step 301: Create a prototype low-pass filter that meets the target frequency response specifications using the frequency sampling method or the window function method.

[0036] In this embodiment, this step involves first obtaining an initial set of coefficients for an FIR low-pass filter, referred to as a prototype. The prototype is a filter with linear phase. The reason for starting with linear phase is that standard FIR design methods (such as the window function method and frequency sampling method) most easily and directly generate linear phase filters. The target frequency response specifications include at least one of the following: a cutoff frequency of 200Hz, passband ripple less than or equal to 0.1dB, and stopband attenuation greater than or equal to 60dB.

[0037] The process of creating it using the window function method is described below.

[0038] First, the principle of this scheme is to create an "ideal" low-pass filter whose frequency response is 1 in the passband, 0 in the stopband, and a step at the cutoff frequency. This ideal filter is infinitely long and non-causal in the time domain. To obtain a finite-length, causal FIR filter, the infinitely long impulse response of this ideal filter is captured through a finite-length "window".

[0039] Then, given the cutoff frequency (e.g., 200Hz) and sampling frequency Calculate the unit impulse response of an ideal low-pass filter. See formula (1) below.

[0040] (1).

[0041] in, N It is the filter order (length), and the sinc function is sinc( x )=sin( x ) / x .

[0042] Next, select a window function. The rectangular window has the narrowest transition band, but its stopband attenuation is poor (approximately 21 dB), failing to meet the ≥60 dB requirement. The Hanning window has a stopband attenuation of approximately 44 dB, still insufficient. The Hamming window has a stopband attenuation of approximately 53 dB, close but possibly still insufficient. The Blackman window has a stopband attenuation of up to 74 dB, meeting the requirement, but has the widest transition band. To achieve the narrowest possible transition band while meeting the 60 dB attenuation requirement, either the Kaiser window or the Dolph-Chebyshev window can be selected. The Kaiser window is more commonly used because its shape parameter β can be continuously adjusted, achieving the narrowest main lobe width for a given stopband attenuation.

[0043] Then, the window function With unit impulse response Multiplying them yields the causal FIR filter coefficients: .

[0044] Finally, calculate The actual frequency response is used to verify whether the passband ripple and stopband attenuation meet the requirements. If not, the filter order N needs to be increased or the window parameters adjusted, and the filter recreated.

[0045] The process of creating it using the frequency sampling method is described below.

[0046] First, the principle of this scheme is directly in the frequency domain, in N Equally spaced frequency points Above, specify the frequency response amplitude we want. For example, the passband frequency is set to 1 and the stopband frequency is set to 0, and then the time-domain coefficients are obtained by inverse discrete Fourier transform.

[0047] Then, based on the cutoff frequency, determine which frequency belongs to the passband. Set it to 1; this will affect the stopband. Set it to 0. In the transition band, 1-2 samples can be set to an intermediate value (e.g., 0.5) to optimize performance. Specifically, to obtain a linear-phase filter, the frequency response amplitude... It needs to satisfy conjugate symmetry: that is In addition, in order to obtain a Type I linear phase filter (with symmetric coefficients and N being odd), the amplitude response must also satisfy even symmetry.

[0048] Finally, for Perform an N-point inverse discrete Fourier transform to obtain the time-domain impulse response. .

[0049] It should be noted that frequency sampling can precisely control the response at a specific frequency point, offering greater flexibility for designs requiring specific frequency response shapes. However, it typically requires more iterations and optimizations to achieve the desired overall passband flatness and stopband attenuation.

[0050] It's important to note that the choice of a cutoff frequency of 200Hz is based on the division of the Bark band and the perceptual characteristics of human hearing. According to the Bark band model, the human ear's frequency discrimination is not linear. The center frequency of the second Bark band is approximately 150Hz, with its boundary range roughly between 100Hz and 200Hz. Setting the cutoff frequency at 200Hz essentially places it near the upper boundary of this Bark band. This design allows the separated low-frequency components (below 200Hz) to be largely controlled within the same Bark band. While the phase distortion of the minimum phase filter increases with frequency in the passband, the phase distortion generated near 200Hz falls precisely within the auditory masking effect dominated by the center frequency of 150Hz, making it difficult for the human ear to perceive. Furthermore, 200Hz is also close to the upper limit of the energy concentration area of ​​most "pure" low frequencies (such as bass and kick drum) in music. Crossover at this point effectively separates low frequencies while preserving the integrity of the mid-to-high frequency sound field information to the maximum extent. Of course, 200Hz is not the only solution. It represents an optimal balance between auditory psychoacoustic models and engineering practice, and can be adjusted at other Bark frequency band boundaries (such as 100Hz or 300Hz) depending on the specific application scenario.

[0051] Step 302: Process the prototype low-pass filter using Hilbert transform or spectral decomposition to obtain the first low-pass filter.

[0052] In this embodiment, the prototype low-pass filter generated by the above steps is linear-phase. By performing a phase characteristic transformation from "linear phase" to "minimum phase" while keeping the amplitude-frequency characteristics (i.e., the filter's "filtering" capability) essentially unchanged, a new low-pass filter is obtained, called the first low-pass filter. That is, the first low-pass filter is a finite impulse response low-pass filter with minimum phase property.

[0053] Due to frequency response It can be decomposed into amplitude spectrum and phase spectrum A linear-phase prototype has a known amplitude spectrum and a linear phase spectrum. A minimum-phase system is defined as having the smallest phase delay among all causal stable systems with the same amplitude spectrum, and all its zeros lie within the unit circle. Finding this minimum-phase system with the same amplitude spectrum yields the first low-pass filter.

[0054] The following describes the processing procedure using the Hilbert transform method.

[0055] First, the principle of this scheme is that the complex cepstrum of a minimum-phase sequence is causal. The essence of this scheme is to set the negative time part of the complex cepstrum of the original sequence (which may be non-minimum-phase) to zero in the complex cepstrum domain, thereby forcing a minimum-phase sequence to be obtained.

[0056] Then, for the linear phase prototype Perform a discrete Fourier transform to obtain .

[0057] Then, calculate its complex logarithm: Unwrap is a phase unwinding operation that ensures phase continuity.

[0058] Then, the minimum phase component is extracted. A minimum phase complex cepstrum is constructed. This operation forces the complex cepstral to be causal, thus corresponding to a minimum-phase system. Wherein: .

[0059] Finally, the minimum phase sequence is recovered. Perform a discrete Fourier transform to obtain Calculate the complex exponent: .right Performing the inverse discrete Fourier transform yields the final minimum phase filter coefficients. .

[0060] The following describes the processing procedure using spectral decomposition.

[0061] First, the principle of this scheme is that the system corresponding to the autocorrelation sequence of a linear-phase FIR filter (or its inverse DFT of zero-phase frequency response) is a zero-phase system. By performing polynomial factorization on the system function of this zero-phase system and selecting all zeros within the unit circle, the resulting new system is a minimum-phase system with the same amplitude response.

[0062] Then, the autocorrelation sequence / zero-phase response of the prototype is calculated. Linear phase prototype. It is real symmetric. Calculate its autocorrelation sequence. This is equivalent to taking The inverse DFT of the amplitude spectrum yields an even-symmetric sequence of length 2N-1. The frequency response corresponding to this sequence... It is zero-phase and non-negative.

[0063] Then, the autocorrelation sequence r [ m Consider it as a z Transformation R ( z The coefficients of the polynomial. Finding the polynomial. R (z The root of ). Because and r [ m Symmetrical, the roots always appear in the form of reciprocal pairs of conjugates, that is, if It is a root, then , , They are all roots.

[0064] Then, construct the minimum-phase system. From each set of reciprocals of the roots, select all roots with a modulus less than or equal to 1 (i.e., roots inside or on the unit circle). Multiplying the factors corresponding to these roots together constitutes the minimum-phase system function. .

[0065] Finally, the system function Expand into z The polynomial form of 1, whose coefficients are the coefficients of the minimum phase filter we are looking for. .

[0066] Step 303: Filter the left and right channel signals of the original audio through the first low-pass filter to obtain low-frequency left and right channel signals.

[0067] In this embodiment, the purpose of this step is to accurately separate the low-frequency components (e.g., below 200Hz) from the original stereo signal, forming an independent, unprocessed, and uncontaminated low-frequency signal stream. This is because in stereo sound field processing (such as MS processing), directly performing a "middle-side" transformation and scaling on a full-range signal containing strong low-frequency energy leads to a serious problem: an unbalanced distribution of low-frequency energy. Since most of the low-frequency energy is concentrated in the "middle" signal, scaling the "side" signal disproportionately alters the overall listening experience of the low frequencies, making the sound muddy or thin. Therefore, separating and "protecting" the low frequencies is a key strategy in audio processing.

[0068] The left and right channel signals of the original audio are two time-aligned discrete-time sequences L[n] and R[n], representing the pressure changes of the sound at the left and right spatial locations. The signals are typically digitized at high sampling rates (e.g., 44.1kHz, 48kHz, 96kHz) and depths (e.g., 24-bit) to ensure sufficient frequency and dynamic range. Filtering operations are performed on these discrete sample points.

[0069] Filtering is a convolution operation, which can be viewed as a weighted moving average of the input signal and the filter's "impulse response" (i.e., its coefficient sequence). The shape of the filter determines which frequency components are preserved (passband) or suppressed (stopband). The following explanation uses the left channel as an example, see formula (2) below.

[0070] (2).

[0071] in, These are the coefficients of a minimum-phase FIR low-pass filter of length N. L[n] is the original left channel signal. The output low-frequency signal is from the left channel. (Right channel) Obtained through the exact same convolution operation, using the same filter. .

[0072] Finally, the output is the low-frequency left and right channel signal. The frequency domain of this signal primarily contains frequency components below the cutoff frequency (e.g., 200Hz). Components above this frequency are significantly attenuated (reaching or exceeding -60dB). Compared to the low-frequency portion of the original signal, the envelope and transients of this signal's time-domain waveform are largely preserved, but due to the band-limiting characteristics of the filter, its waveform becomes smoother (high-frequency details are removed).

[0073] It should be noted that the reason for using a minimum-phase FIR low-frequency filter to separate low frequencies, instead of directly using a linear-phase FIR filter, is that although a linear-phase filter guarantees no phase distortion at all frequencies, it inherently has a constant group delay (typically N). A 1 / 2 sampling point represents a considerable absolute time delay for low frequencies. This delay must be precisely compensated and aligned when subsequently mixed with high-frequency signals; otherwise, it will lead to phase distortion across the entire frequency range. More importantly, the "pre-echo" generated by its symmetrical impulse response appears before the low-frequency transient, which is counterintuitive and may cause auditory discomfort. For low-frequency signals that are bypassed, not involved in complex processing, and have low phase sensitivity in the human ear, prioritizing the integrity of their temporal energy (minimum phase) is a better solution.

[0074] For example, consider processing a stereo drum recording containing a kick drum (around 60Hz) and a snare drum (complex high frequencies). The inputs L[n] and R[n] contain the booming sound of the kick drum and the crackling sound of the snare drum. A 200Hz minimum phase low-pass filter is used. When the bass drum signal is input, its main 60Hz component passes smoothly. Because the filter is minimum phase, the thumping sound of the drum is heard at the output. , The sound remains crisp and powerful, without any muddiness. When the snare drum signal is input, its rich high-frequency components (above 1kHz) are strongly attenuated. The crisp snare drum sound is almost inaudible at the output, leaving only a very faint, muffled low-frequency remnant. (Final output) , The signal is almost a "pure" bass drum signal, along with the low-frequency portions of instruments such as the bass. This signal retains all the energy and dynamics of the original low frequencies, with natural (non-linear but minimized) phase characteristics.

[0075] Step 304: Perform symmetrical processing on the coefficients of the first low-pass filter to obtain the second low-pass filter.

[0076] In this embodiment, this step involves a mathematical transformation from a minimum-phase system to a linear-phase system. The input to the transformation is the coefficient sequence of the minimum-phase FIR low-pass filter. (Of length N), the output is a new sequence of FIR low-pass filter coefficients with linear phase. (Length is 2N-1). The coefficients are used to indicate the impulse response of the first low-pass filter, and the second low-pass filter is a finite impulse response low-pass filter with linear phase properties.

[0077] In some embodiments, symmetrical processing is achieved through zero-phase frequency response. Accordingly, symmetrical processing is performed on the coefficients of the first low-pass filter to obtain the second low-pass filter, including: obtaining the zero-phase frequency response of the first low-pass filter; and generating a second low-pass filter with linear phase and a length of 2N-1 based on the zero-phase frequency response by frequency sampling or inverse Fourier transform, where N is the order of the first low-pass filter.

[0078] in, It is inherently asymmetric, and its Fourier transform... The phase is non-linear. First, calculate... autocorrelation sequence .

[0079] .

[0080] because It is finite in length and causal; in actual calculations, n ranges from 0 to N-1. Autocorrelation sequence. r [ m ] is an even symmetric sequence ( r [ m ]= r [- m ]), and its length becomes 2 N 1. More importantly, its discrete-time Fourier transform That is, the square of the amplitude spectrum of the original minimum-phase system. Since Its corresponding phase is zero, so this is the frequency response of a zero-phase filter.

[0081] A zero-phase system is non-causal (because its impulse response r[m] is symmetric about m=0) and cannot be directly used for real-time filtering. To obtain a causal, realizable linear-phase filter, the zero-phase impulse response r[m] is shifted so that it lies entirely on the non-negative time axis. r[m] is shifted to the right by (N... A new sequence can be obtained from 1 sample: .

[0082] This new sequence These are the coefficients of the second low-pass filter. Since they originate from a shift of the even-symmetric sequence r[m], therefore... It satisfies symmetry: This is precisely the coefficient symmetry condition for a Type I linear-phase FIR filter. At this point, The frequency response is Its amplitude response is the square of the amplitude response of the original minimum-phase filter (meaning that the passband ripple and stopband attenuation will change, usually becoming better), and its phase response is strictly linear. .

[0083] In some embodiments, the above autocorrelation operation is mathematically completely equivalent to... The first low-pass filter is convolved with its time-reversed sequence. Correspondingly, the coefficients of the first low-pass filter are symmetrically processed to obtain the second low-pass filter, including: convolving the coefficients of the first low-pass filter with the time-reversed sequence of the coefficients to obtain the second low-pass filter. The length of the second low-pass filter is 2N-1, where N is the order of the first low-pass filter.

[0084] The convolution operation is as follows: This is consistent with the autocorrelation formula. Therefore, "symmetry processing" can be implemented through convolution operations, which provides a straightforward computational path for digital signal processors.

[0085] Step 305: Perform spectrum inversion on the second low-pass filter based on the unit pulse sequence to obtain the first high-pass filter.

[0086] In this embodiment of the application, in digital signal processing, "spectral inversion" of a low-pass filter means "flipping" its frequency response around the Nyquist frequency (i.e., half the sampling frequency). Mathematically, this is equivalent to multiplying the impulse response of the low-pass filter by an alternating sequence of symbols in the time domain (filter coefficient domain). And superimpose a unit pulse (DC component adjustment).

[0087] The formula for spectrum inversion is: .

[0088] in, These are the coefficients of the linear phase low-pass filter (i.e., the coefficients of the second low-pass filter). ).

[0089] These are the coefficients of the linear phase high-pass filter to be determined (i.e., the coefficients of the first high-pass filter).

[0090] It is a unit impulse sequence (Kronecker delta function), which is 1 when n=0 and 0 otherwise.

[0091] It is a key time offset, equal to the center point of the low-pass filter group delay, typically 1 / 2. , where L is Length (2N) 1) Select this This is to ensure that the coefficient sequence of the high-pass filter still maintains symmetry, thereby inheriting the linear phase property.

[0092] The following is a brief introduction to the spectrum inversion process.

[0093] First, calculate the length L of the second low-pass filter. Since it is linear-phase (type I), it is an odd-length sequence and symmetric. Its group delay is constant and has a value of [value missing]. One sampling point. This The 'd' in the formula ensures that the unit pulse is placed at the center of symmetry of the filter coefficients.

[0094] Then, a sequence is generated where the value is 1 only at the position n=d, and 0 at all other positions. This is... .

[0095] Then, the constructed unit pulse sequence is combined with the low-pass filter coefficient sequence. Perform point-by-point subtraction: .

[0096] The first high-pass filter is derived by inverting the spectrum of a unit pulse sequence, eliminating the complex process of redesigning a high-performance linear-phase high-pass filter. Instead, a linear-phase high-pass filter with complementary amplitude-frequency characteristics and consistent phase characteristics is directly derived from an existing linear-phase low-pass filter with known performance through a simple time-domain subtraction and precise delay alignment. This ensures that the high-pass and low-pass branches have identical group delays, meaning that all frequency components experience the same time delay in both branches. When the high- and low-frequency signals are finally mixed, they are perfectly aligned in the time domain, preventing waveform distortion or image drift caused by phase differences introduced by the frequency division itself.

[0097] Step 306: Filter the left and right channel signals of the original audio through the first high-pass filter to obtain high-frequency left and right channel signals.

[0098] In the embodiments of this application, the purpose of this step is to non-destructively separate the mid-to-high frequency components above a specific cutoff frequency (such as 200Hz) from the original stereo signal to form an independent high-frequency signal stream without causing any contamination to the "spatial imaging" of the sound.

[0099] Human spatial localization of sound primarily relies on the binaural time difference and binaural intensity difference of mid-to-high frequency (above 500Hz) signals. These subtle differences are encoded in the phase and amplitude relationships of the left and right channel signals. Stereo sound field adjustment techniques (such as widening and narrowing) essentially perform mathematical transformations and reweighting on these encoded spatial relationships between the left and right channels. Therefore, the processed signals must contain these spatial cues. If full-band signals (including strong, omnidirectional low frequencies) are directly processed, the powerful low-frequency energy will interfere with the processing algorithm, leading to an overall energy imbalance and blurred sound image after processing. Therefore, the mid-to-high frequency components, which are crucial for sound field processing, must be extracted first.

[0100] The high-pass filtering process is the same convolution operation as the low-frequency path, but the objects and purposes it applies to are completely different.

[0101] For convolution operations (temporal domain), please refer to the following formula: ; .

[0102] in, It is of length M (usually 2N) 1) The coefficients of the linear phase FIR high-pass filter, which are also mentioned above. It should be noted that the left and right channels use the exact same filters, which ensures consistent processing.

[0103] In the frequency domain, for components of the input signal with frequencies below the cutoff frequency (e.g., 200Hz), the filter's amplitude response is very small (e.g., ≤-60dB), and these components are greatly suppressed in the output. For components of the input signal with frequencies above the cutoff frequency, the amplitude response is close to 1 (passband ripple ≤0.1dB), and these components pass through with almost no attenuation. Simultaneously, all components passing through the frequency experience the exact same time delay.

[0104] Finally, the output high-frequency left and right channel signals mainly contain all frequency components above the cutoff frequency (e.g., 200Hz). Components below this frequency are deeply attenuated. and These two sequences, while not containing the low-frequency energy of the original signal, completely and without distortion preserve all the spatial information carried by the high-frequency components of the original signal. The clarity, texture, and precise positioning of instruments and vocals in the stereo field are all preserved faithfully.

[0105] For example, consider processing a complex symphony recording containing double bass (low frequencies), a ensemble of violins (mid-to-high frequencies, wide spatial distribution), and a solo flute (mid-to-high frequencies, precise positioning). The input raw audio contains a mixture of all instruments. High-pass filtering is performed using a 200Hz linear-phase high-pass filter (the first high-pass filter). The deep sound of the double bass (primarily below 100Hz) is almost completely filtered out and is inaudible in the output. The overtones and spatial spread of the violins are fully preserved. Due to the linear-phase characteristics, the distribution of the violins on the left, center, and right of the soundstage sounds exactly the same as the original recording, without any blurring or drift. The clear, precisely positioned notes of the solo flute are extracted completely. It can be clearly determined that it is positioned slightly to the right because the precise time / amplitude relationship between the left and right channels is not disrupted. The final result... and It is a pair of "pure" mid-to-high frequency stereo signals that contain all the details, airiness, spatial atmosphere, and precise sound image positioning in the music.

[0106] Step 307: Convert the high-frequency left and right channel signals into signals on the middle two sides to obtain the first signal.

[0107] In this embodiment of the application, the purpose of this step is to refine the high-frequency stereo signal. and Perform a lossless, reversible coordinate transformation to map the sound from the "left-right" channel domain to the "middle-side" signal domain, thereby separating the "common" and "central" components from the "differentiated" and "spatial" components in the sound.

[0108] The following is a brief explanation of why the high-frequency left and right channel signals are converted into signals from the middle two sides.

[0109] Traditional stereo recording or mixing directly corresponds to the voltage driving the left and right speakers. However, the human brain's interpretation of a sound field is not simply a matter of "the left ear hears the left speaker, and the right ear hears the right speaker." The sound field we perceive is the result of the brain's sophisticated processing of the mixed signals received by both ears.

[0110] MS encoding simulates and simplifies this process. The intermediate signal simulates the components of all sounds that arrive at both ears simultaneously, in phase, and with the same amplitude. This typically corresponds to solo vocals, lead instruments, or mono mix components located at the center of the sound field. Physically, it approximates the superposition of sound pressure levels at the listening position after adding the signals from the left and right speakers. The side signals simulate the components of all sounds that have time differences, intensity differences, or out-of-phase. This encodes the width of the sound, spatial atmosphere, environmental reflections, and off-center sound image localization. Physically, it approximates the difference between the left and right speaker signals. Therefore, converting LR signals to MS signals essentially decouples "physical channel" information into "auditory perception" information, making subsequent adjustments to the sound field width (essentially adjusting the ratio of spatial information to center information) extremely intuitive and efficient.

[0111] The principle of the conversion is explained below.

[0112] The conversion from LR to MS is a system of linear equations, performed independently at each discrete time point n.

[0113] ; .

[0114] in, It represents the intermediate signal, which is a mono-compatible, centered component in the sound field. This represents the side signal, indicating the stereo components in the sound field that determine width and direction. The coefficient 1 / 2 is a normalization factor used to ensure that the total power (energy) remains approximately conserved before and after the transformation.

[0115] It should be noted that the above transformation process is linear and reversible. There exists a unique inverse transformation, which can be derived from... and Restore the original and .

[0116] ; .

[0117] This reversibility ensures that the processing can theoretically be lossless (ignoring quantization errors).

[0118] It should be noted that under ideal stereo conditions, and They are irrelevant. This means that central information and spatial information are statistically separate. For example, a voice that is purely centered only has spatial information... There is energy in it, In an ideal scenario, it is zero; while a pair of completely inverted signals (the ultimate sense of width) is... The center is zero, and all the energy is in In the middle. Because the previous high-pass filtering step was linear phase, and The original relative phase relationship between them is perfectly preserved. This addition and subtraction transformation does not destroy this relationship; it simply repackages it. The in-phase components of the left and right signals are preserved. Then their inverse components were extracted.

[0119] It should be noted that the effectiveness of this conversion process depends entirely on the input signal. and The quality of the signal. If the input signal contains severe phase distortion (such as from an IIR filter), the relative phase relationship between the left and right channels is distorted. Therefore, the calculated... The signal will contain a large number of false, non-musical, inverted components. These false... When energy is subsequently amplified (when the sound field is widened), it will lead to a chaotic sound field, positioning drift, and hollowing of the sound quality. A pure, phase-consistent input signal, however, ensures... The extracted information is real and meaningful spatial information.

[0120] Furthermore, the first signal output in this step is not a single signal, but a pair of signals. First signal = { , This signal will undergo subsequent scaling processing to narrow or widen the sound field.

[0121] For example, let's take processing a high-frequency signal from a piece of pop music, including centered vocals, a guitar positioned to the left, keyboards to the right, and a wide-ranging background harmony, as an example. Input It contains strong vocals, strong guitar, weak keyboards, and background harmonies. It includes strong vocals, soft guitar, strong keyboards, background, and right-side harmonics. The vocals (equal on both sides) are fully preserved and have a focused amplitude. The guitar (strong on the left, soft on the right) and the keyboard (soft on the left, strong on the right) each receive a portion of their energy. (Average). The left and right components of the background harmony are in... The values ​​are added together. The vocal parts (left and right sides are the same) are subtracted to zero, completely disappearing from the S-signal. The guitar (left side positive, right side negative difference) produces a positive result. The component encodes "left-leaning" information. The keyboard input (left negative, right positive difference) produces a negative value. The components, encoding "right-leaning" information, were extracted. The left-right differences in the background harmony were... In the first signal, the width sense is encoded. It sounds like a focused, slightly narrow mono mix that includes all vocals, as well as some of the centralized energy from the guitar and keyboard. It sounds like a "ghost" signal with a sense of space extracted, mainly containing leftward information from the guitar, rightward information from the keyboard, and width information from the background, but without any human voices. and They collectively and uniquely encode all the information of the original stereo sound field. By adjusting their proportions later, one can intuitively control whether the sound image is "widened" from the center to the sides or "tightened" from the sides to the center, just like adjusting a knob.

[0122] In summary, converting high-frequency left and right channel signals into signals from the center and two sides is a smart mapping from physical coordinates to perceptual coordinates. Through simple addition and subtraction operations, the sound field information is decoupled from "left / right" to "center / space". This decoupling is the foundation for all subsequent advanced sound field processing.

[0123] Step 308: Scale the first signal according to the sound field adjustment coefficient to obtain the second signal.

[0124] In this embodiment of the application, this step scales the first signal using two independent, typically associated gain coefficients (referred to as the first coefficient and the second coefficient), changing the relative levels of the center component and the side components, thereby achieving an audible expansion, narrowing, or preservation of the sound field.

[0125] In some embodiments, this gain coefficient is the sound field adjustment coefficient. When the sound field adjustment coefficient is greater than 1, it is used to expand the sound field; when the sound field adjustment coefficient is less than 1, it is used to narrow the sound field. Define a main sound field adjustment coefficient. . At this time, the sound field remains unchanged, and both the first and second coefficients are 1. At this time, the sound field expands, the first coefficient is less than 1, and the second coefficient is greater than 1. Optionally, according to the scaling model, the first coefficient = The second coefficient = This model maintains The total power is (approximately) constant to avoid significant changes in overall loudness. The second signal is represented as { According to the linear model, the first coefficient = The second coefficient = . When the sound field narrows, the first coefficient is greater than 1, and the second coefficient is less than 1. The mapping relationship is the opposite of that during expansion, with the first coefficient = The second coefficient = .

[0126] Optionally, scaling the first signal according to the sound field adjustment coefficient to obtain the second signal includes: scaling the intermediate signal component in the first signal according to the first coefficient to obtain the intermediate signal component in the second signal; and scaling the two side signal components in the first signal according to the second coefficient to obtain the two side signal components in the second signal.

[0127] The mathematical expression for scaling is as follows: ; .

[0128] in, Indicates the first coefficient. This indicates the second coefficient.

[0129] The following explains why scaling is necessary.

[0130] Expanding the sound field: increasing This means that the difference signal between the left and right channels is amplified. When inverting back to LR ( L = M + S , R = M - S This will lead to L and R The difference between them increases. Decrease. This means weakening the common signal between the left and right channels. This reduces the "focus" and prominence of the sound image located in the center of the sound field. From an auditory perspective, sound images that were originally closer to the center (such as vocals slightly to the left) will move outwards. Elements that already have width, such as reverberation, ambient sounds, and underlay synthesizers, will become more expansive, creating a grander, more immersive sense of space. Of course, if this is excessively weakened... This can lead to a "hollow" sound in the center of the sound field, making the lead vocals or main instruments sound weaker, more receding, and disconnected from the accompaniment. Therefore, in practice, a lower limit can be set to prevent this. Too small.

[0131] Narrowing the sound field: reducing This means that the difference between the left and right channels is suppressed, making L and R The signals become more similar. Increase. This means that the shared central component is strengthened. From an auditory perspective, the sound images dispersed on both sides will move closer to the center. The overall listening experience becomes more compact and impactful, similar to the process of "mono-ifying" stereo, but usually retaining a little more width than pure mono. Due to The signal is reduced, and the final signal is closer to that of a mono-compatible signal. This can improve the clarity and loudness of sound on mono playback devices (such as mobile phone speakers and small Bluetooth speakers).

[0132] Step 309: Convert the second signal into left and right channel signals to obtain the third signal.

[0133] In this embodiment of the application, this step uses the precisely scaled "center-side" signal pairs, which represent the desired sound field effect. The second signal, through a lossless, mathematically completely symmetrical inverse transformation relative to the encoding process, is remapped back to the standard "left-right" channel domain, generating high-frequency left and right channel signals that can be directly used for subsequent mixing with low-frequency signals and final output. , (Third signal).

[0134] MS decoding is the inverse operation of encoding. At each discrete time point n, the following linear operation is performed: ; .

[0135] in, and It is the third output signal, namely the high-frequency left and right channel signals after sound field adjustment. This refers to the second input signal, namely the scaled middle and side signals. It should be noted that this assumes the encoding uses M=(L+R) / 2, S=(L... The form is R) / 2.

[0136] For example, continuing to use the high-frequency MS signal of popular music from the previous example: Includes strong vocals and some guitar / keyboard elements. Includes left guitar information, right keyboard information, and background width. When applying sound field extension, use... Using the proportional model as an example, the first coefficient Second coefficient . The volume of the middle voice is reduced by about 3dB, making it sound slightly receding. The left-leaning information of the guitar, the right-leaning information of the keyboard, and the sense of background width are all improved by about 3dB. After the inverse transformation, the guitar sound image is more to the left, the keyboard more to the right, and the background soundstage is significantly wider, expanding the entire musical space from a small room to a large hall. Vocals remain clear, but are no longer as "close," and are more integrated into the expanded space. When applying soundstage narrowing, with... For example, the first coefficient Second coefficient . The volume of the human voice is increased by about 3dB, making it more prominent. Spatial information is attenuated by approximately 3dB. The sound signature after the inverse transformation is as follows: the guitar and keyboard sound images move closer to the center, and the background sounds become narrower and closer. The overall sound becomes more compact and direct, with increased impact, more like listening to a nearby speaker or headphones, making it suitable for obtaining a clearer listening experience in noisy environments.

[0137] Step 310: Mix the third signal with the low-frequency left and right channel signals to obtain a stereo output signal.

[0138] In this embodiment of the application, this step combines the complexly processed high-frequency path (with flexible sound field adjustment capability) with the bypassed low-frequency path (with original impact and energy) into one through a time-domain addition operation, and reconstructs the full-band stereo audio without loss, while solving the core defect of "low-frequency energy imbalance" in traditional sound field processing.

[0139] The following describes the mixing method.

[0140] In mathematics, mixing is simply adding samples one by one: ; .

[0141] in, It is the final stereo output signal. This indicates the third signal, namely the high-frequency left and right channel signals after sound field adjustment. It is the original low-frequency left and right channel signal extracted above, without any processing.

[0142] It should be noted that time delay compensation is required before mixing. Accordingly, the third signal is mixed with the low-frequency left and right channel signals to obtain the stereo output signal, including: processing the low-frequency left and right channel signals based on the gain compensation coefficient to obtain the fourth signal; and mixing the fourth signal with the third signal to obtain the stereo output signal.

[0143] The high-frequency path underwent linear phase high-pass filtering (fixed delay). This includes MS encoding / decoding, scaling, and other processing. All these linear time-invariant processes introduce a constant total group delay, denoted as . Among them, the delay of the linear phase filter This is the main part. The low-frequency path only undergoes minimum-phase low-pass filtering. The group delay of a minimum-phase system is frequency-dependent and usually smaller, and its energy center delay is much smaller than that of a linear-phase filter with equivalent performance. The arrival time of the main energy in the low-frequency part is denoted as... Time difference: That is, the high-frequency signal arrives later than the low-frequency signal. Each sampling point. Without compensation, this will result in severe temporal distortion: for example, the "impact" (low frequency) and "drumhead sound" (high frequency) of the bass drum will be separated in time, making the sound loose, blurry, and lacking impact.

[0144] The compensation method involves inserting a digital delay line (gain compensation factor) into the low-frequency path to delay its signal. The number of sampling points, i.e., the signal used in the actual mixing, is... In this way, the main energies of the two are precisely aligned in time.

[0145] This application provides an audio processing solution that extracts low-frequency signals using an FIR low-pass filter with minimum phase properties. This ensures accurate frequency division while maintaining the transient response of the low-frequency signal, significantly reducing processing delay in the low-frequency path and minimizing signal distortion. By processing the mid-to-high frequency signals after high-pass filtering, while directly mixing the low-frequency signals after minimum phase filtering, the original low-frequency energy distribution and phase characteristics are fully preserved. This solution achieves flexible sound field adjustment while completely resolving the problems of blurred sound image positioning caused by nonlinear phase distortion and auditory imbalance caused by changes in low-frequency energy distribution in traditional solutions, thus achieving high fidelity and high-fidelity stereo sound field processing effects.

[0146] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0147] It should be noted that the order of the method steps provided in the embodiments of this application can be appropriately adjusted, and the steps can also be added or removed as appropriate. Any method variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and therefore will not be elaborated further.

[0148] Figure 4 This is a schematic diagram of the structure of an audio processing device provided according to an embodiment of this application. For example... Figure 4As shown, the device includes: a low-pass filter module 401, a high-pass filter module 402, and a mixing module 403.

[0149] The low-pass filter module 401 is used to filter the left and right channel signals of the original audio through a first low-pass filter to obtain low-frequency left and right channel signals. The first low-pass filter is a finite impulse response low-pass filter with minimum phase property. The high-pass filter module 402 is used to filter the left and right channel signals of the original audio through the first high-pass filter to obtain high-frequency left and right channel signals. The mixing module 403 is used to adjust the sound field of the high-frequency left and right channel signals and mix them with the low-frequency left and right channel signals to obtain a stereo output signal.

[0150] In some embodiments, the apparatus further includes: The filter processing module is used to perform symmetrical processing on the coefficients of the first low-pass filter to obtain the second low-pass filter. The coefficients are used to indicate the impulse response of the first low-pass filter. The second low-pass filter is a finite impulse response low-pass filter with linear phase properties. The second low-pass filter is then subjected to spectral inversion based on the unit impulse sequence to obtain the first high-pass filter. The first high-pass filter is also a finite impulse response high-pass filter with linear phase properties.

[0151] In some embodiments, the filter processing module is used to perform a convolution operation on the coefficients of the first low-pass filter and the time-reversed sequence of the coefficients to obtain a second low-pass filter, wherein the length of the second low-pass filter is 2N-1, and N is the order of the first low-pass filter.

[0152] In some embodiments, the filter processing module is used to obtain the zero-phase frequency response of the first low-pass filter; based on the zero-phase frequency response, a second low-pass filter with linear phase and a length of 2N-1 is generated by frequency sampling or inverse Fourier transform, where N is the order of the first low-pass filter.

[0153] In some embodiments, the apparatus further includes: The filter processing module is used to create a prototype low-pass filter that meets the target frequency response specifications using the frequency sampling method or the window function method. The target frequency response specifications include at least one of the following: cutoff frequency of 200Hz, passband ripple of less than or equal to 0.1dB, and stopband attenuation of greater than or equal to 60dB. The prototype low-pass filter is processed by the Hilbert transform or spectral decomposition method to obtain the first low-pass filter.

[0154] In some embodiments, the mixing module 403 is used to convert the high-frequency left and right channel signals into signals for the middle two sides to obtain a first signal; scale the first signal according to the sound field adjustment coefficient to obtain a second signal; convert the second signal into left and right channel signals to obtain a third signal; and mix the third signal with the low-frequency left and right channel signals to obtain a stereo output signal.

[0155] In some embodiments, the mixing module 403 is configured to scale the intermediate signal component in the first signal according to a first coefficient to obtain the intermediate signal component in the second signal; and to scale the two side signal components in the first signal according to a second coefficient to obtain the two side signal components in the second signal.

[0156] In some embodiments, when the sound field adjustment coefficient is greater than 1, it is used to expand the sound field; when the sound field adjustment coefficient is less than 1, it is used to narrow the sound field.

[0157] In some embodiments, the mixing module 403 is used to process the low-frequency left and right channel signals based on the gain compensation coefficient to obtain a fourth signal; and to mix the fourth signal with the third signal to obtain a stereo output signal.

[0158] This application provides an audio processing device that extracts low-frequency signals using an FIR low-pass filter with minimum phase properties. This ensures accurate frequency division while maintaining the transient response of the low-frequency signal, significantly reducing processing delay in the low-frequency path and minimizing signal distortion. By processing the mid-to-high frequency signals after high-pass filtering, while directly mixing the low-frequency signals after minimum phase filtering, the original low-frequency energy distribution and phase characteristics are fully preserved. This solution achieves flexible sound field adjustment while completely resolving the problems of blurred sound image positioning caused by nonlinear phase distortion and auditory imbalance caused by changes in low-frequency energy distribution in traditional solutions, thus achieving high fidelity and high-fidelity stereo sound field processing effects.

[0159] It should be noted that the audio processing device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device and the audio processing method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0160] Embodiments of this application also provide a computer device, which includes a processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the audio processing methods provided in the above-described method embodiments.

[0161] Figure 5 This is a schematic diagram of the structure of a computer device provided according to an embodiment of this application.

[0162] Computer device 500 includes a central processing unit (CPU) 501, a system memory 504 including random access memory (RAM) 502 and read-only memory (ROM) 503, and a system bus 505 connecting the system memory 504 and the CPU 501. Computer device 500 also includes a basic input / output system (I / O system) 506 that facilitates information transfer between various components within the computer device, and a mass storage device 507 for storing the operating system 513, application programs 514, and other program modules 515.

[0163] The basic input / output system 506 includes a display 508 for displaying information and an input device 509 for user input, such as a mouse or keyboard. Both the display 508 and the input device 509 are connected to the central processing unit 501 via an input / output controller 510 connected to the system bus 505. The basic input / output system 506 may also include the input / output controller 510 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 510 also provides output to a display screen, printer, or other types of output devices.

[0164] Mass storage device 507 is connected to central processing unit 501 via a mass storage controller (not shown) connected to system bus 505. Mass storage device 507 and its associated computer-readable storage media provide non-volatile storage for computer device 500. That is, mass storage device 507 may include computer-readable storage media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.

[0165] Without loss of generality, computer-readable storage media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technique for storing information such as computer-readable storage instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage devices, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 504 and mass storage device 507 described above can be collectively referred to as memory.

[0166] The memory stores one or more programs, which are configured to be executed by one or more central processing units 501. The one or more programs contain instructions for implementing the above method embodiments, and the central processing unit 501 executes the one or more programs to implement the methods provided by the various method embodiments described above.

[0167] According to various embodiments of this application, the computer device 500 can also be connected to a remote computer device on a network, such as the Internet. That is, the computer device 500 can be connected to a network 512 via a network interface unit 511 connected to the system bus 505, or the network interface unit 511 can be used to connect to other types of networks or remote computer device systems (not shown).

[0168] The memory also includes one or more programs stored in the memory, and the one or more programs include steps performed by a computer device in the methods provided in the embodiments of this application.

[0169] This application also provides a computer-readable storage medium storing a computer program that is loaded and executed by a processor to implement the audio processing methods provided in the above-described method embodiments.

[0170] This application also provides a computer program product, which includes a computer program executed by a processor to implement the audio processing methods provided in the above-described method embodiments.

[0171] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0172] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent switching, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, The method includes: The left and right channel signals of the original audio are filtered by a first low-pass filter to obtain low-frequency left and right channel signals. The first low-pass filter is a finite impulse response low-pass filter with minimum phase property. The left and right channel signals of the original audio are filtered by a first high-pass filter to obtain high-frequency left and right channel signals. After the high-frequency left and right channel signals are adjusted for sound field, they are mixed with the low-frequency left and right channel signals to obtain a stereo output signal.

2. The method according to claim 1, characterized in that, The method further includes: The coefficients of the first low-pass filter are symmetrically processed to obtain a second low-pass filter. The coefficients are used to indicate the impulse response of the first low-pass filter. The second low-pass filter is a finite impulse response low-pass filter with linear phase properties. The first high-pass filter is obtained by performing spectral inversion on the second low-pass filter based on the unit pulse sequence. The first high-pass filter is a finite impulse response high-pass filter with linear phase properties.

3. The method according to claim 2, characterized in that, The step of performing symmetrical processing on the coefficients of the first low-pass filter to obtain the second low-pass filter includes: The coefficients of the first low-pass filter are convolved with the time-reversed sequence of the coefficients to obtain the second low-pass filter. The length of the second low-pass filter is 2N-1, where N is the order of the first low-pass filter.

4. The method according to claim 2, characterized in that, The step of performing symmetrical processing on the coefficients of the first low-pass filter to obtain the second low-pass filter includes: Obtain the zero-phase frequency response of the first low-pass filter; Based on the zero-phase frequency response, a second low-pass filter with linear phase and a length of 2N-1 is generated by frequency sampling or inverse Fourier transform, where N is the order of the first low-pass filter.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: A prototype low-pass filter that meets the target frequency response specifications is created using the frequency sampling method or the window function method. The target frequency response specifications include at least one of the following: cutoff frequency of 200Hz, passband ripple of less than or equal to 0.1dB, and stopband attenuation of greater than or equal to 60dB. The prototype low-pass filter is processed using Hilbert transform or spectral decomposition to obtain the first low-pass filter.

6. The method according to any one of claims 1 to 5, characterized in that, The process of adjusting the sound field of the high-frequency left and right channel signals and then mixing them with the low-frequency left and right channel signals to obtain a stereo output signal includes: The high-frequency left and right channel signals are converted into signals on the middle two sides to obtain the first signal; The first signal is scaled according to the sound field adjustment coefficient to obtain the second signal; The second signal is converted into left and right channel signals to obtain the third signal; The third signal is mixed with the low-frequency left and right channel signals to obtain the stereo output signal.

7. The method according to claim 6, characterized in that, The scaling of the first signal according to the sound field adjustment coefficient to obtain the second signal includes: The intermediate signal component in the first signal is scaled according to the first coefficient to obtain the intermediate signal component in the second signal; The two signal components in the first signal are scaled according to the second coefficient to obtain the two signal components in the second signal.

8. The method according to claim 6, characterized in that, When the sound field adjustment coefficient is greater than 1, it is used to expand the sound field; when the sound field adjustment coefficient is less than 1, it is used to narrow the sound field.

9. The method according to claim 6, characterized in that, The step of mixing the third signal with the low-frequency left and right channel signals to obtain the stereo output signal includes: The low-frequency left and right channel signals are processed based on the gain compensation coefficient to obtain the fourth signal; The fourth signal is mixed with the third signal to obtain the stereo output signal.

10. An audio processing apparatus, characterized in that, The device includes: The low-pass filter module is used to filter the left and right channel signals of the original audio through the first low-pass filter to obtain low-frequency left and right channel signals. The first low-pass filter is a finite impulse response low-pass filter with minimum phase property. A high-pass filter module is used to filter the left and right channel signals of the original audio through the first high-pass filter to obtain high-frequency left and right channel signals; The mixing module is used to adjust the sound field of the high-frequency left and right channel signals and mix them with the low-frequency left and right channel signals to obtain a stereo output signal.

11. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the audio processing method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The readable storage medium stores a computer program, which is loaded and executed by a processor to implement the audio processing method as described in any one of claims 1 to 9.

13. A computer program product, characterized in that, The computer program product includes a computer program executed by a processor to implement the audio processing method as described in any one of claims 1 to 9.