Stereo processing method and device

By performing dry and wet sound separation and filtering on the stereo audio signal, the limitations of existing stereo processing technologies are solved, achieving efficient sound field zoning and isolation while preserving stereo effects, thus improving the user experience.

CN121815184APending Publication Date: 2026-04-07IFLYTEK (SUZHOU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing sound field zoning control technology has limitations when processing stereo audio signals, making it difficult to effectively achieve independent and non-interfering audio zoning control.

Method used

The stereo audio signal is decomposed into dry sound, left channel wet sound, and right channel wet sound by dry and wet sound separation processing. Three control filters are used to filter these signals respectively, and the filter coefficients are calculated by optimization algorithm to generate multi-channel output signal, so as to realize accurate reconstruction of sound field and suppression of sound pressure in dark area.

Benefits of technology

Without relying on special hardware, it achieves a high level of sound zone isolation, maintains the sound field width and spatial positioning of stereo audio, and enhances the user's listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815184A_ABST
    Figure CN121815184A_ABST
Patent Text Reader

Abstract

The invention provides a stereophonic sound processing method and device, and relates to the technical field of sound processing, and the method comprises the steps: carrying out the dry and wet sound separation processing of a stereophonic sound audio signal comprising a left sound channel signal and a right sound channel signal, and obtaining a dry sound, a left channel wet sound, and a right channel wet sound; wherein the left channel wet sound is a signal component irrelevant to a right channel signal in the left channel signal, the right channel wet sound is a signal component irrelevant to the left channel signal in the right channel signal, and the dry sound is a signal component with the same amplitude and phase in the left channel signal and the right channel signal; performing filtering processing on the dry sound, the left channel wet sound and the right channel wet sound by using a preset first control filter, a preset second control filter and a preset third control filter respectively to obtain three paths of corresponding filtering signals; and superposing the three paths of filtering signals to generate a multi-channel output signal for driving the loudspeaker array.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound processing technology, and in particular to a stereo processing method and apparatus. Background Technology

[0002] With the development of smart cockpit technology, zoned control of the in-vehicle sound field, providing independent and non-interfering audio content for drivers and passengers in different positions, such as drivers listening to navigation and passengers listening to music, has become a key requirement for improving user experience.

[0003] Existing sound field zoning control technologies are mainly designed for single sound sources, employing traditional algorithms such as acoustic contrast control and sound pressure matching to achieve zoning isolation by maximizing the sound energy ratio between bright and dark areas or minimizing the reconstruction error in the bright area. However, these technologies have significant limitations when processing stereo audio signals.

[0004] Therefore, how to perform stereo processing more effectively has become an urgent problem to be solved in the industry. Summary of the Invention

[0005] This invention provides a stereo processing method and apparatus to solve the problem of how to perform stereo processing more effectively in the prior art.

[0006] This invention provides a stereo processing method, comprising: The stereo audio signal, including the left channel signal and the right channel signal, is subjected to dry and wet sound separation processing to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The dry sound, left channel wet sound, and right channel wet sound are filtered by a preset first control filter, a second control filter, and a third control filter, respectively, to obtain the corresponding three-channel filtered signals. The three filtered signals are superimposed to generate a multi-channel output signal for driving the speaker array.

[0007] According to a stereo processing method provided by the present invention, the method for determining the first control filter, the second control filter, and the third control filter includes: Obtain the bright area transfer function matrix from each loudspeaker in the loudspeaker array to the target bright area, and the dark area transfer function matrix to the target dark area; Different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, respectively; Based on the bright area transfer function matrix, the dark area transfer function matrix, and their respective target sound field distributions, the coefficients of each control filter are calculated using an optimization algorithm.

[0008] According to a stereo processing method provided by the present invention, different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, including: The target sound field of the first control filter is the sound field generated in the target bright area when the left and right channel speakers in a standard stereo configuration simultaneously play the same signal. The target sound field of the second control filter is the sound field generated in the target bright area when only the left channel speaker is working in a standard stereo configuration; The target sound field of the third control filter is the sound field generated in the target bright area when only the right channel speaker is working in a standard stereo configuration.

[0009] According to a stereo processing method provided by the present invention, the step of calculating the coefficients of each control filter through an optimization algorithm includes: Construct an objective function, which includes a bright area sound field reconstruction error term and a dark area sound pressure energy suppression term; By minimizing the objective function, the filter coefficients that maximize the ratio of sound pressure energy in the bright area to sound pressure energy in the dark area are obtained, and the calculated filter coefficients are stored as the coefficients of each control filter.

[0010] According to a stereo processing method provided by the present invention, the step of obtaining the bright area transfer function matrix from each speaker in the speaker array to the target bright area and the dark area transfer function matrix to the target dark area includes: Multiple discrete target control points are set within the target's light area and target's dark area, respectively; The bright zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all bright zone target control points; The dark zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all dark zone target control points.

[0011] According to a stereo processing method provided by the present invention, the step of performing dry and wet sound separation processing on a stereo audio signal including a left channel signal and a right channel signal includes: Calculate the complex correlation coefficient between the left channel signal and the right channel signal; Based on the complex correlation coefficient, the signal components with the same amplitude and phase in the left channel signal and the right channel signal are extracted as the dry sound; The left channel wet signal is obtained by subtracting the component corresponding to the left channel from the dry signal. The right channel wet signal is obtained by subtracting the component corresponding to the right channel from the dry signal.

[0012] According to a stereo processing method provided by the present invention, the dry sound, the left channel wet sound, and the right channel wet sound satisfy the orthogonality condition, the cross-correlation coefficient between the left channel wet sound and the dry sound is zero; the cross-correlation coefficient between the right channel wet sound and the dry sound is zero; and the cross-correlation coefficient between the left channel wet sound and the right channel wet sound is zero.

[0013] The present invention also provides a stereo processing device, comprising: The separation module is used to perform dry and wet sound separation processing on the stereo audio signal including the left channel signal and the right channel signal to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The filtering module is used to filter the dry sound, the left channel wet sound and the right channel wet sound respectively using a preset first control filter, a second control filter and a third control filter to obtain the corresponding three filtered signals; The output module is used to superimpose the three filtered signals to generate a multi-channel output signal for driving the speaker array.

[0014] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Obtain the bright area transfer function matrix from each loudspeaker in the loudspeaker array to the target bright area, and the dark area transfer function matrix to the target dark area; Different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, respectively; Based on the bright area transfer function matrix, the dark area transfer function matrix, and their respective target sound field distributions, the coefficients of each control filter are calculated using an optimization algorithm.

[0015] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: The target sound field of the first control filter is the sound field generated in the target bright area when the left and right channel speakers in a standard stereo configuration simultaneously play the same signal. The target sound field of the second control filter is the sound field generated in the target bright area when only the left channel speaker is working in a standard stereo configuration; The target sound field of the third control filter is the sound field generated in the target bright area when only the right channel speaker is working in a standard stereo configuration.

[0016] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Construct an objective function, which includes a bright area sound field reconstruction error term and a dark area sound pressure energy suppression term; By minimizing the objective function, the filter coefficients that maximize the ratio of sound pressure energy in the bright area to sound pressure energy in the dark area are obtained, and the calculated filter coefficients are stored as the coefficients of each control filter.

[0017] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Multiple discrete target control points are set within the target's light area and target's dark area, respectively; The bright zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all bright zone target control points; The dark zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all dark zone target control points.

[0018] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Calculate the complex correlation coefficient between the left channel signal and the right channel signal; Based on the complex correlation coefficient, the signal components with the same amplitude and phase in the left channel signal and the right channel signal are extracted as the dry sound; The left channel wet signal is obtained by subtracting the component corresponding to the left channel from the dry signal. The right channel wet signal is obtained by subtracting the component corresponding to the right channel from the dry signal.

[0019] According to a stereo processing apparatus provided by the present invention, the dry sound, the left channel wet sound, and the right channel wet sound satisfy an orthogonal condition, wherein the cross-correlation coefficient between the left channel wet sound and the dry sound is zero; the cross-correlation coefficient between the right channel wet sound and the dry sound is zero; and the cross-correlation coefficient between the left channel wet sound and the right channel wet sound is zero.

[0020] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the stereo processing method described above.

[0021] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the stereo processing method as described above.

[0022] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the stereo processing method as described above.

[0023] The stereo processing method and apparatus provided by this invention innovatively separates the wet and dry sound of stereo audio signals, decomposing the signal into wet sound components for the left and right channels and a common dry sound component. These components are then differentiated and superimposed using three independent control filters. By preserving and accurately reconstructing the wet sound signal that constitutes the sense of space, while effectively controlling overall sound energy leakage, a high level of sound zone isolation is achieved using only ordinary speaker arrays without relying on special hardware. This ensures effective zone isolation while maintaining the original stereo audio's sound field width and spatial positioning, significantly enhancing the user's listening experience. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating the stereo processing method provided by the present invention; Figure 2 The audio source input / output flowchart provided by this invention; Figure 3 A schematic diagram of the stereo processing device provided by the present invention; Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] Figure 1 This is a flowchart illustrating the stereo processing method provided by the present invention, as shown below. Figure 1 As shown, the method includes the following: Step 110: Perform dry and wet sound separation processing on the stereo audio signal including the left channel signal and the right channel signal to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; In this application, the stereo audio signal includes at least a left channel signal and a right channel signal. These two signals contain both common components and their own independent components. In the vehicle audio system, the stereo signal creates a sense of spatial positioning through the amplitude difference, phase difference and time difference between the left and right channels, bringing users a more immersive listening experience.

[0028] Dry and wet sound separation processing can decompose the original signal into different components. In this embodiment, it can separate different components related to spatial perception in the stereo audio signal for subsequent independent control.

[0029] Specifically, after dry and wet sound separation processing, the original stereo audio signal is decomposed into three signals: Dry sound is a signal component in both the left and right channel signals that has the same amplitude and phase. In auditory perception, dry sound typically corresponds to the central sound image in a stereo sound field, such as a solo vocal performance. It is a highly correlated portion between the two channels.

[0030] Left channel wet sound refers to the signal components in the left channel signal that are unrelated to the right channel signal. This means that after removing the dry sound components shared with the right channel signal from the original left channel signal, what remains is the portion unique to the left channel. In auditory perception, it primarily contributes to the sense of location and space on the left side of the sound field.

[0031] Right channel wet sound refers to the signal components in the right channel signal that are unrelated to the left channel signal. Similarly, this refers to the portion unique to the right channel remaining after removing the dry sound components shared with the right channel signal from the original right channel signal. In auditory perception, it mainly contributes to the sense of direction and space on the right side of the sound field.

[0032] Through dry and wet sound separation processing, the original stereo signal is resolved into dry sound representing the central sound image and wet sound representing the left and right channels representing the sound images on both sides, laying the foundation for subsequent refined sound field zoning control for different sound image components.

[0033] Step 120: The dry sound, left channel wet sound and right channel wet sound are filtered by the preset first control filter, second control filter and third control filter respectively to obtain three filtered signals corresponding to the dry sound, left channel wet sound and right channel wet sound. In this embodiment, three preset control filters are used: a first control filter, a second control filter, and a third control filter. Each control filter processes one separated signal; the first control filter processes the dry signal, the second control filter processes the wet signal of the left channel, and the third control filter processes the wet signal of the right channel.

[0034] During the filtering process, the three single-channel signals—dry sound, left channel wet sound, and right channel wet sound—are used as inputs to the corresponding control filters. After filtering, each input will generate a set of multi-channel output signals.

[0035] For example, if the loudspeaker array contains N loudspeakers, the second control filter receives a single-channel left-channel wet acoustic signal input and outputs N channels of filtered signals, each channel driving one loudspeaker. The first and third control filters perform similar processing. This results in a three-channel multi-filtered signal.

[0036] Step 130: The three filtered signals are superimposed to generate a multi-channel output signal for driving the speaker array.

[0037] In this application, the three filtered signals are superimposed. For example, for the i-th speaker in the speaker array, its final drive signal is the sum of the signals from the i-th output of the first control filter, the i-th output of the second control filter, and the i-th output of the third control filter.

[0038] The number of channels in the multi-channel output signal generated after superposition is the same as the number of speakers in the speaker array. The superimposed signal is then sent to the power amplifier and finally drives each speaker in the speaker array to produce sound.

[0039] In this application, an innovative wet and dry sound separation process is applied to the stereo audio signal, decomposing the signal into wet sound components for the left and right channels and a common dry sound component. Three independent control filters are then used to differentiate and superimpose these components. By preserving and accurately reconstructing the wet sound signal that constitutes the sense of space, while effectively controlling overall sound energy leakage, a high level of sound zone isolation is achieved using only a standard speaker array without relying on special hardware. This ensures effective zone isolation while maintaining the original stereo audio's sound field width and spatial positioning, significantly enhancing the user's listening experience.

[0040] Optionally, the method for determining the first control filter, the second control filter, and the third control filter includes: Obtain the bright area transfer function matrix from each loudspeaker in the loudspeaker array to the target bright area, and the dark area transfer function matrix to the target dark area; Different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, respectively; Based on the bright area transfer function matrix, the dark area transfer function matrix, and their respective target sound field distributions, the coefficients of each control filter are calculated using an optimization algorithm.

[0041] In this application, a speaker array refers to a collection of multiple speaker units arranged inside a vehicle, which may include front door speakers, rear door speakers, center console speakers, headrest speakers, etc. Each speaker in the speaker array can be controlled independently, and the spatial distribution control of the sound field is achieved through collaborative work.

[0042] The target clear area refers to the area where clear audio is expected to be produced, such as the driver's seat area or the rear passenger area.

[0043] A target dark zone refers to the area where audio suppression is desired, i.e., the area where you don't want to hear any sound or where you want a low sound pressure level. In automotive applications, when the driver's area is designated as the bright zone, the rear passenger area can be set as the dark zone, and vice versa.

[0044] The transfer function matrix describes the acoustic transfer characteristics of each loudspeaker in a loudspeaker array to a specific location in space.

[0045] The bright area transfer function matrix contains the frequency response functions of all loudspeakers to each measurement point in the target bright area, and each element of the matrix represents the complex transfer function of a specific loudspeaker to a specific measurement point.

[0046] The dark zone transfer function matrix contains the frequency response functions of all loudspeakers to each measurement point within the target dark zone.

[0047] The transfer function can be obtained through either actual measurement or acoustic simulation. In actual measurement, a measuring microphone is placed within the target area, and test signals are played sequentially by driving each speaker. The response signals received by the microphones are recorded, and the transfer function is calculated using a system identification method.

[0048] Acoustic simulation, on the other hand, involves establishing an in-vehicle acoustic model and calculating the transfer function using the finite element method or the boundary element method.

[0049] In this application, setting different target sound field distributions for the three control filters is the key to achieving stereo zoning.

[0050] The target sound field distribution defines the desired sound pressure distribution pattern within the target bright area. By setting the target sound field corresponding to each component of stereo for different control filters, stereo effects can be reconstructed in the bright area.

[0051] Based on the obtained transfer function matrix and the set target sound field distribution, the coefficients of each control filter are calculated using an optimization algorithm. The goal of the optimization process is to find a set of filter coefficients that can accurately reconstruct the target sound field in the target bright area, while suppressing sound pressure energy to the greatest extent in the target dark area.

[0052] This embodiment, through systematic transfer function measurement and optimization algorithm design, can obtain the optimal control filter suitable for specific vehicle models and cabin layouts, thereby achieving high-performance stereo processing.

[0053] Optionally, different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, including: The target sound field of the first control filter is the sound field generated in the target bright area when the left and right channel speakers in a standard stereo configuration simultaneously play the same signal. The target sound field of the second control filter is the sound field generated in the target bright area when only the left channel speaker is working in a standard stereo configuration; The target sound field of the third control filter is the sound field generated in the target bright area when only the right channel speaker is working in a standard stereo configuration.

[0054] In this application, the standard stereo configuration refers to a traditional two-channel audio playback configuration, which typically includes two main speakers, one on the left front and one on the right front.

[0055] In a standard stereo configuration, the left channel signal is played only through the left speaker, and the right channel signal is played only through the right speaker. The stereo sound field is generated by the coordinated work of the two speakers.

[0056] The input to the second control filter is the left channel wet sound, which corresponds to the component in stereo that belongs only to the left channel.

[0057] Therefore, the target sound field of the second control filter is set to the sound field generated in the target bright area when only the left channel speaker is working in a standard stereo configuration. This sound field distribution reflects the spatial characteristics of the sound source being to the left, and can reconstruct the spatial positioning information of the left channel in the target bright area.

[0058] The first control filter receives dry sound as its input, corresponding to the common component of the left and right channels. The target sound field of the first control filter is set to the sound field generated in the target bright area when the left and right channel speakers in a standard stereo configuration simultaneously play the same signal.

[0059] This sound field distribution exhibits central positioning characteristics, with the sound image located in the center of the left and right speakers, making it suitable for reproducing audio content that requires central positioning, such as vocals and solo instruments.

[0060] The input to the third control filter is the right channel wet sound, which corresponds to the component in stereo that belongs only to the right channel.

[0061] The target sound field of the third control filter is set to the sound field generated in the target bright area when only the right channel speaker is working in a standard stereo configuration. This sound field distribution reflects the spatial characteristics of the sound source being to the right, and can reconstruct the spatial positioning information of the right channel in the target bright area.

[0062] By setting this target sound field distribution, the three control filters work together to accurately reconstruct the complete stereo sound field in the target bright area, including all spatial information of left-side positioning, center positioning, and right-side positioning, while achieving effective sound pressure suppression in the target dark area.

[0063] The target sound field setting method in this embodiment fully considers the spatial characteristics of stereo sound. Through separate processing and independent control, it not only ensures the integrity of the stereo effect but also achieves efficient sound field zoning and isolation.

[0064] Optionally, the calculation of the coefficients of each control filter using an optimization algorithm includes: Construct an objective function, which includes a bright area sound field reconstruction error term and a dark area sound pressure energy suppression term; By minimizing the objective function, the filter coefficients that maximize the ratio of sound pressure energy in the bright area to sound pressure energy in the dark area are obtained, and the calculated filter coefficients are stored as the coefficients of each control filter.

[0065] In this application, the objective function is a mathematical expression that needs to be minimized or maximized in the optimization problem, used to quantify the quality of the control effect. In sound field zoning control, the objective function typically includes two main components: a bright zone sound field reconstruction error term and a dark zone sound pressure energy suppression term.

[0066] The calculation of the control filter parameters is implemented using the VAST framework in the ACC-PM algorithm. The general objective function of ACC-PM is: ; in, For the bright area transfer function, For dark area transfer function, For the desired control filter, For the target sound field, and and These are the weighting parameters for the reconstruction error in the bright area and the energy in the dark area, respectively.

[0067] Since the input stereo is decomposed into three parts, the corresponding three target sound fields also need to be changed when solving these three control filters. The target sound field of the first control filter corresponds to the wet sound of the left channel input. Therefore, the target sound field is the sum of the transmission paths of all speakers that emit left channel sound sources in the standard stereo mode. Similarly, the target sound field of the third control filter corresponds to the wet sound of the right channel as input. Therefore, the target sound field is the sum of the transmission paths of all loudspeakers emitting right channel sound sources in standard stereo mode. Furthermore, the corresponding input of the first control filter is the dry sound shared by the left and right channels, so the target sound field is the sum of the transmission paths of all speakers in the standard stereo mode.

[0068] In this application, by minimizing the objective function, the filter coefficients that maximize the ratio of sound pressure energy in the bright area to that in the dark area can be obtained. This ratio, also known as acoustic contrast, is an important indicator for evaluating the performance of sound field zoning. The optimization solution can be obtained using numerical optimization methods such as gradient descent, Newton's method, or interior-point method.

[0069] For frequency domain processing, the optimization problem needs to be solved independently at each frequency point to obtain the frequency-dependent filter coefficients. The calculated filter coefficients are then transformed to the time domain using an inverse fast Fourier transform to obtain the time-domain finite impulse response filter coefficients. These coefficients are stored as preset control filters for use in real-time processing.

[0070] This embodiment, by constructing a reasonable objective function and optimizing its solution, can obtain a control filter that achieves the best balance between sound field reconstruction in the bright area and sound pressure suppression in the dark area, thus realizing high-performance sound field zoning control.

[0071] Optionally, obtaining the bright area transfer function matrix from each speaker in the speaker array to the target bright area and the dark area transfer function matrix to the target dark area includes: Multiple discrete target control points are set within the target's light area and target's dark area, respectively; The bright zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all bright zone target control points; The dark zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all dark zone target control points.

[0072] In this application, the target control point is a discrete spatial location selected within the target area to characterize the sound field properties. In practical applications, since it is impossible to completely sample a continuous space, it is necessary to select a finite number of discrete points to approximate the sound field distribution of the entire area.

[0073] When setting multiple discrete target control points within a target bright area, the distribution of the control points should take into account the actual position of the human ear and listening habits. For example, in the driver's seat area, multiple control points can be set near the average head position, including the left ear, right ear, and center of the head. The number of control points needs to be balanced between computational complexity and sound field control accuracy; typically, 3-9 control points are set for each seat area.

[0074] When setting multiple discrete target control points within the target dark area, they should cover the main area where sound suppression is required. The control points can be evenly distributed or non-uniformly distributed depending on actual needs. The number of control points in the dark area is usually equal to or slightly more than in the bright area to ensure effective sound pressure suppression throughout the entire dark area.

[0075] The acoustic transfer function describes the acoustic transmission characteristics from the loudspeaker to the control point, including amplitude and phase response information. Each transfer function is a complex function of frequency, reflecting the attenuation and phase shift characteristics of different frequency components during transmission.

[0076] The transfer function matrix for the bright area is constructed as follows: the rows of the matrix correspond to each target control point in the bright area, and the columns correspond to each loudspeaker in the loudspeaker array.

[0077] The matrix element Hij represents the transfer function from the j-th loudspeaker to the i-th bright-zone control point. If there are M bright-zone control points and N loudspeakers, then the dimension of the bright-zone transfer function matrix is ​​M×N.

[0078] The dark zone transfer function matrix is ​​constructed similarly: the rows of the matrix correspond to the target control points in the dark zone, and the columns correspond to the speakers in the speaker array. If there are P dark zone control points and N speakers, then the dimension of the dark zone transfer function matrix is ​​P×N.

[0079] The transfer function can be measured in a real vehicle environment or an acoustic laboratory. The effects of environmental noise, temperature, humidity, and other factors must be considered during measurement. Using multiple measurements and averaging the results can improve measurement accuracy.

[0080] This embodiment provides a reliable acoustic model foundation for subsequent filter optimization by reasonably setting the target control points and accurately measuring the transfer function.

[0081] Optionally, the step of performing dry and wet sound separation processing on the stereo audio signal including the left channel signal and the right channel signal includes: Calculate the complex correlation coefficient between the left channel signal and the right channel signal; Based on the complex correlation coefficient, the signal components with the same amplitude and phase in the left channel signal and the right channel signal are extracted as the dry sound; The left channel wet signal is obtained by subtracting the component corresponding to the left channel from the dry signal. The right channel wet signal is obtained by subtracting the component corresponding to the right channel from the dry signal.

[0082] In this application, the time-domain stereo signal is first converted to the time-frequency domain using a short-time Fourier transform.

[0083] Time-frequency domain representation can simultaneously reflect the time variation and frequency components of a signal, making it suitable for processing non-stationary signals such as music.

[0084] During the transformation, a window function is used to divide the signal into frames. Each frame is typically 20 to 50 milliseconds long, and there is some overlap between adjacent frames.

[0085] For each time-frequency point, the complex correlation coefficient between the left and right channels needs to be calculated. The complex correlation coefficient considers not only the amplitude correlation but also the phase correlation of the signals. During calculation, the signals within a certain time range need to be statistically averaged to obtain a stable correlation coefficient estimate. The magnitude of the correlation coefficient represents the degree of correlation, and the phase represents the phase relationship between the two signals.

[0086] Dry sound extraction is based on correlation analysis. Highly correlated time-frequency points are considered to primarily contain dry sound components. The amplitude of the dry sound is determined by the geometric mean of the left and right channels, and the phase is the average phase of the two channels. This method preserves the natural characteristics of the original signal's central localization component.

[0087] Wet sound extraction is achieved by removing dry sound components. The dry sound component is subtracted from the left channel signal to obtain the left channel wet sound; the dry sound component is subtracted from the right channel signal to obtain the right channel wet sound. In this application, the similarity between channels is quantified by calculating the complex correlation coefficient in the time-frequency domain, and signal decomposition is performed accordingly. This allows for the accurate extraction of the central common component and the independent components on both sides of the stereo signal. This separation provides a high-quality input signal for subsequent differentiated filtering control and is a prerequisite for the overall technical solution to achieve good results.

[0088] Optionally, the dry sound, the left channel wet sound, and the right channel wet sound satisfy the orthogonality condition, and the cross-correlation coefficient between the left channel wet sound and the dry sound is zero; the cross-correlation coefficient between the right channel wet sound and the dry sound is zero; and the cross-correlation coefficient between the left channel wet sound and the right channel wet sound is zero.

[0089] In this application, orthogonality refers to the absence of a linear correlation between signals. In the technical solution of this application, the three signals after dry and wet sound separation—dry sound, left channel wet sound, and right channel wet sound—satisfy the condition of mutual orthogonality. This orthogonality is an inevitable result of the separation algorithm design.

[0090] The orthogonality of the left channel wet and dry audio stems from their definitions. Left channel wet audio is the remaining component in the left channel after removing the portion shared with the right channel, while dry audio is precisely this shared portion. Therefore, left channel wet audio contains no dry audio components, and their cross-correlation is zero. This is analogous to decomposing a vector into two orthogonal components with no overlap.

[0091] The orthogonality principle of the right channel wet sound and dry sound is the same. The right channel wet sound is a unique component of the right channel and has no overlap with the dry sound, which represents the common component. This orthogonality ensures that processing the right channel wet sound will not affect the center localization image.

[0092] The orthogonality between the wet sound from the left and right channels reflects a fundamental characteristic of stereo recording. In professional stereo recordings, the unique information from the left (such as instruments on the left) and the unique information from the right (such as instruments on the right) are typically independent sound sources, naturally without any correlation between them. The separation algorithm preserves this independence.

[0093] In this application, by ensuring the orthogonality of the separated signals, a theoretical guarantee is provided for achieving high-quality stereo processing, effectively avoiding mutual interference and information redundancy in the signal processing process.

[0094] In an alternative embodiment, the advantages of separating dry and wet audio from stereo sources and calculating the advantages of three control filters for different audio sources will be described in detail.

[0095] The most significant characteristic of stereo is that the left and right channels are correlated but not identical. Most current sound field partitioning algorithms are based on the premise that all speakers use the same audio source. Therefore, when a stereo audio source is input, a common solution is to independently solve for the control filters of the left and right channels. and At this point, under stereo input, the formula for calculating the actual isolation of a single frequency point in the sound field partition is as follows: ; Among them, the transfer function matrix of the bright area Transfer function matrix of dark area Let M be an M*N matrix, where M is the number of microphones in the corresponding area and N is the total number of speakers in the vehicle. and They are respectively and The autocorrelation matrix, i.e. and , and The left and right channel audio sources are represented by an N*N diagonal matrix, with all diagonal elements being identical. and Let N*1 be the control filter vector. The formula can be derived and simplified to obtain: ; Where A represents the isolation of the left channel and B represents the isolation of the right channel. The left and right channel input complex correlation coefficients represent the isolation between two VASTs under stereo input conditions, assuming the left and right channel control filters are known. Depending on the input stereo source, the final isolation is determined by the left and right channel input amplitude ratio and the left and right channel input complex correlation coefficients. Re ( ) takes the real part of the complex number, which represents the mutual interference / coupling term between the left and right channel signals.

[0096] This leads to an uncontrollable overall isolation when playing different stereo songs in real-world scenarios, as the phase relationship between the left and right channels changes in real time. In the worst case, the isolation may deteriorate to a level lower than that of the left and right channels. Therefore, after separating the dry and wet audio in stereo, there is no correlation between the dry and wet audio, nor between the wet audio in the left and right channels. Thus, the influence of the third term in the formula can be approximated as 0, and the final actual isolation must be between the isolation levels of the three VASTs.

[0097] In this application, by separating the dry and wet sound sources of the stereo sound source, the relevant and unrelated parts of the left and right channels are separated and filtered by control filters, which effectively reduces the influence of the correlation of the stereo sound source on the final actual isolation.

[0098] Figure 2 The audio source input / output flowchart provided by this invention is as follows: Figure 2 As shown, the processing flow of this embodiment includes three main stages: dry and wet sound separation stage, control filtering processing stage, and signal superposition output stage.

[0099] First, the stereo audio input signal enters the dry / wet sound separation module. The stereo audio signal contains two signals, the left channel and the right channel, carrying spatial positioning information of the audio. The dry / wet sound separation module analyzes and processes the input stereo signal, decomposing it into three mutually orthogonal signal components: dry sound, left channel wet sound, and right channel wet sound.

[0100] The left channel wet sound contains unique information from the left channel, primarily corresponding to sound source components located on the left side of the sound field. The right channel wet sound contains unique information from the right channel, primarily corresponding to sound source components located on the right side of the sound field. The dry sound contains shared information from both channels, corresponding to sound source components located in the center of the sound field, such as lead vocals or solo instruments.

[0101] The three separated signals are processed by their respective control filters. Control filter 1 processes the wet sound from the left channel; its filter coefficients are pre-optimized to reconstruct the left sound field distribution in the target bright area while suppressing sound pressure in the target dark area. Control filter 2 processes the dry sound signal; its filter coefficients produce a centrally located sound field effect in the target bright area. Control filter 3 processes the wet sound from the right channel, reconstructing the right sound field distribution in the target bright area.

[0102] The three control filters operate in parallel, each independently adjusting the frequency response and spatial distribution of the input signal. Each control filter outputs a multi-channel signal, with the number of channels corresponding to the number of speakers in the speaker array. For example, in a vehicle system equipped with eight speakers, each control filter outputs eight signals.

[0103] Finally, the output signals of the three control filters are superimposed on their corresponding channels to generate the final speaker drive signal. The superposition process is linear, meaning that the three signals corresponding to the same speaker are directly added together. Since the three separated signals satisfy the orthogonality condition, there is no mutual interference after superposition, enabling accurate reconstruction of the complete stereo sound field in the target bright area.

[0104] In practical applications, the entire processing can be completed in real time within a digital signal processor. The input digital audio signal is buffered and processed frame by frame. Each frame of signal undergoes dry and wet sound separation, filtering, and superposition before being output to drive the speaker array to produce the desired sound field distribution.

[0105] The processing method in this embodiment decomposes the stereo signal into orthogonal components and controls them independently, preserving the stereo spatial characteristics of the original audio while achieving efficient sound field zoning and isolation. In in-vehicle scenarios, this can provide drivers and passengers with a personalized audio experience, enhancing driving comfort and safety.

[0106] The stereo processing apparatus provided by the present invention will be described below. The stereo processing apparatus described below and the stereo processing method described above can be referred to in correspondence.

[0107] Figure 3 This is a schematic diagram of the stereo processing device provided by the present invention, as shown below. Figure 3 As shown, it includes: The separation module 310 is used to perform dry and wet sound separation processing on the stereo audio signal including the left channel signal and the right channel signal to obtain dry sound, left channel wet sound and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The filtering module 320 is used to filter the dry sound, the left channel wet sound and the right channel wet sound respectively using a preset first control filter, a second control filter and a third control filter to obtain the corresponding three-channel filtered signals; The output module 330 is used to superimpose the three filtered signals to generate a multi-channel output signal for driving the speaker array.

[0108] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Obtain the bright area transfer function matrix from each loudspeaker in the loudspeaker array to the target bright area, and the dark area transfer function matrix to the target dark area; Different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, respectively; Based on the bright area transfer function matrix, the dark area transfer function matrix, and their respective target sound field distributions, the coefficients of each control filter are calculated using an optimization algorithm.

[0109] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: The target sound field of the first control filter is the sound field generated in the target bright area when the left and right channel speakers in a standard stereo configuration simultaneously play the same signal. The target sound field of the second control filter is the sound field generated in the target bright area when only the left channel speaker is working in a standard stereo configuration; The target sound field of the third control filter is the sound field generated in the target bright area when only the right channel speaker is working in a standard stereo configuration.

[0110] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Construct an objective function, which includes a bright area sound field reconstruction error term and a dark area sound pressure energy suppression term; By minimizing the objective function, the filter coefficients that maximize the ratio of sound pressure energy in the bright area to sound pressure energy in the dark area are obtained, and the calculated filter coefficients are stored as the coefficients of each control filter.

[0111] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Multiple discrete target control points are set within the target's light area and target's dark area, respectively; The bright zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all bright zone target control points; The dark zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all dark zone target control points.

[0112] According to a stereo processing apparatus provided by the present invention, the apparatus is further used for: Calculate the complex correlation coefficient between the left channel signal and the right channel signal; Based on the complex correlation coefficient, the signal components with the same amplitude and phase in the left channel signal and the right channel signal are extracted as the dry sound; The left channel wet signal is obtained by subtracting the component corresponding to the left channel from the dry signal. The right channel wet signal is obtained by subtracting the component corresponding to the right channel from the dry signal.

[0113] According to a stereo processing apparatus provided by the present invention, the dry sound, the left channel wet sound, and the right channel wet sound satisfy an orthogonal condition, wherein the cross-correlation coefficient between the left channel wet sound and the dry sound is zero; the cross-correlation coefficient between the right channel wet sound and the dry sound is zero; and the cross-correlation coefficient between the left channel wet sound and the right channel wet sound is zero.

[0114] In this application, an innovative wet and dry sound separation process is applied to the stereo audio signal, decomposing the signal into wet sound components for the left and right channels and a common dry sound component. Three independent control filters are then used to differentiate and superimpose these components. By preserving and accurately reconstructing the wet sound signal that constitutes the sense of space, while effectively controlling overall sound energy leakage, a high level of sound zone isolation is achieved using only a standard speaker array without relying on special hardware. This ensures effective zone isolation while maintaining the original stereo audio's sound field width and spatial positioning, significantly enhancing the user's listening experience.

[0115] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a stereo processing method, which includes: The stereo audio signal, including the left channel signal and the right channel signal, is subjected to dry and wet sound separation processing to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The dry sound, left channel wet sound, and right channel wet sound are filtered by a preset first control filter, a second control filter, and a third control filter, respectively, to obtain the corresponding three-channel filtered signals. The three filtered signals are superimposed to generate a multi-channel output signal for driving the speaker array.

[0116] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0117] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to execute the stereo processing method provided by the above methods, the method comprising: including: The stereo audio signal, including the left channel signal and the right channel signal, is subjected to dry and wet sound separation processing to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The dry sound, left channel wet sound, and right channel wet sound are filtered by a preset first control filter, a second control filter, and a third control filter, respectively, to obtain the corresponding three-channel filtered signals. The three filtered signals are superimposed to generate a multi-channel output signal for driving the speaker array.

[0118] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the stereo processing methods provided by the above methods, the method comprising: including: The stereo audio signal, including the left channel signal and the right channel signal, is subjected to dry and wet sound separation processing to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The dry sound, left channel wet sound, and right channel wet sound are filtered by a preset first control filter, a second control filter, and a third control filter, respectively, to obtain the corresponding three-channel filtered signals. The three filtered signals are superimposed to generate a multi-channel output signal for driving the speaker array.

[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0120] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A stereo processing method, characterized in that, include: The stereo audio signal, including the left channel signal and the right channel signal, is subjected to dry and wet sound separation processing to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The dry sound, left channel wet sound and right channel wet sound are filtered by the preset first control filter, second control filter and third control filter respectively to obtain three filtered signals corresponding to the dry sound, left channel wet sound and right channel wet sound; The three filtered signals are superimposed to generate a multi-channel output signal for driving the speaker array.

2. The stereo processing method according to claim 1, characterized in that, The method for determining the first control filter, the second control filter, and the third control filter includes: Obtain the bright area transfer function matrix from each loudspeaker in the loudspeaker array to the target bright area, and the dark area transfer function matrix to the target dark area; Different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, respectively; Based on the bright area transfer function matrix, the dark area transfer function matrix, and their respective target sound field distributions, the coefficients of each control filter are calculated using an optimization algorithm.

3. The stereo processing method according to claim 2, characterized in that, Different target sound field distributions are set for the first control filter, the second control filter, and the third control filter, including: The target sound field of the first control filter is the sound field generated in the target bright area when the left and right channel speakers in a standard stereo configuration simultaneously play the same signal. The target sound field of the second control filter is the sound field generated in the target bright area when only the left channel speaker is working in a standard stereo configuration; The target sound field of the third control filter is the sound field generated in the target bright area when only the right channel speaker is working in a standard stereo configuration.

4. The stereo processing method according to claim 2, characterized in that, The calculation of the coefficients of each control filter using an optimization algorithm includes: Construct an objective function, which includes a bright area sound field reconstruction error term and a dark area sound pressure energy suppression term; By minimizing the objective function, the filter coefficients that maximize the ratio of sound pressure energy in the bright area to sound pressure energy in the dark area are obtained, and the calculated filter coefficients are stored as the coefficients of each control filter.

5. The stereo processing method according to claim 2, characterized in that, The acquisition of the bright area transfer function matrix from each loudspeaker in the loudspeaker array to the target bright area, and the dark area transfer function matrix to the target dark area, includes: Multiple discrete target control points are set within the target's light area and target's dark area, respectively; The bright zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all bright zone target control points; The dark zone transfer function matrix is ​​composed of the acoustic transfer functions of each loudspeaker in the loudspeaker array to all dark zone target control points.

6. The stereo processing method according to claim 1, characterized in that, The process of separating dry and wet audio signals from stereo audio signals, including left and right channel signals, includes: Calculate the complex correlation coefficient between the left channel signal and the right channel signal; Based on the complex correlation coefficient, the signal components with the same amplitude and phase in the left channel signal and the right channel signal are extracted as the dry sound; The left channel wet signal is obtained by subtracting the component corresponding to the left channel from the dry signal. The right channel wet signal is obtained by subtracting the component corresponding to the right channel from the dry signal.

7. The stereo processing method according to claim 6, characterized in that, The dry sound, left channel wet sound, and right channel wet sound satisfy the orthogonality condition, and the cross-correlation coefficient between the left channel wet sound and the dry sound is zero; the cross-correlation coefficient between the right channel wet sound and the dry sound is zero; and the cross-correlation coefficient between the left channel wet sound and the right channel wet sound is zero.

8. A stereo processing device, characterized in that, include: The separation module is used to perform dry and wet sound separation processing on the stereo audio signal including the left channel signal and the right channel signal to obtain dry sound, left channel wet sound, and right channel wet sound; wherein, the left channel wet sound is the signal component in the left channel signal that is not related to the right channel signal, the right channel wet sound is the signal component in the right channel signal that is not related to the left channel signal, and the dry sound is the signal component in the left channel signal and the right channel signal that has the same amplitude and phase; The filtering module is used to filter the dry sound, left channel wet sound, and right channel wet sound using a preset first control filter, a second control filter, and a third control filter, respectively, to obtain three filtered signals corresponding to the dry sound, left channel wet sound, and right channel wet sound. The output module is used to superimpose the three filtered signals to generate a multi-channel output signal for driving the speaker array.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the stereo processing method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the stereo processing method as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the stereo processing method as described in any one of claims 1 to 7.