An audio processing method, device, system and storage medium

By using a pre-generated filter matrix in the in-vehicle sound playback algorithm to determine the compensation filter and compensate and fuse the audio signals, the problems of low robustness and poor sound image positioning in the prior art are solved, and better tone performance and compensation effects are achieved.

CN115460513BActive Publication Date: 2025-06-27GUOGUANG ELECTRIC COMPANY LIMITED
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211280996.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-19
Publication Date
2025-06-27
Estimated Expiration
2042-10-19

AI Technical Summary

Technical Problem

The existing in-vehicle sound playback algorithms have shortcomings in terms of robustness and sound image positioning effects, resulting in timbre deterioration and compensation effects that are not obvious.

Method used

By decoding the audio file to be played, the left channel signal and the right channel signal are obtained, and the compensation filter is determined based on the pre-generated filter matrix to compensate and filter the signal. Then, the compensated signal is fused and delay output is performed according to the target delay time of the speaker.

Benefits of technology

It improves the robustness of sound playback and the sound image positioning effect in the car, enhances the expressiveness of the tone, and significantly improves the compensation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115460513B_ABST
    Figure CN115460513B_ABST
Patent Text Reader

Abstract

The present application discloses an audio processing method, apparatus, system and storage medium. The method includes: decoding an audio file to be played to obtain a replay signal, where the replay signal includes a left channel signal and a right channel signal; obtaining the target speaker category and the target delay duration of each speaker in the speaker group; respectively determining the compensation filters for the left channel signal and the right channel signal from a pre-generated filter matrix, and using the compensation filters to respectively perform compensation filtering on the corresponding left channel signal and right channel signal; fusing the compensated left channel signal and the right channel signal to generate a fused signal; determining the output signals corresponding to each speaker based on the fused signal and the target speaker category, and delaying the output signals according to the target delay duration of each speaker and then outputting them to the corresponding speakers, so as to realize two-channel or multi-channel sound replay on the left and right sides of the vehicle and improve the sound replay effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio analysis, and in particular, to an audio processing method, an audio processing device, an audio signal processing system, and a computer-readable storage medium. Background Art

[0002] Good audio equipment can improve the audio effect of a vehicle. One of the evaluation criteria for good audio equipment is whether it can reproduce the real live music effect. The playback effect of the audio equipment depends on the in-vehicle sound playback algorithm used by the audio equipment.

[0003] In the in-vehicle sound playback algorithms mentioned in the related art, a filter is used to compensate the impulse response from the speaker to the ears, so that the binaural signals at this time are the desired playback signals, but the playback area is limited to the measurement points. When this method is applied to in-vehicle sound playback, the algorithm robustness (i.e., stability) is significantly reduced, the sound image localization effect and timbre in sound playback become worse, and the compensation effect becomes insignificant. And using a microphone array to measure the binaural impulse response and compensating the array data can improve the stability of sound image localization during playback, but the processing process is complicated and the timbre improvement is not obvious. Summary of the Invention

[0004] The present application provides an audio processing method, device, system, and storage medium to solve the problems of low robustness, poor sound image localization effect and timbre in the existing in-vehicle sound playback algorithm, and insignificant compensation effect.

[0005] According to the first aspect of the present application, an audio processing method is provided. The method is applied to an audio signal processing system located in a vehicle. A speaker group is provided in the vehicle, and each speaker in the speaker group has a corresponding speaker category;

[0006] The method includes:

[0007] Decoding an audio file to be played to obtain a playback signal, where the playback signal includes a left-channel signal and a right-channel signal;

[0008] Obtaining the target speaker category and target delay duration of each speaker in the speaker group;

[0009] Respectively determining compensation filters for the left-channel signal and the right-channel signal from a pre-generated filter matrix, and using the compensation filters to respectively perform compensation filtering on the corresponding left-channel signal and right-channel signal;

[0010] Fusing the compensated and filtered left-channel signal and right-channel signal to generate a fused signal;

[0011] Determine output signals corresponding to each speaker based on the fused signal and the target speaker category, and delay the output signals according to the target delay duration of each speaker and then output them to the corresponding speakers.

[0012] According to a second aspect of the present application, there is provided an audio processing device, which is arranged in an audio signal processing system located in a vehicle. A speaker group is arranged in the vehicle, and each speaker in the speaker group has a corresponding speaker category;

[0013] The device includes:

[0014] A playback signal acquisition module, configured to decode an audio file to be played to obtain a playback signal, where the playback signal includes a left channel signal and a right channel signal;

[0015] A speaker information acquisition module, configured to acquire the target speaker category and the target delay duration of each speaker in the speaker group;

[0016] A compensation filtering module, configured to respectively determine compensation filters for the left channel signal and the right channel signal from a pre-generated filter matrix, and use the compensation filters to respectively perform compensation filtering on the corresponding left channel signal and right channel signal;

[0017] A signal fusion module, configured to fuse the compensated and filtered left channel signal and right channel signal to generate a fused signal;

[0018] A signal output module, configured to determine output signals corresponding to each speaker based on the fused signal and the target speaker category, and delay the output signals according to the target delay duration of each speaker and then output them to the corresponding speakers.

[0019] According to a third aspect of the present application, there is provided an audio signal processing system, which includes:

[0020] At least one processor and

[0021] A memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can be used to implement the method of the first aspect above.

[0023] According to a fourth aspect of the present application, there is provided a computer-readable storage medium, which stores computer instructions for causing a processor to implement the method of the first aspect above when executed.

[0024] In this embodiment, by determining the compensation filters corresponding to the left-channel signal and the right-channel signal, and using the compensation filters to perform compensation filtering on the corresponding left-channel signal and right-channel signal, the left-channel signal and the right-channel signal after compensation filtering are made closer to the signals in an anechoic chamber. Then, after the left-channel signal and the right-channel signal after compensation filtering are fused into a fused signal, an output signal is further obtained according to the fused signal. Then, the output signal can be delayed by the target delay duration corresponding to each speaker and then fed to the corresponding speaker, so that each speaker can receive the output signal simultaneously and play the received output signal simultaneously, realizing two-channel or multi-channel sound reproduction on both sides of the vehicle and improving the sound reproduction effect.

[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1 is a flowchart of an audio processing method provided in Embodiment 1 of the present application;

[0028] Figure 2 is a schematic diagram of the distribution of an in-vehicle speaker group provided in Embodiment 1 of the present application;

[0029] Figure 3 is a flowchart of a method for generating a compensation filter matrix provided in Embodiment 1 of the present application;

[0030] Figure 4 is a schematic structural diagram of an audio processing device provided in Embodiment 2 of the present application;

[0031] Figure 5 is a schematic structural diagram of an audio signal processing system provided in Embodiment 3 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] Embodiment 1

[0035] Figure 1 It is a flowchart of an audio processing method provided for Embodiment 1 of this application. This embodiment can be applied to an audio signal processing system located inside a vehicle. This audio signal processing system is a DSP (Digital Signal Processing) system for performing sound image positioning and sound reproduction processing.

[0036] In this embodiment, a speaker group can be arranged inside the vehicle. Exemplarily, the speaker group can be arranged in front of the driver's seat and the passenger's seat. Each speaker in the speaker group has a corresponding speaker category. Exemplarily, the speaker category can include a high-frequency category, a mid-frequency category, a low-frequency category, and a full-frequency category. The speaker group can include a left speaker group located on the left side of the vehicle, a right speaker group located on the right side of the vehicle and symmetrically distributed with the left speaker group, and a full-frequency speaker located in the middle of the vehicle and at the same distance from the left speaker group and the right speaker group.

[0037] For example, as Figure 2As shown in the figure, the left speaker group is located in the front left of the driver's seat and includes a left high-frequency speaker S1H, a left mid-frequency speaker S1M, and a left low-frequency speaker S1L; the right speaker group is located in the front right of the passenger seat and includes a right high-frequency speaker S2H, a right mid-frequency speaker S2M, and a right low-frequency speaker S2L. Among them, S1H and S2H are symmetrically distributed, S1M and S2M are symmetrically distributed, and S1L and S2L are symmetrically distributed; the full-frequency speaker SO is located in the middle or the front middle of the left speaker group and the right speaker group, and thus can also be called a center full-frequency speaker. It can be set that the distance from S0 to S1H and S2H is equal, the distance from S0 to S1M and S2M is equal, and the distance from S0 to S1L and S2L is equal.

[0038] As Figure 1 shown, the method may include the following steps:

[0039] Step 101: Decode the audio file to be played to obtain a replay signal, where the replay signal includes a left-channel signal and a right-channel signal.

[0040] This replay signal is an audio signal that needs to be replayed in the vehicle. The source of the replay signal can be an audio file from the network, or an audio file from other devices, or an audio file from the vehicle's local storage. This embodiment does not limit this.

[0041] By decoding the audio file, a replay signal can be obtained, and this replay signal can include a left-channel signal and a right-channel signal.

[0042] If the replay signal is a stereo signal, then a left-channel stereo signal and a right-channel stereo signal are decoded.

[0043] If the replay signal is a channel signal from a 5.1-channel or 7.1-channel source, then the middle-channel signal in the channel signal and preset correction data are used to correct the left-channel signal and the right-channel signal. For example, if the preset correction data is then the corrected left-channel signal X1 is:

[0044]

[0045] The corrected right-channel signal X2 is:

[0046]

[0047] Among them, L is the left-channel signal in the channel signal, R is the right-channel signal in the channel signal, and C is the middle-channel signal in the channel signal.

[0048] Step 102: Obtain the target speaker category and the target delay duration of each speaker in the speaker group.

[0049] When implemented, each speaker in the speaker group has a corresponding speaker category and a delay duration. The delay duration of each speaker can be a preset delay duration, or can be pre-calculated and stored in the audio signal processing system. This delay duration is used to balance the durations of the signals from each speaker reaching both ears, so that the signals emitted by each speaker can reach both ears as simultaneously as possible.

[0050] Step 103: Determine the compensation filters for the left-channel signal and the right-channel signal respectively from a pre-generated filter matrix, and use the compensation filters to perform compensation filtering on the corresponding left-channel signal and right-channel signal respectively.

[0051] Wherein, if the left-channel signal and the right-channel signal are corrected in step 101, the left-channel signal and the right-channel signal in step 103 are the corrected left-channel signal and right-channel signal.

[0052] Wherein, the compensation filter matrix is a matrix containing multiple compensation filters pre-generated and stored in the audio signal processing system. In the compensation filter matrix, each element has corresponding category information and channel information. The category information is used to constrain the speaker categories to which the element can be applied, and the channel information is used to constrain whether the channel signal to which the element can be applied is a left-channel signal or a right-channel signal. For example, assume the compensation filter matrix is If the category information of the filter matrix h 1f is high-frequency category, medium-frequency category and low-frequency category, and the channel information is left-channel signal, it means that this h 1f is applicable to the compensation filtering of the left-channel signals of high-frequency category, medium-frequency category and low-frequency category; if the category information of the filter matrix h 4f is full-frequency category, and the channel information is right-channel signal, it means that this h 4f is applicable to the compensation filtering of the right-channel signals of the full-frequency category.

[0053] In one implementation, the compensation filter may include a first compensation filter corresponding to the left-channel signal and a second compensation filter corresponding to the right-channel signal; then the following steps can be used in step 103 to determine the compensation filters for the left-channel signal and the right-channel signal:

[0054] According to the channel information, obtain a first filter set corresponding to the left-channel signal, and a second filter set corresponding to the right-channel information; according to the speaker category of the speaker to which the signal needs to be fed, select a matching compensation filter from the first filter set as the first compensation filter, and select a matching compensation filter from the second filter set as the second compensation filter.

[0055] For example, in the above-mentioned compensation filter matrix h 1f and h 3f correspond to the left-channel signal, h 2f and h 4f correspond to the right-channel signal, and the category information of h 1f and h 2f includes high-frequency category, medium-frequency category, and low-frequency category, and the category information of h 3f and h 4f includes full-frequency category. Then the first filter set includes h 1f and h 3f , and the second filter set includes h 2f and h 4f . For the left-channel signal, it corresponds to the first filter set {h 1f , h 3f}. If the speakers to be fed are the left speaker group and the right speaker group in Figure 2 , then the first compensation filter is h 1f ; if the speakers to be fed are the full-frequency speakers in Figure 2 , then the second compensation filter is h 3f . For the right-channel signal, it corresponds to the second filter set {h 2f , h 4f}. If the speakers to be fed are the left speaker group and the right speaker group in Figure 2 , then the first compensation filter is h 2f ; if the speakers to be fed are the full-frequency speakers in Figure 2 , then the second compensation filter is h 4f .

[0056] Next, the first compensation filter can be used to perform compensation filtering on the left-channel signal, and the second compensation filter can be used to perform compensation filtering on the right-channel signal.

[0057] Step 104: Fuse the compensated and filtered left-channel signal and the right-channel signal to generate a fused signal.

[0058] In one implementation, one way of fusion can be: using an adder in the audio signal processing system to add the left-channel signal filtered by the first compensation filter and the right-channel signal filtered by the second compensation filter to obtain a fused signal.

[0059] Step 105: Determine the output signals corresponding to each speaker based on the fused signal and the target speaker category, and delay the output signals according to the target delay duration of each speaker and then output them to the corresponding speakers.

[0060] In this step, after obtaining the fused signal, the output signal fed to each speaker can be determined according to the speaker category of each speaker.

[0061] In one implementation, if the speaker category is a high-frequency category, the fused signal is filtered by a high-pass filter to obtain the output signal; if the speaker category is a mid-frequency category, the fused signal is filtered by a band-pass filter to obtain the output signal; if the speaker category is a low-frequency category, the fused signal is filtered by a low-pass filter to obtain the output signal; if the speaker category is a full-frequency category, the fused signal can be directly used as the output signal.

[0062] After obtaining the output signals of the speakers in the speaker group, the output signals can be delayed by the target delay duration corresponding to each speaker and then fed to the corresponding speaker, so that each speaker can receive the output signal simultaneously and play the received output signal simultaneously, realizing two-channel or multi-channel sound reproduction on both sides of the vehicle.

[0063] In one embodiment, refer to Figure 3 FIG. shows a flowchart of a method for generating a compensation filter matrix, which can be generated by a control platform connected to a test vehicle. The test vehicle includes a speaker group as shown in Figure 2 To achieve better effects, a pickup group can also be placed in the test vehicle. The pickup group is used to simulate the human ear effect and includes a left ear pickup and a right ear pickup, simulating the left ear and the right ear respectively. The speaker group and the pickup group are respectively connected to the control platform.

[0064] In a preferred example, the pickup group can be placed in the middle of the driver's seat and at a height close to that of the human ear. For example, an artificial head mold can be placed at a specified height in the middle of the driver's seat, and the pickup group can be placed at the binaural positions of the artificial head mold.

[0065] As Figure 3 shown, the compensation filter matrix is generated using the following steps:

[0066] S1, determine the delay duration of each speaker in the speaker group through a first test signal.

[0067] Among them, the first test signal is a signal used to feed each speaker, and the delay duration of each speaker is obtained by sequentially feeding the first test signal to each speaker.

[0068] The first test signal can be a test signal sent by the control platform or a signal obtained after processing the test signal sent to the test platform, which is specifically determined according to the speaker category of the speaker. Specifically, after obtaining the first original test signal sent by the control platform, if the speaker category of the current speaker is the high-frequency category, a high-pass filter is used to filter the first original test signal to obtain the first test signal; if the speaker category of the current speaker is the mid-frequency category, a band-pass filter is used to filter the first original test signal to obtain the first test signal; if the speaker category of the current speaker is the low-frequency category, a low-pass filter is used to filter the first original test signal to obtain the first test signal; if the speaker category of the current speaker is the full-frequency category, the first original test signal is directly used as the first test signal.

[0069] When determining the delay duration of each speaker, since the left speaker group and the right speaker group are symmetrically distributed and the delay of the right speaker group is designed based on the left speaker group, the delay duration can be calculated based on the left speaker group in this embodiment. In one implementation, each speaker in the left speaker group is traversed in sequence. For the currently traversed speaker, the first test signal is sent to the current speaker for the speaker to play the first test signal. Then, the binaural impulse response corresponding to the first test signal is obtained, where the binaural impulse response includes the left-ear impulse response and the right-ear impulse response. For example, the impulse response can be expressed as S0 i 、S1H i 、S1M i 、S1L i , where i = 1, 2 respectively represent the sound transmission paths from the speaker to the left and right ears. Then, based on the binaural impulse response and the first test signal, the delay duration of the current speaker is determined, and the next speaker is continued to be traversed.

[0070] In one embodiment, the step of determining the delay duration of each speaker based on the binaural impulse response and the first test signal may further include the following steps:

[0071] Respectively obtain the time differences between the first test signal sent by the left high-frequency speaker, the left mid-frequency speaker, and the left low-frequency speaker and the first test signal received by the left-ear pickup, which are respectively denoted as Obtain the time difference between the first test signal sent by the mid-frequency speaker and the first test signal received by the right-ear pickup, which is denoted as Calculate the set constant values respectively with the The differences are respectively denoted as t1, t2, t3, and t4, and t1 is used as the delay duration of the left and right high-frequency speakers, t2 is used as the delay duration of the left and right mid-frequency speakers, t3 is used as the delay duration of the left and right low-frequency speakers, and t4 is used as the delay duration of the full-frequency speaker.

[0072] Specifically, for the left speaker group, let:

[0073]

[0074]

[0075]

[0076]

[0077] where T is a set constant value.

[0078] According to the above formula, the delay durations t1 to t4 of each speaker in the left speaker group and the full-frequency speaker can be obtained. According to the principle of symmetric design, the delay durations of each speaker in the right speaker group are also t1 to t3.

[0079] S2. Determine the first binaural impulse response of the left speaker group, the second binaural impulse response of the right speaker group, and the third binaural impulse response of the full-frequency speaker through the second test signal and the delay duration.

[0080] After obtaining the binaural delay durations of each speaker in the speaker group, a second measurement is performed. Different from sending the first test signal to each speaker in turn during the first measurement, during the second measurement, the second test signal is sent to the speakers on each side simultaneously in the same time period.

[0081] In one embodiment, step S2 may further include the following steps:

[0082] S21. According to the delay durations of each speaker in the left speaker group, after delaying for the time corresponding to the delay duration, send the second test signal to the corresponding speaker, so that each speaker in the left speaker group plays the second test signal simultaneously, and obtain the binaural impulse response corresponding to the second test signal as the first binaural impulse response.

[0083] Similar to the first measurement, before sending the second test signal to each speaker in the left speaker group, it is necessary to filter the test signal through the corresponding band-pass filter. Specifically, after obtaining the second original test signal sent by the control platform, if the speaker category of the current speaker is the high-frequency category, the second original test signal is filtered by a high-pass filter to obtain the second test signal; if the speaker category of the current speaker is the medium-frequency category, the second original test signal is filtered by a band-pass filter to obtain the second test signal; if the speaker category of the current speaker is the low-frequency category, the second original test signal is filtered by a low-pass filter to obtain the second test signal.

[0084] After obtaining the second test signal, increase the corresponding speaker and then send the second test signal to the corresponding speaker to enable each speaker to play the second test signal simultaneously, and measure the equivalent first binaural impulse response h 1i 。

[0085] S22. According to the delay duration of each speaker in the right speaker group, after delaying for the time corresponding to the delay duration, send the second test signal to the corresponding speaker so that each speaker in the right speaker group plays the second test signal simultaneously, and obtain the binaural impulse response corresponding to the second test signal as the second binaural impulse response.

[0086] Similarly, before sending the second test signal to each speaker in the right speaker group, it is necessary to filter the test signal through the corresponding band-pass filter. Specifically, after obtaining the second original test signal sent by the control platform, if the speaker category of the current speaker is the high-frequency category, the second original test signal is filtered by a high-pass filter to obtain the second test signal; if the speaker category of the current speaker is the medium-frequency category, the second original test signal is filtered by a band-pass filter to obtain the second test signal; if the speaker category of the current speaker is the low-frequency category, the second original test signal is filtered by a low-pass filter to obtain the second test signal.

[0087] After obtaining the second test signal, increase the corresponding speaker and then send the second test signal to the corresponding speaker to enable each speaker to play the second test signal simultaneously, and measure the equivalent second binaural impulse response h 2i 。

[0088] S23. According to the delay duration of the full-frequency speaker, after delaying for the time corresponding to the delay duration, send the second test signal to the full-frequency speaker so that the full-frequency speaker plays the second test signal, and obtain the binaural impulse response corresponding to the second test signal as the third binaural impulse response.

[0089] For the full - range speaker, after directly delaying the corresponding delay duration, the second original test signal is sent to the full - range speaker, thereby measuring the third binaural impulse response h 0i .

[0090] Through step S2, the first binaural impulse response of the left speaker group includes h 11 and h 12 , the second binaural impulse response of the right speaker group includes h 21 and h 22 , and the third binaural impulse response of the full - range speaker includes h 01 and h 02 .

[0091] S3. Respectively perform stability processing on the first binaural impulse response, the second binaural impulse response, and the third binaural impulse response to obtain the corresponding first stable binaural impulse response, second stable binaural impulse response, and third stable binaural impulse response.

[0092] In order to obtain a more stable binaural amplitude - frequency response at the measurement point, this step can perform stability processing on the binaural impulse response.

[0093] Taking the first binaural impulse response as an example, the method of stability processing includes the following steps:

[0094] S31. Perform Fourier transform on each impulse response in the first binaural impulse response in the time domain to obtain the binaural transfer function in the frequency domain.

[0095] Among them, this Fourier transform can include an N - point Fourier transform, expressed as:

[0096] H(k)=FFT(h(n),N), N is an even number

[0097] In the above formula, h(n) is the binaural impulse response (including the first binaural impulse response), and H(k) is the binaural transfer function obtained after performing the N - point Fourier transform.

[0098] S32. In a specified interval, perform weighted summation on the binaural transfer function using a preset window function.

[0099] In one implementation, the specified interval can be [k - m(k)+1,k + m(k)], and a preset window function w can be used to perform weighted summation on the binaural transfer function H(k) to obtain H1(k), as shown in the following formula:

[0100]

[0101] m(k)=roundn[(fu - fl) / 2,0], and the process roundn is an integer;

[0102] fu = f(k)·2 0.5·oct ,fl = f(k)·0.5 0.5·oct ,oct is a constant

[0103]

[0104] where f s is the sampling rate for measuring the binaural impulse signal.

[0105] S33. Keep the phase of the binaural transfer function after weighted summation unchanged and correct its amplitude.

[0106] After obtaining H1(k), keep the phase of H1(k) unchanged and correct its amplitude to obtain the corrected amplitude-frequency response H3(k) as follows:[[]]

[0107]

[0108]

[0109] S34. Perform an inverse Fourier transform on the corrected binaural transfer function to obtain a stable impulse response.

[0110] In this step, perform an inverse Fourier transform on H3(k), that is, obtain the stable binaural impulse response h s (n) after robust design as follows:[[]]

[0111] h s (n) = IFFT(H3(k))

[0112] Through step S3, the first stable binaural impulse response of the left speaker group includes h 11s and h 12s , the second stable binaural impulse response of the right speaker group includes h 21s and h 22s , and the third stable binaural impulse response of the full-frequency speaker includes h 01s and h 02s .

[0113] S4. Generate a compensation filter matrix according to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response.

[0114] In one embodiment, after obtaining the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response by the method of step S3, the following steps can be used to generate a compensation filter matrix:[[]]

[0115] S41, respectively obtaining convolution matrices corresponding to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response.

[0116] Among them, each stable binaural impulse response has a corresponding convolution matrix, then h 01s 、h 02s ,h 11s 、h 12s ,h 21s 、h 22s The corresponding convolution matrices are: h 01s_convm 、h 02s_convm ,h 11s_convm 、h 12s_convm ,h 21s_convm 、h 22s_convm .

[0117] In one implementation, the convolution matrix can be expressed as:

[0118] h s_convm =convmtx(h s (n),L1)

[0119] Wherein, L1 is the length of the compensation filter, and the compensation filter may be a FIR filter.

[0120] For example, for h 01s For , its convolution matrix is ​​expressed as:

[0121] h 01s_convm =convmtx(h 01s ,L)

[0122] The other stable binaural impulse responses are similar.

[0123] S42, using each convolution matrix to generate a speaker matrix.

[0124] In one embodiment, the speaker matrix can be represented as follows:

[0125]

[0126] S43, obtaining a preset expected playback matrix.

[0127] The expected playback matrix is ​​a matrix generated based on an anechoic room environment. The anechoic room environment may exemplarily include an azimuth angle θ of a loudspeaker arranged in the anechoic room and a binaural impulse response hrir measured in the anechoic room.

[0128] In one implementation, the expected replay matrix can be expressed as:

[0129]

[0130] Among them, B represents the desired playback matrix, and hrir θ (:, 1) represents the impulse response collected by the left ear pickup after the speaker at the θ azimuth angle emits a signal; hrir θ (:, 2) represents the impulse response collected by the right ear pickup after the speaker at the θ azimuth angle emits a signal; hrir -θ (:, 1) represents the impulse response collected by the left ear pickup after the speaker at the -θ azimuth angle emits a signal; hrir -θ (:, 2) represents the impulse response collected by the right ear pickup after the speaker at the -θ azimuth angle emits a signal. And θ and -θ are the azimuth angles of two speakers symmetrically arranged in an anechoic chamber.

[0131] S44. Obtain the transpose matrix of the speaker matrix.

[0132] S45. Determine the compensation filter matrix according to the speaker matrix, the transpose matrix, and the desired playback matrix.

[0133] In one implementation, the relationship between the speaker matrix A, the compensation filter matrix h f , and the desired playback matrix B is as follows:

[0134] A·h f = B

[0135] Through matrix pseudo-inverse, the compensation filter h f can be obtained:

[0136] h f = (A′·A + β·E) -1 ·A′·B

[0137] Where A′ is the transpose matrix of A, β is a regularization factor, which is a user-defined constant; E is the identity matrix.

[0138] In this embodiment, stable binaural impulse responses are obtained through two measurements and stability processing, and a compensation filter matrix is generated based on the stable binaural impulse responses, thereby reducing the difference in the binaural amplitude-frequency response at a position deviated from the measurement point, improving the robustness of the algorithm, and the processed binaural amplitude-frequency response curve is smooth, the amplitude-frequency curve of the compensation filter becomes smooth, and the timbre is significantly improved. The entire processing process is simpler and the effect is obvious.

[0139] Embodiment 2

[0140] Figure 4Schematic structural diagram of an audio processing device provided in the second embodiment of the present application. The device is disposed in an audio signal processing system located inside a vehicle. A speaker group is provided inside the vehicle, and each speaker in the speaker group has a corresponding speaker category;

[0141] The device may include the following modules:

[0142] A playback signal acquisition module 201, configured to decode an audio file to be played to obtain a playback signal, where the playback signal includes a left-channel signal and a right-channel signal;

[0143] A speaker information acquisition module 202, configured to acquire the target speaker category and the target delay duration of each speaker in the speaker group;

[0144] A compensation filtering module 203, configured to respectively determine a compensation filter for the left-channel signal and the right-channel signal from a pre-generated filter matrix, and use the compensation filter to respectively perform compensation filtering on the corresponding left-channel signal and right-channel signal;

[0145] A signal fusion module 204, configured to fuse the compensated and filtered left-channel signal and right-channel signal to generate a fused signal;

[0146] A signal output module 205, configured to determine an output signal corresponding to each speaker based on the fused signal and the target speaker category, and delay the output signal according to the target delay duration of each speaker and then output it to the corresponding speaker.

[0147] In one embodiment, the speaker group includes a left speaker group located on the left side of the vehicle, a right speaker group located on the right side of the vehicle and symmetrically distributed with the left speaker group, and a full-frequency speaker located at an intermediate position in front of the left speaker group and the right speaker group and at an equal distance from the left speaker group and the right speaker;

[0148] The audio signal processing system is connected to a control platform, and the control platform includes a compensation filter matrix generation module. The compensation filter matrix module may further include the following modules:

[0149] A delay duration determination module, configured to determine the delay duration of each speaker in the speaker group through a first test signal;

[0150] A pulse response determination module, configured to determine a first binaural pulse response of the left speaker group, a second binaural pulse response of the right speaker group, and a third binaural pulse response of the full-frequency speaker through a second test signal and the delay duration;

[0151] A stability processing module, configured to perform stability processing on the first binaural impulse response, the second binaural impulse response, and the third binaural impulse response respectively, to obtain corresponding first stable binaural impulse response, second stable binaural impulse response, and third stable binaural impulse response;

[0152] A matrix generation module, configured to generate a compensation filter matrix according to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response.

[0153] In one embodiment, the delay duration determination module is configured to:

[0154] Traverse each speaker in the left speaker group in sequence. For the currently traversed speaker, send a first test signal to the current speaker to play the first test signal by the speaker;

[0155] Obtain the binaural impulse response corresponding to the first test signal;

[0156] Determine the delay duration of the current speaker according to the binaural impulse response and the first test signal, and continue to traverse the next speaker.

[0157] In one embodiment, the speaker categories include high-frequency category, medium-frequency category, low-frequency category, and full-frequency category; the compensation filter matrix module further includes a first test signal acquisition module, configured to:

[0158] Obtain a first original test signal;

[0159] If the speaker category of the current speaker is the high-frequency category, filter the first original test signal with a high-pass filter to obtain a first test signal;

[0160] If the speaker category of the current speaker is the medium-frequency category, filter the first original test signal with a band-pass filter to obtain a first test signal;

[0161] If the speaker category of the current speaker is the low-frequency category, filter the first original test signal with a low-pass filter to obtain a first test signal;

[0162] If the speaker category of the current speaker is the full-frequency category, use the first original test signal as the first test signal.

[0163] In one embodiment, the left speaker group includes a left high-frequency speaker, a left medium-frequency speaker, and a left low-frequency speaker; the right speaker group includes a right high-frequency speaker, a right medium-frequency speaker, and a right low-frequency speaker; a pickup group is further arranged in the vehicle, and the pickup group includes a left ear pickup and a right ear pickup;

[0164] The delay time determination module is further configured to:

[0165] Obtain the time differences between the first test signal emitted by the left high-frequency speaker, the left mid-frequency speaker, and the left low-frequency speaker respectively, and the first test signal received by the left ear pickup, and denote them as

[0166] Obtain the time difference between the first test signal emitted by the mid-frequency speaker and the first test signal received by the right ear pickup, and denote it as

[0167] Calculate the differences between the set constant values and the respectively, and denote them as t1, t2, t3, and t4. Take t1 as the delay time of the left high-frequency speaker and the right high-frequency speaker, take t2 as the delay time of the left mid-frequency speaker and the right mid-frequency speaker, take t3 as the delay time of the left low-frequency speaker and the right low-frequency speaker, and take t4 as the delay time of the full-frequency speaker.

[0168] In one embodiment, the impulse response determination module is specifically configured to:

[0169] According to the delay time of each speaker in the left speaker group, after delaying for the time corresponding to the delay time, send a second test signal to the corresponding speaker, so that each speaker in the left speaker group plays the second test signal simultaneously, and obtain the binaural impulse response corresponding to the second test signal as the first binaural impulse response;

[0170] According to the delay time of each speaker in the right speaker group, after delaying for the time corresponding to the delay time, send a second test signal to the corresponding speaker, so that each speaker in the right speaker group plays the second test signal simultaneously, and obtain the binaural impulse response corresponding to the second test signal as the second binaural impulse response;

[0171] According to the delay time of the full-frequency speaker, after delaying for the time corresponding to the delay time, send a second test signal to the full-frequency speaker, so that the full-frequency speaker plays the second test signal, and obtain the binaural impulse response corresponding to the second test signal as the third binaural impulse response.

[0172] In one embodiment, the stability processing module is specifically configured to:

[0173] Perform Fourier transform on each impulse response in the first binaural impulse response in the time domain to obtain the binaural transfer function in the frequency domain;

[0174] Within a specified interval, perform weighted summation on the binaural transfer function using a preset window function;

[0175] Keep the phase of the binaural transfer function after weighted summation unchanged, and correct its amplitude;

[0176] Perform inverse Fourier transform on the corrected binaural transfer function to obtain a stable impulse response.

[0177] In one embodiment, the matrix generation module is specifically configured to:

[0178] Obtain the convolution matrices corresponding to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response respectively;

[0179] Generate a speaker matrix using each convolution matrix;

[0180] Obtain a preset desired playback matrix, where the desired playback matrix is a matrix generated based on an anechoic chamber environment;

[0181] Obtain the transpose matrix of the speaker matrix;

[0182] Determine a compensation filter matrix according to the speaker matrix, the transpose matrix, and the desired playback matrix.

[0183] In one embodiment, the compensation filter includes a first compensation filter and a second compensation filter; the compensation filtering module 203 is specifically configured to:

[0184] According to the channel information, obtain a first filter set corresponding to the left channel signal and a second filter set corresponding to the right channel information;

[0185] Select a matching compensation filter from the first filter set as the first compensation filter and select a matching compensation filter from the second filter set as the second compensation filter according to the speaker category of the speaker to which feeding is required.

[0186] In one embodiment, the device may further include:

[0187] A signal correction module, configured to, if the playback signal is a channel signal from a 5.1-channel or 7.1-channel, correct the left channel signal and the right channel signal using the middle channel signal in the channel signal and preset correction data.

[0188] An audio processing device provided by an embodiment of the present application can execute an audio processing device provided by Embodiment 1 of the present application, and has corresponding functional modules and beneficial effects for executing the method in Embodiment 1 above.

[0189] Embodiment III

[0190] Figure 5 FIG. shows a schematic structural diagram of an audio signal processing system 10 that can be used to implement the method embodiments of the present application. As Figure 5 shown, the audio signal processing system 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the audio signal processing system 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0191] Multiple components in the audio signal processing system 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the audio signal processing system 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0192] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method described in Embodiment I.

[0193] In some embodiments, the method described in Embodiment 1 can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the audio signal processing system 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method described in Embodiment 1 above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the method described in Embodiment 1 by any other suitable means (e.g., by means of firmware).

[0194] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0195] The computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0196] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0197] To provide for interaction with a user, the systems and techniques described herein can be implemented on an audio signal processing system having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the audio signal processing system. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0198] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0199] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0200] It should be understood that various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this application can be achieved, and no limitations are imposed herein.

[0201] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. An audio processing method, characterized in that The method is applied to an audio signal processing system located inside a vehicle. A speaker group is provided inside the vehicle, and each speaker in the speaker group has a corresponding speaker category; The method includes: Decoding the audio file to be played to obtain a replay signal, where the replay signal includes a left-channel signal and a right-channel signal; Obtaining the target speaker category and target delay duration of each speaker in the speaker group; Respectively determining the compensation filters for the left-channel signal and the right-channel signal from a pre-generated filter matrix, and using the compensation filters to respectively perform compensation filtering on the corresponding left-channel signal and right-channel signal; Fusing the compensated and filtered left-channel signal and right-channel signal to generate a fused signal; Based on the fused signal and the target speaker category, determining the output signal corresponding to each speaker, and delaying the output signal according to the target delay duration of each speaker and outputting it to the corresponding speaker; Wherein, the speaker group includes a left speaker group located on the left side of the vehicle, a right speaker group located on the right side of the vehicle and symmetrically distributed with the left speaker group, and a full-frequency speaker located at the middle position in front of the left speaker group and the right speaker group and at an equal distance from the left speaker group and the right speaker group; The method further includes: Determining the delay duration of each speaker in the speaker group through a first test signal; Determining a first binaural impulse response of the left speaker group, a second binaural impulse response of the right speaker group, and a third binaural impulse response of the full-frequency speaker through a second test signal and the delay duration; Respectively performing stability processing on the first binaural impulse response, the second binaural impulse response, and the third binaural impulse response to obtain corresponding first stable binaural impulse responses, second stable binaural impulse responses, and third stable binaural impulse responses; Generating a compensation filter matrix according to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response.

2. The method according to claim 1, wherein The determining the delay duration of each speaker in the speaker group through the first test signal includes: Sequentially traversing each speaker in the left speaker group. For the currently traversed speaker, sending a first test signal to the current speaker so that the speaker plays the first test signal; Obtaining the binaural impulse response corresponding to the first test signal; Determining the delay duration of the current speaker according to the binaural impulse response and the first test signal, and continuing to traverse the next speaker.

3. The method according to claim 2, wherein The speaker category includes a high-frequency category, a mid-frequency category, a low-frequency category, and a full-frequency category; Before sending the first test signal to the current speaker, the method further includes: Obtaining a first original test signal; If the speaker category of the current speaker is the high-frequency category, filtering the first original test signal with a high-pass filter to obtain a first test signal; If the speaker category of the current speaker is the mid-frequency category, filtering the first original test signal with a band-pass filter to obtain a first test signal; If the speaker category of the current speaker is a low-frequency category, a low-pass filter is used to filter the first original test signal to obtain a first test signal; If the speaker category of the current speaker is a full-frequency category, the first original test signal is used as the first test signal.

4. The method according to claim 2 or 3, characterized in that, The left speaker group includes a left high-frequency speaker, a left mid-frequency speaker, and a left low-frequency speaker; the right speaker group includes a right high-frequency speaker, a right mid-frequency speaker, and a right low-frequency speaker; a pickup group is further provided in the vehicle, and the pickup group includes a left ear pickup and a right ear pickup; Determining the delay duration of each speaker according to the binaural impulse response and the first test signal includes: Obtain the time differences between the first test signal emitted by the left high-frequency speaker, the left mid-frequency speaker, and the left low-frequency speaker respectively, and the first test signal received by the left ear pickup, and record them respectively as Obtain the time difference between the medium-frequency speaker emitting the first test signal and the right ear pickup receiving the first test signal, denoted as Calculate the differences between the calculated set constant values and the respectively, denoted as t1, t2, t3, and t4, and use the t1 as the delay duration of the left and right high-frequency speakers, use the t2 as the delay duration of the left and right mid-frequency speakers, use the t3 as the delay duration of the left and right low-frequency speakers, and use the t4 as the delay duration of the full-frequency speaker.

5. The method according to any one of claims 1 to 4, characterized in that, Determining the first binaural impulse response of the left speaker group, the second binaural impulse response of the right speaker group, and the third binaural impulse response of the full-frequency speaker through the second test signal and the delay duration includes: According to the delay duration of each speaker in the left speaker group, after delaying for a time corresponding to the delay duration, a second test signal is sent to the corresponding speaker so that each speaker in the left speaker group plays the second test signal simultaneously, and the binaural impulse response corresponding to the second test signal is obtained as the first binaural impulse response; According to the delay duration of each speaker in the right speaker group, after delaying for a time corresponding to the delay duration, a second test signal is sent to the corresponding speaker so that each speaker in the right speaker group plays the second test signal simultaneously, and the binaural impulse response corresponding to the second test signal is obtained as the second binaural impulse response; According to the delay duration of the full-frequency speaker, after delaying for a time corresponding to the delay duration, a second test signal is sent to the full-frequency speaker so that the full-frequency speaker plays the second test signal, and the binaural impulse response corresponding to the second test signal is obtained as the third binaural impulse response.

6. The method according to any one of claims 1-4, characterized in that The following method is used to perform stability processing on the first binaural impulse response to obtain a corresponding first stable binaural impulse response: Perform Fourier transform on each impulse response in the first binaural impulse response in the time domain to obtain a binaural transfer function in the frequency domain; In a specified interval, a preset window function is used to perform weighted summation on the binaural transfer function; Keep the phase of the binaural transfer function after weighted summation unchanged and correct its amplitude; Perform inverse Fourier transform on the corrected binaural transfer function to obtain a stable impulse response.

7. The method according to any one of claims 1 to 4, characterized in that Generating a compensation filter matrix according to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response includes: Respectively obtain the convolution matrices corresponding to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response; Use each convolution matrix to generate a speaker matrix; Obtain a preset desired playback matrix, and the desired playback matrix is a matrix generated based on an anechoic chamber environment; Obtain the transpose matrix of the speaker matrix; Determine a compensation filter matrix according to the speaker matrix, the transpose matrix, and the desired playback matrix.

8. The method according to claim 1, wherein The compensation filter includes a first compensation filter and a second compensation filter; Determining the compensation filters for the left-channel signal and the right-channel signal respectively from a pre-generated filter matrix includes: According to the channel information, obtaining a first filter set corresponding to the left-channel signal and a second filter set corresponding to the right-channel signal; According to the speaker category of the speakers to be fed, selecting a matching compensation filter from the first filter set as the first compensation filter, and selecting a matching compensation filter from the second filter set as the second compensation filter.

9. The method according to claim 1 or 8, characterized in that After decoding the audio file to be played to obtain a replay signal, where the replay signal includes a left-channel signal and a right-channel signal, the method further includes: If the replay signal is a channel signal from a 5.1-channel or 7.1-channel, using the center-channel signal in the channel signal and preset correction data to correct the left-channel signal and the right-channel signal.

10. An audio processing device, characterized in that, The device is arranged in an audio signal processing system in a vehicle, and a speaker group is arranged in the vehicle, and each speaker in the speaker group has a corresponding speaker category; The device includes: A replay signal acquisition module, configured to decode the audio file to be played to obtain a replay signal, where the replay signal includes a left-channel signal and a right-channel signal; A speaker information acquisition module, configured to acquire the target speaker category and the target delay duration of each speaker in the speaker group; A compensation filtering module, configured to determine the compensation filters for the left-channel signal and the right-channel signal respectively from a pre-generated filter matrix, and use the compensation filters to perform compensation filtering on the corresponding left-channel signal and right-channel signal respectively; A signal fusion module, configured to fuse the compensated and filtered left-channel signal and right-channel signal to generate a fusion signal; A signal output module, configured to determine the output signal corresponding to each speaker based on the fusion signal and the target speaker category, and delay the output signal according to the target delay duration of each speaker and then output it to the corresponding speaker; Wherein, the speaker group includes a left speaker group on the left side of the vehicle, a right speaker group symmetrically distributed with the left speaker group on the right side of the vehicle, and a full-frequency speaker at the middle position in front of the left speaker group and the right speaker group and at an equal distance from the left speaker group and the right speaker; The audio signal processing system is connected to a control platform, and the control platform includes a compensation filter matrix generation module, and the compensation filter matrix generation module includes: A delay duration determination module, configured to determine the delay duration of each speaker in the speaker group through a first test signal; An impulse response determination module, configured to determine a first binaural impulse response of the left speaker group, a second binaural impulse response of the right speaker group, and a third binaural impulse response of the full-frequency speaker through a second test signal and the delay duration; A stability processing module, configured to perform stability processing on the first binaural impulse response, the second binaural impulse response, and the third binaural impulse response respectively, to obtain corresponding first stable binaural impulse response, second stable binaural impulse response, and third stable binaural impulse response; A matrix generation module, configured to generate a compensation filter matrix according to the first stable binaural impulse response, the second stable binaural impulse response, and the third stable binaural impulse response.

11. An audio signal processing system, characterized in that, The audio signal processing system includes: At least one processor and a memory communicatively connected to the at least one processor; Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor is configured to execute the method according to any one of claims 1-9.

12. A computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the method according to any one of claims 1-9 when executed by a processor.

Citation Information

Patent Citations

  • Vehicle-mounted sound playback signal delay method adaptive to listening center position

    CN113115199A

  • Apparatus and method of reproducing virtual sound

    CN1630434A