Signal processing method and device, electronic equipment and storage medium

By selecting the target audio signal in the headphones to determine the first downmixing matrix and performing symmetrical processing, the second downmixing matrix adapted to the other ear is directly obtained. This solves the problem of large computational load for headphone stereo signals, improves spatial perception, and reduces the in-head effect.

CN121815155APending Publication Date: 2026-04-07UNISOC CHONGQING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies require a large amount of computation when processing stereo signals in headphones, which can cause the in-head effect when the headphones are too close to the ears, affecting the user experience.

Method used

By selecting the audio signal from the left or right ear as the target, a first downmixing matrix for channel adaptation is determined, and a second downmixing matrix for adapting the other ear is obtained directly through symmetric processing, reducing the repeated matrix solving operations for the downmixing matrices of the left and right ears.

Benefits of technology

It reduces the computational load of signal processing, improves the spatial perception of the headphone stereo signal, reduces the in-head effect, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815155A_ABST
    Figure CN121815155A_ABST
Patent Text Reader

Abstract

The invention provides a signal processing method and device, electronic equipment and a storage medium, and relates to the technical field of signal processing. The method comprises the following steps: receiving an audio signal; determining a first down-mixing matrix of the audio signal; the first down-mixing matrix is a single-channel matrix for performing channel adaptation on the target audio signal; the target audio signal is any one of a left ear audio signal and a right ear audio signal; performing symmetric processing on the first down-mixed matrix to obtain a second down-mixed matrix, the second down-mixed matrix being a down-mixed matrix of the left ear audio signal or a down-mixed matrix of the right ear audio signal; and generating a binaural output signal according to the first downmix matrix, the second downmix matrix and the audio signal. According to the scheme, the calculation amount of signal processing is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing technology, and in particular to a signal processing method, apparatus, electronic device and storage medium. Background Technology

[0002] When users listen to standard stereo recordings with headphones, the close proximity of the headphones to the ears results in a shallow vertical sound field, easily creating a "head-in-the-head effect," where the sound seems to be playing inside the head. In some scenarios, this can enhance the spatial feel of the sound in the headphones.

[0003] In related technologies, headphones can enhance the spatial sound of audio signals by processing them using channel-based spatial audio technology. Specifically, the headphones can perform spherical harmonic transformation on the input audio signal to obtain a virtual channel signal in the spherical harmonic domain. Then, the virtual channel signal is convolved and superimposed with the Head-Related Impulse Response (HRIR) to obtain the output signal. However, processing the input audio signal in this way involves a large amount of computation. Summary of the Invention

[0004] This application provides a signal processing method, apparatus, electronic device, and storage medium to reduce the computational load of signal processing.

[0005] In a first aspect, this application provides a signal processing method, the method comprising:

[0006] Receive audio signals;

[0007] Determine the first downmixing matrix of the audio signal; the first downmixing matrix is ​​a single-channel matrix for channel adaptation of the target audio signal; the target audio signal is either the left ear audio signal or the right ear audio signal;

[0008] The first downmixing matrix is ​​symmetrically processed to obtain the second downmixing matrix, which is either the downmixing matrix of the left ear audio signal or the downmixing matrix of the right ear audio signal.

[0009] Based on the first submixing matrix, the second submixing matrix, and the audio signal, a binaural output signal is generated.

[0010] In one possible implementation, determining the first downmixing matrix of the audio signal includes:

[0011] Acquire at least one impulse response signal, each of which corresponds to a different virtual channel;

[0012] At least one impulse response signal is subjected to dimensionality reduction processing to obtain at least one target impulse signal;

[0013] The first downmixing matrix is ​​determined based on at least one target pulse signal.

[0014] In one possible implementation, at least one impulse response signal is subjected to dimensionality reduction processing to obtain at least one target impulse signal, including:

[0015] Determine at least one pulse signal pair from at least one target pulse signal, wherein each pulse signal pair includes a left channel pulse signal and a right channel pulse signal;

[0016] For each pulse signal pair in at least one pulse signal pair, the time delay difference of the pulse signal pair is determined based on the time delay of the left channel pulse signal and the time delay of the right channel pulse signal;

[0017] At least one target pulse signal is determined based on the time delay difference between at least one pulse signal pair and the frequency of at least one pulse signal pair.

[0018] In one possible implementation, determining at least one target pulse signal based on the time delay difference between at least one pulse signal pair and the frequency of each of the at least one pulse signal includes:

[0019] For each pulse signal in at least one pulse signal, perform the following operation:

[0020] Based on the frequency of the pulse signal, determine the low-frequency signal and the high-frequency signal of the pulse signal;

[0021] Based on the frequency of the pulse signal and the time delay difference of the corresponding pulse signal pairs, the high-frequency signal of the pulse signal is filtered to obtain the filtered signal.

[0022] The target pulse signal corresponding to the pulse signal includes the low-frequency signal of the pulse signal and the filtered signal.

[0023] In one possible implementation, determining the first downmixing matrix based on at least one target pulse signal includes:

[0024] Based on the spatial orientation of each of the at least one target pulse signal, determine the spherical harmonic coefficients corresponding to each of the at least one target pulse signal;

[0025] The first undermixing matrix is ​​determined based on at least one target pulse signal and the spherical harmonic coefficients corresponding to each target pulse signal.

[0026] In one possible implementation, generating binaural output signals based on a first downmixing matrix, a second downmixing matrix, and an audio signal includes:

[0027] The audio signal is processed by spherical harmonic domain conversion to obtain an analog audio signal;

[0028] Based on the first downmixing matrix, the second downmixing matrix, and the analog audio signal, generate the left ear channel signal and the right ear channel signal;

[0029] The signals from the left and right ear channels are processed to generate binaural output signals.

[0030] In one possible implementation, the left ear channel signal and the right ear channel signal are processed to generate a binaural output signal, including:

[0031] In response to the selection operation of the spatial type, a signal generation network is determined; wherein the generation network is used to generate a signal with spatial reverberation corresponding to the spatial type;

[0032] The left ear channel signal is input into the signal generation network to obtain the left ear output signal;

[0033] The right ear channel signal is input into the signal generation network to obtain the right ear output signal;

[0034] The binaural output signal includes the left ear output signal and the right ear output signal.

[0035] Secondly, this application provides a data processing apparatus, comprising:

[0036] The receiving module is used to receive audio signals;

[0037] The determination module is used to determine the first downmixing matrix of the audio signal; the first downmixing matrix is ​​a single-channel matrix for channel adaptation of the target audio signal; the target audio signal is either the left ear audio signal or the right ear audio signal.

[0038] The processing module is used to perform symmetrical processing on the first downmixing matrix to obtain the second downmixing matrix, which is either the downmixing matrix of the left ear audio signal or the downmixing matrix of the right ear audio signal.

[0039] The generation module is used to generate binaural output signals based on the first downmixing matrix, the second downmixing matrix, and the audio signal.

[0040] In one possible implementation, the determining module is specifically used for:

[0041] Acquire at least one impulse response signal, each of which corresponds to a different virtual channel;

[0042] At least one impulse response signal is subjected to dimensionality reduction processing to obtain at least one target impulse signal;

[0043] The first downmixing matrix is ​​determined based on at least one target pulse signal.

[0044] In one possible implementation, the determining module is specifically used for:

[0045] Determine at least one pulse signal pair from at least one target pulse signal, wherein each pulse signal pair includes a left channel pulse signal and a right channel pulse signal;

[0046] For each pulse signal pair in at least one pulse signal pair, the time delay difference of the pulse signal pair is determined based on the time delay of the left channel pulse signal and the time delay of the right channel pulse signal;

[0047] At least one target pulse signal is determined based on the time delay difference between at least one pulse signal pair and the frequency of at least one pulse signal pair.

[0048] In one possible implementation, the determining module is specifically used for:

[0049] For each pulse signal in at least one pulse signal, perform the following operation:

[0050] Based on the frequency of the pulse signal, determine the low-frequency signal and the high-frequency signal of the pulse signal;

[0051] Based on the frequency of the pulse signal and the time delay difference of the corresponding pulse signal pairs, the high-frequency signal of the pulse signal is filtered to obtain the filtered signal.

[0052] The target pulse signal corresponding to the pulse signal includes the low-frequency signal of the pulse signal and the filtered signal.

[0053] In one possible implementation, the determining module is specifically used for:

[0054] Based on the spatial orientation of each of the at least one target pulse signal, determine the spherical harmonic coefficients corresponding to each of the at least one target pulse signal;

[0055] The first undermixing matrix is ​​determined based on at least one target pulse signal and the spherical harmonic coefficients corresponding to each target pulse signal.

[0056] In one possible implementation, the generation module is specifically used for:

[0057] The audio signal is processed by spherical harmonic domain conversion to obtain an analog audio signal;

[0058] Based on the first downmixing matrix, the second downmixing matrix, and the analog audio signal, generate the left ear channel signal and the right ear channel signal;

[0059] The signals from the left and right ear channels are processed to generate binaural output signals.

[0060] In one possible implementation, the generation module is specifically used for:

[0061] In response to the selection operation of the spatial type, a signal generation network is determined; wherein the generation network is used to generate a signal with spatial reverberation corresponding to the spatial type;

[0062] The left ear channel signal is input into the signal generation network to obtain the left ear output signal;

[0063] The right ear channel signal is input into the signal generation network to obtain the right ear output signal;

[0064] The binaural output signal includes the left ear output signal and the right ear output signal.

[0065] Thirdly, this application provides an electronic device, comprising:

[0066] At least one processor; and

[0067] A memory that is communicatively connected to at least one processor; wherein,

[0068] The memory stores instructions that can be executed by at least one processor to cause the at least one processor to perform the methods involved in the first aspect and any possible implementation.

[0069] Fourthly, this application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods involved in the first aspect and any possible implementation.

[0070] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods involved in the first aspect and any possible implementation.

[0071] In a sixth aspect, this application provides a chip including at least one processor for executing program instructions to perform the methods involved in the first aspect and any possible implementation.

[0072] The signal processing method, apparatus, electronic device, and storage medium provided in this application select either the left ear audio signal or the right ear audio signal as the target audio signal, determine a first downmixing matrix for channel adaptation, and then directly obtain a second downmixing matrix for the other ear through symmetric processing. This eliminates the need to derive complete downmixing matrices for the left and right ears independently, saving the repetitive matrix solving operations when determining the downmixing matrices for each ear, thereby reducing the overall computational load of signal processing. Attached Figure Description

[0073] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0074] Figure 1 This is a schematic diagram of the system architecture provided for an embodiment of this application;

[0075] Figure 2 A schematic flowchart of a signal processing method provided in an embodiment of this application;

[0076] Figure 3 A flowchart illustrating the process of determining a first undermixing matrix provided in an embodiment of this application;

[0077] Figure 4 A schematic diagram of a process for generating binaural output signals is provided for an embodiment of this application;

[0078] Figure 5 This is a schematic diagram of the structure of a signal generation network provided in an embodiment of this application;

[0079] Figure 6 A schematic flowchart illustrating another signal processing method provided in an embodiment of this application;

[0080] Figure 7 This application provides a schematic diagram of the structure of a signal processing device according to an embodiment of the present application;

[0081] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0082] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0083] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0084] The collection, storage, use, processing, transmission, provision, and disclosure of financial data or user data involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0085] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0086] When users listen to live-recorded stereo sound using headphones, the close proximity of the headphones to the ears results in a shallower vertical sound field, creating the "in-head effect." This means the user perceives the stereo sound as originating inside their head, leading to a poor listening experience. While some recording methods, such as head recording, can be used, a significant portion of stereo sound is still recorded using standard techniques.

[0087] For audio recorded using ordinary recording methods, the audio signal can be processed to increase the user's sense of space when hearing the sound, thereby improving the user experience.

[0088] In related technologies, headphones can process input audio signals using spatial audio technology to obtain output audio signals with a sense of space. Spatial audio technology can include channel-based audio technology, object-based audio technology, and scene-based audio technology.

[0089] For channel-based audio technology, headphones can perform spherical harmonic domain (SHV) conversion on the input audio signal to obtain virtual channel signals in the SHV domain. These virtual channel signals can be in a three-dimensional spatial audio coding format (ambisonic format). Then, the headphones convolve and superimpose the virtual channel signals with the corresponding HRIR signals to obtain the output audio signal.

[0090] Furthermore, the HRIR library contains a large amount of impulse response data from different spatial orientations. The convolution operation between the virtual channel signal and the HRIR signal involves numerous iterative calculations, and the HRIR signal is typically quite long, further increasing the computational complexity of the convolution operation. Therefore, processing audio signals using the above method involves a significant amount of computation.

[0091] Based on this, this application provides a signal processing method that selects either the left ear audio signal or the right ear audio signal as the target audio signal, determines a first downmixing matrix for channel adaptation, and then directly obtains a second downmixing matrix for the other ear through symmetric processing. This eliminates the need to derive complete downmixing matrices for the left and right ears independently, saving the repetitive matrix solving operations when determining the downmixing matrices for each ear, thereby reducing the overall computational load of signal processing.

[0092] To facilitate understanding, the following will be combined with... Figure 1The system architecture applicable to the embodiments of this application will be described.

[0093] Figure 1 This is a schematic diagram of the system architecture provided for an embodiment of this application. Figure 1 As shown, it includes a terminal device 11 and an earphone 12, wherein the terminal device 11 stores the audio signal of the live recording.

[0094] In practical applications, terminal device 11 can establish a communication connection with headset 12. Terminal device 11 and headset 12 can transmit data through this communication connection. For example, terminal device 11 can send audio signals to headset 12 via the communication connection. After receiving the audio signal, headset 12 can process the audio signal, obtain an output signal, and output the output signal.

[0095] It should be noted that, Figure 1 This is merely an example to illustrate a system architecture diagram, and is not a limitation on system architecture diagrams.

[0096] It should be noted that the execution subject in each embodiment of this application can be a chip, chip module, processor, microprocessor, etc., or it can be a device integrating the above-mentioned chips, chip modules, processors, or microprocessors, such as a server. The specific execution subject in each embodiment of this application is not limited, and it can be selected and set according to actual needs. In the following embodiments, a server integrating the above-mentioned chips, chip modules, processors, or microprocessors is used as an example for description, which does not constitute a limitation on the actual execution subject.

[0097] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0098] Figure 2 This is a schematic flowchart illustrating a signal processing method provided in an embodiment of this application. Figure 2 As shown, the method may include the following steps:

[0099] S201, Receive audio signals.

[0100] Audio signals are electrical or digital signals that carry sound information and can be received, processed, transmitted, and played back by electronic devices (such as headphones, mobile phones, and audio decoders). They serve as the conversion medium between physical sound waves and data that can be recognized by electronic devices. Audio signals can be dual-channel or multi-channel signals.

[0101] In some embodiments, the headphones and the terminal device establish a communication connection via wired or wireless means. The terminal device can then send audio signals to the headphones through this communication connection.

[0102] S202. Determine the first downmixing matrix of the audio signal; the first downmixing matrix is ​​a single-channel matrix for channel adaptation of the target audio signal; the target audio signal is either the left ear audio signal or the right ear audio signal.

[0103] The left-ear audio signal is specifically designed for the left ear's sensory channel. It carries sound information about the intensity, phase, and frequency of sounds from a specific location, tailored for the left ear. The right-ear audio signal is specifically designed for the right ear's sensory channel. It carries sound information about the intensity, phase, and frequency of sounds from a specific location, tailored for the right ear.

[0104] In some embodiments, either the left ear audio signal or the right ear audio signal can be determined as the target audio signal.

[0105] The first downmixing matrix is ​​designed for the target audio signal and is used to achieve "dimensionality reduction and channel adaptation of high-dimensional signals". The core function of the first downmixing matrix is ​​to establish the mapping relationship between the target audio signal and a single output channel.

[0106] In some embodiments, an HRIR database may be acquired, which includes at least one HRIR signal. The headphones may process the at least one HRIR signal using a preset algorithm and, in conjunction with the characteristics and spatial orientation information of the target audio signal, determine a first downmixing matrix for the communication adaptation requirements of the target audio signal.

[0107] S203. Perform symmetrical processing on the first downmixing matrix to obtain the second downmixing matrix, which is either the downmixing matrix of the left ear audio signal or the downmixing matrix of the right ear audio signal.

[0108] When the target audio signal is the left ear audio signal, the second downmixing matrix is ​​the downmixing matrix of the right ear audio signal; when the target audio signal is the right ear audio signal, the second downmixing matrix is ​​the downmixing matrix of the left ear audio signal.

[0109] Taking the second downmixing matrix as an example, which is the downmixing matrix for the left ear audio signal, the second downmixing matrix is ​​designed for the left ear audio signal and is used to achieve "dimensionality reduction of high-dimensional signals and channel adaptation". The core function of the second downmixing matrix is ​​to establish the mapping relationship between the left ear audio signal and the single output channel.

[0110] In some embodiments, with the head as the central origin, the azimuth angles of the left and right ears differ by 180 degrees, and the pitch angles of both are 0 degrees. Assuming perfect symmetry in the head, and with the subharmonic number m of the spherical harmonic function less than 0, the parameters of the first and second lower mixing matrices are opposites of each other. Therefore, the first lower mixing matrix can be symmetrically processed; that is, for the parameters in the first lower mixing matrix where the subharmonic m is less than 0, the parameters are updated to their opposites to obtain the second lower mixing matrix.

[0111] For example, when the subharmonic wavenumber m of the spherical harmonic function is less than 0, assume the transpose of the first undermixing matrix. By performing symmetric processing on the first lower mixing matrix, the second lower mixing matrix is ​​obtained as follows: Where NA is the pre-defined order of the spherical harmonic function, representing the overall resolution of the spherical harmonic domain in representing three-dimensional space. n is the sub-order of the spherical harmonic function. For example, The core weights used to indicate the mapping of spherical harmonic spatial features with sub-order and sub-harmonic number of 0 to a single-channel audio signal are directly determined by their numerical value and sign, which in turn determine the degree of contribution of such high-resolution spatial features to the target channel output signal.

[0112] In some embodiments, taking the target audio signal as the left ear audio signal as an example, considering the memory and computational requirements of binaural rendering, a first downmixing matrix corresponding to the left ear can be calculated, and the left ear output signal can be determined based on the first downmixing matrix and the left ear audio signal. When determining the right ear output signal, a second downmixing matrix is ​​determined by performing symmetrical processing on the first downmixing matrix, and then the right ear output signal is determined based on the second downmixing matrix and the right ear audio signal. This saves computational resources and memory required to calculate the second downmixing matrix, thereby reducing the computational load of signal processing.

[0113] S204. Generate binaural output signals based on the first downmixing matrix, the second downmixing matrix, and the audio signal.

[0114] Binaural output signal is the output signal with spatial quality obtained by the headphones after processing the audio signal. Binaural output signal can include left ear output signal and right ear output signal.

[0115] Taking the target audio signal as the left ear audio signal as an example, the left ear output signal refers to the output signal adapted to the left ear obtained by processing the audio signal through the first downmixing matrix; the right ear output signal refers to the output signal adapted to the right ear obtained by processing the right ear audio signal through the second downmixing matrix.

[0116] exist Figure 2In the embodiment shown, either the left ear audio signal or the right ear audio signal is selected as the target audio signal, a first downmixing matrix for channel adaptation is determined, and then a second downmixing matrix for the other ear is directly obtained through symmetric processing. There is no need to derive complete downmixing matrices for the left and right ears separately, which saves the repeated matrix solving operations when determining the downmixing matrices for the left and right ears, thereby reducing the overall computational load of signal processing.

[0117] exist Figure 2 Based on the illustrated embodiment, the following is combined with Figure 3 The method for determining the first downmixing matrix of the audio signal in the embodiments of this application will be further explained.

[0118] Figure 3 This is a schematic diagram illustrating a process for determining a first undermixing matrix, provided as an embodiment of this application. Figure 3 As shown, the process may include the following steps:

[0119] S301. Obtain at least one impulse response signal, wherein each of the at least one impulse response signal corresponds to a different virtual channel.

[0120] At least one impulse response signal is a set of time-series signals that record the instantaneous impulse responses at one or more specific locations in space, and which, after reflection, diffraction, and attenuation by physiological structures such as the human head, auricle, and tragus, reach the left and right ears respectively. Each location corresponds to one impulse response signal for the left ear and one for the right ear. At least one impulse response signal can be, for example, an HRIR signal.

[0121] In some embodiments, the headphones store an HRIR database, which includes at least one impulse response signal. The headphones can retrieve at least one impulse response signal from the HRIR database.

[0122] S302. Perform dimensionality reduction processing on at least one impulse response signal to obtain at least one target impulse signal.

[0123] In some embodiments, the method for performing dimensionality reduction processing on at least one pulse signal to obtain at least one target pulse signal can be as follows: determining at least one pulse signal pair in the at least one target pulse signal, each pulse signal pair including a left channel pulse signal and a right channel pulse signal; for each pulse signal pair in the at least one pulse signal pair, determining the time delay difference of the pulse signal pair based on the time delay of the left channel pulse signal and the time delay of the right channel pulse signal; determining at least one target pulse signal based on the time delay difference of each of the at least one pulse signal pair and the frequency of each of the at least one pulse signal.

[0124] The pulse signal pair includes two pulse response signals, which are the pulse response signals of the left ear and the right ear, respectively, located in the same direction.

[0125] For each pulse signal pair in at least one pulse signal pair, the left channel pulse signal in the pulse signal pair is the pulse response signal of the left ear in the corresponding orientation. The right channel pulse signal in the pulse signal pair is the pulse response signal of the right ear in the corresponding orientation.

[0126] In some embodiments, for any one of the at least one pulse signals, the orientation and channel type corresponding to the pulse signal can be determined, and the corresponding opposite pulse signal can be determined from the at least one pulse signal based on the orientation of the pulse signal. Then, the pulse signal and the corresponding opposite pulse signal are determined as a pulse signal pair.

[0127] Meanwhile, if the channel type of the pulse signal is left channel, the pulse signal is identified as the left channel pulse signal in the pulse signal pair, and the corresponding pulse signal on the opposite side is identified as the right channel pulse signal in the pulse signal pair.

[0128] If the channel type of the pulse signal is the right channel, the pulse signal is identified as the right channel pulse signal in the pulse signal pair, and the corresponding pulse signal on the opposite side is identified as the left channel pulse signal in the pulse signal pair.

[0129] In some embodiments, for at least one pulse signal pair, the delay of the left channel pulse signal and the delay of the right channel pulse signal in the pulse signal pair can be determined. Then, the delay of the left channel pulse signal and the delay of the right channel pulse signal are processed by calculating the correlation between the left and right channels (Impulse Response, IR) to determine the delay difference of the pulse signal pair.

[0130] In some embodiments, the method for determining at least one target pulse signal based on the time delay difference between at least one pulse signal pair and the frequency of at least one pulse signal can be as follows: For each pulse signal among the at least one pulse signal, the following operations are performed: determining the low-frequency signal and the high-frequency signal of the pulse signal based on the frequency of the pulse signal; filtering the high-frequency signal of the pulse signal based on the frequency of the pulse signal and the time delay difference between the corresponding pulse signal pair to obtain a filtered signal; wherein, the target pulse signal corresponding to the pulse signal includes the low-frequency signal of the pulse signal and the filtered signal.

[0131] In some embodiments, the low-frequency signal and high-frequency signal of the pulse signal can be determined according to the frequency of the pulse signal as follows: the signal with a frequency less than a preset frequency in the pulse signal frequency is determined as the frequency signal of the pulse signal; the signal with a frequency greater than or equal to the preset frequency in the pulse signal frequency is determined as the high-frequency signal of the pulse signal.

[0132] Then, the high-frequency signal of the pulse signal can be filtered to align with the binaural time delay of the high-frequency signal. In this way, the interaural time difference (ITD) information in the high-frequency signal can be removed, while retaining the feature information in the high-frequency signal related to spatial orientation perception.

[0133] Specifically, the high-frequency signal H of the pulse signal is filtered to obtain the filtered signal. The method can be as shown in formula (1):

[0134] (1)

[0135] Where A is the filter function, and A = . Here, w is the frequency of the high-frequency pulse signal, which is a complex exponential term. For preset frequency, This represents the time delay difference between the corresponding pulse signal pairs.

[0136] In some embodiments, after determining the filtered signal, the high-frequency signal of the pulse signal can be replaced with the filtered signal to obtain the target pulse signal corresponding to the pulse signal. Therefore, the target pulse signal corresponding to the pulse signal includes the low-frequency signal of the pulse signal and the filtered signal.

[0137] S303. Determine the first downmixing matrix based on at least one target pulse signal.

[0138] In some embodiments, the method for determining the first downmixing matrix based on at least one target pulse signal may be as follows: determining the spherical harmonic coefficients corresponding to each of the at least one target pulse signal based on their respective spatial orientations; and determining the first downmixing matrix based on the at least one target pulse signal and the spherical harmonic coefficients corresponding to each of the at least one target pulse signal.

[0139] Spherical harmonic coefficients are real-valued coefficients derived from spherical harmonic functions. The spherical harmonic coefficients of each target pulse signal can be uniquely determined by its corresponding spatial orientation. The essence of spherical harmonic coefficients is to convert the three-dimensional spatial orientation characteristics of the target pulse signal into quantized index values.

[0140] In some embodiments, for each target pulse signal in at least one target pulse signal, the spherical harmonic coefficients of the target pulse signal can be obtained by calculating the trigonometric function corresponding to the azimuth angle and the sine value corresponding to the elevation angle of the target pulse signal along with the Legendre function.

[0141] In some embodiments, a first undermixing matrix can be determined by processing at least one target pulse signal and its corresponding spherical harmonic coefficients using the least squares method. Specifically, the first undermixing matrix is ​​determined. The method can be as shown in formula (2):

[0142] (2)

[0143] Determined by formula (2) The optimal solution can be shown below:

[0144]

[0145] in, It is used to solve The spherical harmonic coefficient matrix is ​​essentially a matrix formed by concatenating the column-wise spherical harmonic coefficient vectors corresponding to P target pulse signals. P is the total number of at least one target pulse signal. W is a frequency-independent diagonal weighting matrix with orthogonal weights. This is the transpose of the first undermixing matrix. For at least one target pulse signal, It is the spatial orientation corresponding to a single target pulse signal. The spherical harmonic coefficient vector has a dimension of (NA+1). 2 ×1, where each element in the vector corresponds to a set of coefficient values ​​for a spherical harmonic function (defined by the spherical harmonic order NA and the subharmonic wave number m). These are the three-dimensional spatial orientation parameters corresponding to a single target pulse signal. It is the angular frequency of the audio signal. It is the square of the Frobenius norm. Used to quantify the magnitude of the error between the "mapping result" and at least one target pulse signal (the smaller the norm, the smaller the error).

[0146] yes transpose, It is the target pulse signal. It is the transpose of the target pulse signal. This is the transpose of the first undermixed matrix.

[0147] exist Figure 3In the illustrated embodiment, by acquiring the impulse response signals corresponding to different virtual channels and performing targeted dimensionality reduction processing, redundant signal data is eliminated, reducing the computational load of subsequent signal processing. Furthermore, by determining the time delay difference of the impulse signal pairs and processing them according to frequency and scene (directly retaining key positioning information for low frequencies and filtering and aligning high frequencies based on time delay difference), the core features of the entire frequency band related to spatial orientation perception (such as low-frequency ITD and high-frequency interaural level difference (ILD)) in the impulse response signals are preserved, thus retaining key information affecting hearing in the high-frequency signals of the impulse signals. Simultaneously, the spherical harmonic coefficients are determined based on the spatial orientation of each target impulse signal, and then the first downmixing matrix is ​​solved in conjunction with the target impulse signal. This ensures that the downmixing matrix can accurately establish the mapping relationship between high-dimensional spatial features and single-channel output, maximizing the restoration of realistic sound field information while achieving channel adaptation.

[0148] Based on the above embodiments, the following is combined with Figure 4 The method of generating binaural output signals in the embodiments of this application will be further explained.

[0149] Figure 4 This is a schematic diagram illustrating a process for generating binaural output signals, provided as an embodiment of this application. Figure 4 As shown, the process may include the following steps:

[0150] S401. Perform spherical harmonic domain conversion on the audio signal to obtain an analog audio signal.

[0151] The analog audio signal is a virtual audio signal obtained by performing spherical harmonic domain conversion on the left ear audio signal.

[0152] In some embodiments, the audio signal can be processed by spherical harmonic domain conversion using spatial audio coding rules, based on the standard configuration of the recorded audio signal (e.g., stereo dual channels are generally located at a position of ±30 degrees in front), to obtain the analog audio signal.

[0153] Specifically, the audio signal is subjected to spherical harmonic domain conversion to obtain an analog audio signal. The method can be as shown in formula (3):

[0154] (3)

[0155] in, The signal is an audio signal, and w is the frequency of the audio signal. It is the spatial orientation corresponding to a single target pulse signal. The vector of spherical harmonic coefficients. It is the three-dimensional spatial orientation parameter corresponding to a single target pulse signal, and , These are real-valued spherical harmonic coefficients. .

[0156] in, It is associated with the Legendre function. It is the azimuth angle. The range of values ​​for is [-π, π]. Angle of elevation The range of values ​​for is [-π / 2, π / 2]. for The cosine of , m is the second harmonic number of the spherical harmonic function, NA is the preset order of the left ear analog signal in the spherical harmonic domain, and n is the sub-order of the left ear analog signal in the spherical harmonic domain.

[0157] S402. Generate the left ear channel signal and the right ear channel signal based on the first downmixing matrix, the second downmixing matrix and the analog audio signal.

[0158] The left ear channel signal is the audio signal of the left ear output channel of the compatible headphones, and the right ear channel signal is the audio signal of the right ear output channel of the compatible headphones.

[0159] Taking the target audio signal as the left ear audio signal and determining the left ear channel signal as an example, the left ear channel signal can be determined based on the first downmixing matrix and the analog audio signal, and the right ear channel signal can be determined based on the second downmixing matrix and the analog audio signal.

[0160] In some embodiments, the left ear channel signal X can be determined based on the first downmixing matrix and the analog audio signal as shown in formula (4):

[0161] (4)

[0162] in, Let a(w) be the transpose of the first undermixing matrix, and let a(w) be the analog signal of the left ear.

[0163] It should be noted that the method for determining the right ear channel signal is similar to the method for determining the left ear channel signal, and will not be repeated here.

[0164] S403. Process the left and right ear channel signals to generate binaural output signals.

[0165] In some embodiments, the binaural output signal can be generated as follows: in response to a spatial type selection operation, a signal generation network is determined; wherein the generation network is used to generate a signal with spatial reverberation corresponding to the spatial type; the left ear channel signal is input to the signal generation network to obtain the left ear output signal; the right ear channel signal is input to the signal generation network to obtain the right ear output signal; wherein the binaural output signal includes the left ear output signal and the right ear output signal.

[0166] A signal generation network is a signal processing model specifically designed to simulate spatial reverberation effects. For example, a signal generation network can be a feedback delay network (FDN).

[0167] It should be noted that an FDN network includes at least one reverberation line, a feedback matrix, and a delay unit. To better simulate actual spatial reverberation, attenuation filters can be added after each reverberation line. These attenuation filters can be composed of equalizers (EQ).

[0168] In some embodiments, the frequency-dependent reverberation time curve T can be determined based on the actual room impulse response (RIR). 60 (w). Specifically, T can be determined based on empirical values. 60 (w).

[0169] In some embodiments, the amplitude-frequency response of the attenuation filter can be determined as follows: the amplitude-frequency response of the attenuation filter to be approximated can be calculated based on the reverberation time, wherein the relationship between the reverberation time and the amplitude-frequency response of the attenuation filter can be as shown in formula (5):

[0170] (5)

[0171] in, It is the frequency domain complex-valued transfer function of the feedback path in the signal generation network. It is the frequency domain transfer function of the feedback path. The amplitude, M is the number of delay sampling points in the delay unit of the signal generation network, fs is the sampling frequency, and T is the amplitude. 60 (w) represents the reverberation time curve.

[0172] In some embodiments, human hearing can be divided into different levels, and the center frequency of the attenuation filter can be determined based on the mapping relationship between human hearing levels and center frequencies.

[0173] Taking the signal generation network as an example, we can combine... Figure 5 To understand, Figure 5 This is a schematic diagram of a signal generation network provided in an embodiment of this application. Figure 5 As shown, the FDN network is deployed with a pre-adder, a delay unit, a feedback matrix, an attenuation filter, and an output adder.

[0174] The pre-adder is a signal fusion node in the FDN network, located between the input and each delay unit. Its function is to linearly superimpose the input signal with the feedback signal returned by the feedback matrix to generate the input signal for each delay unit.

[0175] The delay unit is the core component of the FDN network for simulating reverberation time delay characteristics. Different delay units are configured with different delay lengths in order to simulate the dispersed time delay of sound reflected through different paths in real space (such as the multipath delay of early reflected sound), so that the reverberation signal generated later is more in line with the acoustic timing characteristics of real space, and avoids the harsh and distorted reverberation effect caused by a single delay.

[0176] The feedback matrix is ​​the core logic module for implementing multipath cyclic reflection in FDN networks, and is typically a square matrix matching the number of delay units. The function of the feedback matrix is ​​to receive the signals output from each delay unit, perform weighted and mixed operations on these signals, and then feed the result back to each pre-adder, forming a cyclic iterative link of signals. This multipath feedback mechanism can simulate the complex superposition effect of sound after multiple reflections from walls and objects in real space, generating dense and natural post-reverberation.

[0177] The attenuation filter is a component in FDN networks that simulates the reverberation energy attenuation characteristics. It can perform frequency-dependent or amplitude-dependent attenuation processing on the signal output from the delay unit. Based on the fact that sound weakens with each reflection in a real space due to energy loss (the degree of attenuation varies at different frequencies), the attenuation filter controls the attenuation amplitude of each delayed signal, ensuring the reverberation effect matches the acoustic parameters of the target space type (such as a concert hall or conference room).

[0178] The output adder is the signal aggregation node of the FDN network. It first performs preliminary superposition of the multipath reverberation signals processed by each attenuation unit, and then combines them with other link signals (if any) to generate the final output signal. The output adder can fuse the multipath signals after delay, feedback, and attenuation into a complete reverberation processed signal, ensuring that the output signal contains the acoustic information of all reflection paths, and ultimately achieves a natural reverberation effect that conforms to the target space type.

[0179] In some embodiments, users can select a space type through an interactive interface. The space type refers to a scene category with specific acoustic reflection characteristics. For example, the space type can be a small room, a large room, etc.

[0180] In response to the selection of a spatial type, the signal generation network corresponding to the spatial type selected by the user is determined based on the mapping relationship between the spatial type and the signal generation network, as well as the spatial type selected by the user.

[0181] Then, the left ear channel signal is input to the signal generation network. The signal generation network processes the input left ear channel signal and outputs the left ear output signal. The right ear channel signal is then input to the signal generation network. The signal generation network processes the input right ear channel signal and outputs the right ear output signal, thus obtaining the binaural output signals.

[0182] exist Figure 4 In the illustrated embodiment, the second submixing matrix is ​​obtained by symmetrically processing the first submixing matrix, eliminating the need to independently solve for the complete submixing matrix for each ear. This significantly simplifies the matrix calculation process and reduces the overall computational load of signal processing. Furthermore, the signal processing method provided in this application supports response space type selection operations and generates corresponding spatial reverberation output signals through a signal generation network, meeting users' immersive listening needs for different scenarios and further enhancing the realism and presence of the sound field while preserving accurate spatial positioning.

[0183] Based on the above embodiments, the following is combined with Figure 6 The signal processing method provided in this application will be further described.

[0184] Figure 6 This is a schematic flowchart illustrating another signal processing method provided in an embodiment of this application. Figure 6 As shown, a signal processing method may include the following steps:

[0185] S601, Receives audio signals.

[0186] For an introduction to receiving audio signals, please refer to [link / reference]. Figure 2 S201 in the illustrated embodiment will not be described in detail here.

[0187] S602. Determine the first and second submixing matrices.

[0188] For an introduction to determining the first lower mixing matrix, please refer to [link / reference needed]. Figure 3 The embodiments shown will not be described in detail here.

[0189] For an introduction to determining the second lower mixing matrix, please refer to [link / reference needed]. Figure 2 S203 in the illustrated embodiment will not be described in detail here.

[0190] S603. Perform spherical harmonic domain conversion on the audio signal to obtain an analog audio signal.

[0191] For an introduction to obtaining analog audio signals, please refer to [link / reference]. Figure 4 S401 in the illustrated embodiment will not be described in detail here.

[0192] S604. Generate the left ear channel signal and the right ear channel signal based on the first downmixing matrix, the second downmixing matrix and the analog audio signal.

[0193] For an introduction to generating left and right ear channel signals, please refer to [link / reference needed]. Figure 4 S402 in the illustrated embodiment will not be described in detail here.

[0194] S605, in response to the selection operation of the space type, determine the signal generation network.

[0195] For an introduction to determining signal generation networks, please refer to [link / reference]. Figure 4 S403 in the illustrated embodiment will not be described in detail here.

[0196] S606. Input the left ear channel signal into the signal generation network to obtain the left ear output signal.

[0197] For an explanation of obtaining the left ear output signal, please refer to [link / reference]. Figure 4 S403 in the illustrated embodiment will not be described in detail here.

[0198] S607. Input the right ear channel signal into the signal generation network to obtain the right ear output signal.

[0199] For an explanation of obtaining the right ear output signal, please refer to [link / reference needed]. Figure 4 S403 in the illustrated embodiment will not be described in detail here.

[0200] Figure 6 In the illustrated embodiment, the headphones only need to determine either the left-ear audio signal or the right-ear audio signal as the target audio signal, and only calculate the first downmixing matrix of the target audio signal. Then, the output audio signal is determined based on the first downmixing matrix, the left-ear audio signal, and the right-ear audio signal. This eliminates the need to calculate the downmixing matrices of the left-ear and right-ear audio signals separately, reducing the computational load of calculating the downmixing matrix of the other audio signal besides the target audio signal, thereby reducing the computational load in the data processing process.

[0201] Figure 7 This is a schematic diagram of the structure of a signal processing device provided in an embodiment of this application. For example... Figure 7 As shown, the signal processing device 70 includes: a receiving module 71, a determining module 72, a processing module 73, and a generating module 74, wherein...

[0202] Receiver module 71 is used to receive audio signals;

[0203] The determination module 72 is used to determine the first downmixing matrix of the audio signal; the first downmixing matrix is ​​a single-channel matrix for channel adaptation of the target audio signal; the target audio signal is either the left ear audio signal or the right ear audio signal;

[0204] Processing module 73 is used to perform symmetrical processing on the first downmixing matrix to obtain a second downmixing matrix, which is either the downmixing matrix of the left ear audio signal or the downmixing matrix of the right ear audio signal.

[0205] The generation module 74 is used to generate binaural output signals based on the first downmixing matrix, the second downmixing matrix, and the audio signal.

[0206] In one possible implementation, the determining module 72 is specifically used for:

[0207] Acquire at least one impulse response signal, each of which corresponds to a different virtual channel;

[0208] At least one impulse response signal is subjected to dimensionality reduction processing to obtain at least one target impulse signal;

[0209] The first downmixing matrix is ​​determined based on at least one target pulse signal.

[0210] In one possible implementation, the determining module 72 is specifically used for:

[0211] Determine at least one pulse signal pair from at least one target pulse signal, wherein each pulse signal pair includes a left channel pulse signal and a right channel pulse signal;

[0212] For each pulse signal pair in at least one pulse signal pair, the time delay difference of the pulse signal pair is determined based on the time delay of the left channel pulse signal and the time delay of the right channel pulse signal;

[0213] At least one target pulse signal is determined based on the time delay difference between at least one pulse signal pair and the frequency of at least one pulse signal pair.

[0214] In one possible implementation, the determining module 72 is specifically used for:

[0215] For each pulse signal in at least one pulse signal, perform the following operation:

[0216] Based on the frequency of the pulse signal, determine the low-frequency signal and the high-frequency signal of the pulse signal;

[0217] Based on the frequency of the pulse signal and the time delay difference of the corresponding pulse signal pairs, the high-frequency signal of the pulse signal is filtered to obtain the filtered signal.

[0218] The target pulse signal corresponding to the pulse signal includes the low-frequency signal of the pulse signal and the filtered signal.

[0219] In one possible implementation, the determining module 72 is specifically used for:

[0220] Based on the spatial orientation of each of the at least one target pulse signal, determine the spherical harmonic coefficients corresponding to each of the at least one target pulse signal;

[0221] The first undermixing matrix is ​​determined based on at least one target pulse signal and the spherical harmonic coefficients corresponding to each target pulse signal.

[0222] In one possible implementation, the generation module 74 is specifically used for:

[0223] The audio signal is processed by spherical harmonic domain conversion to obtain an analog audio signal;

[0224] Based on the first downmixing matrix, the second downmixing matrix, and the analog audio signal, generate the left ear channel signal and the right ear channel signal;

[0225] The signals from the left and right ear channels are processed to generate binaural output signals.

[0226] In one possible implementation, the generation module 74 is specifically used for:

[0227] In response to the selection operation of the spatial type, a signal generation network is determined; wherein the generation network is used to generate a signal with spatial reverberation corresponding to the spatial type;

[0228] The left ear channel signal is input into the signal generation network to obtain the left ear output signal;

[0229] The right ear channel signal is input into the signal generation network to obtain the right ear output signal;

[0230] The binaural output signal includes the left ear output signal and the right ear output signal.

[0231] The signal processing device 70 provided in this application embodiment can execute the technical solution of the signal processing method in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0232] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 80 includes:

[0233] At least one processor 82; and

[0234] Memory 81 is communicatively connected to at least one processor 82; wherein,

[0235] The memory 81 stores instructions that can be executed by at least one processor 82, which in turn executes the signal processing method described in the above method embodiments.

[0236] Optionally, the aforementioned processor can be a central processing unit (CPU), or it can be a GPU, other general-purpose processors, digital signal processors (DSPs), or application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0237] The electronic device 80 provided in this application embodiment can execute the signal processing method involved in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0238] This application provides a non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, which are used to cause the computer to execute the signal processing method involved in the above method embodiments.

[0239] This application provides a computer program product, including a computer program that, when executed by an electronic device, implements the signal processing method involved in the above method embodiments.

[0240] This application provides a chip, which includes at least one processor. The processor is used to run program instructions to perform the signal processing methods involved in the above method embodiments.

[0241] This application provides a chip module on which a computer program is stored. When the computer program is executed by the chip module, it implements the signal processing method involved in the above method embodiments.

[0242] All or part of the steps in the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.

[0243] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable terminal device to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0244] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0245] These computer program instructions can also be loaded onto a computer or other programmable terminal device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0246] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

[0247] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0248] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A signal processing method, characterized in that, The method includes: Receive audio signals; A first downmixing matrix is ​​determined for the audio signal; the first downmixing matrix is ​​a single-channel matrix for channel adaptation of the target audio signal; the target audio signal is either the left ear audio signal or the right ear audio signal. The first downmixing matrix is ​​symmetrically processed to obtain a second downmixing matrix, which is either the downmixing matrix of the left ear audio signal or the downmixing matrix of the right ear audio signal. Based on the first downmixing matrix, the second downmixing matrix, and the audio signal, a binaural output signal is generated.

2. The method according to claim 1, characterized in that, Determining the first downmixing matrix of the audio signal includes: Acquire at least one impulse response signal, wherein each of the at least one impulse response signal corresponds to a different virtual channel; The at least one impulse response signal is subjected to dimensionality reduction processing to obtain at least one target impulse signal; The first downmixing matrix is ​​determined based on the at least one target pulse signal.

3. The method according to claim 2, characterized in that, The step of performing dimensionality reduction processing on the at least one impulse response signal to obtain at least one target impulse signal includes: At least one pulse signal pair is determined from the at least one target pulse signal, and each pulse signal pair includes a left channel pulse signal and a right channel pulse signal; For each pulse signal pair in the at least one pulse signal pair, the time delay difference of the pulse signal pair is determined based on the time delay of the left channel pulse signal and the time delay of the right channel pulse signal; At least one target pulse signal is determined based on the time delay difference between the at least one pulse signal pair and the frequency of the at least one pulse signal pair.

4. The method according to claim 3, characterized in that, Determining at least one target pulse signal based on the time delay difference between the at least one pulse signal pairs and the frequency of the at least one pulse signal pairs includes: For each of the at least one pulse signal, the following operations are performed: Based on the frequency of the pulse signal, determine the low-frequency signal and the high-frequency signal of the pulse signal; Based on the frequency of the pulse signal and the time delay difference of the corresponding pulse signal pair, the high-frequency signal of the pulse signal is filtered to obtain the filtered signal. The target pulse signal corresponding to the pulse signal includes the low-frequency signal of the pulse signal and the filtered signal.

5. The method according to any one of claims 2-4, characterized in that, Determining the first downmixing matrix based on the at least one target pulse signal includes: Based on the spatial orientation of each of the at least one target pulse signal, determine the spherical harmonic coefficients corresponding to each of the at least one target pulse signal; The first downmixing matrix is ​​determined based on the at least one target pulse signal and the spherical harmonic coefficients corresponding to each of the at least one target pulse signal.

6. The method according to any one of claims 1-3, characterized in that, The step of generating binaural output signals based on the first downmixing matrix, the second downmixing matrix, and the audio signal includes: The audio signal is subjected to spherical harmonic domain conversion to obtain an analog audio signal; Based on the first downmixing matrix, the second downmixing matrix, and the analog audio signal, generate the left ear channel signal and the right ear channel signal; The left ear channel signal and the right ear channel signal are processed to generate the binaural output signal.

7. The method according to claim 6, characterized in that, The process of processing the left ear channel signal and the right ear channel signal to generate the binaural output signal includes: In response to a spatial type selection operation, a signal generation network is determined; wherein the generation network is used to generate a signal with spatial reverberation corresponding to the spatial type; The left ear channel signal is input into the signal generation network to obtain the left ear output signal; The right ear channel signal is input into the signal generation network to obtain the right ear output signal; The binaural output signal includes the left ear output signal and the right ear output signal.

8. A signal processing apparatus, characterized in that, The device includes: The receiving module is used to receive audio signals; The determining module is used to determine a first downmixing matrix of the audio signal; the first downmixing matrix is ​​a single-channel matrix for channel adaptation of the target audio signal; the target audio signal is either a left-ear audio signal or a right-ear audio signal; The processing module is used to perform symmetrical processing on the first downmixing matrix to obtain a second downmixing matrix, wherein the second downmixing matrix is ​​the downmixing matrix of the left ear audio signal or the downmixing matrix of the right ear audio signal; The generation module is used to generate a binaural output signal based on the first downmixing matrix, the second downmixing matrix, and the audio signal.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to cause the at least one processor to perform the method of any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.