Audio signal alignment method and device, computer storage medium and terminal

Through frequency response processing of short-time Fourier transform and time-varying linear prediction error filter, the problem of poor audio signal quality in multi-person conference scenarios is solved, and the accuracy of delay estimation and the alignment effect of audio signals are improved.

CN117275508BActive Publication Date: 2025-10-21GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210671975.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-10-21
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

In multi-person conference scenarios, when using the array microphone of the conference terminal all-in-one machine for long-distance sound pickup, the audio signal quality is poor, and the relative delay between multiple audio signals is not negligible and changes dynamically, making subsequent audio processing difficult.

Method used

The short-time frequency domain representation of the audio signal is determined by short-time Fourier transform, the frequency response of the time-varying linear prediction error filter is calculated, the low-frequency coefficients are extracted and recombined, the short-time spectrum of the linear prediction error envelope is determined, the frame delay estimation result of the audio signal is calculated, and speed-varying but pitch-invariant processing is performed to align the audio signal.

Benefits of technology

The accuracy of delay estimation is improved, the mutual conversion operations between time and frequency domains are reduced, the amount of calculation is reduced, and the alignment effect of audio signals is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117275508B_ABST
    Figure CN117275508B_ABST
Patent Text Reader

Abstract

The application provides an audio signal alignment method and device, a medium and a terminal, and relates to the technical field of audio processing. The method comprises the following steps: calculating a short-time frequency domain representation of an audio signal, wherein the audio signal comprises a first audio signal and a second audio signal; calculating a frequency response of a time-varying linear prediction error filter according to the short-time frequency domain representation of the audio signal, and calculating a short-time spectrum of a linear prediction error; extracting a low-frequency coefficient from the short-time spectrum and calculating a short-time spectrum of a corresponding linear prediction error envelope; calculating a short-time cross-power spectrum of the linear prediction error envelope corresponding to the audio signal according to the short-time spectrum of the linear prediction error envelope; calculating a frame-based time delay between the two audio signals according to the short-time cross-power spectrum of the linear prediction error envelope; and performing a variable-speed constant-pitch processing on the first audio signal or the second audio signal according to the calculation result of the frame-based time delay, so as to obtain an aligned audio signal. The present scheme can improve the accuracy of time delay estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to an audio signal alignment method and device, a computer-readable storage medium, and a terminal. Background Art

[0002] In a multi-person conference scenario, the quality of the sound pickup signal obtained by long-distance sound pickup using the array microphone on the all-in-one conference terminal (usually installed against a wall) is usually poor. To improve the microphone pickup effect, a feasible solution is to place one or more wireless microphones closer to the speaker and jointly process the sound pickup signals of the wireless microphones and the sound pickup signals of the all-in-one conference terminal. The relative delay between multiple audio signals cannot be ignored and is in a dynamically changing state, which will bring great difficulties to the subsequent audio processing process. Therefore, when performing multi-device joint processing, it is first necessary to estimate the relative delay of multiple audio signals, align them according to the estimated results, and then perform subsequent processing.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0004] The purpose of this application is to provide an audio signal alignment method and apparatus, a computer-readable storage medium, and a device, which can at least improve the accuracy of delay estimation to a certain extent.

[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.

[0006] According to a first aspect of the present application, a method for aligning audio signals is provided, the method comprising: determining a short-time frequency domain representation of an audio signal by short-time Fourier transform, wherein the audio signal comprises a first audio signal collected by a first device and a second audio signal collected by a second device; determining a frequency response of a time-varying linear prediction error filter of the audio signal based on the short-time frequency domain representation of the audio signal, and determining a short-time spectrum of a linear prediction error corresponding to the audio signal based on the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter; extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine a short-time spectrum of a linear prediction error envelope corresponding to the audio signal; determining a short-time cross-power spectrum of a linear prediction error envelope corresponding to the audio signal based on the short-time spectrum of the linear prediction error envelope; determining a frame delay estimation result between the first audio signal and the second audio signal based on the short-time cross-power spectrum of the linear prediction error envelope; and performing speed-varying, pitch-invariant processing on the first audio signal or the second audio signal based on the frame delay estimation result to obtain an aligned audio signal.

[0007] In one embodiment of the present application, the short-time frequency domain representation of the audio signal determined by short-time Fourier transform includes: framing the audio signal to obtain a time domain framed signal corresponding to the audio signal; windowing and fast Fourier transforming the time domain framed signal to obtain a short-time frequency domain representation of the audio signal.

[0008] In one embodiment of the present application, determining the frequency response of the time-varying linear prediction error filter of the audio signal based on the short-time frequency domain representation of the audio signal includes: determining the sum of the squares of the real part and the imaginary part of each frequency point in a target frame of the short-time frequency domain representation of the audio signal to obtain a power spectrum of the audio signal in the target frame; performing an inverse fast Fourier transform on the power spectrum of the audio signal in the target frame to obtain an autocorrelation function of the audio signal in the target frame; determining the time-varying linear prediction error filter coefficients of the audio signal based on the autocorrelation function of the audio signal in the target frame; and performing a fast Fourier transform on the time-varying linear prediction error filter coefficients of the audio signal to obtain the frequency response of the time-varying linear prediction error filter of the audio signal.

[0009] In one embodiment of the present application, determining the time-varying linear prediction error filter coefficients of the audio signal based on the autocorrelation function of the audio signal in the target frame includes: selecting the first p+1 values ​​of the autocorrelation function of the audio signal in the target frame, and determining the p-order linear prediction coefficients of the audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; and taking the opposite of the p-order linear prediction coefficients and adding the leading term 1 to obtain the time-varying linear prediction error filter coefficients of the audio signal with a length of p+1.

[0010] In one embodiment of the present application, determining the short-time frequency spectrum of the linear prediction error corresponding to the audio signal based on the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter includes: multiplying the complex coefficients of the target frequency point in the short-time frequency domain representation of the audio signal with the complex coefficients of the target frequency point in the frequency response of the time-varying linear prediction error filter to obtain the short-time frequency spectrum of the linear prediction error of the audio signal.

[0011] In one embodiment of the present application, the extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine the short-time spectrum of the linear prediction error envelope corresponding to the audio signal includes: determining the downsampling rate required for approximating the envelope, and determining the subscripts of the frequency points to be extracted in the short-time spectrum of the linear prediction error based on the number of frequency points in the short-time spectrum of the linear prediction error and the downsampling rate; extracting the coefficients of the corresponding frequency points from the short-time spectrum of the linear prediction error based on the subscripts of the frequency points to be extracted, and recombining the short-time spectrum of the linear prediction error envelope.

[0012] In one embodiment of the present application, determining the short-time cross-power spectrum of the linear prediction error envelope corresponding to the audio signal based on the short-time spectrum of the linear prediction error envelope includes: performing a conjugate transformation on the short-time spectrum of the linear prediction error envelope of the second audio signal to obtain a conjugate spectrum; and multiplying the conjugate spectrum with the short-time spectrum of the linear prediction error envelope of the first audio signal to obtain the short-time cross-power spectrum of the linear prediction error envelope.

[0013] In one embodiment of the present application, determining the frame delay estimation result between the first audio signal and the second audio signal based on the short-time cross-power spectrum of the linear prediction error envelope includes: performing an inverse fast Fourier transform on the short-time cross-power spectrum of the linear prediction error envelope to obtain a short-time cross-correlation function of the linear prediction error envelope corresponding to the audio signal; searching for a maximum value of the short-time cross-correlation function of the linear prediction error envelope, and using a sampling point offset corresponding to the maximum value as the frame delay estimation result between the first audio signal and the second audio signal.

[0014] In one embodiment of the present application, the above-mentioned first audio signal or the above-mentioned second audio signal is subjected to speed-changing but not pitch-changing processing based on the above-mentioned frame delay estimation result to obtain an aligned audio signal, including: if the maximum value of the short-time cross-correlation function of the above-mentioned linear prediction error envelope is greater than a preset threshold, then, based on the above-mentioned frame delay estimation result, the above-mentioned first audio signal or the above-mentioned second audio signal is subjected to speed-changing but not pitch-changing processing to obtain an aligned audio signal.

[0015] According to a second aspect of the present application, an audio signal alignment device is provided, the device comprising: a first determination module, configured to determine a short-time frequency domain representation of an audio signal by short-time Fourier transform, wherein the audio signal comprises a first audio signal collected by a first device and a second audio signal collected by a second device; a second determination module, configured to determine a frequency response of a time-varying linear prediction error filter of the audio signal based on the short-time frequency domain representation of the audio signal, and determine a short-time spectrum of a linear prediction error corresponding to the audio signal based on the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter; a third determination module, configured to determine a frequency response of a time-varying linear prediction error filter of the audio signal from the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter; and a third determination module, configured to determine a frequency response of a time-varying linear prediction error filter of the audio signal from the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter. Low-frequency coefficients are extracted from the short-time spectrum of the linear prediction error and recombined to determine the short-time spectrum of the linear prediction error envelope corresponding to the audio signal; a fourth determination module is used to determine the short-time cross-power spectrum of the linear prediction error envelope corresponding to the audio signal based on the short-time spectrum of the linear prediction error envelope; a fifth determination module is used to determine the frame delay estimation result between the first audio signal and the second audio signal based on the short-time cross-power spectrum of the linear prediction error envelope; and an alignment module is used to perform speed-changing but pitch-invariant processing on the first audio signal or the second audio signal based on the frame delay estimation result to obtain an aligned audio signal.

[0016] According to a third aspect of the present application, a terminal is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the audio signal alignment method described in the first aspect when executing the computer program.

[0017] According to a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the audio signal alignment method described in the first aspect is implemented.

[0018] The audio signal alignment method and device, computer storage medium, and terminal provided in the embodiments of the present application have the following technical effects:

[0019] A short-time frequency domain representation of an audio signal is calculated using a short-time Fourier transform, where the audio signal includes a first audio signal collected by a first device and a second audio signal collected by a second device. A frequency response of a time-varying linear prediction error filter of the audio signal is calculated based on the short-time frequency domain representation of the audio signal, and a short-time spectrum of the linear prediction error corresponding to the audio signal is calculated based on the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter. Low-frequency coefficients are then extracted from the short-time spectrum of the linear prediction error and recombined to determine a short-time spectrum of the linear prediction error envelope corresponding to the audio signal. Based on the short-time spectrum of the linear prediction error envelope, a short-time cross-power spectrum of the linear prediction error envelope corresponding to the audio signal is calculated. A frame-by-frame delay estimation result is calculated based on the short-time cross-power spectrum of the linear prediction error envelope. Finally, based on the frame-by-frame delay estimation result, the first audio signal or the second audio signal is subjected to speed-variable processing without pitch-variable processing to obtain an aligned audio signal. This solution can improve the accuracy of delay estimation, reduce time-frequency domain conversion operations, and reduce computational complexity.

[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0022] Figure 1 The following schematically shows a flow chart of an audio signal alignment method provided by an exemplary embodiment of the present application;

[0023] Figure 2 A flowchart of determining a delay estimation result provided by an exemplary embodiment of the present application is shown;

[0024] Figure 3 A flowchart of determining a linear prediction error frequency response according to an exemplary embodiment of the present application is shown;

[0025] Figure 4 The following schematically shows a structural diagram of an audio signal alignment device provided by an embodiment of the present application;

[0026] Figure 5 A schematic diagram illustrating the structure of an audio signal alignment device provided by another embodiment of the present application is shown;

[0027] Figure 6 A block diagram of a terminal provided in an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0028] In order to make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the accompanying drawings.

[0029] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0030] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. In addition, in the description of this application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0031] In related technologies, the delay estimation method for two audio signals is typically to calculate their correlation function and use the sampling point corresponding to the maximum value in the correlation function as the delay estimate. However, in typical conference room scenarios, reverberation is often quite pronounced, and the sound waves propagate from the speaker to the microphone through multiple reflections. Therefore, the delay calculated solely based on the correlation function often suffers from significant errors.

[0032] In response to the problems existing in the above-mentioned related technologies, the present application proposes an audio signal alignment method and device, a computer storage medium and a terminal to improve the accuracy of delay estimation.

[0033] Below, each step of the audio signal alignment method in this exemplary implementation will be described in more detail with reference to the accompanying drawings and embodiments.

[0034] in, Figure 1 The following schematically shows a flow chart of an audio signal alignment method according to an exemplary embodiment of the present application. Figure 2 FIG1 shows a flow chart of determining the delay estimation result according to an exemplary embodiment of the present application. Figure 2 right Figure 1 The illustrated embodiment will be described in detail.

[0035] S110 , determining a short-time frequency domain representation of an audio signal by short-time Fourier transform, wherein the audio signal includes a first audio signal collected by a first device and a second audio signal collected by a second device.

[0036] In an exemplary embodiment, as Figure 2 As shown, it is assumed that two microphones are used: a first device (device 1) and a second device (device 2) to collect audio signals and form two audio signals. In this embodiment, device 1 is taken as an example to illustrate processing process A. First, the collected first-channel audio signal is framed, and the first-channel audio signal after framing is subjected to short-time Fourier transform (STFT) processing to divide the first-channel audio signal (long-time signal) into several shorter equal-length time-domain framed signals. The framed signal is then subjected to windowing processing and fast Fourier transform (FFT) processing, thereby converting the original representation of each audio signal in the time domain into a representation in the time-frequency domain, that is, obtaining the short-time-frequency domain representation corresponding to the first-channel audio signal.

[0037] S120: Determine a frequency response of a time-varying linear prediction error filter of the audio signal based on the short-time frequency domain representation of the audio signal, and determine a short-time frequency spectrum of the linear prediction error corresponding to the audio signal based on the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter.

[0038] As a specific implementation of "determining the frequency response of the time-varying linear prediction error filter of the audio signal according to the short-time frequency domain representation of the audio signal" in step S120, Figure 3 The flow chart for determining the linear prediction error frequency response is shown below. Figure 3 For detailed description of the embodiment shown, please refer to Figure 3 .

[0039] S310 , determining the sum of the squares of the real part and the imaginary part of each frequency point in a target frame represented by the short-time frequency domain of the audio signal, and obtaining a power spectrum of the audio signal in the target frame.

[0040] In an exemplary embodiment, the time domain signal is usually represented by a real number, and after the short-time Fourier transform, it is converted into a short-time frequency domain signal represented by a complex number. Figure 2 As shown, the sum of the squares of the real part and the imaginary part of each frequency point in the target frame corresponding to the short-time frequency domain signal of the first audio signal is calculated, thereby obtaining the power spectrum corresponding to the first audio signal in the target frame.

[0041] S320 , performing an inverse fast Fourier transform on the power spectrum of the audio signal in the target frame to obtain an autocorrelation function of the audio signal in the target frame.

[0042] In the exemplary embodiment, continue to refer to Figure 2 After obtaining the power spectrum of the audio signal in the target frame, an inverse fast Fourier transform (IFFT) is performed on it to obtain the autocorrelation function (autocorrelation function, used to characterize the degree of correlation between the values ​​of the same sequence at different times) of the first audio signal corresponding to the target frame, which is recorded as r(n).

[0043] S330 : Determine a time-varying linear prediction error filter coefficient of the audio signal according to an autocorrelation function of the audio signal in the target frame.

[0044] In an exemplary embodiment, the first p+1 values ​​r(0), r(1) ... r(p) of the autocorrelation function r(n) are selected, and the Levinson-Durbin algorithm is used to solve the selected autocorrelation function, ultimately obtaining the corresponding p-order linear prediction coefficients c, which can be denoted as c(1), c(2) ... c(p).

[0045] In an exemplary embodiment, after obtaining the linear prediction coefficients, a time-varying linear prediction error filter coefficient h of length p+1 can be obtained. The first term of the time-varying linear prediction error filter coefficient h is 1, and the remaining terms are the negative values ​​(opposite numbers) of the p-order linear prediction coefficients, i.e., h(0) = 1, h(1) = -c(1), ..., h(p) = -c(p).

[0046] S340 , performing fast Fourier transform on the time-varying linear prediction error filter coefficients of the audio signal to obtain a frequency response of the time-varying linear prediction error filter of the audio signal.

[0047] In an exemplary embodiment, it is assumed that in the first audio signal after framing, the length of each frame of the audio signal is n, that is, there are n frequency points in total. Figure 2 As shown, the p+1 order time-varying linear prediction error filter coefficients are subjected to n-point fast Fourier transform to obtain the frequency response of the time-varying linear prediction error filter, which is denoted as H(f), and the number of its frequency points is also n.

[0048] In an exemplary embodiment, as Figure 2 As shown, after executing Figure 3After obtaining the frequency response H(f) of the time-varying linear prediction error filter in the steps shown, the complex coefficients of the target frequency point in the short-time frequency domain representation of the audio signal are multiplied by the complex coefficients of the target frequency point in H(f) to obtain the short-time spectrum of the linear prediction error corresponding to the first audio signal. The audio signal collected by each microphone device is processed through process A to obtain the corresponding short-time spectrum of the linear prediction error. That is, for device 2, the process for obtaining the short-time spectrum of the linear prediction error corresponding to the second audio signal is the same as for device 1 and is not further described here.

[0049] S130 , extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine a short-time spectrum of the linear prediction error envelope corresponding to the audio signal.

[0050] In an exemplary embodiment, as Figure 2 As shown in the figure, after obtaining the short-time spectrum of the linear prediction error corresponding to the two audio signals, it is necessary to determine the downsampling rate required for calculating the approximate envelope. Based on the number of frequency points in the short-time spectrum of the linear prediction error and the downsampling rate, the subscripts of the frequency points to be extracted from the short-time spectrum of the linear prediction error are calculated. Then, based on the subscripts of the frequency points to be extracted, the coefficients of the corresponding frequency points are extracted from the short-time spectrum of the linear prediction error, and the short-time spectrum of the linear prediction error envelope is reassembled.

[0051] S140 : Determine a short-time cross-power spectrum of the linear prediction error envelope corresponding to the audio signal according to the short-time frequency spectrum of the linear prediction error envelope.

[0052] In an exemplary embodiment, as Figure 2 As shown, after obtaining the short-time spectrum of the linear prediction error envelope, the short-time spectrum of the linear prediction error envelope of the second audio signal is conjugated, and the conjugated spectrum is multiplied with the short-time spectrum of the linear prediction error envelope of the first audio signal to obtain the short-time cross-power spectrum of the linear prediction error envelope between the two audio signals.

[0053] S150 , determining a frame delay estimation result between the first audio signal and the second audio signal according to the short-time cross power spectrum of the linear prediction error envelope.

[0054] In the exemplary embodiment, continue to refer to Figure 2 After determining the short-time cross-power spectrum of the linear prediction error envelope, an inverse fast Fourier transform is performed on it to obtain the short-time cross-correlation function of the linear prediction error envelope between the two audio signals. A peak search is then performed on the short-time cross-correlation function, and the sampling point offset corresponding to the maximum value in the short-time cross-correlation function is used as a preliminary estimate of the frame delay between the two audio signals.

[0055] S160 , performing speed-changing but pitch-invariant processing on the first audio signal or the second audio signal according to the frame delay estimation result to obtain an aligned audio signal.

[0056] In an exemplary embodiment, as Figure 2 As shown, after obtaining the preliminary estimate of the frame delay, it still needs to be verified and judged. For example, the maximum value of the short-time cross-correlation function can be compared with a preset threshold. If the maximum value is greater than the preset threshold, it indicates that the correlation between the first audio signal and the second audio signal is sufficiently strong, and the preliminary estimate of the frame delay can be output as the final estimate of the frame delay.

[0057] In an exemplary embodiment, the final estimation result of the frame delay and the two original time domain audio signals are input into the expansion and compression module, such as Figure 2 As shown, the expansion and compression module performs frame-level time scale modification (TSM) on the first and second audio signals based on the final estimated frame delay. This process maintains the pitch and semantics of the audio signals while speeding up or slowing down the speech rate. The module then outputs the final synchronized signals: the aligned audio signal corresponding to device 1 and the aligned audio signal corresponding to device 2. This module can reduce listener discomfort caused by time-advancing or delaying the audio signals.

[0058] Since sound waves have multiple reflections when propagating in an actual room, it is equivalent to the convolution of the original audio signal and the room impulse response. When the reverberation lasts for a long time, the signals received by microphones at different positions in the room will be severely distorted. If the cross-correlation function of the audio signals collected by the two devices is directly calculated, the peak position will produce a large error from the actual time delay. In the audio signal alignment method provided in this application, part of the environmental influence is eliminated by the envelope of the linear prediction error signal, making the audio signal closer to the speaker's vocal cord vibration pattern, which is conducive to improving the accuracy of the time delay estimation.

[0059] In addition, when calculating the linear prediction coefficient, the method commonly used in related technologies is to calculate the autocorrelation function of the audio signal in the time domain and, after obtaining the linear prediction error filter, convolve it in the time domain. However, the audio signal alignment method provided in this application utilizes the conversion relationship between the power spectrum and the autocorrelation function to convert the above processing method to the frequency domain. The final result is the short-time spectrum of the linear prediction error, which can be directly used to calculate the short-time cross-correlation function of the two devices, thereby reducing the operation of converting between the time and frequency domains and reducing the amount of calculation.

[0060] In addition, a deep learning model can be designed and trained using the audio signal fusion method provided in this application. The trained model can calculate the time delay of the two input audio signals and output the alignment results of the two audio signals.

[0061] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0062] in, Figure 4 The figure shows a structural diagram of an audio signal alignment device according to an exemplary embodiment of the present application.

[0063] The audio signal alignment apparatus 400 in the embodiment of the present application includes: a first determination module 410, a second determination module 420, a third determination module 430, a fourth determination module 440, a fifth determination module 450, and an alignment module 460, wherein:

[0064] The first determining module 410 is configured to determine a short-time frequency domain representation of an audio signal by short-time Fourier transform, wherein the audio signal includes a first audio signal collected by the first device and a second audio signal collected by the second device;

[0065] a second determining module 420 configured to determine a frequency response of a time-varying linear prediction error filter of the audio signal based on the short-time-frequency-domain representation of the audio signal, and determine a short-time spectrum of the linear prediction error corresponding to the audio signal based on the short-time-frequency-domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter;

[0066] The third determining module 430 is configured to extract low-frequency coefficients from the short-time spectrum of the linear prediction error and recombine them to determine a short-time spectrum of the linear prediction error envelope corresponding to the audio signal;

[0067] The fourth determining module 440 is configured to determine a short-time cross power spectrum of the linear prediction error envelope corresponding to the audio signal based on the short-time spectrum of the linear prediction error envelope.

[0068] A fifth determining module 450 is configured to determine a frame delay estimation result between the first audio signal and the second audio signal based on a short-time cross power spectrum of a linear prediction error envelope;

[0069] The alignment module 460 is configured to perform speed-changing but pitch-invariant processing on the first audio signal or the second audio signal according to the frame delay estimation result to obtain an aligned audio signal.

[0070] Figure 5 A structural diagram of an audio signal alignment device according to another exemplary embodiment of the present application is shown.

[0071] The first determination module 410 includes a framing unit 4101 and a first transform unit 4102. The framing unit 4101 is configured to frame the audio signal to obtain a time-domain framed signal corresponding to the audio signal; and the transform unit 4102 is configured to perform windowing and fast Fourier transform on the time-domain framed signal to obtain a short-time-frequency domain representation of the audio signal.

[0072] The second determination module 420 includes: a first determination unit 4201, a second transformation unit 4202, a second determination unit 4203, and a third transformation unit 4204. The first determination unit 4201 is configured to determine the sum of the squares of the real and imaginary parts of each frequency point in a target frame represented by the short-time frequency domain of the audio signal to obtain a power spectrum of the audio signal in the target frame; the second transformation unit 4202 is configured to perform an inverse fast Fourier transform on the power spectrum of the audio signal in the target frame to obtain an autocorrelation function of the audio signal in the target frame; the second determination unit 4203 is configured to determine a time-varying linear prediction error filter coefficient of the audio signal based on the autocorrelation function of the audio signal in the target frame; and the third transformation unit 4204 is configured to perform a fast Fourier transform on the time-varying linear prediction error filter coefficient of the audio signal to obtain a frequency response of the time-varying linear prediction error filter of the audio signal.

[0073] The above-mentioned second determination unit 4203 is specifically used to: select the first p+1 values ​​of the autocorrelation function of the audio signal in the target frame, and determine the p-order linear prediction coefficient of the audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; take the opposite of the p-order linear prediction coefficient and add the leading term 1 to obtain the time-varying linear prediction error filter coefficient of the audio signal with a length of p+1.

[0074] The second determination module 420 includes a third determination unit 4205. The third determination unit 4205 is configured to multiply the complex coefficients of the target frequency point in the short-time frequency domain representation of the audio signal by the complex coefficients of the target frequency point in the frequency response of the time-varying linear prediction error filter to obtain a short-time frequency spectrum of the linear prediction error of the audio signal.

[0075] The third determination module 430 includes a fourth determination unit 4301 and a fifth determination unit 4302. The fourth determination unit 4301 is configured to determine the downsampling rate required for approximating the envelope, and determine the subscripts of the frequency points to be extracted in the short-time spectrum of the linear prediction error based on the number of frequency points in the short-time spectrum of the linear prediction error and the downsampling rate. The fifth determination unit 4302 is configured to extract the coefficients of the corresponding frequency points from the short-time spectrum of the linear prediction error based on the subscripts of the frequency points to be extracted, and reassemble the short-time spectrum of the linear prediction error envelope.

[0076] The fourth determination module 440 includes a conjugation unit 4401 and a sixth determination unit 4402. The conjugation unit 4401 is configured to perform a conjugate transform on the short-time spectrum of the linear prediction error envelope of the second audio signal to obtain a conjugate spectrum; and the sixth determination unit 4402 is configured to multiply the conjugate spectrum by the short-time spectrum of the linear prediction error envelope of the first audio signal to obtain a short-time cross-power spectrum of the linear prediction error envelope.

[0077] The fifth determination module 450 includes a fourth transform unit 4501 and a search unit 4502. The fourth transform unit 4501 is configured to perform an inverse fast Fourier transform on the short-time cross-power spectrum of the linear prediction error envelope to obtain a short-time cross-correlation function of the linear prediction error envelope corresponding to the audio signal; and the search unit 4502 is configured to search for a maximum value of the short-time cross-correlation function of the linear prediction error envelope and use a sampling point offset corresponding to the maximum value as a frame delay estimation result between the first audio signal and the second audio signal.

[0078] The alignment module 460 includes an alignment unit 4601. The alignment unit 4601 is configured to: if the maximum value of the short-time cross-correlation function of the linear prediction error envelope is greater than a preset threshold, perform speed-shifting without pitch-invariance processing on the first audio signal or the second audio signal based on the frame delay estimation result to obtain an aligned audio signal.

[0079] It should be noted that the audio signal alignment device provided in the above embodiment only uses the division of the above functional modules as an example when executing the audio signal alignment method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio signal alignment device provided in the above embodiment and the audio signal alignment method embodiment belong to the same concept. Therefore, for details not disclosed in the device embodiment of the present application, please refer to the embodiment of the audio signal alignment method of the present application, which will not be repeated here.

[0080] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0081] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the aforementioned methods. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive, a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.

[0082] An embodiment of the present application also provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of any of the above-mentioned method embodiments are implemented.

[0083] Figure 6 The schematic diagram shows the structure of the terminal according to an exemplary embodiment of the present application. Figure 6 As shown, the terminal 600 includes a processor 601 and a memory 602 .

[0084] In the embodiment of the present application, the processor 601 is the control center of the computer system, which can be the processor of a physical machine or the processor of a virtual machine. The processor 601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state.

[0085] In the embodiment of the present application, the processor 601 is specifically configured to:

[0086] Determining a short-time frequency domain representation of an audio signal through short-time Fourier transform, wherein the audio signal includes a first audio signal collected by a first device and a second audio signal collected by a second device; determining a frequency response of a time-varying linear prediction error filter of the audio signal based on the short-time frequency domain representation of the audio signal, and determining a short-time spectrum of the linear prediction error corresponding to the audio signal based on the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter; extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine a short-time spectrum of a linear prediction error envelope corresponding to the audio signal; determining a short-time cross-power spectrum of the linear prediction error envelope corresponding to the audio signal based on the short-time spectrum of the linear prediction error envelope; determining a frame delay estimation result between the first audio signal and the second audio signal based on the short-time cross-power spectrum of the linear prediction error envelope; and performing speed-varying, pitch-invariant processing on the first audio signal or the second audio signal based on the frame delay estimation result to obtain an aligned audio signal.

[0087] Furthermore, in one embodiment of the present application, the above-mentioned determination of the short-time frequency domain representation of the audio signal by short-time Fourier transform includes: framing the above-mentioned audio signal to obtain a time domain frame signal corresponding to the above-mentioned audio signal; windowing and fast Fourier transforming the above-mentioned time domain frame signal to obtain the short-time frequency domain representation of the above-mentioned audio signal.

[0088] Optionally, the above-mentioned determining the frequency response of the time-varying linear prediction error filter of the audio signal based on the short-time frequency domain representation of the audio signal includes: determining the sum of the squares of the real part and the imaginary part of each frequency point in the target frame of the short-time frequency domain representation of the audio signal to obtain the power spectrum of the audio signal in the target frame; performing an inverse fast Fourier transform on the power spectrum of the audio signal in the target frame to obtain the autocorrelation function of the audio signal in the target frame; determining the time-varying linear prediction error filter coefficients of the audio signal based on the autocorrelation function of the audio signal in the target frame; and performing a fast Fourier transform on the time-varying linear prediction error filter coefficients of the audio signal to obtain the frequency response of the time-varying linear prediction error filter of the audio signal.

[0089] Optionally, determining the time-varying linear prediction error filter coefficient of the audio signal based on the autocorrelation function of the audio signal in the target frame includes: selecting the first p+1 values ​​of the autocorrelation function of the audio signal in the target frame, and determining the p-order linear prediction coefficient of the audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; taking the opposite of the p-order linear prediction coefficient and adding the leading term 1 to obtain the time-varying linear prediction error filter coefficient of the audio signal with a length of p+1.

[0090] Optionally, the above-mentioned determining the short-time frequency spectrum of the linear prediction error corresponding to the above-mentioned audio signal based on the short-time frequency domain representation of the above-mentioned audio signal and the frequency response of the above-mentioned time-varying linear prediction error filter includes: multiplying the complex coefficients of the target frequency point in the short-time frequency domain representation of the above-mentioned audio signal with the complex coefficients of the above-mentioned target frequency point in the frequency response of the above-mentioned time-varying linear prediction error filter to obtain the short-time frequency spectrum of the linear prediction error of the above-mentioned audio signal.

[0091] Optionally, the above-mentioned extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine the short-time spectrum of the linear prediction error envelope corresponding to the audio signal includes: determining the downsampling rate required for the approximate envelope, and determining the subscripts of the frequency points to be extracted in the short-time spectrum of the linear prediction error based on the number of frequency points of the short-time spectrum of the linear prediction error and the above-mentioned downsampling rate; extracting the coefficients of the corresponding frequency points from the short-time spectrum of the linear prediction error based on the subscripts of the frequency points to be extracted, and recombining the short-time spectrum of the linear prediction error envelope.

[0092] Optionally, the above-mentioned determining the short-time cross-power spectrum of the linear prediction error envelope corresponding to the above-mentioned audio signal based on the short-time spectrum of the above-mentioned linear prediction error envelope includes: performing a conjugate transformation on the short-time spectrum of the linear prediction error envelope of the above-mentioned second audio signal to obtain a conjugate spectrum; and multiplying the above-mentioned conjugate spectrum with the short-time spectrum of the linear prediction error envelope of the above-mentioned first audio signal to obtain the short-time cross-power spectrum of the above-mentioned linear prediction error envelope.

[0093] Optionally, the above-mentioned determining the frame delay estimation result between the above-mentioned first audio signal and the above-mentioned second audio signal based on the short-time cross-power spectrum of the above-mentioned linear prediction error envelope includes: performing an inverse fast Fourier transform on the short-time cross-power spectrum of the above-mentioned linear prediction error envelope to obtain the short-time cross-correlation function of the linear prediction error envelope corresponding to the above-mentioned audio signal; searching for the maximum value of the short-time cross-correlation function of the above-mentioned linear prediction error envelope, and taking the sampling point offset corresponding to the above-mentioned maximum value as the frame delay estimation result between the above-mentioned first audio signal and the above-mentioned second audio signal.

[0094] Optionally, the above-mentioned first audio signal or the second audio signal is subjected to speed-changing but not pitch-changing processing based on the above-mentioned frame delay estimation result to obtain an aligned audio signal, including: if the maximum value of the short-time cross-correlation function of the above-mentioned linear prediction error envelope is greater than a preset threshold, then the above-mentioned first audio signal or the second audio signal is subjected to speed-changing but not pitch-changing processing based on the above-mentioned frame delay estimation result to obtain an aligned audio signal.

[0095] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage terminals and flash memory storage terminals. In some embodiments of the present application, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one instruction, which is used to be executed by the processor 601 to implement the method in the embodiment of the present application.

[0096] In some embodiments, terminal 600 further includes a peripheral terminal interface 603 and at least one peripheral terminal. Processor 601, memory 602, and peripheral terminal interface 603 may be connected via a bus or signal lines. Each peripheral terminal may be connected to peripheral terminal interface 603 via a bus, signal lines, or circuit boards. Specifically, the peripheral terminal includes at least one of a display screen 604, a camera 606, and an audio circuit 606.

[0097] The peripheral terminal interface 603 can be used to connect at least one peripheral terminal related to input / output (I / O) to the processor 601 and the memory 602. In some embodiments of the present application, the processor 601, the memory 602, and the peripheral terminal interface 603 are integrated on the same chip or circuit board; in some other embodiments of the present application, any one or two of the processor 601, the memory 602, and the peripheral terminal interface 603 can be implemented on separate chips or circuit boards. This embodiment of the present application is not specifically limited to this.

[0098] Display screen 604 is used to display a user interface (UI). The UI may include graphics, text, icons, videos, or any combination thereof. When display screen 604 is a touch screen display, it is also capable of collecting touch signals on or above the surface of display screen 604. The touch signals can be input as control signals to processor 601 for processing. In this case, display screen 604 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments of the present application, there may be one display screen 604, provided on the front panel of terminal 600; in other embodiments of the present application, there may be at least two display screens 604, provided on different surfaces of terminal 600 or in a foldable design; in still other embodiments of the present application, display screen 604 may be a flexible display, provided on a curved or foldable surface of terminal 600. Furthermore, display screen 604 can be configured as a non-rectangular irregular shape, i.e., a special-shaped screen. Display screen 604 can be made of materials such as liquid crystal display (LCD) and organic light-emitting diode (OLED).

[0099] Camera 605 is used to capture images or videos. Optionally, camera 605 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and virtual reality (VR) shooting function or other fusion shooting functions. In some embodiments of the present application, camera 605 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0100] Audio circuit 606 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals for input to processor 601 for processing. For the purpose of stereo sound collection or noise reduction, multiple microphones may be provided, respectively, at different locations on terminal 600. The microphone may also be an array microphone or an omnidirectional microphone.

[0101] Power supply 605 is used to power various components in terminal 600. Power supply 605 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 605 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0102] The terminal structure block diagram shown in the embodiment of the present application does not constitute a limitation on the terminal 600. The terminal 600 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0103] In this application, the terms "first," "second," etc. are used for descriptive purposes only and are not to be construed as indicating or implying relative importance or order; the term "plurality" refers to two or more, unless expressly limited otherwise. Terms such as "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, "connected" can mean a fixed connection, a detachable connection, or an integral connection; "connected" can mean a direct connection or an indirect connection through an intermediary. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.

[0104] In the description of this application, it should be understood that the terms "upper" and "lower" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or unit referred to must have a specific direction, be constructed and operated in a specific orientation. Therefore, they should not be understood as limitations on this application.

[0105] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, equivalent modifications made according to the claims of the present application are still within the scope of protection of the present application.

Claims

1. A method for aligning an audio signal, characterized in that: include: Determine a short-time frequency domain representation of an audio signal by short-time Fourier transform, wherein the audio signal includes a first audio signal collected by the first device and a second audio signal collected by the second device; determining a frequency response of a time-varying linear prediction error filter of the audio signal based on a short-time frequency domain representation of the audio signal, and determining a short-time frequency spectrum of a linear prediction error corresponding to the audio signal based on the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter; Extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine a short-time spectrum of the linear prediction error envelope corresponding to the audio signal; Determining a short-time cross-power spectrum of the linear prediction error envelope corresponding to the audio signal according to the short-time spectrum of the linear prediction error envelope; determining a frame delay estimation result between the first audio signal and the second audio signal according to the short-time cross power spectrum of the linear prediction error envelope; According to the frame delay estimation result, speed-changing but pitch-invariant processing is performed on the first audio signal or the second audio signal to obtain an aligned audio signal.

2. The audio signal alignment method according to claim 1, characterized in that: Determining the short-time frequency domain representation of the audio signal by short-time Fourier transform includes: framing the audio signal to obtain a time-domain frame signal corresponding to the audio signal; Windowing and fast Fourier transform are performed on the time-domain framed signal to obtain a short-time frequency-domain representation of the audio signal.

3. The audio signal alignment method according to claim 1, characterized in that: The determining the frequency response of the time-varying linear prediction error filter of the audio signal according to the short-time frequency domain representation of the audio signal comprises: determining the sum of the squares of the real part and the imaginary part of each frequency point in a target frame represented by the short-time frequency domain of the audio signal to obtain a power spectrum of the audio signal in the target frame; performing an inverse fast Fourier transform on a power spectrum of the audio signal in the target frame to obtain an autocorrelation function of the audio signal in the target frame; Determining a time-varying linear prediction error filter coefficient of the audio signal according to an autocorrelation function of the audio signal in the target frame; Performing a fast Fourier transform on the time-varying linear prediction error filter coefficients of the audio signal to obtain a frequency response of the time-varying linear prediction error filter of the audio signal.

4. The audio signal alignment method according to claim 3, characterized in that: The step of determining the time-varying linear prediction error filter coefficient of the audio signal according to the autocorrelation function of the audio signal in the target frame includes: Selecting first p+1 values ​​of the autocorrelation function of the audio signal in the target frame, and determining a p-order linear prediction coefficient of the audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; The inverse of the p-order linear prediction coefficient is taken and the leading term 1 is added to obtain the time-varying linear prediction error filter coefficient of the audio signal with a length of p+1.

5. The audio signal alignment method according to any one of claims 1 to 4, characterized in that: The determining of the short-time frequency spectrum of the linear prediction error corresponding to the audio signal according to the short-time frequency domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter comprises: The complex coefficient of the target frequency point in the short-time frequency domain representation of the audio signal is multiplied by the complex coefficient of the target frequency point in the frequency response of the time-varying linear prediction error filter to obtain a short-time frequency spectrum of the linear prediction error of the audio signal.

6. The audio signal alignment method according to claim 1, characterized in that: The extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine the short-time spectrum of the linear prediction error envelope corresponding to the audio signal includes: Determining a downsampling rate required for approximating the envelope, and determining the subscript of the frequency point to be extracted in the short-time spectrum of the linear prediction error according to the number of frequency points of the short-time spectrum of the linear prediction error and the downsampling rate; According to the subscripts of the frequency points to be extracted, coefficients of the corresponding frequency points are extracted from the short-time spectrum of the linear prediction error, and the short-time spectrum of the linear prediction error envelope is reassembled.

7. The audio signal alignment method according to claim 1 or 6, characterized in that: Determining the short-time cross power spectrum of the linear prediction error envelope corresponding to the audio signal according to the short-time spectrum of the linear prediction error envelope includes: performing a conjugate transform on the short-time spectrum of the linear prediction error envelope of the second audio signal to obtain a conjugate spectrum; The conjugate spectrum is multiplied by the short-time spectrum of the linear prediction error envelope of the first audio signal to obtain the short-time cross-power spectrum of the linear prediction error envelope.

8. The audio signal alignment method according to claim 1, characterized in that: The determining, based on the short-time cross power spectrum of the linear prediction error envelope, a frame delay estimation result between the first audio signal and the second audio signal includes: Performing an inverse fast Fourier transform on the short-time cross power spectrum of the linear prediction error envelope to obtain a short-time cross-correlation function of the linear prediction error envelope corresponding to the audio signal; A maximum value of the short-time cross-correlation function of the linear prediction error envelope is searched, and a sampling point offset corresponding to the maximum value is used as a frame delay estimation result between the first audio signal and the second audio signal.

9. The audio signal alignment method according to claim 8, characterized in that: The step of performing speed-changing but pitch-invariant processing on the first audio signal or the second audio signal according to the frame delay estimation result to obtain an aligned audio signal includes: If the maximum value of the short-time cross-correlation function of the linear prediction error envelope is greater than a preset threshold, the first audio signal or the second audio signal is subjected to speed-changing processing without changing pitch according to the frame delay estimation result to obtain an aligned audio signal.

10. An audio signal alignment device, characterized in that: include: A first determining module is configured to determine a short-time frequency domain representation of an audio signal by short-time Fourier transform, wherein the audio signal includes a first audio signal collected by the first device and a second audio signal collected by the second device; a second determining module, configured to determine a frequency response of a time-varying linear prediction error filter of the audio signal based on a short-time frequency-domain representation of the audio signal, and determine a short-time frequency spectrum of a linear prediction error corresponding to the audio signal based on the short-time frequency-domain representation of the audio signal and the frequency response of the time-varying linear prediction error filter; A third determining module is configured to extract low-frequency coefficients from the short-time spectrum of the linear prediction error and recombine them to determine a short-time spectrum of the linear prediction error envelope corresponding to the audio signal; a fourth determining module, configured to determine a short-time cross power spectrum of the linear prediction error envelope corresponding to the audio signal based on the short-time spectrum of the linear prediction error envelope; a fifth determining module, configured to determine a frame delay estimation result between the first audio signal and the second audio signal according to a short-time cross power spectrum of the linear prediction error envelope; The alignment module is configured to perform speed-changing but pitch-invariant processing on the first audio signal or the second audio signal according to the frame delay estimation result to obtain an aligned audio signal.

11. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the audio signal alignment method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio signal alignment method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Sound signal time delay estimation method and device

    CN104700842A

  • Audio spatial localization apparatus and methods

    US6078669A