Method and apparatus for fusing audio signals, computer storage medium and terminal

By performing short-time frequency domain representation, time-varying linear prediction error filter response, and nonlinear compression processing on multiple audio signals, and determining the weights, the problem of poor audio signal quality in multi-person conferences is solved, and speech clarity and intelligibility are improved.

CN117351989BActive Publication Date: 2026-05-19GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU SHIYUAN ELECTRONICS CO LTD
Filing Date
2022-06-28
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In multi-person conference scenarios, the sound pickup signal obtained by long-distance sound pickup through the array microphone on the conference terminal is of poor quality, especially the signal-to-noise ratio and signal-to-mixing ratio of the microphone signal that is far away from the speaker, which affects the quality of the audio signal.

Method used

By determining the short-time frequency domain representation of each audio signal corresponding to multiple devices, the frequency response of the time-varying linear prediction error filter is calculated, low-frequency coefficients are extracted and the linear prediction error envelope is recombined, nonlinear compression processing is performed, weights are determined based on the variance of the nonlinear compressed signal, and the audio signals are fused.

Benefits of technology

It improves the quality of the fused audio signal, enhances speech clarity and intelligibility, adapts to changes in speaker position, reduces reverberation, lowers the probability of weight assignment errors, and reduces computational load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351989B_ABST
    Figure CN117351989B_ABST
Patent Text Reader

Abstract

The application provides a fusion method and device of an audio signal, a storage medium and a terminal, and relates to the technical field of audio processing. The method comprises the following steps: determining a short-time frequency domain representation of audio signals of a plurality of devices; determining a frequency response of a time-varying linear prediction error filter according to the short-time frequency domain representation of the audio signals, and calculating a short-time frequency spectrum of a linear prediction error; extracting low-frequency coefficients from the short-time frequency spectrum and determining a linear prediction error envelope; performing amplitude compensation and nonlinear compression on the linear prediction error envelope, and extracting features; determining weights corresponding to the audio signals of each channel according to the variance of the nonlinear compressed signals, and fusing the audio signals of each channel according to the weights. The scheme can improve the quality of the fused audio signals and the intelligibility and clarity of the speech by determining the weights of the multi-channel audio in a strong reverberation environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio signal fusion method and apparatus, a computer-readable storage medium and terminal. Background Technology

[0002] In multi-person conference scenarios, the quality of audio signals obtained from long-distance sound pickup using array microphones on a conference terminal (typically mounted against a wall) is often poor. To improve microphone pickup performance, a feasible solution is to place one or more wireless microphones closer to the speaker and combine their signals with those from the conference terminal. Generally, microphones closer to the speaker produce better quality signals, while microphones farther away have lower signal-to-noise ratios and signal-to-mixing ratios, resulting in poorer quality. Therefore, to leverage the capabilities of multiple wireless microphones, it is necessary to fuse the multiple audio signals.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] The purpose of this application is to provide a method and apparatus for fusing audio signals, a computer-readable storage medium and device, which can at least improve the quality of the fused audio signals to a certain extent.

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0006] According to a first aspect of this application, an audio signal fusion method is provided, the method comprising: determining short-time frequency domain representations of audio signals corresponding to multiple devices; determining the frequency response of a time-varying linear prediction error filter for each audio signal based on the short-time frequency domain representations of the audio signals, and calculating the short-time spectrum of the linear prediction error of each audio signal based on the short-time frequency domain representations and the frequency response of the time-varying linear prediction error filter; extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal, and determining the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope; performing nonlinear compression processing on the linear prediction error envelope to obtain a nonlinear compressed signal of each audio signal; determining the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal, and fusing the audio signals based on the weights.

[0007] In one embodiment of this application, the above-mentioned determination of the short-time frequency domain representation of each audio signal corresponding to multiple devices includes: dividing the audio signals into frames to obtain time-domain framed signals corresponding to each audio signal; and performing windowing and fast Fourier transform on the time-domain framed signals to obtain the short-time frequency domain representation of each audio signal.

[0008] In one embodiment of this application, determining the frequency response of the time-varying linear prediction error filter for each audio signal based on its short-time frequency domain representation includes: determining the sum of squares of the real and imaginary parts of each frequency point in the target frame of the short-time frequency domain representation to obtain the power spectrum of each audio signal in the target frame; performing an inverse fast Fourier transform on the power spectrum to obtain the autocorrelation function of each audio signal in the target frame; determining the coefficients of the time-varying linear prediction error filter for each audio signal based on the autocorrelation function; and performing a fast Fourier transform on the coefficients of the time-varying linear prediction error filter to obtain the frequency response of the time-varying linear prediction error filter for each audio signal.

[0009] In one embodiment of this application, the determination of the time-varying linear prediction error filter coefficients of each audio signal based on the autocorrelation function includes: selecting the first p+1 values ​​of the autocorrelation function of each audio signal in the target frame, and determining the p-th order linear prediction coefficients of each audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; taking the negative of the p-th order linear prediction coefficients and adding the first term 1 to obtain the time-varying linear prediction error filter coefficients of each audio signal with a length of p+1.

[0010] In one embodiment of this application, the above-mentioned calculation of the short-time spectrum of the linear prediction error of each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter includes: multiplying the complex coefficients of each frequency point in the frequency domain representation with the corresponding complex coefficients of each frequency point in the frequency response of the time-varying linear prediction error filter to obtain the short-time spectrum of the linear prediction error of each audio signal.

[0011] In one embodiment of this application, the above-mentioned extraction and recombination of low-frequency coefficients from the short-time spectrum of the linear prediction error to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal includes: determining the downsampling rate required for the approximate envelope; determining the index of the frequency point to be extracted in the short-time spectrum of the linear prediction error based on the number of frequency points in the short-time spectrum of the linear prediction error and the downsampling rate; extracting the coefficients of the corresponding frequency point from the short-time spectrum of the linear prediction error based on the index of the frequency point to be extracted; and recombinating the short-time spectrum of the linear prediction error envelope.

[0012] In one embodiment of this application, determining the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope includes: performing an inverse fast Fourier transform on the short-time spectrum of the linear prediction error envelope to obtain the linear prediction error envelope.

[0013] In one embodiment of this application, the nonlinear compression processing of the linear prediction error envelope to obtain the nonlinear compressed signal of each audio signal includes: calculating the average energy of the linear prediction error envelope frame by frame, and performing exponential smoothing on the average energy of the linear prediction error envelope to obtain the updated average energy of the current frame; subtracting the updated average energy of the current frame from the logarithmic transform of the linear prediction error envelope signal of the current frame to obtain the subtraction result, and calculating the exponential function of the subtraction result to obtain the amplitude-compensated linear prediction error envelope; and calculating the cube root of the amplitude-compensated linear prediction error envelope to obtain the nonlinear compressed signal.

[0014] In one embodiment of this application, the above-mentioned method of determining the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal and fusing the audio signals based on the weights includes: calculating the variance of the nonlinear compressed signal to obtain the variance of each audio signal; inputting the variance of each audio signal into a weight adjuster, and fusing the audio signals based on the weights output by the weight adjuster.

[0015] In one embodiment of this application, the method further includes: using the variance of the nonlinear compressed signal as a feature of each audio signal; the weight adjuster is used to: assign corresponding weights to each audio signal according to the distribution of the features of each audio signal in each audio signal; and limit the rate of change of the weights of each audio signal.

[0016] According to a second aspect of this application, an audio signal fusion apparatus is provided, the apparatus comprising: a first determining module, configured to: determine short-time frequency domain representations of audio signals corresponding to multiple devices; a second determining module, configured to: determine the frequency response of a time-varying linear prediction error filter for each audio signal based on the short-time frequency domain representations of the audio signals, and calculate the short-time spectrum of the linear prediction error of each audio signal based on the short-time frequency domain representations and the frequency response of the time-varying linear prediction error filter; a third determining module, configured to: extract low-frequency coefficients from the short-time spectrum of the linear prediction error and recombine them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal, and determine the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope; a nonlinear compression module, configured to: perform nonlinear compression processing on the linear prediction error envelope to obtain a nonlinear compressed signal of each audio signal; and a fusion module, configured to: determine the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal, and fuse the audio signals based on the weights.

[0017] According to a third aspect of this application, a terminal is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the audio signal fusion method described in the first aspect.

[0018] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the audio signal fusion method described in the first aspect.

[0019] The audio signal fusion method, apparatus, computer storage medium, and terminal provided in the embodiments of this application have the following technical effects: Determining the short-time frequency domain representation of each audio signal corresponding to multiple devices. Based on the short-time frequency domain representation of each audio signal, determining the frequency response of the time-varying linear prediction error filter for each audio signal, and calculating the short-time spectrum of the linear prediction error for each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter. Extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombinating them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal, and determining the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope. Performing nonlinear compression processing on the linear prediction error envelope to obtain the nonlinear compressed signal of each audio signal. Determining the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal, and fusing the audio signals according to the weights. This solution can improve the quality of the fused audio signal and enhance speech clarity and intelligibility by determining the weights of multi-channel audio in environments with strong reverberation.

[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0022] Figure 1 The flowchart illustrating an exemplary embodiment of the audio signal fusion method provided in this application is shown in the illustration.

[0023] Figure 2 A flowchart illustrating the determination of weights for multiple audio signals is shown in an exemplary embodiment of this application;

[0024] Figure 3 A flowchart illustrating the determination of the frequency response of a linear prediction error is shown in an exemplary embodiment of this application;

[0025] Figure 4 This schematic diagram illustrates the structure of an audio signal fusion apparatus according to an embodiment of this application;

[0026] Figure 5 This schematic diagram illustrates the structure of an audio signal fusion apparatus according to another embodiment of this application;

[0027] Figure 6 The diagram illustrates a block diagram of a terminal provided in one embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0030] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0031] In related technologies, one method for fusing multiple audio signals is to directly add them together. Since the amplitude of the audio signal from the microphone closest to the speaker is larger, the signal from the microphone closest to the speaker has the largest energy proportion among the multiple audio signals, which helps improve the quality of the fused audio signal. However, this method is actually equivalent to averaging multiple audio signals with different signal-to-noise ratios, thus its improvement on the quality of the fused audio signal is limited. Furthermore, in practical applications, the speaker's position constantly changes, and correspondingly, the quality of the multiple audio signals also changes. Therefore, it is necessary to adjust the weights in real time based on the acquired audio signals.

[0032] To address the problems existing in the aforementioned related technologies, this application proposes an audio signal fusion method and apparatus, a computer storage medium and a terminal, to increase the proportion of the microphone pickup signal closest to the speaker in multiple audio signals, and to adjust the weights in real time according to the collected multiple audio signals, thereby achieving a significant improvement in pickup quality.

[0033] The following will describe in more detail the steps of the audio signal fusion method in this exemplary embodiment, with reference to the accompanying drawings and embodiments.

[0034] in, Figure 1 A flowchart illustrating an audio signal fusion method according to an exemplary embodiment of this application is shown schematically. Figure 2 A flowchart illustrating the determination of weights for multiple audio signals according to an exemplary embodiment of this application is shown. The following is in conjunction with... Figure 2 right Figure 1 The embodiments shown will be described in detail.

[0035] S110, determine the short-time frequency domain representation of each audio signal corresponding to multiple devices.

[0036] In an exemplary embodiment, such as Figure 2 As shown, assume there are m microphone devices collecting audio signals, and these m microphone devices are: device 1, device 2, ..., device m, each device corresponding to one audio signal. The process of fusing the various audio signals is described in detail below.

[0037] In an exemplary embodiment, such as Figure 2 As shown, the audio signal processing for each device is independent. In this embodiment, the processing of the audio signal collected by device 1 is taken as an example. First, the audio signal is segmented into frames, and then the segmented audio signal is subjected to a Short Time Fourier Transform (STFT) to divide each audio signal (long-duration signal) into several shorter time-domain frame signals of equal length. Then, the segmented signals are subjected to windowing processing and Fast Fourier Transform (FFT) processing, thereby converting each audio signal from its original representation in the time domain to its representation in the time-frequency domain, that is, obtaining the short time-frequency domain representation of each audio signal.

[0038] S120: Based on the short-time frequency domain representation of each audio signal, determine the frequency response of the time-varying linear prediction error filter for each audio signal, and calculate the short-time spectrum of the linear prediction error for each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter.

[0039] As a specific implementation of step S120, "determine the frequency response of the time-varying linear prediction error filter for each audio signal based on the short-time frequency domain representation of each audio signal", Figure 3 A flowchart for determining the frequency response of linear prediction error is shown below. Figure 3 The illustrated embodiments are described in detail below; please refer to them. Figure 3 .

[0040] S310: Determine the sum of squares of the real and imaginary parts of each frequency point in the target frame represented in the short time frequency domain to obtain the power spectrum of each audio signal in the target frame.

[0041] In an exemplary embodiment, the time-domain signal is typically represented by real numbers, and after a Fast Fourier Transform, it is converted into a time-frequency domain signal represented by complex numbers. For example... Figure 2 As shown, the next step is to calculate the complex value corresponding to each frequency point in the target frame of the time-frequency domain signal, the sum of the squares of the real and imaginary parts, and thus obtain the power spectrum corresponding to the audio signal.

[0042] S320 performs an inverse fast Fourier transform on the power spectrum to obtain the autocorrelation function of each audio signal in the target frame.

[0043] In the exemplary embodiments, reference continues to be made to... Figure 2 After obtaining the power spectrum of each audio signal, we perform inverse fast fourier transform (IFFT) processing on them to obtain the autocorrelation function (used to characterize the degree of correlation between the values ​​of the same sequence at different times) of each audio signal in the target frame, which is denoted as r(n).

[0044] S330 determines the time-varying linear prediction error filter coefficients for each audio signal based on the autocorrelation function.

[0045] In an exemplary embodiment, the first p+1 values ​​of the autocorrelation function r(n) in the target frame are selected as r(0), r(1)...r(p). The Levinson-Durbin algorithm is then used to solve for the selected autocorrelation function, and the corresponding p-order linear prediction coefficients c are obtained, which can be denoted as c(1), c(2)...c(p).

[0046] In an exemplary embodiment, after obtaining the linear prediction coefficients, the time-varying linear prediction error filter coefficients h of length p+1 can be obtained. The first term of the time-varying linear prediction error filter coefficient h is 1, and the remaining terms take the negative values ​​(opposite numbers) of the above p-order linear prediction coefficients, that is, it can be written as h(0)=1, h(1)=-c(1), ..., h(p)=-c(p).

[0047] S340 performs a fast Fourier transform on the coefficients of the time-varying linear prediction error filter to obtain the frequency response of the time-varying linear prediction error filter for each audio signal.

[0048] In an exemplary embodiment, it is assumed that the length of each frame of the segmented audio signal is n, meaning there are n frequency points in total. For example... Figure 2As shown, the n-point fast Fourier transform of the coefficients of the p+1 order time-varying linear prediction error filter can be performed to obtain the frequency response of the time-varying linear prediction error filter corresponding to each audio signal, which is denoted as H(f), and the number of its frequency points is also n.

[0049] In an exemplary embodiment, such as Figure 2 As shown, after execution Figure 3 After obtaining the frequency response H(f) of the time-varying linear prediction error filter through the steps shown, the complex coefficients of each frequency point in the short-time frequency domain representation of the audio signal are multiplied correspondingly with the complex coefficients of each frequency point in the frequency response of the time-varying linear prediction error filter to obtain the short-time spectrum of the linear prediction error corresponding to each audio signal. The audio signal collected by each microphone device can obtain the corresponding short-time spectrum of the linear prediction error through processing step A. That is, for device 2, device 3, ..., device m, the process of obtaining the short-time spectrum of the linear prediction error is the same as that for device 1, and will not be repeated here.

[0050] S130: Extract low-frequency coefficients from the short-time spectrum of the linear prediction error and recombine them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal, and determine the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope.

[0051] In an exemplary embodiment, after obtaining the short-time spectra of the linear prediction errors corresponding to the two audio signals, it is necessary to determine the downsampling rate required to calculate the approximate envelope, and calculate the indices of the frequency points to be extracted from the short-time spectrum of the linear prediction error based on the number of frequency points and the downsampling rate. Then, based on the indices of the frequency points to be extracted, the coefficients of the corresponding frequency points are extracted from the short-time spectrum of the linear prediction error, and the short-time spectrum of the linear prediction error envelope is reassembled.

[0052] In an exemplary embodiment, after obtaining the short-time spectrum of the linear prediction error envelope, performing an inverse fast Fourier transform on it yields the linear prediction error envelope. For example, since the short-time spectrum of the linear prediction error envelope is frame-by-frame, the time-domain signal of the m-th linear prediction error envelope obtained after performing an inverse fast Fourier transform on the audio signal of the m-th device can be expressed as:

[0053] e m [k]=∑ l e m [k-Cl,l],

[0054] in,

[0055]

[0056] The short-time spectrum S corresponding to the linear prediction error envelope recombined in the l-th frame m The time-domain signal is [k, l], where C represents the number of short-time spectral points extracted above, and is also the time-domain signal e after the inverse fast Fourier transform. m The effective length of each frame in [k, l]. When k takes values ​​from 1 to C, for S... m Perform an inverse fast Fourier transform on [k, l] to obtain the time-domain signal of the linear prediction error envelope of the l-th frame; when the value of k is not in the range of 1 to C, the time-domain signal of the linear prediction error envelope of the l-th frame is 0.

[0057] In this exemplary embodiment, the operation process is more computationally efficient than traditional envelope extraction methods. Linear prediction error is used to reconstruct the original input audio signal. However, the original audio signal undergoes multiple reflections during propagation, which can introduce errors into the reconstruction process. Linear prediction error minimizes these errors.

[0058] S140 performs nonlinear compression processing on the linear prediction error envelope to obtain nonlinear compressed signals for each audio signal.

[0059] In an exemplary embodiment, during the nonlinear compression processing stage, the average energy μ of the linear prediction error envelope is first calculated frame by frame. m [l], where

[0060] μ m [l] = mean(e m [k, l]), k = 1, ..., C,

[0061] The mean function is used to calculate the average of a sample. Then, for μ... m [l] Perform exponential smoothing to obtain the updated average energy of each frame. in,

[0062]

[0063] λ is the forgetting factor, taking the value of a positive decimal slightly less than 1, such as 0.9. The first term in the summation above... This represents the forgetting effect on the value at the previous moment. The second term (1-λ)μ in the summation above represents this effect. m [l] represents the update effect of the initial observation at the current moment. The exponential smoothing process is completed by adding the two items together.

[0064] In an exemplary embodiment, assuming the current frame is the l-th frame, the linear prediction error envelope signal e for the l-th frame is... m Perform a logarithmic transformation on [k, l], i.e.

[0065]

[0066] and from Subtract the updated average energy of the current frame from the middle. Then, the exponential function of the subtraction result is calculated to obtain the amplitude-compensated linear prediction error envelope signal. That is, the amplitude-compensated linear prediction error envelope signal. The calculation method is as follows

[0067]

[0068] exp represents an exponential function with the natural constant e as its base.

[0069] In an exemplary embodiment, the amplitude-compensated linear prediction error envelope signal Compared with the original linear prediction error envelope signal e m Compared to [k], this eliminates differences in microphone devices and distance from the speaker, thus allowing for the extraction of more robust metrics.

[0070] In an exemplary embodiment, the amplitude-compensated linear prediction error envelope signal is finally calculated. The cube root of the given value yields the nonlinear compressed signal, i.e., the nonlinear compressed signal y. m The calculation method for [k] is as follows:

[0071]

[0072] y m [k]=∑ l y m [k-Cl,l].

[0073] S150 determines the weights of each audio signal based on the variance of the nonlinear compressed signal, and then fuses the audio signals according to the weights.

[0074] In an exemplary embodiment, the nonlinear compressed signal y is... m In the variance input weight adjuster of [k], the variance of each audio signal is...

[0075] V m [l] = var(y m [k, l]), k = 1, ..., C,

[0076] The `var` function calculates the variance of a given sample and then fuses the audio signals based on the weights output by the weight adjuster.

[0077] In an exemplary embodiment, such as Figure 2As shown, after processing step A, the audio signals acquired by each device produce a variance of the corresponding nonlinear compressed signal (the variance of the nonlinear compressed signal has been normalized). The variance of the nonlinear compressed signal can be used as a feature of each audio signal. The larger the variance of the nonlinear compressed signal, the higher the quality of the original audio signal, i.e., the higher the signal-to-noise ratio and signal-to-mixing ratio. After inputting the variances of these m nonlinear compressed signals into the weight adjuster, the weight adjuster can assign higher weights to devices with larger variances of the nonlinear compressed signals (features of each audio signal) to increase the proportion of high-quality audio signals in the weighted fused audio signal, thereby improving the fusion effect of multiple audio signals.

[0078] In the exemplary embodiment, two problems exist in reality: first, rapid changes in weights can introduce noise; second, the speaker may change at any time. Therefore, in addition to the aforementioned role of assigning weights, the comparison decision unit can also add two corresponding functions: first, controlling the rate of change of weights, that is, changing the weights relatively smoothly when the acquired audio signals change; second, in the stage when no one is speaking, averaging the audio signals acquired by all devices to prevent a serious deterioration in the quality of the fused audio due to the weights not updating in time when the speaker switches.

[0079] Compared to related technologies that directly add up the audio signals and assign weights based on the signal-to-noise ratio of multiple audio signals, the audio signal fusion method provided in this application can reflect the quality differences of audio signals collected by different devices. By calculating a set of weight values ​​through linear prediction error, higher weights are assigned to audio signals with high signal-to-noise ratio and high signal-to-mixing ratio, which can significantly improve the quality of the fused signal, is less susceptible to the effects of reverberation, reduces the probability of weight assignment errors, and requires less computation.

[0080] Furthermore, in distributed wireless microphone applications, the transmission delay of multiple devices is not negligible and needs to be estimated. Using linear prediction error to estimate signal delay is a reliable and low-cost solution. Therefore, the estimation results of linear prediction error can be reused when estimating the delay of the acquired audio signal.

[0081] Furthermore, a deep learning model can be designed and trained using the audio signal fusion method provided in this application. The trained model can achieve end-to-end audio signal enhancement and fusion, that is, it can directly output a fused audio signal with improved quality from the input audio signals collected by multiple devices.

[0082] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0083] in, Figure 4 A structural diagram of an audio signal fusion apparatus according to an exemplary embodiment of this application is shown.

[0084] The audio signal fusion device 400 in this embodiment includes: a first determining module 410, a second determining module 420, a third determining module 430, a nonlinear compression module 440, and a fusion module 450, wherein:

[0085] The first determining module 410 is used to: determine the short-time frequency domain representation of each audio signal corresponding to multiple devices.

[0086] The second determining module 420 is used to: determine the frequency response of the time-varying linear prediction error filter for each audio signal based on the short-time frequency domain representation of each audio signal, and calculate the short-time spectrum of the linear prediction error of each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter.

[0087] The third determining module 430 is used to: extract low-frequency coefficients from the short-time spectrum of the linear prediction error and recombine them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal, and determine the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope.

[0088] The nonlinear compression module 440 is used to perform nonlinear compression processing on the linear prediction error envelope to obtain nonlinear compressed signals for each audio signal.

[0089] The fusion module 450 is used to: determine the weights of each audio signal based on the variance of the nonlinear compressed signal, and fuse the audio signals according to the weights.

[0090] Figure 5 A structural diagram of an audio signal fusion apparatus according to another exemplary embodiment of this application is shown.

[0091] The aforementioned first determining module 410 includes: a framing unit 4101 and a first transformation unit 4102. The framing unit 4101 is used to: framing each audio signal to obtain a time-domain framing signal corresponding to each audio signal; the transformation unit 4102 is used to: window and perform fast Fourier transform on the time-domain framing signal to obtain a short-time frequency domain representation of each audio signal.

[0092] The second determining module 420 includes: a first determining unit 4201, a second transforming unit 4202, a second determining unit 4203, and a third transforming unit 4204. The first determining unit 4201 is used to: determine the sum of squares of the real and imaginary parts of each frequency point in the target frame represented in the short time-frequency domain, thereby obtaining the power spectrum of each audio signal in the target frame; the second transforming unit 4202 is used to: perform an inverse fast Fourier transform on the power spectrum to obtain the autocorrelation function of each audio signal in the target frame; the second determining unit 4203 is used to: determine the time-varying linear prediction error filter coefficients of each audio signal based on the autocorrelation function; the third transforming unit 4204 is used to: perform a fast Fourier transform on the time-varying linear prediction error filter coefficients to obtain the frequency response of the time-varying linear prediction error filter for each audio signal.

[0093] The second determining unit 4203 is specifically used to: select the first p+1 values ​​of the autocorrelation function of each audio signal in the target frame, and determine the p-th order linear prediction coefficients of each audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; take the negative of the p-th order linear prediction coefficients and add the first term 1 to obtain the time-varying linear prediction error filter coefficients of each audio signal with a length of p+1.

[0094] The second determining module 420 mentioned above includes a third determining unit 4205. The third determining unit 4205 is used to: multiply the complex coefficients of each frequency point in the short-time frequency domain representation with the corresponding complex coefficients of each frequency point in the frequency response of the time-varying linear prediction error filter to obtain the short-time spectrum of the linear prediction error of each audio signal.

[0095] The aforementioned third determining module 430 includes: a fourth determining unit 4301 and a coefficient extraction unit 4302. The fourth determining unit 4301 is used to: determine the downsampling rate required for the approximate envelope, and determine the index of the frequency point to be extracted in the short-time spectrum of the linear prediction error based on the number of frequency points and the downsampling rate of the short-time spectrum of the linear prediction error; the coefficient extraction unit 4302 is used to: extract the coefficients of the corresponding frequency points from the short-time spectrum of the linear prediction error based on the index of the frequency points to be extracted, and recombine them to form the short-time spectrum of the linear prediction error envelope.

[0096] The third determining module 430 mentioned above includes a fourth transformation unit 4303. The fourth transformation unit 4303 is used to perform an inverse fast Fourier transform on the short-time spectrum of the linear prediction error envelope to obtain the linear prediction error envelope.

[0097] The aforementioned nonlinear compression module 440 includes: an exponential smoothing unit 4401, an amplitude compensation unit 4402, and a nonlinear compressed signal determination unit 4403. The exponential smoothing unit 4401 is used to: calculate the average energy of the linear prediction error envelope frame by frame, and perform exponential smoothing on the average energy of the linear prediction error envelope to obtain the updated average energy of the current frame; the amplitude compensation unit 4402 is used to: subtract the updated average energy of the current frame from the logarithmic transform of the linear prediction error envelope of the current frame to obtain the subtraction result, and calculate the exponential function of the subtraction result to obtain the amplitude-compensated linear prediction error envelope; the nonlinear compressed signal determination unit 4403 is used to: calculate the cube root of the amplitude-compensated linear prediction error envelope to obtain the nonlinear compressed signal.

[0098] The aforementioned fusion module 450 includes: a variance determination unit 4501 and a fusion unit 4502. The variance determination unit 4501 is used to: calculate the variance of the nonlinear compressed signal to obtain the variance of each audio signal; the fusion unit 4502 is used to: input the variance of each audio signal into a weight adjuster, and fuse the audio signals according to the weights output by the weight adjuster.

[0099] The aforementioned device also includes: a feature determination module 460.

[0100] The aforementioned feature determination module 460 is used to: use the variance of the nonlinear compressed signal as a feature of each audio signal. In an exemplary embodiment, based on the aforementioned scheme, the weight adjuster is used to: assign corresponding weights to each audio signal according to the distribution of the features of each audio signal in each audio signal; and limit the rate of change of the weights of each audio signal.

[0101] It should be noted that the audio signal fusion apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when performing the audio signal fusion method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio signal fusion apparatus and the audio signal fusion method embodiments provided in the above embodiments belong to the same concept. Therefore, for details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the audio signal fusion method of this application, which will not be repeated here.

[0102] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0103] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, as well as magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0104] This application also provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of any of the methods described above.

[0105] Figure 6 This diagram schematically illustrates the structure of a terminal according to an exemplary embodiment of this application. Please refer to... Figure 6 As shown, terminal 600 includes a processor 601 and a memory 602.

[0106] In this embodiment, processor 601 is the control center of the computer system, and can be a processor of a physical machine or a processor of a virtual machine. Processor 601 may include one or more processing cores, such as a 4-core processor or an 8-core processor. Processor 601 can be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 601 may also include a main processor and coprocessors. The main processor is used to process data in the wake-up state, also known as the Central Processing Unit (CPU); the coprocessor is a low-power processor used to process data in the standby state.

[0107] In this embodiment of the application, the processor 601 is specifically used for:

[0108] The process involves: determining the short-time frequency domain representation of each audio signal corresponding to multiple devices; determining the frequency response of the time-varying linear prediction error filter for each audio signal based on the short-time frequency domain representation; calculating the short-time spectrum of the linear prediction error for each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter; extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal; determining the linear prediction error envelope for each audio signal based on the short-time spectrum of the linear prediction error envelope; performing nonlinear compression processing on the linear prediction error envelope to obtain the nonlinear compressed signal for each audio signal; determining the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal; and fusing the audio signals based on the weights.

[0109] Furthermore, in one embodiment of this application, the above-mentioned audio signals are framed to obtain time-domain framed signals corresponding to the above-mentioned audio signals; the time-domain framed signals are windowed and subjected to fast Fourier transform to obtain short-time frequency domain representations of the above-mentioned audio signals.

[0110] Optionally, determining the frequency response of the time-varying linear prediction error filter for each audio signal based on its short-time frequency domain representation includes: determining the sum of squares of the real and imaginary parts of each frequency point in the target frame of the short-time frequency domain representation to obtain the power spectrum of each audio signal in the target frame; performing an inverse fast Fourier transform on the power spectrum to obtain the autocorrelation function of each audio signal in the target frame; determining the coefficients of the time-varying linear prediction error filter for each audio signal based on the autocorrelation function; and performing a fast Fourier transform on the coefficients of the time-varying linear prediction error filter to obtain the frequency response of the time-varying linear prediction error filter for each audio signal.

[0111] Optionally, determining the time-varying linear prediction error filter coefficients of each audio signal based on the autocorrelation function includes: selecting the first p+1 values ​​of the autocorrelation function of each audio signal in the target frame, and determining the p-th order linear prediction coefficients of each audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; taking the negative of the p-th order linear prediction coefficients and adding the first term 1 to obtain the time-varying linear prediction error filter coefficients of each audio signal with a length of p+1.

[0112] Optionally, the above calculation of the short-time spectrum of the linear prediction error of each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter includes: multiplying the complex coefficients of each frequency point in the frequency domain representation with the corresponding complex coefficients of each frequency point in the frequency response of the time-varying linear prediction error filter to obtain the short-time spectrum of the linear prediction error of each audio signal.

[0113] Optionally, the above-mentioned extraction and recombination of low-frequency coefficients from the short-time spectrum of the linear prediction error to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal includes: determining the downsampling rate required for the approximate envelope, and determining the index of the frequency point to be extracted in the short-time spectrum of the linear prediction error based on the number of frequency points in the short-time spectrum of the linear prediction error and the downsampling rate; extracting the coefficients of the corresponding frequency points from the short-time spectrum of the linear prediction error based on the index of the frequency points to be extracted, and recombinating the short-time spectrum of the linear prediction error envelope.

[0114] Optionally, determining the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope includes: performing an inverse fast Fourier transform on the short-time spectrum of the linear prediction error envelope to obtain the linear prediction error envelope.

[0115] Optionally, the nonlinear compression processing of the linear prediction error envelope to obtain the nonlinear compressed signal of each audio signal includes: calculating the average energy of the linear prediction error envelope frame by frame, and performing exponential smoothing on the average energy of the linear prediction error envelope to obtain the updated average energy of the current frame; subtracting the updated average energy of the current frame from the logarithmic transform of the linear prediction error envelope of the current frame to obtain the subtraction result, and calculating the exponential function of the subtraction result to obtain the amplitude-compensated linear prediction error envelope; and calculating the cube root of the amplitude-compensated linear prediction error envelope to obtain the nonlinear compressed signal.

[0116] Optionally, the above-mentioned determination of the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal, and the fusion of the audio signals based on the weights, includes: calculating the variance of the nonlinear compressed signal to obtain the variance of each audio signal; inputting the variance of each audio signal into a weight adjuster, and fusing the audio signals based on the weights output by the weight adjuster.

[0117] Optionally, the above method further includes: using the variance of the nonlinear compressed signal as a feature of each audio signal; the weight adjuster is used to: assign corresponding weights to each audio signal according to the distribution of the features of each audio signal in each audio signal; and limit the rate of change of the weights of each audio signal.

[0118] Memory 602 may include one or more computer-readable storage media, which may be non-transitory. Memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage terminals or flash memory terminals. In some embodiments of this application, the non-transitory computer-readable storage media in memory 602 is used to store at least one instruction, which is executed by processor 601 to implement the method in the embodiments of this application.

[0119] In some embodiments, terminal 600 further includes a peripheral terminal interface 603 and at least one peripheral terminal. The processor 601, memory 602, and peripheral terminal interface 603 can be connected via a bus or signal line. Each peripheral terminal can be connected to peripheral terminal interface 603 via a bus, signal line, or circuit board. Specifically, the peripheral terminal includes at least one of a display screen 604, a camera 605, and an audio circuit 606.

[0120] The peripheral terminal interface 603 can be used to connect at least one input / output (I / O) related peripheral terminal to the processor 601 and the memory 602. In some embodiments of this application, the processor 601, memory 602, and peripheral terminal interface 603 are integrated on the same chip or circuit board; in some other embodiments of this application, any one or two of the processor 601, memory 602, and peripheral terminal interface 603 can be implemented on separate chips or circuit boards. This application does not specifically limit this aspect.

[0121] Display screen 604 is used to display a user interface (UI). The UI may include graphics, text, icons, videos, and any combination thereof. When display screen 604 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 601 for processing. In this case, display screen 604 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments of this application, there may be one display screen 604, which serves as the front panel of terminal 600; in other embodiments, there may be at least two display screens 604, respectively disposed on different surfaces of terminal 600 or in a folded design; in still other embodiments, display screen 604 may be a flexible display screen, disposed on a curved or folded surface of terminal 600. Furthermore, display screen 604 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 604 may be made of materials such as liquid crystal display (LCD) or organic light-emitting diode (OLED).

[0122] Camera 605 is used to capture images or videos. Optionally, camera 605 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, virtual reality (VR) shooting, or other fusion shooting functions. In some embodiments of this application, camera 605 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.

[0123] The audio circuit 606 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input to the processor 601 for processing. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 600. The microphone may also be an array microphone or an omnidirectional microphone.

[0124] Power supply 607 is used to power the various components in terminal 600. Power supply 607 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 607 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0125] The terminal structure block diagram shown in the embodiments of this application does not constitute a limitation on the terminal 600. The terminal 600 may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0126] In this application, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or order; the term "multiple" refers to two or more unless otherwise expressly defined. The terms "install," "connect," "link," "fix," etc., should be interpreted broadly. For example, "connect" can be a fixed connection, a detachable connection, or an integral connection; "link" can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0127] In the description of this application, it should be understood that the terms "upper" and "lower" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or unit referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0128] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, equivalent variations made in accordance with the claims of this application still fall within the scope of this application.

Claims

1. A method for fusing audio signals, characterized in that, include: Determine the short-time frequency domain representation of each audio signal corresponding to multiple devices; Based on the short-time frequency domain representation of each audio signal, the frequency response of the time-varying linear prediction error filter for each audio signal is determined, and the short-time spectrum of the linear prediction error of each audio signal is calculated based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter. Low-frequency coefficients are extracted from the short-time spectrum of the linear prediction error and recombined to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal, and the linear prediction error envelope of each audio signal is determined based on the short-time spectrum of the linear prediction error envelope. The linear prediction error envelope is subjected to nonlinear compression processing to obtain the nonlinear compressed signal of each audio signal. The weights corresponding to each audio signal are determined based on the variance of the nonlinear compressed signal, and the audio signals are fused according to the weights.

2. The audio signal fusion method according to claim 1, characterized in that, The determination of the short-time frequency domain representation of each audio signal corresponding to multiple devices includes: Each audio signal is divided into frames to obtain the time-domain framed signal corresponding to each audio signal. Windowing and Fast Fourier Transform are applied to the time-domain framed signals to obtain the short-time frequency domain representation of each audio signal.

3. The audio signal fusion method according to claim 1, characterized in that, The step of determining the frequency response of the time-varying linear prediction error filter for each audio signal based on its short-time frequency domain representation includes: In the target frame represented by the short time frequency domain, the sum of squares of the real and imaginary parts of each frequency point is determined to obtain the power spectrum of each audio signal in the target frame; Perform an inverse fast Fourier transform on the power spectrum to obtain the autocorrelation function of each audio signal in the target frame; The time-varying linear prediction error filter coefficients for each audio signal are determined based on the autocorrelation function. The frequency response of the time-varying linear prediction error filter for each audio signal is obtained by performing a fast Fourier transform on the coefficients of the time-varying linear prediction error filter.

4. The audio signal fusion method according to claim 3, characterized in that, The step of determining the time-varying linear prediction error filter coefficients of each audio signal based on the autocorrelation function includes: Select the first p+1 values ​​of the autocorrelation function of each audio signal in the target frame, and determine the p-th order linear prediction coefficient of each audio signal based on the first p+1 values ​​of the autocorrelation function, where p is a positive integer; By taking the negative of the p-th order linear prediction coefficients and adding the first term 1, the time-varying linear prediction error filter coefficients with a length of p+1 for each audio signal are obtained.

5. The method for fusing audio signals according to any one of claims 1 to 4, characterized in that, The step of calculating the short-time spectrum of the linear prediction error of each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter includes: The short-time spectrum of the linear prediction error of each audio signal is obtained by multiplying the complex coefficients of each frequency point in the short-time frequency domain representation with the corresponding complex coefficients of each frequency point in the frequency response of the time-varying linear prediction error filter.

6. The method for fusing audio signals according to any one of claims 1 to 4, characterized in that, The step of extracting low-frequency coefficients from the short-time spectrum of the linear prediction error and recombining them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal includes: Determine the downsampling rate required for the approximate envelope, and determine the index of the frequency point to be extracted in the short-time spectrum of the linear prediction error based on the number of frequency points in the short-time spectrum of the linear prediction error and the downsampling rate; Based on the index of the frequency point to be extracted, the coefficients of the corresponding frequency point are extracted from the short-time spectrum of the linear prediction error, and the short-time spectrum of the linear prediction error envelope is recombined.

7. The method for fusing audio signals according to any one of claims 1 to 4, characterized in that, Determining the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope includes: The linear prediction error envelope is obtained by performing an inverse fast Fourier transform on the short-time spectrum of the linear prediction error envelope.

8. The method for fusing audio signals according to any one of claims 1 to 4, characterized in that, The nonlinear compression processing of the linear prediction error envelope to obtain the nonlinear compressed signal of each audio signal includes: The average energy of the linear prediction error envelope is calculated frame by frame, and the average energy of the linear prediction error envelope is exponentially smoothed to obtain the updated average energy of the current frame. Subtract the updated average energy of the current frame from the logarithmic transform of the current frame of the linear prediction error envelope to obtain the subtraction result, and calculate the exponential function of the subtraction result to obtain the amplitude-compensated linear prediction error envelope. The cube root of the linear prediction error envelope of the amplitude compensation is calculated to obtain the nonlinear compressed signal of each audio signal.

9. The audio signal fusion method according to claim 1, characterized in that, The step of determining the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal, and fusing the audio signals according to the weights, includes: Calculate the variance of the nonlinear compressed signal for each audio signal; The variance of the nonlinear compressed signal is input into a weight adjuster, and the audio signals are fused according to the weights output by the weight adjuster.

10. The audio signal fusion method according to claim 9, characterized in that, The method further includes: using the variance of the nonlinear compressed signal as a feature of each audio signal; The weight adjuster is used to assign corresponding weights to each audio signal based on the distribution of the characteristics of each audio signal in each audio signal. Limit the rate of change of the weights of the audio signals.

11. An audio signal fusion device, characterized in that, include: The first determining module is used to: determine the short-time frequency domain representation of each audio signal corresponding to multiple devices; The second determining module is used to: determine the frequency response of the time-varying linear prediction error filter of each audio signal based on the short-time frequency domain representation of each audio signal, and calculate the short-time spectrum of the linear prediction error of each audio signal based on the short-time frequency domain representation and the frequency response of the time-varying linear prediction error filter. The third determining module is used to: extract low-frequency coefficients from the short-time spectrum of the linear prediction error and recombine them to determine the short-time spectrum of the linear prediction error envelope corresponding to each audio signal, and determine the linear prediction error envelope of each audio signal based on the short-time spectrum of the linear prediction error envelope. A nonlinear compression module is used to: perform nonlinear compression processing on the linear prediction error envelope to obtain nonlinear compressed signals for each audio signal; The fusion module is used to: determine the weights corresponding to each audio signal based on the variance of the nonlinear compressed signal, and fuse the audio signals according to the weights.

12. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for fusing audio signals as described in any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for fusing audio signals as described in any one of claims 1 to 10.