Audio signal processing method, device and storage medium

By judging the reversibility of the covariance matrix in the blind source separation technology, and using the current or previous frame matrix for sound source separation, the problem of insufficient sound source separation effect and stability of blind source separation technology is solved, and the speech recognition rate is improved.

CN114724578BActive Publication Date: 2025-08-26BEIJING XIAOMI PINECONE ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110015417.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-04
Publication Date
2025-08-26
Estimated Expiration
2041-01-04

AI Technical Summary

Technical Problem

The existing blind source separation technology has shortcomings in the sound source separation effect and stability, especially in complex environments, the speech recognition rate is not high.

Method used

By obtaining the covariance matrix of the multi-frame audio time domain signal, judging its reversibility degree, and using the current frame matrix when it meets the set degree. Otherwise, the current frame matrix is ​​updated with the previous frame matrix and the intermediate matrix is ​​calculated for sound source separation.

Benefits of technology

It improves the separation effect of blind source signals, enhances the robustness and stability of the algorithm, reduces voice damage, and improves recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724578B_ABST
    Figure CN114724578B_ABST
Patent Text Reader

Abstract

This article discloses an audio signal processing method, device and storage medium, which includes: determining the first covariance matrix of the first sound source at each frequency point of the current frame audio time domain signal and the second covariance matrix of the second sound source at each frequency point; judging whether the reversibility of the first covariance matrix meets the set degree; when the reversibility of the first covariance matrix meets the set degree, determining the first covariance matrix as the first covariance matrix of the current frame audio time domain signal; when the reversibility of the first covariance matrix does not meet the set degree, updating the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal. The present disclosure can improve the separation effect of blind source signals, improve the robustness and stability of the algorithm, and enhance the separation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article relates to the field of mobile terminal data processing technology, and in particular to an audio signal processing method, device and storage medium. Background Art

[0002] In the era of the Internet of Things and AI, intelligent voice, as one of the core technologies of artificial intelligence, can effectively improve the mode of human-computer interaction and greatly improve the convenience of using smart products.

[0003] Currently, the sound collection devices of smart products and equipment mostly use microphone arrays, and apply microphone beamforming technology to improve the quality of voice signal processing, so as to improve the voice recognition rate in real environments.

[0004] Blind source separation technology uses the independence between different sound source signals to separate the sound sources, thereby separating the target signal and the noise source signal and improving the signal-to-noise ratio.

[0005] How to improve the performance of blind source separation technology is a technical problem that needs to be solved. Summary of the Invention

[0006] To overcome the problems existing in the related art, this article provides an audio signal processing method, device and storage medium.

[0007] According to a first aspect of the embodiments of this document, a method for processing an audio signal is provided, the method comprising:

[0008] Acquire aliased audio signals of at least two sound sources collected by at least two microphones;

[0009] Performing frame processing on the mixed audio signal to obtain a multi-frame audio time domain signal;

[0010] The following processing is performed on each frame of audio time domain signal:

[0011] Determine a first covariance matrix of a first sound source at each frequency point and a second covariance matrix of a second sound source at each frequency point of a current frame of audio time domain signal;

[0012] Determining whether the reversibility of the first covariance matrix meets a set degree;

[0013] When the reversibility of the first covariance matrix satisfies a set degree, determining the first covariance matrix as the first covariance matrix of the current frame audio time domain signal; when the reversibility of the first covariance matrix does not satisfy the set degree, updating the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal;

[0014] Calculate an intermediate matrix using the inverse matrix of the first covariance matrix and the second covariance matrix;

[0015] Calculating a separation matrix based on the intermediate matrix;

[0016] The separation matrix is ​​used to separate audio signals of different sound sources from the current frame audio time domain signal.

[0017] In one embodiment, updating the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal includes one of the following:

[0018] Using the first covariance matrix of the previous frame of audio time domain signal as the first covariance matrix of the current frame of audio time domain signal;

[0019] Determine a product matrix of the first covariance matrix of the previous frame audio time domain signal and the coefficient matrix, and use the product matrix as the first covariance matrix of the current frame audio time domain signal.

[0020] In one embodiment, determining whether the reversibility of the first covariance matrix satisfies a set degree includes:

[0021] Determine the auxiliary matrix corresponding to the first covariance matrix using an inversion formula;

[0022] Determining a product matrix of the first covariance matrix and the auxiliary matrix;

[0023] Determining a first difference value between the product matrix and the identity matrix;

[0024] When the first gap value is less than or equal to a set threshold, it is determined that the reversibility degree of the first covariance matrix meets the set degree.

[0025] In one embodiment, determining the auxiliary matrix corresponding to the first covariance matrix using an inversion formula includes:

[0026] determining an adjoint matrix of the first covariance matrix, and determining a determinant of the first covariance matrix;

[0027] Determining a ratio of the adjoint matrix to the determinant;

[0028] The ratio result is used as the auxiliary matrix corresponding to the first covariance matrix.

[0029] In one embodiment, determining a first difference value between the product matrix and the identity matrix includes:

[0030] Determine the absolute value of the difference between each element on the main diagonal of the product matrix and 1,

[0031] determining the absolute value of each element of the product matrix that is outside the main diagonal;

[0032] Determine the sum of the absolute values;

[0033] The sum is used as a first difference value between the product matrix and the identity matrix.

[0034] In one embodiment, the method further comprises:

[0035] Determine a first gap value corresponding to multiple historical frame audio time domain signals before the current frame audio time domain signal, determine a first coefficient based on the first gap values ​​corresponding to the multiple historical frame audio time domain signals, and determine that the set threshold is the product of the first fixed value and the first coefficient.

[0036] In one embodiment, determining the first coefficient according to the first gap values ​​corresponding to the audio time domain signals of the plurality of historical frames includes:

[0037] Determine the difference between a first gap value corresponding to each historical frame audio time domain signal and a first fixed value, determine an average value of the difference corresponding to each historical frame audio time domain signal, and determine a first coefficient based on the average value, where the average value is positively correlated with the first coefficient.

[0038] According to a second aspect of the embodiments of this document, there is provided an audio signal processing apparatus, including:

[0039] an acquisition module configured to acquire aliased audio signals of at least two sound sources collected by at least two microphones;

[0040] A framing module is configured to perform framing processing on the mixed audio signal to obtain a multi-frame audio time domain signal;

[0041] A processing module, configured to process each frame of audio time domain signal;

[0042] The processing module includes:

[0043] A first determining module is configured to determine a first covariance matrix of a first sound source at each frequency point and a second covariance matrix of a second sound source at each frequency point of a current frame of audio time domain signal;

[0044] a judging module configured to judge whether the invertibility degree of the first covariance matrix satisfies a set degree;

[0045] A second determining module is configured to, when the reversibility degree of the first covariance matrix satisfies a set degree, determine that the first covariance matrix is ​​the first covariance matrix of the current frame audio time domain signal; when the reversibility degree of the first covariance matrix does not satisfy the set degree, update the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal;

[0046] A third determining module is configured to calculate an intermediate matrix using the inverse matrix of the first covariance matrix and the second covariance matrix; and calculate a separation matrix according to the intermediate matrix;

[0047] The separation module is configured to use the separation matrix to separate audio signals of different sound sources from the current frame audio time domain signal.

[0048] In one embodiment, the second determining module is further configured to update the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal using one of the following methods:

[0049] Using the first covariance matrix of the previous frame of audio time domain signal as the first covariance matrix of the current frame of audio time domain signal;

[0050] Determine a product matrix of the first covariance matrix of the previous frame audio time domain signal and the coefficient matrix, and use the product matrix as the first covariance matrix of the current frame audio time domain signal.

[0051] In one embodiment, the judgment module includes:

[0052] a fourth determining module, configured to determine an auxiliary matrix corresponding to the first covariance matrix using an inversion formula;

[0053] a fifth determining module, configured to determine a product matrix of the first covariance matrix and the auxiliary matrix;

[0054] a sixth determining module, configured to determine a first difference value between the product matrix and the identity matrix;

[0055] The seventh determination module is configured to determine whether the reversibility of the first covariance matrix meets a set degree when the first gap value is less than or equal to a set threshold.

[0056] In one embodiment, the fourth determining module is further configured to determine the auxiliary matrix corresponding to the first covariance matrix using an inversion formula using the following method:

[0057] determining an adjoint matrix of the first covariance matrix, and determining a determinant of the first covariance matrix;

[0058] Determining a ratio of the adjoint matrix to the determinant;

[0059] The ratio result is used as the auxiliary matrix corresponding to the first covariance matrix.

[0060] In one embodiment, the sixth determining module is configured to determine the first difference value between the product matrix and the identity matrix using the following method:

[0061] Determine the absolute value of the difference between each element on the main diagonal of the product matrix and 1,

[0062] determining the absolute value of each element of the product matrix that is outside the main diagonal;

[0063] Determine the sum of the absolute values;

[0064] The sum is used as a first difference value between the product matrix and the identity matrix.

[0065] In one embodiment, the device further comprises:

[0066] The eighth determination module is configured to determine the first gap value corresponding to multiple historical frame audio time domain signals before the current frame audio time domain signal, determine the first coefficient based on the first gap value corresponding to the multiple historical frame audio time domain signals, and determine that the set threshold is the product of the first fixed value and the first coefficient.

[0067] In one embodiment, the eighth determination module is further configured to determine the first coefficient according to the first gap values ​​corresponding to the audio time domain signals of multiple historical frames using the following method:

[0068] Determine the difference between a first gap value corresponding to the audio time domain signal of each historical frame and a first fixed value, determine an average value of the difference corresponding to each historical frame, and determine a first coefficient based on the average value, where the average value is positively correlated with the first coefficient.

[0069] According to a third aspect of the embodiments of this document, there is provided an audio signal processing apparatus, comprising:

[0070] processor;

[0071] a memory for storing processor-executable instructions;

[0072] The processor is configured to execute the executable instructions in the memory to implement the steps of the method.

[0073] According to a fourth aspect of the embodiments of this document, a non-transitory computer-readable storage medium is provided, on which executable instructions are stored, and when the executable instructions are executed by a processor, the steps of the method are implemented.

[0074] The technical solution provided by the embodiments of this article may include the following beneficial effects: after calculating the first covariance matrix and the second covariance matrix of each frame of audio time domain signal, the reversibility of the first covariance matrix is ​​judged; when the reversibility meets the set degree, this first covariance matrix is ​​used as the first covariance matrix of the current frame of audio time domain signal; when the reversibility does not meet the set degree, the first covariance matrix of the previous frame of audio time domain signal is used to determine the first covariance matrix of the current frame of audio time domain signal, thereby improving the separation effect of blind source signals, improving the robustness and stability of the algorithm, improving the separation performance, reducing the degree of speech damage after separation, and improving the recognition performance.

[0075] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the present invention.

[0077] Figure 1 is a flowchart of an audio signal processing method according to an exemplary embodiment;

[0078] Figure 2 is a structural diagram of an audio signal processing device according to an exemplary embodiment;

[0079] Figure 3 is a structural diagram of an audio signal processing apparatus according to an exemplary embodiment. DETAILED DESCRIPTION

[0080] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this disclosure, as detailed in the appended claims.

[0081] The present disclosure provides an audio signal processing method for a terminal, which is an electronic device integrating two or more microphones. For example, the terminal may be a mobile phone, a laptop, a tablet computer, a vehicle-mounted terminal, a computer, or a server; or the terminal may be a device connected to multiple microphones.

[0082] Reference Figure 1 , Figure 1 FIG. 1 is a flow chart of an audio signal processing method according to an exemplary embodiment. Figure 1As shown, this method includes:

[0083] Step S11: Acquire aliased audio signals of at least two sound sources collected by at least two microphones.

[0084] Step S12: performing frame processing on the mixed audio signal to obtain multiple frames of audio time domain signals.

[0085] Step S13: Perform the following processing on each frame of the audio time domain signal:

[0086] Step S141, determining a first covariance matrix of a first sound source at each frequency point and a second covariance matrix of a second sound source at each frequency point of a current frame of audio time domain signal;

[0087] Step S142, determining whether the reversibility of the first covariance matrix meets a set degree;

[0088] Step S143: When the reversibility of the first covariance matrix satisfies a set degree, determining the first covariance matrix as the first covariance matrix of the current frame audio time domain signal; when the reversibility of the first covariance matrix does not satisfy the set degree, updating the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal;

[0089] Step S144, using the inverse matrix of the first covariance matrix and the second covariance matrix to calculate an intermediate matrix; and calculating a separation matrix based on the intermediate matrix;

[0090] Step S145: Use the separation matrix to separate the audio signals of different sound sources from the current frame audio time domain signal.

[0091] In this embodiment, after calculating the first covariance matrix and the second covariance matrix of each frame of audio time domain signal, the reversibility of the first covariance matrix is ​​judged. When the reversibility degree meets the set degree, this first covariance matrix is ​​used as the first covariance matrix of the current frame of audio time domain signal. When the reversibility degree does not meet the set degree, the first covariance matrix of the previous frame of audio time domain signal is used to determine the first covariance matrix of the current frame of audio time domain signal, thereby improving the separation effect of the blind source signal, improving the robustness and stability of the algorithm, improving the separation performance, reducing the degree of speech damage after separation, and improving the recognition performance.

[0092] In this embodiment, there are two or more microphones and two or more sound sources. Generally, the number of sound sources is the same as the number of microphones. In some embodiments, the number of sound sources and the number of microphones may be different.

[0093] In one application scenario, there are two microphones, microphone 1 and microphone 2, and two sound sources, sound source 1 and sound source 2. The aliased audio signal collected by microphone 1 is a mixture of the audio signals from sound source 1 and sound source 2, and the aliased audio signal collected by microphone 2 is also a mixture of the audio signals from sound source 1 and sound source 2.

[0094] In another application scenario, there are three microphones, namely microphone 1, microphone 2 and microphone 3; there are three sound sources, namely sound source 1, sound source 2 and sound source 3; then the aliased audio signals collected by microphone 1, microphone 2 and microphone 3 are the aliased audio signals of sound source 1, sound source 2 and sound source 3.

[0095] When the number of sound sources is greater than 2, it is generally considered that the number of sound sources is 2, that is, the audio signal of one sound source is used as the target audio signal, and the audio signals of other sound sources are used as interference target audio signals. Therefore, when separating the sound source signals in this embodiment, a total of two sound source signals are separated.

[0096] When the number of microphones is greater than two, during sound source separation, the signals collected by the multiple microphones are subjected to redundancy removal processing (or dimensionality reduction processing) to obtain aliased audio signals corresponding to the two microphones.

[0097] The present invention provides an audio signal processing method, which includes: Figure 1 The method shown in FIG5 , and: in step S14, the first covariance matrix of the audio time domain signal of the current frame is updated according to the first covariance matrix of the audio time domain signal of the previous frame, including one of the following:

[0098] 1. Using the first covariance matrix of the previous frame of audio time domain signal as the first covariance matrix of the current frame of audio time domain signal;

[0099] Second, determine a product matrix of the first covariance matrix of the previous frame audio time domain signal and the coefficient matrix, and use the product matrix as the first covariance matrix of the current frame audio time domain signal.

[0100] In this embodiment, when the reversibility degree of the first covariance matrix of the current frame audio time domain signal does not meet the set degree, the first covariance matrix of the current frame audio time domain signal is abandoned, and the first covariance matrix of the previous frame is used as the first covariance matrix of the current frame audio time domain signal, or the first covariance matrix of the previous frame audio time domain signal is corrected and used as the first covariance matrix of the current frame audio time domain signal, thereby obtaining a better separation effect than using the first covariance matrix of the current frame.

[0101] The present invention provides an audio signal processing method, which includes: Figure 1 The method shown, and:

[0102] The determining whether the reversibility of the first covariance matrix satisfies a set degree includes:

[0103] Step 1, using an inversion formula to determine the auxiliary matrix corresponding to the first covariance matrix;

[0104] Step 2, determining a product matrix of the first covariance matrix and the auxiliary matrix;

[0105] Step 3, determining a first difference value between the product matrix and the identity matrix;

[0106] Step 4: When the first gap value is less than or equal to a set threshold, determine whether the reversibility of the first covariance matrix meets the set degree.

[0107] In one embodiment, step 1 uses an inversion formula to determine the auxiliary matrix corresponding to the first covariance matrix, including:

[0108] determining an adjoint matrix of the first covariance matrix, and determining a determinant of the first covariance matrix;

[0109] Determining a ratio of the adjoint matrix to the determinant;

[0110] The ratio result is used as the auxiliary matrix corresponding to the first covariance matrix.

[0111] For example:

[0112] The first covariance matrix is ​​V1(k,n), where k refers to k=1,..,K, k represents the position identifier of the frequency point, the number of frequency points is K, where K=Nfft / 2+1, the system frame length is Nfft, and n represents the frame number.

[0113] When V1(k,n) is a 2*2 matrix,

[0114]

[0115] The auxiliary matrix calculated using the inversion formula shown below is invWtmp(k,n):

[0116]

[0117] In one embodiment, determining the first difference between the product matrix and the identity matrix in step 3 includes:

[0118] Determine the absolute value of the difference between each element on the main diagonal of the product matrix and 1,

[0119] determining the absolute value of each element of the product matrix that is outside the main diagonal;

[0120] Determine the sum of the absolute values;

[0121] The sum is used as a first difference value between the product matrix and the identity matrix.

[0122] For example:

[0123] The product matrix is ​​V dot (k,n), the first difference between it and the unit matrix is ​​amp1_V dot (k,n):

[0124] amp_V dot (k,n)=abs(V dot (1,1,k,n)-1)+abs(V dot (1,2,k,n))+abs(V dot (2,1,k,n))+abs(V dot (2,2,k,n)-1)

[0125] In this embodiment, the auxiliary matrix corresponding to the first covariance matrix is ​​calculated using an inversion formula. When the first covariance matrix is ​​highly reversible, the product matrix of the first covariance matrix and the corresponding auxiliary matrix is ​​closer to the identity matrix. When the first covariance matrix is ​​less reversible, the product matrix of the first covariance matrix and the corresponding auxiliary matrix is ​​more different from the identity matrix. The method in this embodiment can effectively determine the reversibility of the first covariance matrix.

[0126] In an embodiment of the present disclosure, a method for processing an audio signal is provided. The method includes the method shown in the previous embodiment, and the threshold value is set in one of the following ways:

[0127] Method 1: Set the threshold to a fixed value, for example, set the threshold to 1e-2.

[0128] Method 2: Setting the threshold value is an adjustable dynamic value.

[0129] For example:

[0130] Determine a first gap value corresponding to multiple historical frame audio time domain signals before the current frame audio time domain signal, determine a first coefficient based on the first gap values ​​corresponding to the multiple historical frame audio time domain signals, and determine that the set threshold is the product of the first fixed value and the first coefficient.

[0131] In one embodiment, a first coefficient is determined based on first gap values ​​corresponding to multiple historical frame audio time domain signals, including: determining the difference between the first gap value corresponding to each historical frame audio time domain signal and a first fixed value, determining the average value of the difference corresponding to each historical frame audio time domain signal, and determining the first coefficient based on the average value, wherein the average value is positively correlated with the first coefficient.

[0132] In this embodiment, the set threshold of the current frame is adjusted accordingly according to the first gap values ​​corresponding to multiple historical frames between the current frame, so that the set threshold is closely related to the reversibility of the historical first covariance matrix, thereby achieving a better overall separation effect.

[0133] The following describes it in detail through specific embodiments. Specific embodiment:

[0135] Two sound sources are set up, and the speaker is provided with two microphones. Each microphone receives aliased sound data of the two sound sources, and the data of the two sound sources are distinguished according to the aliased sound data received by the two microphones.

[0136] Step 1: Set parameter values.

[0137] Step 1.1: Set the system frame length to Nfft and the number of frequency points to K, where K = Nfft / 2+1.

[0138] Step 1.2: Set the initial value of the separation matrix corresponding to each frequency point according to formula (1):

[0139]

[0140] in, is the unit matrix, k=1,..,K, where k represents the position identifier of the frequency point.

[0141] The H here stands for conjugate transpose.

[0142] w1(k, 0) is the initial value matrix of the separation matrix for the first sound source, and w2(k, 0) is the initial value matrix of the separation matrix for the second sound source. The 0 in w1(k, 0) and w2(k, 0) represents frame 0. After the sound data is framed, the corresponding frames 1 and 2 are obtained, and so on. Subsequent calculations use the separation matrix of the previous frame for each current frame. Therefore, to facilitate processing of the first frame, the value representing the current frame number in the initial value matrix is ​​set to 0.

[0143] Step 1.3, set the weighted covariance matrix V corresponding to each frequency point according to formula (2) i Initial value of (k):

[0144]

[0145] in, is a zero matrix, k = 1, .., K, k represents the position identifier of the frequency point, and i = 1, 2, i represents the identifier of the sound source.

[0146] Step 2: Calculate frequency domain data.

[0147] The aliased sound data collected by each microphone is framed to obtain a frame of the sound signal collected by each microphone.

[0148] by The discrete sequence representing the time domain signal of the nth frame of the pth microphone, p = 1, 2; m = 1, ..., Nfft.

[0149] According to formula (3), Perform FFT transformation of the windowed Nfft point to obtain the corresponding frequency domain signal X p (k,n),

[0150]

[0151] According to the X of each microphone p (k,n) constructs the observation signal matrix as:

[0152] X(k,n)=[X1(k,n),X2(k,n)] T

[0153] Wherein, k=1,..,K; T represents transpose.

[0154] Step 3: Calculate the frequency band estimate.

[0155] According to formula (4), the prior frequency domain estimation of all sound source signals in the current frame is calculated using the separation matrix W(k,n-1) of the previous frame and the observation signal matrix.

[0156] Y(k,n)=W(k,n-1)X(k,n) (4)

[0157] Where k = 1,..,K.

[0158] Let Y(k,n)=[Y1(k,n),Y2(k,n)] T , k=1,..,K.

[0159] Y1(k,n) and Y2(k,n) are two elements in Y(k,n). They are determined based on Y(k,n). Y1(k,n) and Y2(k,n) are the estimated values ​​of sound sources s1 and s2 at the time-frequency point (k,n).

[0160] Determine the frequency domain estimate of each sound source in the entire frequency band of the current frame as:

[0161]

[0162] Where i=1,2.

[0163] Step 4: Update the corresponding weighted covariance matrix V according to the frequency domain estimation of each sound source in the entire frequency band of the current frame i (k,n).

[0164] The weighted covariance matrix of each sound source at the (k, n)th time-frequency point is updated according to formula (6).

[0165]

[0166] Here, β is a weighting coefficient, for example, the value of β is 0.98.

[0167] Determined by formula (7):

[0168]

[0169] in

[0170] is the contrast function, which is determined by formula (9):

[0171]

[0172] in, Represents the multidimensional super-Gaussian prior probability density distribution model of the i-th sound source based on the entire frequency band.

[0173] In general, According to formula (10), we can get:

[0174]

[0175] at this time,

[0176]

[0177] thereby,

[0178]

[0179]

[0180] In the existing algorithm, the degree of reversibility of the obtained first weighted covariance matrix V1(k,n) is not judged here, and the first weighted covariance matrix V1(k,n) is directly used to calculate the intermediate matrix. Use the intermediate matrix to solve for the eigenvalues.

[0181] However, in actual scenarios, when the reversibility of the first weighted covariance matrix V1(k,n) is poor, directly calculating the intermediate matrix based on the auxiliary matrix obtained by inverting the matrix and performing subsequent separation based on the intermediate matrix will destroy the stability of the algorithm and lead to deterioration of the separation performance.

[0182] In view of this, the present application proposes to judge the reversibility of the first weighted covariance matrix V1(k,n). When the reversibility of the first covariance matrix V1(k,n) meets the set degree, the first covariance matrix V1(k,n) is determined as the first covariance matrix of the current frame; when the reversibility of the first covariance matrix V1(k,n) does not meet the set degree, the first covariance matrix of the current frame is determined according to the first covariance matrix V1(k,n-1) of the previous frame, thereby improving the robustness of the algorithm, ensuring the stability of the algorithm convergence, and improving the voice quality.

[0183] Step 5: Determine the first covariance matrix.

[0184] Use the inverse formula to calculate the auxiliary matrix corresponding to V1(k,n):

[0185] For example:

[0186] For example: When V1(k,n) is a 2*2 matrix,

[0187]

[0188] Then the auxiliary matrix calculated using the inverse formula shown in formula (11) is invWtmp(k,n):

[0189]

[0190] Here, det(V1(k,n)) represents the determinant of V1(k,n).

[0191] During the calculation process using the operation program, the value of det(V1(k,n)) will not be 0. If the value of the determinant of det(V1(k,n)) is 0, a correction value will be automatically added so that the corrected value of det(V1(k,n)) is not 0.

[0192] Calculate the product of the two and get V dot (k,n):

[0193]

[0194] Calculate V dot The first difference between (k,n) and the identity matrix is ​​amp1_V dot (k,n):

[0195] amp_V dot (k,n)=abs(V dot (1,1,k,n)-1)+abs(V dot (1,2,k,n))+abs(V dot (2,1,k,n))+abs(V dot (2,2,k,n)-1)

[0196] If amp_V dot (k,n)≤TH, TH is the setting level, for example, TH is 1e-10, then

[0197] Use V1(k,n) as the first covariance matrix of the current frame.

[0198] If amp_V dot (k,n)>TH, the first covariance matrix of the previous frame is used as the first covariance matrix of the current frame:

[0199] V1(k,n)=V1(k,n-1)

[0200] Step 6: Solve the eigenvalue.

[0201] Calculate the intermediate matrix

[0202] Solve the eigenvalue according to formula (11):

[0203] V2(k,n)e i (k,n)=λ i (k,n)V1(k,n)e i (k,n) (12)

[0204] Where i=1,2.

[0205] The solution is:

[0206]

[0207]

[0208]

[0209]

[0210] Where tr is the trace function. tr(A) sums the elements on the main diagonal of matrix A. det(A) finds the determinant of matrix A, and λ1, λ2, e1, and e2 are the eigenvalues.

[0211] Among them, H 22 (k,n) represents the element in the second row and second column of the H(k,n) matrix. 12 (k,n) represents the element in the first row and second column of the H(k,n) matrix. 11 (k,n) represents the element in the first row and first column of the H(k,n) matrix.

[0212] Step 7. Calculate the separation matrix of all sound sources at each frequency point in the current frame based on the eigenvalues:

[0213] W(k,n)=[w1(k,n),w2(k,n)] H , k=1,..,K. (17)

[0214]

[0215] i=1,2.

[0216] Step 8: Use W(k,n) to separate the aliased audio signal and obtain the posterior frequency domain estimate of the sound source signal:

[0217] Y(k,n)=[Y1(k,n),Y2(k,n)] T =W(k,n)X(k,n) (19)

[0218] Step 9: Perform IFFT and overlap-add to obtain the separated time domain sound source signal s i (m,n).

[0219]

[0220] Where, i=1,2; m=1,...,Nfft.

[0221] The present disclosure provides an audio signal processing device for use in a terminal, which is an electronic device integrating two or more microphones. For example, the terminal may be a mobile phone, a laptop, a tablet computer, a vehicle-mounted terminal, a computer, or a server; or the terminal may be a device connected to multiple microphones.

[0222] Reference Figure 2 , Figure 2 FIG. 1 is a structural diagram of an audio signal processing device according to an exemplary embodiment. Figure 2 As shown, this device includes:

[0223] An acquisition module 21 is configured to acquire aliased audio signals of at least two sound sources collected by at least two microphones;

[0224] The framing module 22 is configured to perform framing processing on the mixed audio signal to obtain a multi-frame audio time domain signal;

[0225] The processing module 23 is configured to process each frame of the audio time domain signal;

[0226] The processing module 23 includes:

[0227] The first determining module 231 is configured to determine a first covariance matrix of a first sound source at each frequency point and a second covariance matrix of a second sound source at each frequency point of the current frame audio time domain signal;

[0228] A judging module 232 is configured to judge whether the invertibility of the first covariance matrix satisfies a set degree;

[0229] The second determining module 233 is configured to, when the reversibility degree of the first covariance matrix satisfies a set degree, determine that the first covariance matrix is ​​the first covariance matrix of the current frame audio time domain signal; when the reversibility degree of the first covariance matrix does not satisfy the set degree, update the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal;

[0230] A third determining module 234 is configured to calculate an intermediate matrix using the inverse matrix of the first covariance matrix and the second covariance matrix; and calculate a separation matrix based on the intermediate matrix;

[0231] The separation module 235 is configured to use the separation matrix to separate the audio signals of different sound sources from the current frame audio time domain signal.

[0232] An embodiment of the present disclosure provides an audio signal processing device, which includes Figure 2 the device shown, and:

[0233] The second determining module 233 is further configured to update the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal using one of the following methods:

[0234] Using the first covariance matrix of the previous frame of audio time domain signal as the first covariance matrix of the current frame of audio time domain signal;

[0235] Determine a product matrix of the first covariance matrix of the previous frame audio time domain signal and the coefficient matrix, and use the product matrix as the first covariance matrix of the current frame audio time domain signal.

[0236] An embodiment of the present disclosure provides an audio signal processing device, which includes Figure 2 the device shown, and:

[0237] The judgment module 232 includes:

[0238] a fourth calculation module, configured to determine an auxiliary matrix corresponding to the first covariance matrix using an inversion formula;

[0239] a fifth computing module, configured to determine a product matrix of the first covariance matrix and the auxiliary matrix;

[0240] A sixth calculation module is configured to determine a first difference value between the product matrix and the identity matrix;

[0241] The seventh determination module is configured to determine whether the reversibility degree of the first covariance matrix meets a set degree when the first gap value is less than or equal to a set threshold.

[0242] In one embodiment, the fourth determining module is further configured to determine the auxiliary matrix corresponding to the first covariance matrix using an inversion formula using the following method:

[0243] determining an adjoint matrix of the first covariance matrix, and determining a determinant of the first covariance matrix;

[0244] Determining a ratio of the adjoint matrix to the determinant;

[0245] The ratio result is used as the auxiliary matrix corresponding to the first covariance matrix.

[0246] In one embodiment, the sixth determining module is configured to determine the first difference value between the product matrix and the identity matrix using the following method:

[0247] Determine the absolute value of the difference between each element on the main diagonal of the product matrix and 1,

[0248] determining the absolute value of each element of the product matrix that is outside the main diagonal;

[0249] Determine the sum of the absolute values;

[0250] The sum is used as a first difference value between the product matrix and the identity matrix.

[0251] In one embodiment, the device further comprises:

[0252] The eighth determination module is configured to determine the first gap value corresponding to multiple historical frame audio time domain signals before the current frame audio time domain signal, determine the first coefficient based on the first gap value corresponding to the multiple historical frame audio time domain signals, and determine that the set threshold is the product of the first fixed value and the first coefficient.

[0253] In one embodiment, the eighth determination module is further configured to determine the first coefficient according to the first gap values ​​corresponding to the audio time domain signals of multiple historical frames using the following method:

[0254] Determine the difference between a first gap value corresponding to the audio time domain signal of each historical frame and a first fixed value, determine an average value of the difference corresponding to each historical frame, and determine a first coefficient based on the average value, where the average value is positively correlated with the first coefficient.

[0255] An embodiment of the present disclosure provides an audio signal processing device, the device comprising:

[0256] processor;

[0257] a memory for storing processor-executable instructions;

[0258] The processor is configured to execute the executable instructions in the memory to implement the steps of the method.

[0259] In an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, on which executable instructions are stored. When the executable instructions are executed by a processor, the steps of the method are implemented.

[0260] Figure 3 FIG3 is a block diagram of an audio signal processing apparatus 300 according to an exemplary embodiment. For example, the apparatus 300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0261] Reference Figure 3 , apparatus 300 may include one or more of the following components: a processing component 302 , a memory 304 , a power component 306 , a multimedia component 308 , an audio component 310 , an input / output (I / O) interface 312 , a sensor component 314 , and a communication component 316 .

[0262] The processing component 302 generally controls the overall operation of the device 300, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 302 may include one or more processors 320 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 302 may include one or more modules to facilitate interaction between the processing component 302 and other components. For example, the processing component 302 may include a multimedia module to facilitate interaction between the multimedia component 308 and the processing component 302.

[0263] The memory 304 is configured to store various types of data to support operations on the device 300. Examples of such data include instructions for any application or method operating on the device 300, contact data, phone book data, messages, pictures, videos, etc. The memory 304 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0264] The power supply component 306 provides power to the various components of the device 300. The power supply component 306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 300.

[0265] The multimedia component 308 includes a screen that provides an output interface between the device 300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 308 includes a front camera and / or a rear camera. When the device 300 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0266] The audio component 310 is configured to output and / or input audio signals. For example, the audio component 310 includes a microphone (MIC) that is configured to receive external audio signals when the device 300 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 304 or transmitted via the communication component 316. In some embodiments, the audio component 310 further includes a speaker for outputting audio signals.

[0267] I / O interface 312 provides an interface between processing component 302 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0268] The sensor assembly 314 includes one or more sensors for providing various aspects of the status assessment of the device 300. For example, the sensor assembly 314 can detect the open / closed state of the device 300, the relative positioning of components, such as the display and keypad of the device 300. The sensor assembly 314 can also detect changes in the position of the device 300 or a component of the device 300, the presence or absence of user contact with the device 300, the orientation or acceleration / deceleration of the device 300, and temperature changes of the device 300. The sensor assembly 314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 314 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 314 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0269] The communication component 316 is configured to facilitate wired or wireless communication between the device 300 and other devices. The device 300 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 316 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0270] In an exemplary embodiment, the apparatus 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0271] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 304 including instructions, which can be executed by the processor 320 of the apparatus 300 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0272] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0273] Other embodiments of the present invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present invention being indicated by the following claims.

[0274] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof, which is limited only by the appended claims.

Claims

1. A method for processing an audio signal, characterized in that: include: Acquire aliased audio signals of at least two sound sources collected by at least two microphones; Performing frame processing on the mixed audio signal to obtain a multi-frame audio time domain signal; The following processing is performed on each frame of audio time domain signal: Determine a first covariance matrix of a first sound source at each frequency point and a second covariance matrix of a second sound source at each frequency point of a current frame of audio time domain signal; Determining whether the reversibility of the first covariance matrix meets a set degree; When the reversibility of the first covariance matrix satisfies a set degree, determining the first covariance matrix as the first covariance matrix of the current frame audio time domain signal; when the reversibility of the first covariance matrix does not satisfy the set degree, updating the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal; Calculate an intermediate matrix using the inverse matrix of the first covariance matrix and the second covariance matrix; Calculating a separation matrix based on the intermediate matrix; Using the separation matrix to separate audio signals of different sound sources from the current frame audio time domain signal; The step of determining whether the reversibility of the first covariance matrix satisfies a set degree includes: Determine the auxiliary matrix corresponding to the first covariance matrix using an inversion formula; Determining a product matrix of the first covariance matrix and the auxiliary matrix; Determining a first difference value between the product matrix and the identity matrix; When the first gap value is less than or equal to a set threshold, it is determined that the reversibility degree of the first covariance matrix meets the set degree.

2. The method according to claim 1, wherein The updating of the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal includes one of the following: Using the first covariance matrix of the previous frame of audio time domain signal as the first covariance matrix of the current frame of audio time domain signal; Determine a product matrix of the first covariance matrix of the previous frame audio time domain signal and the coefficient matrix, and use the product matrix as the first covariance matrix of the current frame audio time domain signal.

3. The method according to claim 1, wherein The determining the auxiliary matrix corresponding to the first covariance matrix using an inversion formula includes: determining an adjoint matrix of the first covariance matrix, and determining a determinant of the first covariance matrix; Determining a ratio of the adjoint matrix to the determinant; The ratio result is used as the auxiliary matrix corresponding to the first covariance matrix.

4. The method according to claim 1, wherein Determining a first difference value between the product matrix and the identity matrix includes: Determine the absolute value of the difference between each element on the main diagonal of the product matrix and 1, determining the absolute value of each element of the product matrix that is outside the main diagonal; Determine the sum of the absolute values; The sum is used as a first difference value between the product matrix and the identity matrix.

5. The method according to claim 1, wherein The method further comprises: Determine a first gap value corresponding to multiple historical frame audio time domain signals before the current frame audio time domain signal, determine a first coefficient based on the first gap values ​​corresponding to the multiple historical frame audio time domain signals, and determine that the set threshold is the product of the first fixed value and the first coefficient.

6. The method according to claim 5, wherein The determining of the first coefficient according to the first gap values ​​corresponding to the audio time domain signals of the plurality of historical frames includes: Determine the difference between a first gap value corresponding to each historical frame audio time domain signal and a first fixed value, determine an average value of the difference corresponding to each historical frame audio time domain signal, and determine a first coefficient based on the average value, where the average value is positively correlated with the first coefficient.

7. An audio signal processing device, applied to a mobile terminal, characterized in that: include: an acquisition module configured to acquire aliased audio signals of at least two sound sources collected by at least two microphones; A framing module is configured to perform framing processing on the mixed audio signal to obtain a multi-frame audio time domain signal; A processing module is configured to process each frame of audio time domain signal; The processing module includes: A first determining module is configured to determine a first covariance matrix of a first sound source at each frequency point and a second covariance matrix of a second sound source at each frequency point of a current frame of audio time domain signal; a judging module configured to judge whether the invertibility degree of the first covariance matrix satisfies a set degree; A second determining module is configured to, when the reversibility degree of the first covariance matrix satisfies a set degree, determine that the first covariance matrix is ​​the first covariance matrix of the current frame audio time domain signal; when the reversibility degree of the first covariance matrix does not satisfy the set degree, update the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal; A third determining module is configured to calculate an intermediate matrix using the inverse matrix of the first covariance matrix and the second covariance matrix; and calculate a separation matrix according to the intermediate matrix; A separation module is configured to separate audio signals of different sound sources from the current frame audio time domain signal using the separation matrix; Wherein, the judgment module includes: a fourth determining module, configured to determine an auxiliary matrix corresponding to the first covariance matrix using an inversion formula; a fifth determining module, configured to determine a product matrix of the first covariance matrix and the auxiliary matrix; a sixth determining module, configured to determine a first difference value between the product matrix and the identity matrix; The seventh determination module is configured to determine whether the reversibility of the first covariance matrix meets a set degree when the first gap value is less than or equal to a set threshold.

8. The device according to claim 7, wherein The second determining module is further configured to update the first covariance matrix of the current frame audio time domain signal according to the first covariance matrix of the previous frame audio time domain signal using one of the following methods: Using the first covariance matrix of the previous frame of audio time domain signal as the first covariance matrix of the current frame of audio time domain signal; Determine a product matrix of the first covariance matrix of the previous frame audio time domain signal and the coefficient matrix, and use the product matrix as the first covariance matrix of the current frame audio time domain signal.

9. The device according to claim 7, wherein The fourth determining module is further configured to determine the auxiliary matrix corresponding to the first covariance matrix using an inversion formula using the following method: determining an adjoint matrix of the first covariance matrix, and determining a determinant of the first covariance matrix; Determining a ratio of the adjoint matrix to the determinant; The ratio result is used as the auxiliary matrix corresponding to the first covariance matrix.

10. The device according to claim 7, wherein The sixth determining module is configured to determine a first difference value between the product matrix and the identity matrix using the following method: Determine the absolute value of the difference between each element on the main diagonal of the product matrix and 1, determining the absolute value of each element of the product matrix that is outside the main diagonal; Determine the sum of the absolute values; The sum is used as a first difference value between the product matrix and the identity matrix.

11. The device according to claim 7, wherein The device further comprises: The eighth determination module is configured to determine the first gap value corresponding to multiple historical frame audio time domain signals before the current frame audio time domain signal, determine the first coefficient based on the first gap value corresponding to the multiple historical frame audio time domain signals, and determine that the set threshold is the product of the first fixed value and the first coefficient.

12. The device according to claim 11, wherein The eighth determination module is further configured to determine the first coefficient according to the first gap values ​​corresponding to the audio time domain signals of multiple historical frames using the following method: Determine the difference between a first gap value corresponding to the audio time domain signal of each historical frame and a first fixed value, determine an average value of the difference corresponding to each historical frame, and determine a first coefficient based on the average value, where the average value is positively correlated with the first coefficient.

13. An audio signal processing device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium having executable instructions stored thereon, characterized in that: When the executable instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Audio signal processing method and device and storage medium

    CN112863537A