Audio processing method and apparatus therefor

By constructing the speech covariance matrix and the noise covariance matrix, solving the mixing matrix, and outputting the speech signal and noise signal, the problem of high computational complexity in stereo output is solved, and efficient stereo output is achieved.

CN115862651BActive Publication Date: 2026-01-30VIVO MOBILE COMM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211436870.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2026-01-30
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

Stereo output requires multiple spatial filtering steps, resulting in high computational complexity.

Method used

By acquiring audio signals, constructing speech covariance and noise covariance matrices, solving the mixing matrix and determining the unmixing matrix, and directly outputting speech and noise signals, the system avoids multiple spatial filtering steps.

Benefits of technology

It effectively reduces computational complexity, improves algorithm robustness, and achieves efficient stereo output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862651B_ABST
    Figure CN115862651B_ABST
Patent Text Reader

Abstract

This application discloses an audio processing method and apparatus, belonging to the field of communication technology. It includes: acquiring an audio signal, the audio signal comprising a first audio sub-signal and a second audio sub-signal collected by different microphones of an electronic device; constructing a speech covariance matrix and a noise covariance matrix corresponding to the audio signal based on the probability of the presence of a speech signal corresponding to each audio frequency point in the audio signal; obtaining a mixing matrix corresponding to the audio signal based on the speech covariance matrix and the noise covariance matrix, and inverting the mixing matrix to determine the demixing matrix of the audio signal; wherein the mixing matrix includes a first spatial transfer function corresponding to the speech signal channel and a second spatial transfer function corresponding to the noise signal channel in the audio signal; and outputting a first speech signal and a first noise signal corresponding to the first audio sub-signal, a second speech signal corresponding to the second audio sub-signal, and a second noise signal corresponding to the second audio sub-signal based on the demixing matrix and the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technology, specifically relating to an audio processing method and apparatus. Background Technology

[0002] Human perception of sound includes not only the three elements of loudness, pitch, and timbre, but also spatial information of sound, such as direction, distance, and environmental information.

[0003] Compared to mono signals, stereo signals include spatial information. With technological advancements, portable devices such as mobile phones and tablets with multiple microphones have become increasingly common. Consequently, stereo recording has gradually become a fundamental function.

[0004] In related technologies, stereo output requires multiple audio signals containing spatial information, while conventional audio enhancement algorithms only output one audio signal. Therefore, regardless of whether beamforming or blind source separation algorithms are used, stereo output generally requires multiple spatial filtering steps, which have high computational complexity.

[0005] Therefore, how to achieve better stereo output has become an urgent problem to be solved in the industry. Summary of the Invention

[0006] The purpose of this application is to provide an audio processing method and apparatus that can solve the problem that stereo output requires multiple spatial filters, resulting in high computational complexity.

[0007] In a first aspect, embodiments of this application provide an audio processing method, the method comprising:

[0008] Acquire audio signals, the audio signals including a first audio sub-signal and a second audio sub-signal collected by different microphones of the electronic device;

[0009] Based on the probability of the presence of the speech signal corresponding to each audio frequency point in the audio signal, construct the speech covariance matrix and noise covariance matrix corresponding to the audio signal;

[0010] The mixing matrix corresponding to the audio signal is obtained based on the speech covariance matrix and the noise covariance matrix, and the inverse of the mixing matrix is ​​used to determine the demixing matrix of the audio signal; wherein, the mixing matrix includes a first spatial transfer function corresponding to the speech signal channel in the audio signal and a second spatial transfer function corresponding to the noise signal channel in the audio signal;

[0011] Based on the demixing matrix and the audio signal, the first speech signal and the first noise signal corresponding to the first audio sub-signal, and the second speech signal and the second noise signal corresponding to the second audio sub-signal are output respectively.

[0012] Secondly, embodiments of this application provide an audio processing apparatus, including:

[0013] The acquisition module is used to acquire audio signals, which include a first audio sub-signal and a second audio sub-signal collected by different microphones of the electronic device;

[0014] The construction module is used to construct the speech covariance matrix and noise covariance matrix corresponding to the audio signal based on the probability of the existence of the speech signal corresponding to each audio frequency point in the audio signal.

[0015] The processing module is configured to obtain the mixing matrix corresponding to the audio signal based on the speech covariance matrix and the noise covariance matrix, and to invert the mixing matrix to determine the demixing matrix of the audio signal; wherein the mixing matrix includes a first spatial transfer function corresponding to the speech signal channel in the audio signal and a second spatial transfer function corresponding to the noise signal channel in the audio signal;

[0016] The output module is configured to output, based on the demixing matrix and the audio signal, a first speech signal and a first noise signal corresponding to the first audio sub-signal, and a second speech signal and a second noise signal corresponding to the second audio sub-signal, respectively.

[0017] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0018] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0019] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0020] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0021] In this embodiment, after acquiring the audio signal, the probability of the speech signal corresponding to each audio frequency point in the audio signal can be used as supervision information. Then, a speech covariance matrix and a noise covariance matrix are constructed based on the supervision information. This supervision information can help select the speech covariance matrix, which can solve the channel selection problem in the blind source separation algorithm. Furthermore, the mixing matrix corresponding to the audio signal is first calculated through the spatial transfer function, and then the unmixing matrix is ​​determined based on the mixing matrix. Then, the first speech signal, the first noise signal, the second speech signal, and the second noise signal are output based on the unmixing matrix and the audio information, respectively. This eliminates the need for multiple spatial filtering, effectively reduces the computational complexity, and improves the robustness of the algorithm. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of voice enhancement in related technologies;

[0023] Figure 2 This is a schematic diagram of the audio processing method provided in the embodiments of this application;

[0024] Figure 3 This is a schematic diagram of the audio processing device structure provided in the embodiments of this application;

[0025] Figure 4 This is a schematic diagram of the electronic device structure provided in the embodiments of this application;

[0026] Figure 5 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0028] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0029] The following description, in conjunction with the accompanying drawings, details an audio processing method and apparatus provided in this application through specific embodiments and application scenarios.

[0030] In related technologies, stereo voice enhancement is an important application scenario for stereo sound. Figure 1 This is a schematic diagram of voice enhancement in related technologies, such as... Figure 1 As shown, taking two microphones as an example, assuming the sound environment is a noisy and reverberant scene, according to Formula 1, the signals collected by the microphones can be expressed as x1(n) and x2(n).

[0031] x m (n)=a m (n)*s(n)+r m (n), (1)

[0032] Where m = 1, 2, s(n) represents the audio signal source, a m (n) represents the acoustic transfer function (ATF) of the audio signal source relative to the m-th microphone, * denotes convolution, r m (n) represents the noise component corresponding to the m-th microphone. Stereo voice enhancement refers to enhancing the speech of these two signals, removing noise signals from these two signals, and retaining the corresponding speech components and spatial information to obtain the enhanced y1(n) and y2(n) respectively. The goal of voice enhancement is y1(n)≈a1(n)*s(n) and y2(n)≈a2(n)*s(n).

[0033] Figure 2 This is a schematic diagram of the audio processing method provided in the embodiments of this application, such as... Figure 2 As shown, it includes:

[0034] Step 210: Acquire audio signals, the audio signals including first audio sub-signals and second audio sub-signals collected by different microphones of the electronic device;

[0035] Specifically, audio signals are acquired by multiple microphones installed in the electronic device. The microphones can be set together to form a microphone array, or they can be set in different locations in the electronic device. Each microphone can collect audio sub-signals independently. For example, microphone A can collect the first audio sub-signal, and microphone B can collect the second audio sub-signal.

[0036] More specifically, the audio sub-signals collected by different microphones may include speech signals from human voice sources and noise signals from other sound sources.

[0037] It is understood that in the embodiments of this application, after the microphone collects each audio sub-signal x mAfter (n), further frame segmentation, windowing, and Fourier transform can be performed to obtain X. m (k,l), m=1,2, where k represents the audio frequency point in the audio signal and l represents the time frame of the audio signal, and the final audio signal obtained is X(k,l)=[x1(k,l)X2(k,l)] T .

[0038] Step 220: Based on the probability of the presence of the speech signal corresponding to each audio frequency point in the audio signal, construct the speech covariance matrix and noise covariance matrix corresponding to the audio signal;

[0039] Specifically, each audio signal contains multiple audio frequency points. Each audio frequency point may correspond to a speech signal or a noise signal. It can be understood that the probability of a speech signal being present is the probability that the audio frequency point may be the frequency point corresponding to the speech signal.

[0040] In this embodiment, the probability of speech signal existence vad(l) corresponding to each audio frequency point can be obtained by analysis through a related deep learning neural network. That is, the audio signal is input into the deep learning neural network, and the probability of speech signal existence vad(l) corresponding to each audio frequency point in the audio signal can be output.

[0041] It is understood that the probability of the presence of the voice signal vad(l) in the embodiments of this application can also be analyzed by other conventional methods. The embodiments of this application do not limit the method of obtaining it, and the embodiments of this application do not require a very accurate probability of the presence of the voice signal vad(l), which can be a relatively coarse piece of information.

[0042] In this embodiment, the probability of speech signal existence vad(l) can be used as a supervisory information to filter out speech frames from the audio signal in the desired sense, thereby constructing a covariance matrix. Furthermore, it can be used to select which covariance matrix is ​​the speech covariance matrix, thereby solving the channel selection problem in the blind source separation algorithm.

[0043] Furthermore, the speech covariance matrix Φ is constructed based on the existence probability of the speech signal. XX (k,l) and noise covariance matrix Φ NN (k,l), specifically Formula 2 and Formula 3:

[0044] Φ XX (k,l)=(1-α)Φ XX (k,l)+α vad(l)X(k,l)X H (k,l), (2)

[0045] Φ NN(k,l)=(1-α)Φ NN (k,l)+α(1-vad(l))X(k,l)X H (k,l), (3)

[0046] Where α is the smoothing factor, k is the audio frequency, l is the time frame, and vad(l) is the probability of the speech signal.

[0047] Step 230: Obtain the mixing matrix corresponding to the audio signal based on the speech covariance matrix and the noise covariance matrix, and invert the mixing matrix to determine the demixing matrix of the audio signal; wherein, the mixing matrix includes a first spatial transfer function corresponding to the speech signal channel in the audio signal and a second spatial transfer function corresponding to the noise signal channel in the audio signal;

[0048] In related technologies, the demixing matrix corresponding to the audio signal is usually solved directly. However, this calculation method is relatively complex and computationally intensive. In this embodiment, the demixing matrix of the audio signal can be obtained by directly calculating the mixing matrix corresponding to the audio signal and then inverting the mixing matrix.

[0049] More specifically, the column vector a of the hybrid matrix described in the embodiments of this application i (k,l) has a clear physical meaning, namely the spatial transfer function, which can specifically include the first spatial transfer function a1(k,l) corresponding to the speech signal channel and the second spatial transfer function a2(k,l) corresponding to the noise signal channel.

[0050] Therefore, it can be understood that solving the mixture matrix in this application can specifically involve updating the first spatial transfer function and the second spatial transfer function. After updating the first spatial transfer function and the second spatial transfer function, the update and solution of the mixture matrix A(k,l) are completed.

[0051] More specifically, in this embodiment of the application, after updating the mixing matrix A(k,l), it will be further normalized to avoid the amplitude uncertainty problem in the blind source separation algorithm. Finally, the normalized mixing matrix is ​​inverted, and the unmixing matrix W(k,l) can be obtained according to Formula 4.

[0052] W(k,l)=A -1 (k,l) (4).

[0053] In this embodiment of the application, it is assumed that the outputs of the blind source separation algorithm are the speech signal channel Y1(k,l) and the noise signal channel N1(k,l), respectively. According to Formula 5, specifically:

[0054]

[0055] For stereo input, the demixing matrix W(k,l) is a 2x2 matrix, which can be decomposed according to Equation 6 as follows:

[0056] W(k,l)=[w1(k,l) w2(k,l)] H (6)

[0057] Among them, w i (k,l) is a 2-dimensional column vector, i = 1, 2.

[0058] According to Formula 7, the mixture matrix can be decomposed into the following specific components:

[0059] A(k,l)=[a1(k,l) a2(k,l)], (7)

[0060] Among them, a i (k,l) is a 2-dimensional column vector, i = 1, 2.

[0061] Step 240: Based on the demixing matrix and the audio signal, output the first speech signal and the first noise signal corresponding to the first audio sub-signal, and the second speech signal and the second noise signal corresponding to the second audio sub-signal, respectively.

[0062] In this embodiment, the first speech signal and the first noise signal corresponding to the first audio sub-signal can be obtained from the demixing matrix and the audio signal, respectively. The second speech signal and the second noise signal corresponding to the second audio sub-signal can also be obtained from the demixing matrix and the audio signal.

[0063] Furthermore, the first speech signal, the first noise signal, the second speech signal, and the second noise signal are respectively subjected to inverse FFT transformation, windowing, and frame transformation to the time domain, and the first speech signal, the first noise signal, the second speech signal, and the second noise signal corresponding to the first audio sub-signal are respectively output.

[0064] In this embodiment, after acquiring the audio signal, the probability of the speech signal corresponding to each audio frequency point in the audio signal can be used as supervision information. Then, a speech covariance matrix and a noise covariance matrix are constructed based on the supervision information. This supervision information can help select the speech covariance matrix, which can solve the channel selection problem in the blind source separation algorithm. Furthermore, the mixing matrix corresponding to the audio signal is first calculated through the spatial transfer function, and then the unmixing matrix is ​​determined based on the mixing matrix. Then, the first speech signal, the first noise signal, the second speech signal, and the second noise signal are output based on the unmixing matrix and the audio information, respectively. This eliminates the need for multiple spatial filtering, effectively reduces the computational complexity, and improves the robustness of the algorithm.

[0065] Optionally, obtaining the mixing matrix corresponding to the audio signal based on the speech covariance matrix and the noise covariance matrix includes:

[0066] The first spatial transfer function and the second spatial transfer function are updated according to the speech covariance matrix and the noise covariance matrix to obtain the first target spatial transfer function and the second target spatial transfer function.

[0067] Based on the first spatial relative transfer function and the second spatial relative transfer function, the first target spatial transfer function and the second target spatial transfer function are normalized respectively to obtain the mixing matrix corresponding to the audio signal;

[0068] Wherein, the first spatial relative transfer function is determined based on the ratio of the third spatial transfer function to the fourth spatial transfer function, and the second spatial relative transfer function is determined based on the ratio of the fifth spatial transfer function to the sixth spatial transfer function; the third spatial transfer function is the spatial transfer function of the speech signal relative to the first microphone, the fourth spatial transfer function is the spatial transfer function of the speech signal relative to the second microphone, the fifth spatial transfer function is the spatial transfer function of the noise signal relative to the second microphone, and the sixth spatial transfer function is the spatial transfer function of the noise signal relative to the first microphone.

[0069] Specifically, the first spatial transfer function described in the embodiments of this application can be the transfer function of the sound source of the speech signal relative to the microphone, and the second spatial transfer function can be the transfer function of the sound source of the noise signal relative to the microphone. The noise may theoretically come from multiple directions, but in the embodiments of this application, the noise is considered to come from a source in a desired direction.

[0070] More specifically, in the embodiments of this application, when any audio frequency point in the audio signal is detected, the first spatial transfer function and the second spatial transfer function are updated.

[0071] In other embodiments, to reduce the number of updates and computational load, the first spatial transfer function is updated only when the audio frequency point may correspond to a speech signal, and the second spatial transfer function is updated only when the audio frequency point may correspond to a noise signal.

[0072] In this embodiment, after updating the first spatial transfer function and the second spatial transfer function, due to the amplitude uncertainty problem in the blind source separation algorithm, the column vectors of the mixing matrix can be further calibrated to obtain a spatial transfer function with clear physical meaning, thus solving the amplitude uncertainty problem of the blind source separation algorithm.

[0073] More specifically, the column vector correction and calibration can be performed by normalizing the first spatial transfer function using the first spatial relative transfer function, and by normalizing the second spatial transfer function using the second spatial relative transfer function, ultimately obtaining the normalized mixed matrix.

[0074] The first spatial relative transfer function described in this application embodiment refers to the transmission coefficient of the voice signal between the first microphone and the second microphone, and the second spatial relative transfer function refers to the transmission coefficient of the noise signal between the first microphone and the second microphone.

[0075] Understandably, according to Formula 8, the normalization process can be specifically described as follows:

[0076] make

[0077] Among them, according to Formula 9, the first spatial relative transfer function a 1rtf (k,l)=a 21 (k,l) / a 11 (k,l).

[0078]

[0079] Wherein, according to Formula 10, the second space relative transfer function is:

[0080] a 2rtf (k,l)=a 12 (k,l) / a 22 (k,l). (10)

[0081] More specifically, in this embodiment, the spatial transfer coefficient of the voice signal source relative to the first microphone is the third spatial transfer function a. 11 (k,l), the spatial transfer coefficient of the speech signal source relative to the second microphone is the fourth spatial transfer function a. 21 (k,l), the spatial transfer coefficient of the noise signal source relative to the first microphone is the fifth spatial transfer function a. 12 (k,l), the spatial transfer coefficient of the noise signal source relative to the second microphone is the sixth spatial transfer function a. 22 (k,l).

[0082] After the above update and normalization loop processing, and after traversing all audio frequency points in the audio signal, the mixing matrix is ​​updated, and the mixing matrix corresponding to the audio signal is obtained.

[0083] In this embodiment, the update and request of the mixing matrix are achieved by updating the first spatial transfer function and the second spatial transfer function, which effectively simplifies the calculation process. At the same time, the first spatial relative transfer function and the second spatial relative transfer function between the speech signal and the noise signal are used to normalize the first spatial transfer function and the second spatial transfer function, which effectively achieves the calibration of the column vector of the mixing matrix and solves the problem of amplitude uncertainty in the blind source separation algorithm.

[0084] Optionally, updating the first spatial transfer function and the second spatial transfer function based on the speech covariance matrix and the noise covariance matrix to obtain the first target spatial transfer function and the second target spatial transfer function includes:

[0085] If a first target audio frequency point is detected in the audio signal, the first spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed to obtain the first target spatial transfer function.

[0086] If a second target audio frequency point is detected in the audio signal, the second spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed to obtain the second target spatial transfer function.

[0087] Wherein, the first target audio frequency point is the audio frequency point in the audio signal where the probability of the presence of a speech signal exceeds a first preset threshold, and the second target audio frequency point is the audio frequency point in the audio signal where the probability of the presence of a noise signal exceeds a second preset threshold.

[0088] Specifically, in order to effectively reduce the number of updates and the amount of computation, the first spatial transfer function is only updated when speech is present in this embodiment. Correspondingly, the second spatial transfer function can also be updated only when noise is present.

[0089] It is understandable that when the probability of a speech signal corresponding to an audio frequency point exceeds the first preset threshold, it indicates that the audio frequency point may correspond to a speech signal, and the first spatial transfer function is updated at this time.

[0090] When the probability of a noise signal corresponding to an audio frequency point exceeds the second preset threshold, it indicates that the audio frequency point may correspond to a noise signal, and the second spatial transfer function is updated at this time.

[0091] More specifically, according to Formula 11, the first spatial transfer function is updated in this embodiment of the application as follows:

[0092] When vad(l) > thr1,

[0093] ds1=w1 H (k,l)Φ XX (k,l)w1(k,l),

[0094] ds2=w1 H (k,l)Φ NN (k,l)w1(k,l), (11)

[0095] us = w2 H (k,l)Φ NN (k,l)w1(k,l),

[0096] v s =us / ds2,

[0097]

[0098] Where thr1 is the first preset threshold, w1(k,l) is the column vector corresponding to the speech signal channel in the demixing matrix, and ds1, ds2, us and v s These are all intermediate quantities in the calculation process.

[0099] More specifically, according to Formula 12, the second space transfer function is updated in this embodiment of the application as follows:

[0100] When vad(l) > thr2,

[0101] un = w1 H (k,l)Φ XX (k,l)e2(k,l),

[0102] dn1=w2 H (k,l)Φ XX (k,l)w2(k,l), (12)

[0103] dn2=w2 H (k,l)Φ NN (k,l)w2(k,l),

[0104] v n =un / dn1,

[0105]

[0106] Among them, un, dn1, dn2 and v n It is an intermediate quantity in the calculation process, thr2 is the second preset threshold, and w2(k,l) is the column vector corresponding to the noise signal channel in the demixing matrix.

[0107] In this embodiment of the application, after traversing all audio frequency points in the audio signal, the first spatial transfer function and the second spatial transfer function are updated.

[0108] In this embodiment, when a first target audio frequency point is detected in the audio signal, the first spatial transfer function is updated, and when a second target audio frequency point is detected in the audio signal, the second spatial transfer function is updated. This can effectively reduce the number of updates and improve update efficiency while ensuring update efficiency.

[0109] Optionally, based on the demixing matrix and the audio signal, a first speech signal and a first noise signal corresponding to the first audio sub-signal, a second speech signal and a second noise signal corresponding to the second audio sub-signal are output respectively, including:

[0110] Based on the product of the demixing matrix and the audio signal, the first speech signal and the first noise signal corresponding to the first audio sub-signal are obtained;

[0111] Based on the first speech signal, the first noise signal, the first spatial relative transfer function, and the second spatial relative transfer function, the second speech signal and the second noise signal corresponding to the second audio sub-signal are obtained.

[0112] Specifically, the first speech signal and the first noise signal corresponding to the first audio sub-signal are obtained according to Formula 13, as follows:

[0113]

[0114] Wherein, Y1(k,l) is the first speech signal in the first audio sub-signal collected by the first microphone, and N1(k,l) is the first noise signal in the first audio sub-signal.

[0115] Furthermore, after obtaining the first speech signal and the first noise signal, the relative transfer coefficient between the two microphones can be further combined to determine the second speech signal and the second noise signal in the second audio sub-signal acquired by the second microphone, according to Formula 14, specifically as follows:

[0116]

[0117] Where Y2(k,l) is the second speech signal in the second audio sub-signal, N2(k,l) is the second noise signal in the second audio sub-signal, and a 1rtf (k,l) is the first-space relative transfer function, a 2rtf (k,l) is the second-space relative transfer function.

[0118] More specifically, after obtaining the first speech signal, the first noise signal, the second speech signal, and the second noise signal, the first speech signal and the second speech signal can be further enhanced.

[0119] In this embodiment, a first speech signal and a first noise signal are obtained by demixing the matrix and the audio signal. At the same time, a second speech signal and a second noise signal are obtained by using the first spatial relative transfer function and the second spatial relative transfer function. This achieves stereo four-channel output without the need for two spatial filters, thus reducing the complexity of the algorithm.

[0120] Optionally, the first spatial relative transfer function is subject to causal constraints, wherein the causal constraints specifically include:

[0121] The first spatial relative transfer function is transformed to the time domain to obtain the first time domain signal;

[0122] The first time-domain signal is truncated according to a preset time-domain range to obtain a constrained first spatial relative transfer function, wherein the preset time-domain range is determined based on the finite-length impulse response corresponding to the first spatial transfer function.

[0123] Specifically, in this embodiment, only the first spatial relative transfer function corresponding to the audio signal is considered. The first spatial relative transfer function is set to be an effective long impulse response. According to Formula 15, its corresponding time-domain impulse response can be expressed as:

[0124]

[0125] In order to make the first spatial relative transfer function a 1rt (k,l) satisfies h T (n) This structure can be used to represent a 1rtf Transform (k,l) to the time domain, truncate the time-domain signal, and retain [-K]. L ,K R This range establishes a causal constraint on the relative transmission path.

[0126] Therefore, Formula 16 is as follows:

[0127]

[0128] in, Indicates that for a 1rtf The result of (k,l) after being constrained by the effective long impulse response is then assigned to a. 1rtf (k,l).

[0129] In this embodiment, the noise suppression effect of the algorithm can be effectively improved by using the first spatial relative transfer function of the voice channel.

[0130] Optionally, the solution in this application embodiment is applicable to voice enhancement scenarios with multi-channel input and multi-channel output, such as recording and spatial audio scenarios. It can enhance noisy multi-channel speech through spatial filtering, achieving voice enhancement for each channel. Compared to existing blind source and beamforming methods, the solution in this application embodiment has lower complexity and is suitable for time-varying acoustic scenarios.

[0131] The audio processing method provided in this application can be executed by an audio processing device. This application uses an audio processing device to perform the audio processing method as an example to illustrate the audio processing device provided in this application.

[0132] Figure 3 This is a schematic diagram of the audio processing device structure provided in the embodiments of this application, such as... Figure 3 As shown, the system includes: an acquisition module 310, a construction module 320, a processing module 330, and an output module 340. The acquisition module 310 acquires an audio signal, which includes a first audio sub-signal and a second audio sub-signal collected by different microphones of the electronic device. The construction module 320 constructs a speech covariance matrix and a noise covariance matrix corresponding to the audio signal based on the probability of the presence of the speech signal corresponding to each audio frequency point in the audio signal. The processing module 330 obtains a mixing matrix corresponding to the audio signal based on the speech covariance matrix and the noise covariance matrix, and inverts the mixing matrix to determine the demixing matrix of the audio signal. The mixing matrix includes a first spatial transfer function corresponding to the speech signal channel and a second spatial transfer function corresponding to the noise signal channel in the audio signal. The output module 340 outputs a first speech signal, a first noise signal, a second speech signal, and a second noise signal corresponding to the second audio sub-signal, respectively, based on the demixing matrix and the audio signal.

[0133] Optionally, the processing module is specifically used for:

[0134] The first spatial transfer function and the second spatial transfer function are updated according to the speech covariance matrix and the noise covariance matrix to obtain the first target spatial transfer function and the second target spatial transfer function.

[0135] Based on the first spatial relative transfer function and the second spatial relative transfer function, the first target spatial transfer function and the second target spatial transfer function are normalized respectively to obtain the mixing matrix corresponding to the audio signal;

[0136] Wherein, the first spatial relative transfer function is determined based on the ratio of the third spatial transfer function to the fourth spatial transfer function, and the second spatial relative transfer function is determined based on the ratio of the fifth spatial transfer function to the sixth spatial transfer function; the third spatial transfer function is the spatial transfer function of the speech signal relative to the first microphone, the fourth spatial transfer function is the spatial transfer function of the speech signal relative to the second microphone, the fifth spatial transfer function is the spatial transfer function of the noise signal relative to the second microphone, and the sixth spatial transfer function is the spatial transfer function of the noise signal relative to the first microphone.

[0137] Optionally, the processing module is specifically used for:

[0138] If a first target audio frequency point is detected in the audio signal, the first spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed to obtain the first target spatial transfer function.

[0139] If a second target audio frequency point is detected in the audio signal, the second spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed to obtain the second target spatial transfer function.

[0140] Wherein, the first target audio frequency point is the audio frequency point in the audio signal where the probability of the presence of a speech signal exceeds a first preset threshold, and the second target audio frequency point is the audio frequency point in the audio signal where the probability of the presence of a noise signal exceeds a second preset threshold.

[0141] Optionally, the output module is specifically used for:

[0142] Based on the product of the demixing matrix and the audio signal, the first speech signal and the first noise signal corresponding to the first audio sub-signal are obtained;

[0143] Based on the first speech signal, the first noise signal, the first spatial relative transfer function, and the second spatial relative transfer function, the second speech signal and the second noise signal corresponding to the second audio sub-signal are obtained.

[0144] Optionally, the first spatial relative transfer function is subject to causal constraints, wherein the causal constraints specifically include:

[0145] The first spatial relative transfer function is transformed to the time domain to obtain the first time domain signal;

[0146] The first time-domain signal is truncated according to a preset time-domain range to obtain a constrained first spatial relative transfer function, wherein the preset time-domain range is determined based on the finite-length impulse response corresponding to the first spatial transfer function.

[0147] In this embodiment, after acquiring the audio signal, the probability of the speech signal corresponding to each audio frequency point in the audio signal can be used as supervision information. Then, a speech covariance matrix and a noise covariance matrix are constructed based on the supervision information. This supervision information can help select the speech covariance matrix, which can solve the channel selection problem in the blind source separation algorithm. Furthermore, the mixing matrix corresponding to the audio signal is first calculated through the spatial transfer function, and then the unmixing matrix is ​​determined based on the mixing matrix. Then, the first speech signal, the first noise signal, the second speech signal, and the second noise signal are output based on the unmixing matrix and the audio information, respectively. This eliminates the need for multiple spatial filtering, effectively reduces the computational complexity, and improves the robustness of the algorithm.

[0148] The audio processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0149] The audio processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0150] The audio processing device provided in this application embodiment can achieve... Figures 1 to 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0151] Optionally, Figure 4 This is a schematic diagram of the electronic device structure provided in the embodiments of this application, such as... Figure 4 As shown, this application embodiment also provides an electronic device 400, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they implement the various steps of the above-described audio processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0152] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0153] Figure 5 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0154] The electronic device 500 includes, but is not limited to, components such as: radio frequency unit 501, network module 502, audio output unit 503, input unit 504, sensor 505, display unit 506, user input unit 507, interface unit 508, memory 509, and processor 510.

[0155] Those skilled in the art will understand that the electronic device 500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0156] The input unit 504 is used to acquire audio signals, which include a first audio sub-signal and a second audio sub-signal collected by different microphones of the electronic device.

[0157] The processor 510 is used to construct the speech covariance matrix and noise covariance matrix corresponding to the audio signal based on the probability of the presence of the speech signal corresponding to each audio frequency point in the audio signal;

[0158] The processor 510 is configured to obtain a mixing matrix corresponding to the audio signal based on the speech covariance matrix and the noise covariance matrix, and invert the mixing matrix to determine the demixing matrix of the audio signal; wherein, the mixing matrix includes a first spatial transfer function corresponding to the speech signal channel in the audio signal and a second spatial transfer function corresponding to the noise signal channel in the audio signal;

[0159] The audio output unit 503 is used to output, according to the demixing matrix and the audio signal, a first speech signal and a first noise signal corresponding to the first audio sub-signal, a second speech signal and a second noise signal corresponding to the second audio sub-signal, respectively.

[0160] Processor 510 is used to update the first spatial transfer function and the second spatial transfer function according to the speech covariance matrix and the noise covariance matrix to obtain the first target spatial transfer function and the second target spatial transfer function;

[0161] Based on the first spatial relative transfer function and the second spatial relative transfer function, the first target spatial transfer function and the second target spatial transfer function are normalized respectively to obtain the mixing matrix corresponding to the audio signal;

[0162] Wherein, the first spatial relative transfer function is determined based on the ratio of the third spatial transfer function to the fourth spatial transfer function, and the second spatial relative transfer function is determined based on the ratio of the fifth spatial transfer function to the sixth spatial transfer function; the third spatial transfer function is the spatial transfer function of the speech signal relative to the first microphone, the fourth spatial transfer function is the spatial transfer function of the speech signal relative to the second microphone, the fifth spatial transfer function is the spatial transfer function of the noise signal relative to the second microphone, and the sixth spatial transfer function is the spatial transfer function of the noise signal relative to the first microphone.

[0163] The processor 510 is used to update the first spatial transfer function based on the speech covariance matrix and the noise covariance matrix when a first target audio frequency point is detected in the audio signal, until all audio frequency points in the audio signal are traversed to obtain the first target spatial transfer function;

[0164] If a second target audio frequency point is detected in the audio signal, the second spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed to obtain the second target spatial transfer function.

[0165] Wherein, the first target audio frequency point is the audio frequency point in the audio signal where the probability of the presence of a speech signal exceeds a first preset threshold, and the second target audio frequency point is the audio frequency point in the audio signal where the probability of the presence of a noise signal exceeds a second preset threshold.

[0166] The processor 510 is used to obtain a first speech signal and a first noise signal corresponding to the first audio sub-signal based on the product of the demixing matrix and the audio signal;

[0167] Based on the first speech signal, the first noise signal, the first spatial relative transfer function, and the second spatial relative transfer function, the second speech signal and the second noise signal corresponding to the second audio sub-signal are obtained.

[0168] Processor 510 is used to transform the first spatial relative transfer function to the time domain to obtain a first time domain signal;

[0169] The first time-domain signal is truncated according to a preset time-domain range to obtain a constrained first spatial relative transfer function, wherein the preset time-domain range is determined based on the finite-length impulse response corresponding to the first spatial transfer function.

[0170] In this embodiment, after acquiring the audio signal, the probability of the speech signal corresponding to each audio frequency point in the audio signal can be used as supervision information. Then, a speech covariance matrix and a noise covariance matrix are constructed based on the supervision information. This supervision information can help select the speech covariance matrix, which can solve the channel selection problem in the blind source separation algorithm. Furthermore, the mixing matrix corresponding to the audio signal is first calculated through the spatial transfer function, and then the unmixing matrix is ​​determined based on the mixing matrix. Then, the first speech signal, the first noise signal, the second speech signal, and the second noise signal are output based on the unmixing matrix and the audio information, respectively. This eliminates the need for multiple spatial filtering, effectively reduces the computational complexity, and improves the robustness of the algorithm.

[0171] It should be understood that, in this embodiment, the input unit 504 may include a graphics processing unit (GPU) 5041 and a microphone 5042. The GPU 5041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 506 may include a display panel 5061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 507 includes at least one of a touch panel 5071 and other input devices 5072. The touch panel 5071 is also called a touch screen. The touch panel 5071 may include a touch detection device and a touch controller. Other input devices 5072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0172] The memory 509 can be used to store software programs and various data. The memory 509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0173] Processor 510 may include one or more processing units; optionally, processor 510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 510.

[0174] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0175] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0176] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0177] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0178] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0179] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0181] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio processing method, characterized by, The method comprises the following steps: acquiring an audio signal, the audio signal comprising a first audio sub-signal and a second audio sub-signal collected by different microphones of an electronic device; constructing a speech covariance matrix and a noise covariance matrix corresponding to the audio signal according to a speech signal existence probability corresponding to each audio frequency point in the audio signal; obtaining a mixing matrix corresponding to the audio signal according to the speech covariance matrix and the noise covariance matrix, and inverting the mixing matrix to determine a demixing matrix of the audio signal; wherein the mixing matrix comprises a first spatial transfer function corresponding to a speech signal channel in the audio signal and a second spatial transfer function corresponding to a noise signal channel in the audio signal; wherein the first spatial transfer function is a transfer function of a sound source of a speech signal relative to a microphone, and the second spatial transfer function is a transfer function of a sound source of a noise signal relative to a microphone; outputting a first speech signal corresponding to the first audio sub-signal, a first noise signal, a second speech signal corresponding to the second audio sub-signal, and a second noise signal according to the demixing matrix and the audio signal.

2. The audio processing method of claim 1, wherein, obtaining a mixing matrix corresponding to the audio signal according to the speech covariance matrix and the noise covariance matrix comprises: updating the first spatial transfer function and the second spatial transfer function according to the speech covariance matrix and the noise covariance matrix to obtain a first target spatial transfer function and a second target spatial transfer function; normalizing the first target spatial transfer function and the second target spatial transfer function according to a first spatial relative transfer function and a second spatial relative transfer function to obtain the mixing matrix corresponding to the audio signal; wherein the first spatial relative transfer function is determined based on a ratio of a third spatial transfer function and a fourth spatial transfer function, and the second spatial relative transfer function is determined based on a ratio of a fifth spatial transfer function and a sixth spatial transfer function; the third spatial transfer function is a spatial transfer function of the speech signal relative to a first microphone, the fourth spatial transfer function is a spatial transfer function of the speech signal relative to a second microphone, the fifth spatial transfer function is a spatial transfer function of the noise signal relative to the second microphone, and the sixth spatial transfer function is a spatial transfer function of the noise signal relative to the first microphone.

3. The audio processing method of claim 2, wherein, updating the first spatial transfer function and the second spatial transfer function according to the speech covariance matrix and the noise covariance matrix to obtain a first target spatial transfer function and a second target spatial transfer function comprises: in a case where a first target audio frequency point is detected in the audio signal, updating the first spatial transfer function based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed to obtain the first target spatial transfer function; In a case where a second target audio frequency point is detected in the audio signal, the second spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed, to obtain a second target spatial transfer function; The first target audio frequency point is an audio frequency point in the audio signal in which a probability of existence of a speech signal exceeds a first preset threshold, and the second target audio frequency point is an audio frequency point in the audio signal in which a probability of existence of a noise signal exceeds a second preset threshold.

4. The audio processing method of claim 2, wherein, According to the demixing matrix and the audio signal, first speech signals and first noise signals corresponding to the first audio sub-signals and second speech signals and second noise signals corresponding to the second audio sub-signals are respectively output, including: According to a product of the demixing matrix and the audio signal, first speech signals and first noise signals corresponding to the first audio sub-signals are obtained. Based on the first speech signals, the first noise signals, the first spatial relative transfer function, and the second spatial relative transfer function, second speech signals and second noise signals corresponding to the second audio sub-signals are obtained.

5. The audio processing method of claim 2, wherein, The first spatial relative transfer function is subject to a causality constraint, where the causality constraint is specifically: The first spatial relative transfer function is transformed to a time domain to obtain a first time domain signal. According to a preset time domain range, the first time domain signal is truncated to obtain a constrained first spatial relative transfer function, where the preset time domain range is determined based on a finite impulse response corresponding to the first spatial transfer function.

6. An audio processing apparatus, characterized by comprising: including: The obtaining module is configured to obtain an audio signal, the audio signal including first audio sub-signals and second audio sub-signals collected by different microphones of an electronic device. The constructing module is configured to construct a speech covariance matrix and a noise covariance matrix corresponding to the audio signal according to a probability of existence of a speech signal corresponding to each audio frequency point in the audio signal. The processing module is configured to obtain a mixing matrix corresponding to the audio signal according to the speech covariance matrix and the noise covariance matrix, and to determine a demixing matrix of the audio signal by inverting the mixing matrix, where the mixing matrix includes a first spatial transfer function corresponding to a speech signal channel in the audio signal and a second spatial transfer function corresponding to a noise signal channel in the audio signal, the first spatial transfer function is a transfer function of a sound source of a speech signal relative to a microphone, and the second spatial transfer function is a transfer function of a sound source of a noise signal relative to a microphone. The output module is configured to output first speech signals and first noise signals corresponding to the first audio sub-signals and second speech signals and second noise signals corresponding to the second audio sub-signals respectively according to the demixing matrix and the audio signal.

7. The audio processing apparatus of claim 6, wherein, The processing module is specifically configured to: update the first spatial transfer function and the second spatial transfer function based on the speech covariance matrix and the noise covariance matrix to obtain a first target spatial transfer function and a second target spatial transfer function; According to the first spatial relative transfer function and the second spatial relative transfer function, the first target spatial transfer function and the second target spatial transfer function are normalized respectively, and a mixing matrix corresponding to the audio signal is obtained; The first spatial relative transfer function is determined based on a ratio of a third spatial transfer function and a fourth spatial transfer function, and the second spatial relative transfer function is determined based on a ratio of a fifth spatial transfer function and a sixth spatial transfer function; the third spatial transfer function is a spatial transfer function of the speech signal relative to a first microphone, the fourth spatial transfer function is a spatial transfer function of the speech signal relative to a second microphone, the fifth spatial transfer function is a spatial transfer function of the noise signal relative to the second microphone, and the sixth spatial transfer function is a spatial transfer function of the noise signal relative to the first microphone.

8. The audio processing apparatus of claim 7, wherein, The processing module is specifically configured to: In a case where a first target audio frequency point is detected in the audio signal, the first spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed, and a first target spatial transfer function is obtained; In a case where a second target audio frequency point is detected in the audio signal, the second spatial transfer function is updated based on the speech covariance matrix and the noise covariance matrix until all audio frequency points in the audio signal are traversed, and a second target spatial transfer function is obtained; The first target audio frequency point is an audio frequency point in the audio signal in which a probability of existence of the speech signal exceeds a first preset threshold, and the second target audio frequency point is an audio frequency point in the audio signal in which a probability of existence of the noise signal exceeds a second preset threshold.

9. The audio processing apparatus of claim 7, wherein, The output module is specifically configured to: According to a product of the demixing matrix and the audio signal, a first speech signal corresponding to a first audio sub-signal and a first noise signal are obtained; Based on the first speech signal, the first noise signal, the first spatial relative transfer function and the second spatial relative transfer function, a second speech signal corresponding to a second audio sub-signal and a second noise signal are obtained.

10. The audio processing apparatus of claim 7, wherein, The first spatial relative transfer function is subject to a causality constraint, wherein the causality constraint is specifically: The first spatial relative transfer function is transformed to a time domain to obtain a first time domain signal; The first time domain signal is truncated according to a preset time domain range to obtain a constrained first spatial relative transfer function, wherein the preset time domain range is determined based on a finite impulse response corresponding to the first spatial transfer function.

Citation Information

Patent Citations

  • Multi-channel speech enhancement method and device, terminal and readable storage medium

    CN113689870A