Distributed speech enhancement method and speech enhancement device
By constructing the optimal communication probability function and generating the node-to-node communication probability matrix, the problem of high energy consumption in wireless acoustic sensor networks is solved, and a low-energy, high-quality voice enhancement effect is achieved.
Patent Information
- Application Number
- CN202310214884.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-03-02
AI Technical Summary
Wireless acoustic sensor networks consume a lot of energy when transmitting voice information, and traditional microphone arrays have limited spatial acquisition capabilities, especially when the array is far from the speaker, which affects the intelligibility of the target voice signal.
By constructing an optimal communication probability function, using a trade-off factor, updating the second largest eigenvalue of the matrix, and the average transmission energy consumption, a node-pair communication probability matrix is generated. Based on this matrix, beamforming signal model processing is performed to output enhanced speech information.
It effectively reduces the average transmission energy consumption during signal transmission, improves the speech signal-to-noise ratio and subjective perception quality, and achieves speech intelligibility close to the effects of the classic random rumor algorithm and the centralized minimum variance distortionless response beamforming method.
Smart Images

Figure CN116312603B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech signal processing technology, and more specifically, to a distributed speech enhancement method, speech enhancement device, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] In recent years, multi-channel voice enhancement technology has been widely used in scenarios such as smart conferencing and in-vehicle voice systems. Microphone arrays can collect voice signals from specific directions in space. Different application scenarios can select appropriate geometric arrangements and adaptively control the beam direction to enable the array to collect sound source signals from different directions.
[0003] The audio signals captured by microphone arrays include not only the target speaker's voice but also environmental noise, reverberation, and interference from multiple speakers. These interference signals can severely affect the intelligibility of the target speech signal. Therefore, front-end speech enhancement is essential in many speech signal processing systems to improve speech signal quality. However, traditional microphone arrays have very limited spatial signal acquisition capabilities, especially when the array is far from the speaker, and wired microphone arrays restrict the array's scalability.
[0004] In realizing the present invention, the inventors discovered at least the following problems in the related technology: when wireless acoustic sensor networks transmit voice information, the energy consumption of the sensor network is large when obtaining enhanced voice with better voice signal quality. Summary of the Invention
[0005] In view of the above, embodiments of this disclosure provide a distributed speech enhancement method, a speech enhancement device, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] One aspect of this disclosure provides a distributed speech enhancement method, including:
[0007] The optimal communication probability function is constructed based on the speech parameter set, which includes a trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption. The average transmission energy consumption represents the loss of signal energy when transmitting speech between multiple nodes, including a primary speech node and multiple secondary speech nodes.
[0008] Solving the above optimal communication probability function yields the node pair communication probability matrix, where the node pair communication probability matrix represents the probability that the speech master node selects any speech auxiliary node to form a node pair.
[0009] Based on the above node-to-node communication probability matrix, multiple sound signals received by the above voice master node and multiple above voice auxiliary nodes are processed using preset processing rules to obtain a beamforming signal model.
[0010] The input speech information obtained from the aforementioned speech master node is input into the aforementioned beamforming signal model, and the enhanced speech information is output.
[0011] According to embodiments of this disclosure, the above-described construction of the optimal communication probability function based on the speech parameter set includes:
[0012] Based on the sensor network composed of the above nodes, generate an undirected graph of nodes;
[0013] Based on the above undirected graph of nodes, an average communication matrix is generated according to the N-dimensional identity matrix and the N-dimensional column vector.
[0014] Based on the above average communication matrix and the initial column vector of node states, generate an iterative expression for the column vector;
[0015] If the first iteration number of the above column vector iterative expression satisfies the first iteration number threshold, the iteration result of the above column vector iterative expression is taken as the second largest eigenvalue of the above update matrix, wherein the iteration result represents the second largest eigenvalue of the update matrix generated by the above column vector iterative expression in the iteration.
[0016] Based on the aforementioned trade-off factors, the second largest eigenvalue of the aforementioned update matrix, and the aforementioned average transmission energy consumption, the aforementioned optimal communication probability function characterizing the energy-aware random rumor algorithm is constructed.
[0017] According to embodiments of this disclosure, the above-mentioned average transmission power consumption is determined in the following manner:
[0018] Determine the Eulerian distance between the two nodes in each of the above node pairs;
[0019] Based on the above Eulerian distances, generate a squared distance matrix between node pairs;
[0020] Based on the above update matrix and the above node-pair distance squared matrix, the above average transmission energy consumption is generated.
[0021] According to embodiments of this disclosure, the optimal communication probability function described above is given by the following formula:
[0022]
[0023] Where P represents the node-pair communication probability matrix to be solved, and α represents the pre-defined tradeoff factor. The second largest eigenvalue represents the updated matrix, and λ² represents taking the second largest eigenvalue. The characterization update matrix is represented by D, which represents the squared distance matrix between node pairs. Characterizes average transmission energy consumption.
[0024] According to embodiments of this disclosure, the above-mentioned beamforming signal model is obtained by processing multiple audio signals received by the primary voice node and multiple secondary voice nodes based on the node pair communication probability matrix and using preset processing rules, including:
[0025] Perform a short-time Fourier transform on multiple of the above-mentioned sound signals to obtain a set of coefficients, wherein the set of coefficients includes a subset of node Fourier coefficients, a subset of clean speech Fourier coefficients, and a subset of noise Fourier coefficients.
[0026] Based on the above set of coefficients and set of steering vectors, a vector signal model is constructed, wherein the above set of steering vectors is generated based on the initial acoustic transfer function corresponding to each of the above sound signals;
[0027] Based on the above vector signal model and target covariance matrix, the above beamforming signal model is generated.
[0028] According to embodiments of this disclosure, the aforementioned sound signal is generated based on a sound source signal, the aforementioned initial sound transfer function, and noise;
[0029] The above-mentioned short-time Fourier transform of multiple sound signals yields a set of coefficients, including:
[0030] Perform a short-time Fourier transform on multiple of the above sound signals to obtain the node Fourier coefficients, clean speech Fourier coefficients, noise Fourier coefficients, and target sound transfer function corresponding to each of the above sound signals in the frequency domain.
[0031] Based on the aforementioned node Fourier coefficients, the aforementioned clean speech Fourier coefficients, the aforementioned noise Fourier coefficients, and the aforementioned target acoustic transfer functions, the aforementioned node Fourier coefficient subsets, the aforementioned clean speech Fourier coefficient subsets, the aforementioned noise Fourier coefficient subsets, and the aforementioned steering vector set are constructed respectively.
[0032] According to embodiments of this disclosure, generating the beamforming signal model based on the vector signal model and the target covariance matrix includes:
[0033] Based on the above vector signal model, the preset signal covariance matrix is inverted to obtain the initial beamforming iterative expression, wherein the above initial beamforming iterative expression characterizes the above target covariance matrix;
[0034] The initial beamforming iterative expression above is iterated twice to obtain the first instantaneous average estimate and the second instantaneous average estimate, respectively.
[0035] Substituting the average value of the first instantaneous estimate and the average value of the second instantaneous estimate into the initial beamforming iterative expression, the target beamforming expression is obtained;
[0036] Based on the above target beamforming expression and the above steering vector set, the above beamforming signal model is generated.
[0037] According to embodiments of this disclosure, the above-mentioned initial beamforming iterative expression is subjected to two iterations to obtain a first instantaneous estimated average value and a second instantaneous estimated average value, including:
[0038] If the second iteration number of the above initial beamforming iterative expression satisfies the second iteration number threshold, the first instantaneous estimated average value is determined according to the above vector signal model and the first intermediate beamforming expression, wherein the above first intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the second iteration;
[0039] If the third iteration of the above initial beamforming iterative expression satisfies the third iteration threshold, the second instantaneous estimated average value is determined according to the above vector signal model and the second intermediate beamforming expression. The above second intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the third iteration.
[0040] According to embodiments of this disclosure, generating the beamforming signal model based on the target beamforming expression and the steering vector set includes:
[0041] Based on the above target beamforming expression and the above steering vector set, a first consensus vector is generated;
[0042] Based on the first consensus vector and the set of turning vectors mentioned above, a second consensus vector is generated;
[0043] Based on the first consensus vector and the second consensus vector, the beamforming signal model is generated.
[0044] Another aspect of this disclosure provides a voice enhancement device, including:
[0045] The construction module is used to construct the optimal communication probability function based on the speech parameter set, wherein the speech parameter set includes a trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption. The average transmission energy consumption represents the loss of signal energy when transmitting speech between multiple nodes, and the multiple nodes include a primary speech node and multiple secondary speech nodes.
[0046] The solution module is used to solve the above optimal communication probability function to obtain the node pair communication probability matrix, wherein the above node pair communication probability matrix represents the probability that the voice master node selects any voice auxiliary node to form a node pair;
[0047] The processing module is used to process multiple sound signals received by the above-mentioned main voice node and multiple above-mentioned auxiliary voice nodes based on the above-mentioned node pair communication probability matrix and using preset processing rules to obtain a beamforming signal model.
[0048] The enhancement module is used to input the input voice information obtained from the above-mentioned voice master node into the above-mentioned beamforming signal model and output enhanced voice information.
[0049] Another aspect of this disclosure provides an electronic device, including: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method described above.
[0050] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the method described above.
[0051] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, implement the method described above.
[0052] According to embodiments of this disclosure, by introducing a trade-off factor to weight and balance the second largest eigenvalue of the update matrix and the average transmission energy consumption, the average transmission energy consumption during signal transmission can be effectively reduced. The beamforming signal model generated by the optimal communication probability function constructed based on the trade-off factor can greatly reduce the network energy consumption of distributed minimum variance distortionless response beamforming. The speech signal-to-noise ratio, subjective perceived quality, and speech intelligibility of the output enhanced speech information can all converge to the classical random rumor algorithm and the centralized minimum variance distortionless response beamforming method. Attached Figure Description
[0053] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0054] Figure 1 A flowchart illustrating a distributed speech enhancement method according to an embodiment of the present disclosure is shown schematically.
[0055] Figure 2 A schematic diagram of an experimental environment according to an embodiment of the present disclosure is shown;
[0056] Figure 3 A schematic diagram illustrating the output signal-to-noise ratio versus the input signal-to-noise ratio according to an embodiment of the present disclosure is shown.
[0057] Figure 4 The diagram schematically illustrates the cumulative energy consumption curve of the distributed speech enhancement method according to embodiments of the present disclosure with respect to the number of iterations.
[0058] Figure 5 The diagram illustrates the variation of the subjective perceived quality (PESQ) of the output signal with respect to the input signal-to-noise ratio according to an embodiment of the present disclosure.
[0059] Figure 6 The diagram illustrates the variation of speech intelligibility (STOI) with respect to the input signal-to-noise ratio according to embodiments of the present disclosure.
[0060] Figure 7 A block diagram schematically illustrates a speech enhancement apparatus according to an embodiment of the present disclosure; and
[0061] Figure 8 A block diagram of an electronic device suitable for implementing the methods described above, according to embodiments of the present disclosure, is illustrated schematically. Detailed Implementation
[0062] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0063] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0064] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0065] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).
[0066] In Wireless Acoustic Sensor Networks (WASNs) (also known as distributed microphone arrays), each sensor node houses one or more microphones, and multiple microphone nodes can form a distributed microphone network. The free distribution of microphone nodes and flexible network construction effectively overcome the spatial sampling limitations of traditional wired arrays. WASNs consist of independent nodes. Considering the limited battery energy resources of each sensor node, efficient and energy-saving distributed signal algorithms are needed to extend network lifetime and facilitate data transmission and processing between nodes. Randomized Gossip (RG) algorithms are widely used in WASNs. RG is an algorithm for finding average consensus in arbitrarily connected networks. To further improve the convergence speed of the RG algorithm, the fastest distributed linear average (FDLA) algorithm can be established by minimizing the second-largest eigenvalue of the network communication probability matrix. Considering network energy consumption constraints, the RG algorithm can be improved by balancing convergence speed and energy consumption, keeping network energy consumption within a controllable range.
[0067] Adaptive beamforming algorithms are widely used multi-channel speech enhancement algorithms. Among them, the minimum variance distortionless response (MVDR) beamforming algorithm is based on the minimum mean square error criterion. Under the constraint of no distortion of the target sound source signal, it minimizes the output signal power, which is equivalent to minimizing the output noise power.
[0068] With the continuous development of adaptive beamforming technology, distributed beamforming algorithms have also gradually emerged. For example, the MVDR algorithm is implemented using the acoustic transfer function in conjunction with a correlation function used for data estimation. The MVDR beamformer depends on the inverse correlation matrix corresponding to the noise plus the target signal, finding the optimal rate and optimal estimation weights among all sensors in the network. A related technique proposes a method for implementing distributed MVDR beamforming in WASN. This algorithm is based on the assumption that noise signals are uncorrelated between microphone nodes, simplifying the distributed computation of MVDR but sacrificing performance in canceling coherent noise.
[0069] However, during data transmission in WASN, since the sensors and information processors are not owned by a specific user, distributed processing may lead to privacy leaks. In speech enhancement applications, homomorphic encryption can be used to provide the necessary privacy. Considering that one user in the network maintains the privacy of the precise source of interest for the other users, this algorithm is quite complex. Further, a distributed beamforming signal estimation method based on the RG algorithm in WASN has been proposed, which also provides privacy protection for the source of interest. Considering that nodes in the network do not want to share their specific sources, WASN is needed to estimate the signal sources they are interested in. However, this method has a slow convergence speed and does not consider the node energy consumption problem.
[0070] In view of this, embodiments of the present disclosure provide a distributed speech enhancement method and a speech enhancement device. The method includes constructing an optimal communication probability function based on a speech parameter set, wherein the speech parameter set includes a tradeoff factor, the second largest eigenvalue of an update matrix, and average transmission energy consumption, where average transmission energy consumption characterizes the signal energy loss during speech transmission between multiple nodes, and the multiple nodes include a primary speech node and multiple secondary speech nodes; solving the optimal communication probability function to obtain a node pair communication probability matrix, wherein the node pair communication probability matrix characterizes the probability that the primary speech node selects any one of the secondary speech nodes to form a node pair; based on the node pair communication probability matrix, processing multiple sound signals received by the primary speech node and the multiple secondary speech nodes using preset processing rules to obtain a beamforming signal model; and inputting the input speech information obtained from the primary speech node into the beamforming signal model to output enhanced speech information.
[0071] Figure 1 A flowchart illustrating a distributed speech enhancement method according to an embodiment of the present disclosure is shown schematically.
[0072] like Figure 1 As shown, the distributed speech enhancement method includes operations S101 to S104.
[0073] In operation S101, an optimal communication probability function is constructed based on the voice parameter set. The voice parameter set includes a trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption. The average transmission energy consumption characterizes the loss of signal energy when transmitting voice between multiple nodes. The multiple nodes include a voice master node and multiple voice auxiliary nodes.
[0074] In operation S102, the optimal communication probability function is solved to obtain the node pair communication probability matrix, where the node pair communication probability matrix represents the probability that the speech master node selects any speech auxiliary node to form a node pair.
[0075] In operation S103, based on the node-to-node communication probability matrix, multiple audio signals received by the voice master node and multiple voice auxiliary nodes are processed using preset processing rules to obtain a beamforming signal model.
[0076] In operation S104, the input voice information obtained from the voice master node is input into the beamforming signal model, and the enhanced voice information is output.
[0077] According to embodiments of this disclosure, the tradeoff factor can be specifically set according to the specific application scenario. The second largest eigenvalue of the update matrix is the second largest eigenvalue of the update matrix, i.e., the second largest eigenvalue. The magnitude of this eigenvalue determines the convergence speed of the optimal communication probability function based on the energy-aware random rumor algorithm. The smaller the eigenvalue, the faster the convergence speed; typically, the maximum eigenvalue is 1. Average transmission energy consumption characterizes the average signal loss during signal transmission between any two nodes at different locations.
[0078] According to embodiments of this disclosure, before performing speech enhancement processing, an optimal communication probability function is first constructed using a set of speech parameters including a trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption. By introducing a trade-off factor to perform a weighted balance on the second largest eigenvalue of the update matrix and the average transmission energy consumption, an optimal communication probability function with a faster convergence speed is constructed.
[0079] According to embodiments of this disclosure, the optimal communication probability function can be further rewritten as a standard semi-definite programming (SDP) problem using epi-graph, and the node pair communication probability matrix P can be quickly solved using optimization tools such as SeDuMi or CVX.
[0080] According to embodiments of this disclosure, a voice node selected based on a node-to-node communication probability matrix processes multiple audio signals received by a primary voice node and multiple secondary voice nodes using preset processing rules to obtain a beamforming signal model. The input voice information obtained from the primary voice node is then input into the beamforming signal model, and enhanced voice information is output.
[0081] According to embodiments of this disclosure, by introducing a trade-off factor to weight and balance the second largest eigenvalue of the update matrix and the average transmission energy consumption, the average transmission energy consumption during signal transmission can be effectively reduced. The beamforming signal model generated by the optimal communication probability function constructed based on the trade-off factor can greatly reduce the network energy consumption of distributed minimum variance distortionless response beamforming. The speech signal-to-noise ratio, subjective perceived quality, and speech intelligibility of the output enhanced speech information can all converge to the classical random rumor algorithm and the centralized minimum variance distortionless response beamforming method.
[0082] According to embodiments of this disclosure, constructing an optimal communication probability function based on a set of speech parameters includes the following operations:
[0083] Generate an undirected graph of nodes based on a sensor network consisting of multiple nodes;
[0084] Based on the node-undirected graph, an average communication matrix is generated using an N-dimensional identity matrix and an N-dimensional column vector.
[0085] Generate column vector iterative expressions based on the average communication matrix and the initial column vector of node states;
[0086] If the first iteration number of the column vector iterative expression satisfies the first iteration number threshold, the iteration result of the column vector iterative expression is taken as the second largest eigenvalue of the update matrix, where the iteration result represents the second largest eigenvalue of the update matrix generated by the column vector iterative expression in the iteration.
[0087] Based on the trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption, an optimal communication probability function is constructed to characterize the energy-aware random rumor algorithm.
[0088] According to embodiments of this disclosure, consider a wireless acoustic sensor network comprising N nodes, which can be represented by an undirected graph G = (V, E), where V = {1, ..., N} and Let represent the set of fixed points and the set of edges, respectively. When node i and node j can communicate directly, (i, j) ∈ E.
[0089] According to embodiments of this disclosure, it is assumed that the nodes follow a symmetric relationship in communication, and the initial state value of node i is y. i The classic Randomized Gossip (RG) algorithm calculates the mean of all nodes by iteratively selecting two communicating nodes, calculating their average, and then updating the state values of the node pairs. Therefore, it is also called the average consensus problem, where the initial iteration value of each node is...
[0090] In the classic RG algorithm, the master node i in the k-th iteration is randomly selected. Let matrix P = [P ij ] represents the node-to-node communication probability matrix, i.e., P ij Let P represent the probability that the primary node i selects node j as the secondary node. Therefore, when (i, j) ∈ E, P ij >0. For node pair (i, j) at this time, the data update matrix is shown in formula (1).
[0091] W ij (k)=I N -0.5×(e i -e j (e) i -e j ) T (1)
[0092] Where I N Let e be an N-dimensional identity matrix. i Let x be an N-dimensional column vector (the i-th element is 1, and all other elements are 0). Represent all observed node state values as a column vector to obtain the initial column vector of node states x(k) = [x1(k), x...]. N (k), ..., x N (k)] T Then the kth iteration process can be written as the column vector iteration expression shown in formula (2).
[0093] x(k)=W ij (k)x(k-1) (2)
[0094] After k→∞ iterations, the average communication matrix is obtained. This indicates that the spectral radius of the update matrix is less than 1, and its largest eigenvalue is 1, thus ensuring the convergence of the RG algorithm, and the second largest eigenvalue... The convergence speed was measured by the maximum eigenvalue at this time. This involves updating the second largest eigenvalue of the matrix, where λ² represents the second largest eigenvalue. Based on this, the Fast Distributed Linear Average (FDLA) algorithm updates the second largest eigenvalue of the matrix. This improves the convergence speed of the classic RG algorithm. It should be noted that the threshold for the first iteration in this embodiment is ∞, which can be set according to actual needs.
[0095] According to embodiments of this disclosure, in determining the second largest eigenvalue of the update matrix... Then, by combining the trade-off factor and average transmission energy consumption, an optimal communication probability function is constructed to characterize the energy-aware random rumor algorithm.
[0096] According to embodiments of this disclosure, the average transmission power consumption is determined as follows:
[0097] Determine the Eulerian distance between the two nodes in each node pair; generate a squared distance matrix between node pairs based on multiple Eulerian distances; generate the average transmission energy consumption based on the update matrix and the squared distance matrix between node pairs.
[0098] According to embodiments of this disclosure, the free space path-loss (FSPL) of the communication node pair (i, j) in the k-th iteration is considered as shown in Equation (3):
[0099]
[0100] Where, d ij Let λ represent the Eulerian distance between node pairs (i, j), λ1 represent the wavelength, f be the frequency, and c be the speed of sound. The power consumed by communication between nodes is shown in formula (4).
[0101]
[0102] in, and Let represent the transmitted signal energy and the received signal energy, respectively. Assuming the received signal energy remains constant, it can be seen that the transmitted energy is directly proportional to the square of the transmission distance. Let The matrix representing the squared distance between node pairs is the average communication matrix. The average transmission energy consumption can be obtained by applying it to matrix D. Represents the set of real numbers.
[0103] Therefore, a tradeoff factor α is introduced to control the balance between convergence speed and network transmission energy consumption, thereby improving the optimization objective function of the FDLA method and obtaining the optimal communication probability function based on the energy-aware random rumor algorithm as shown in formula (5).
[0104]
[0105] Where P represents the node-pair communication probability matrix to be solved, and α represents the pre-defined tradeoff factor. The second largest eigenvalue represents the updated matrix, and λ² represents taking the second largest eigenvalue. The update matrix represents the average communication matrix described above, and D represents the squared distance matrix between node pairs. Characterizes average transmission energy consumption.
[0106] According to embodiments of this disclosure, based on a node-to-node communication probability matrix, multiple audio signals received by a primary voice node and multiple secondary voice nodes are processed using preset processing rules to obtain a beamforming signal model, including the following operations:
[0107] Short-time Fourier transform is performed on multiple sound signals to obtain a set of coefficients, which includes a subset of nodal Fourier coefficients, a subset of clean speech Fourier coefficients, and a subset of noise Fourier coefficients.
[0108] A vector signal model is constructed based on the set of coefficients and the set of steering vectors, wherein the set of steering vectors is generated based on the initial acoustic transfer function corresponding to each sound signal;
[0109] A beamforming signal model is generated based on the vector signal model and the target covariance matrix.
[0110] According to embodiments of this disclosure, the sound signal is generated based on a sound source signal, an initial sound transfer function, and noise;
[0111] The process involves performing a short-time Fourier transform on multiple sound signals to obtain a set of coefficients, including:
[0112] Perform short-time Fourier transform on multiple sound signals to obtain the node Fourier coefficients, clean speech Fourier coefficients, noise Fourier coefficients, and target sound transfer function corresponding to each sound signal in the frequency domain;
[0113] Based on multiple node Fourier coefficients, multiple clean speech Fourier coefficients, multiple noise Fourier coefficients, and multiple target acoustic transfer functions, subsets of node Fourier coefficients, subsets of clean speech Fourier coefficients, subsets of noise Fourier coefficients, and sets of turning vectors are constructed respectively.
[0114] According to embodiments of this disclosure, m represents the number of each node (e.g., a microphone), s is the sound source signal, and d is the initial sound transfer function. m The noise is n m Then the sound signal y received by microphone m m As shown in formula (6).
[0115] y m =d m *s+n m (6)
[0116] After performing a Short Time Fourier Transform (STFT) on multiple audio signals, Y m (f, t) represents the STFT coefficient of the m-th microphone in the array, i.e., the point Fourier coefficient. Similarly, S m (f, t) and N m(f, t) represent the STFT coefficients of the clean speech and noise signals, respectively, i.e., the Fourier coefficients of the clean speech and the Fourier coefficients of the noise, where f is the frequency index and t is the time frame. m (f, t) represents the target acoustic transfer function in the frequency domain. A vector is used to represent the signal received by the microphones in the array as a whole, i.e., the subset of nodal Fourier coefficients Y = [Y1, ..., Y]. M ] T Similarly, the subset of pure speech Fourier coefficients S and the subset of noise Fourier coefficients N can also be represented by vectors, d = [d1, ..., dn]. M ] T The vector signal model is shown in Equation (7), which is the set of steering vectors.
[0117] Y = dS + N = S + N (7)
[0118] According to embodiments of this disclosure, a beamforming signal model is generated based on a vector signal model and a target covariance matrix.
[0119] According to embodiments of this disclosure, a beamforming signal model is generated based on a vector signal model and a target covariance matrix, including the following operations:
[0120] Based on the vector signal model, the preset signal covariance matrix is inverted to obtain the initial beamforming iterative expression, where the initial beamforming iterative expression characterizes the target covariance matrix.
[0121] The initial beamforming iterative expression is iterated twice to obtain the first instantaneous average estimate and the second instantaneous average estimate, respectively.
[0122] Substituting the average value of the first instantaneous estimate and the average value of the second instantaneous estimate into the initial beamforming iterative expression, we obtain the target beamforming expression;
[0123] A beamforming signal model is generated based on the target beamforming expression and the set of steering vectors.
[0124] According to embodiments of this disclosure, given a signal covariance matrix, the inverse correlation matrix is used to estimate the MVDR beamformer in a distributed manner, resulting in an initial beamforming iterative expression as shown in Equation (8).
[0125]
[0126] Where H represents the transpose, and λ is a forgetting factor close to 1 (greater than 0 and less than 1).
[0127] According to embodiments of this disclosure, beamforming refers to the process of outputting an estimated signal by the inner product of the MVDR filter and the observed signal vector. The inner product of the two vectors can essentially be regarded as a summation or average consensus problem. Therefore, it is necessary to use the optimal communication probability function to iterate and output the estimated signal again. That is, the optimal communication probability matrix obtained by the energy-aware random rumor algorithm is used to select the node, and two rounds of iteration are required to obtain the average value, namely the first instantaneous estimated average value and the second instantaneous estimated average value.
[0128] According to an embodiment of this disclosure, the first instantaneous estimated average value and the second instantaneous estimated average value are substituted into the initial beamforming iterative expression to obtain the target beamforming expression as shown in formula (9).
[0129]
[0130] According to embodiments of this disclosure, based on the target beamforming expression and the steering vector set d = [d1, ..., d2], M ] T Generate a beamforming signal model.
[0131] According to embodiments of this disclosure, the initial beamforming iterative expression is iterated twice to obtain a first instantaneous average estimate and a second instantaneous average estimate, including the following operations:
[0132] If the second iteration number of the initial beamforming iterative expression meets the second iteration number threshold, the first instantaneous estimated average value is determined according to the vector signal model and the first intermediate beamforming expression, wherein the first intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the second iteration;
[0133] If the third iteration of the initial beamforming iterative expression satisfies the threshold of the third iteration, the second instantaneous estimated average value is determined based on the vector signal model and the second intermediate beamforming expression. The second intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the third iteration.
[0134] According to embodiments of this disclosure, the optimal communication probability matrix obtained by solving the energy-optimized RG algorithm is used to select nodes. The calculation process can be broken down into two average consensus subproblems, which require two rounds of iteration to obtain the average value, resulting in the first instantaneous estimated average value and the second instantaneous estimated average value as shown in formula (10).
[0135]
[0136]
[0137] First instantaneous estimated average value and the second instantaneous estimated average value These are the instantaneous estimates of the average values of the i-th microphone node in the k-th iteration and at time frame t, respectively, in the two iterations. It is assumed that the inverse correlation matrix at each microphone node in the sensor network is an identity matrix before iteration. In the first iteration, initial values are assigned to each node in frame t. Where Y m This refers to the STFT coefficient of the received signal at the m-th microphone in the network, corresponding to the i-th node. The average value among the nodes after each iteration can be expressed as... This is the estimated value at node i during the K1th iteration of the first round of iterations. The second round of iterations begins after the sensor network reaches an average consensus. Let represent the initial value of each node. In the t-th iteration, the average value between nodes i and j is calculated in the same way as in the first iteration, letting This represents the estimated value at node i during the K2th iteration of the second round of iterations. Substituting the average value into the expression of the inverse correlation matrix yields another expression of formula (9), as shown in formula (11).
[0138]
[0139] According to embodiments of this disclosure, a beamforming signal model is generated based on the target beamforming expression and the steering vector set, including the following operations:
[0140] Based on the target beamforming expression and the set of steering vectors, a first consensus vector is generated;
[0141] Based on the first consensus vector and the set of turning vectors, a second consensus vector is generated;
[0142] A beamforming signal model is generated based on the first consensus vector and the second consensus vector.
[0143] According to embodiments of this disclosure, after obtaining Subsequently, it can be seen that the centralized MVDR beamformer can also be decomposed into two average consensus problems, as shown in the first consensus vector of formula (12). Second consensus vector
[0144]
[0145] When the first consensus vector Second consensus vector After convergence, the expression for the distributed adaptive MVDR beamformer is shown in Equation (13).
[0146]
[0147] It can be seen that after two rounds of energy-sensing random rumors iteration, the MVDR filter coefficients of each node reached a consensus.
[0148] According to embodiments of this disclosure, based on the target beamforming expression and the steering vector set d = [d1, ..., d2], M ] T The beamforming signal model shown in formula (14) is generated.
[0149]
[0150] According to embodiments of this disclosure, the input speech information Y(t) obtained from the speech master node is input to the beamforming signal model, and enhanced speech information is output.
[0151] Figure 2 A schematic diagram of an experimental environment according to an embodiment of the present disclosure is shown; Figure 3 A schematic diagram illustrating the output signal-to-noise ratio versus the input signal-to-noise ratio according to an embodiment of the present disclosure is shown. Figure 4 The diagram illustrates the cumulative energy consumption curve of the distributed speech enhancement method according to an embodiment of the present disclosure with respect to the number of iterations.
[0152] In one exemplary embodiment, to verify the effectiveness of the distributed speech enhancement method of this disclosure, a three-dimensional room (6m × 8m × 3m) was simulated, in which 10 wireless sensor nodes were randomly distributed. Each node, without loss of generality, contained only one microphone, and a target sound source and an interfering sound source were placed around the node, such as... Figure 2 The diagram shows 10 randomly distributed microphone nodes (hollow circles), 1 target sound source (black dot), and 1 interference sound source (asterisk). The clean audio signal is a 30-second speech signal. Both the speech source and the interference source are from the Timit database, with a sampling rate of 16kHz. The microphone self-noise is independent and identically distributed additive white Gaussian noise (signal-to-noise ratio of 50dB). The energy ratio between the speech source and the interference source is defined as the pure signal-to-noise ratio (controlling the input signal quality). Other relevant parameter settings are as follows: forgetting factor λ = 0.999, short-time Fourier transform frame length of 32ms, and a Hanning window with 50% overlap. Furthermore, to focus on the distributed speech enhancement problem, it is assumed that all sensor nodes use a synchronized sampling frequency.
[0153] According to embodiments of this disclosure, Figure 3 and Figure 4 The output signal-to-noise ratio and cumulative transmission energy consumption of the enhanced speech information were compared between the classic random rumor algorithm and the distributed speech enhancement method disclosed in this paper. Figure 4 The parameters 0.01 and 0.03 represent the energy consumption of the distributed speech enhancement methods of this disclosure under different trade-off factors. From... Figure 3 and Figure 4 It can be seen that the output signal-to-noise ratio (SNR) of both the distributed speech enhancement method of this disclosure and the classical RG method converges to the result of the centralized MVDR beamforming method (this result represents the optimal case). This convergence relationship is independent of the input SNR, and the enhanced speech signal quality improves with increasing input SNR. Furthermore, the distributed speech enhancement method of this disclosure has a significant advantage in energy consumption compared to the classical RG and FDLA algorithms; increasing the energy consumption optimization weight parameters can significantly reduce output energy consumption. Table 1 specifically shows the overall network energy consumption of the classical RG algorithm and the distributed speech enhancement method of this disclosure under different input SNR conditions. As can be seen from Table 1, the distributed speech enhancement method of this disclosure can save approximately 50% of transmission energy consumption.
[0154] Table 1 Overall Network Transmission Energy Consumption
[0155] Input signal-to-noise ratio (dB) Classical RG method The distributed speech enhancement method disclosed herein -5 1.0508e+09 5.2179e+08 -2.5 1.0213e+09 5.1842e+08 0 1.0110e+09 4.9982e+08 2.5 1.0127e+09 5.1200e+08 5 1.0133e+09 5.0415e+08
[0156] Figure 5 The diagram illustrates the variation of the subjective perceived quality (PESQ) of the output signal with respect to the input signal-to-noise ratio according to an embodiment of the present disclosure. Figure 6 The diagram illustrates the variation of speech intelligibility (STOI) with respect to the input signal-to-noise ratio according to an embodiment of the present disclosure.
[0157] According to embodiments of this disclosure, Figure 5 and Figure 6 This study compared the relationship between the subjective perceived quality of speech (PESQ) and short-time objective speech intelligibility (STOI) of several methods and the input signal-to-noise ratio. Higher values for both metrics indicate better speech quality. Figure 5 and Figure 6 As can be seen, the classic RG algorithm (random rumors in the illustration) and the distributed speech enhancement method of this disclosure (energy optimization in the illustration) achieve comparable performance, with the difference from the optimal result obtained by centralized MVDR beamforming (fusion center in the illustration) being negligible, and both can significantly improve the signal quality of the enhanced speech information. Figure 5 and Figure 6 The “denoising” in the middle diagram represents the original speech signal.
[0158] In summary, the distributed speech enhancement method disclosed herein applies the energy-optimized RG model, which can effectively save network energy consumption, to MVDR beamforming. This can significantly reduce the network energy consumption of distributed MVDR beamforming, and the output speech signal-to-noise ratio, subjective perceived quality, and speech intelligibility can all converge to the classical RG algorithm and the centralized MVDR beamforming method.
[0159] Figure 7 A block diagram of a speech enhancement apparatus according to an embodiment of the present disclosure is shown schematically.
[0160] like Figure 4 As shown, the speech enhancement device 700 includes a construction module 710, a solution module 720, a processing module 730, and an enhancement module 740.
[0161] The construction module 710 is used to construct the optimal communication probability function based on the speech parameter set, wherein the speech parameter set includes a trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption. The average transmission energy consumption characterizes the loss of signal energy when transmitting speech between multiple nodes, and the multiple nodes include a primary speech node and multiple secondary speech nodes.
[0162] The solver module 720 is used to solve the optimal communication probability function to obtain the node pair communication probability matrix, where the node pair communication probability matrix represents the probability that the speech master node selects any speech auxiliary node to form a node pair.
[0163] The processing module 730 is used to process multiple audio signals received by the voice master node and multiple voice auxiliary nodes based on the node-to-node communication probability matrix and using preset processing rules to obtain a beamforming signal model.
[0164] The enhancement module 740 is used to input the input speech information obtained from the speech master node into the beamforming signal model and output enhanced speech information.
[0165] According to embodiments of this disclosure, by introducing a trade-off factor to weight and balance the second largest eigenvalue of the update matrix and the average transmission energy consumption, the average transmission energy consumption during signal transmission can be effectively reduced. The beamforming signal model generated by the optimal communication probability function constructed based on the trade-off factor can greatly reduce the network energy consumption of distributed minimum variance distortionless response beamforming. The speech signal-to-noise ratio, subjective perceived quality, and speech intelligibility of the output enhanced speech information can all converge to the classical random rumor algorithm and the centralized minimum variance distortionless response beamforming method.
[0166] According to embodiments of this disclosure, the construction module 710 includes a first generation submodule, a second generation submodule, a third generation submodule, a obtaining submodule, and a first construction submodule.
[0167] The first generation submodule is used to generate an undirected graph of nodes based on a sensor network consisting of multiple nodes.
[0168] The second generation submodule is used to generate an average communication matrix based on the node undirected graph, using an N-dimensional identity matrix and an N-dimensional column vector.
[0169] The third generation submodule is used to generate column vector iterative expressions based on the average communication matrix and the initial column vector of node states.
[0170] The resulting submodule is used to take the iteration result of the column vector iterative expression as the second largest eigenvalue of the update matrix when the first iteration number of the column vector iterative expression meets the first iteration number threshold. Here, the iteration result represents the second largest eigenvalue of the update matrix generated by the column vector iterative expression in the iteration.
[0171] The first construction submodule is used to construct the optimal communication probability function based on the energy-aware random rumor algorithm, according to the trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption.
[0172] According to embodiments of this disclosure, the average transmission energy consumption is determined by a determining submodule, a fourth generating submodule, and a fifth generating submodule.
[0173] The determination submodule is used to determine the Eulerian distance between the two nodes in each node pair.
[0174] The fourth generation submodule is used to generate a squared distance matrix between node pairs based on multiple Eulerian distances.
[0175] The fifth generation submodule is used to generate the average transmission energy consumption based on the update matrix and the squared distance matrix between node pairs.
[0176] According to embodiments of this disclosure, the processing module 730 includes a transformation submodule, a second construction submodule, and a sixth generation submodule.
[0177] The transform submodule is used to perform short-time Fourier transform on multiple audio signals to obtain a set of coefficients, which includes a subset of node Fourier coefficients, a subset of clean speech Fourier coefficients, and a subset of noise Fourier coefficients.
[0178] The second construction submodule is used to construct a vector signal model based on the set of coefficients and the set of steering vectors, wherein the set of steering vectors is generated based on the initial acoustic transfer function corresponding to each sound signal.
[0179] The sixth generation submodule is used to generate a beamforming signal model based on the vector signal model and the target covariance matrix.
[0180] According to embodiments of this disclosure, the sound signal is generated based on a sound source signal, an initial sound transfer function, and noise.
[0181] According to embodiments of this disclosure, the transformation submodule includes a transformation unit and a construction unit.
[0182] The transform unit is used to perform short-time Fourier transform on multiple sound signals to obtain the node Fourier coefficients, clean speech Fourier coefficients, noise Fourier coefficients, and target sound transfer function corresponding to each sound signal in the frequency domain.
[0183] The construction unit is used to construct subsets of node Fourier coefficients, subsets of clean speech Fourier coefficients, subsets of noise Fourier coefficients, and sets of turning vectors based on multiple node Fourier coefficients, multiple clean speech Fourier coefficients, multiple noise Fourier coefficients, and multiple target acoustic transfer functions.
[0184] According to embodiments of this disclosure, the sixth generation submodule includes an inversion unit, an iteration unit, a substitution unit, and a generation unit.
[0185] The inversion unit is used to invert the preset signal covariance matrix based on the vector signal model to obtain the initial beamforming iterative expression, where the initial beamforming iterative expression characterizes the target covariance matrix.
[0186] The iterative unit is used to perform two iterations on the initial beamforming iterative expression to obtain the first instantaneous estimate average and the second instantaneous estimate average, respectively.
[0187] The substitution unit is used to substitute the average value of the first instantaneous estimate and the average value of the second instantaneous estimate into the initial beamforming iterative expression to obtain the target beamforming expression.
[0188] The generation unit is used to generate a beamforming signal model based on the target beamforming expression and the set of steering vectors.
[0189] According to embodiments of this disclosure, the iteration unit includes a second iteration subunit and a third iteration subunit.
[0190] The second iterative subunit is used to determine the first instantaneous estimated average value based on the vector signal model and the first intermediate beamforming expression, provided that the second iteration number of the initial beamforming iterative expression meets the second iteration number threshold. The first intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the second iteration.
[0191] The third iteration subunit is used to determine the second instantaneous estimated average value based on the vector signal model and the second intermediate beamforming expression, provided that the third iteration number of the initial beamforming iterative expression meets the third iteration number threshold. The second intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the third iteration.
[0192] According to embodiments of this disclosure, the generation unit includes a first generation subunit, a second generation subunit, and a third generation subunit.
[0193] The first generating sub-unit is used to generate the first consensus vector based on the target beamforming expression and the set of steering vectors.
[0194] The second generation subunit is used to generate a second consensus vector based on the first consensus vector and the set of turning vectors.
[0195] The third generation subunit is used to generate a beamforming signal model based on the first consensus vector and the second consensus vector.
[0196] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuits, such as Field Programmable Gate Arrays (FPGAs), Programmable Logic Arrays (PLAs), Systems-on-Chip, Systems-on-Substrate, Systems-on-Package, Application-Specific Integrated Circuits (ASICs), or implemented by hardware or firmware through any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, or firmware, or in a suitable combination of any one or more of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0197] For example, any plurality of the construction module 710, solving module 720, processing module 730, and enhancement module 740 can be combined into one module / unit / subunit, or any one of these modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the construction module 710, solving module 720, processing module 730, and enhancement module 740 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the building module 710, solving module 720, processing module 730, and enhancement module 740 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0198] It should be noted that the speech enhancement device part in the embodiments of this disclosure corresponds to the distributed speech enhancement method part in the embodiments of this disclosure. The description of the speech enhancement device part is specifically referred to in the distributed speech enhancement method part, and will not be repeated here.
[0199] Figure 8 A block diagram of an electronic device suitable for implementing the methods described above, according to embodiments of the present disclosure, is illustrated schematically. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0200] like Figure 8As shown, an electronic device 800 according to an embodiment of this disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 802 or a program loaded from storage portion 808 into random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this disclosure.
[0201] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0202] According to embodiments of this disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The system 800 may also include one or more of the following components connected to the I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.
[0203] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by processor 801, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0204] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0205] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0206] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.
[0207] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the distributed speech enhancement method provided in the embodiments of this disclosure.
[0208] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0209] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0210] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0211] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features recited in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not expressly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0212] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A distributed speech enhancement method, comprising: An optimal communication probability function is constructed based on a set of speech parameters, wherein the set of speech parameters includes a trade-off factor, the second largest eigenvalue of the update matrix, and average transmission energy consumption. The average transmission energy consumption characterizes the loss of signal energy when transmitting speech between multiple nodes, and the multiple nodes include a primary speech node and multiple secondary speech nodes. Solving the optimal communication probability function yields a node pair communication probability matrix, where the node pair communication probability matrix represents the probability that the voice master node selects any one of the voice auxiliary nodes to form a node pair. Based on the node-to-node communication probability matrix, multiple audio signals received by the main voice node and multiple auxiliary voice nodes are processed using preset processing rules to obtain a beamforming signal model. The input voice information obtained from the voice master node is input into the beamforming signal model, and the enhanced voice information is output. The step of constructing the optimal communication probability function based on the speech parameter set includes: Based on the sensor network composed of multiple nodes, an undirected graph of nodes is generated; Based on the aforementioned undirected graph of nodes, an average communication matrix is generated according to an N-dimensional identity matrix and an N-dimensional column vector. Based on the average communication matrix and the initial column vector of node states, generate an iterative expression for the column vector; If the first iteration number of the column vector iterative expression satisfies the first iteration number threshold, the iteration result of the column vector iterative expression is taken as the second largest eigenvalue of the update matrix, wherein the iteration result represents the second largest eigenvalue of the update matrix generated by the column vector iterative expression in the iteration; Based on the aforementioned trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption, the optimal communication probability function based on the energy-aware random rumor algorithm is constructed. The optimal communication probability function is shown in formula (1): Where P represents the node-pair communication probability matrix to be solved. Characterizing the pre-defined trade-off factors, Represents the second largest eigenvalue of the update matrix. The representation is taken as the second largest eigenvalue. Representation update matrix, The squared distance matrix representing the node pairs Characterizes average transmission energy consumption.
2. The method according to claim 1, wherein, The average transmission power consumption is determined in the following manner: Determine the Eulerian distance between the two nodes in each node pair; Generate a squared distance matrix between node pairs based on the multiple Eulerian distances described; The average transmission energy consumption is generated based on the update matrix and the squared distance matrix between node pairs.
3. The method according to claim 1, wherein, The process of processing multiple audio signals received by the primary voice node and multiple secondary voice nodes based on the node-to-node communication probability matrix using preset processing rules to obtain a beamforming signal model includes: Perform a short-time Fourier transform on multiple sound signals to obtain a set of coefficients, wherein the set of coefficients includes a subset of node Fourier coefficients, a subset of clean speech Fourier coefficients, and a subset of noise Fourier coefficients. A vector signal model is constructed based on the set of coefficients and the set of steering vectors, wherein the set of steering vectors is generated based on the initial acoustic transfer function corresponding to each of the sound signals; The beamforming signal model is generated based on the vector signal model and the target covariance matrix.
4. The method according to claim 3, wherein the sound signal is generated based on the sound source signal, the initial sound transfer function, and noise; in, The step of performing a short-time Fourier transform on the multiple sound signals to obtain a set of coefficients includes: Perform a short-time Fourier transform on the multiple sound signals to obtain the node Fourier coefficients, clean speech Fourier coefficients, noise Fourier coefficients, and target sound transfer function corresponding to each sound signal in the frequency domain; Based on multiple node Fourier coefficients, multiple clean speech Fourier coefficients, multiple noise Fourier coefficients, and multiple target acoustic transfer functions, the node Fourier coefficient subset, the clean speech Fourier coefficient subset, the noise Fourier coefficient subset, and the steering vector set are constructed respectively.
5. The method according to claim 3, wherein, The step of generating the beamforming signal model based on the vector signal model and the target covariance matrix includes: Based on the vector signal model, the preset signal covariance matrix is inverted to obtain an initial beamforming iterative expression, wherein the initial beamforming iterative expression characterizes the target covariance matrix; The initial beamforming iterative expression is iterated twice to obtain the first instantaneous average estimate and the second instantaneous average estimate, respectively. Substituting the first instantaneous average estimate and the second instantaneous average estimate into the initial beamforming iterative expression, the target beamforming expression is obtained; The beamforming signal model is generated based on the target beamforming expression and the steering vector set.
6. The method according to claim 5, wherein, The step of performing two iterations on the initial beamforming iterative expression to obtain the first instantaneous average estimate and the second instantaneous average estimate, includes: If the second iteration number of the initial beamforming iterative expression satisfies the second iteration number threshold, the first instantaneous estimated average value is determined according to the vector signal model and the first intermediate beamforming expression, wherein the first intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the second iteration; If the third iteration of the initial beamforming iterative expression satisfies the third iteration threshold, the second instantaneous estimated average value is determined according to the vector signal model and the second intermediate beamforming expression, wherein the second intermediate beamforming expression is obtained by inverting the initial beamforming iterative expression after the third iteration.
7. The method according to claim 5, wherein, The step of generating the beamforming signal model based on the target beamforming expression and the steering vector set includes: A first consensus vector is generated based on the target beamforming expression and the set of steering vectors. A second consensus vector is generated based on the first consensus vector and the set of turning vectors; The beamforming signal model is generated based on the first consensus vector and the second consensus vector.
8. A speech enhancement device, comprising: The construction module is used to construct an optimal communication probability function based on a speech parameter set, wherein the speech parameter set includes a trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption. The average transmission energy consumption characterizes the loss of signal energy when transmitting speech between multiple nodes, and the multiple nodes include a primary speech node and multiple secondary speech nodes. The solution module is used to solve the optimal communication probability function to obtain the node pair communication probability matrix, wherein the node pair communication probability matrix represents the probability that the voice master node selects any voice auxiliary node to form a node pair; The processing module is used to process multiple sound signals received by the voice master node and multiple voice auxiliary nodes based on the node pair communication probability matrix and using preset processing rules to obtain a beamforming signal model. The enhancement module is used to input the input voice information obtained from the voice master node into the beamforming signal model and output enhanced voice information. The building modules include: The first generation submodule is used to generate an undirected graph of nodes based on a sensor network consisting of multiple nodes. The second generation submodule is used to generate an average communication matrix based on the node undirected graph, according to an N-dimensional identity matrix and an N-dimensional column vector. The third generation submodule is used to generate column vector iterative expressions based on the average communication matrix and the initial column vector of node states; The submodule is obtained to take the iteration result of the column vector iterative expression as the second largest eigenvalue of the update matrix when the first iteration number of the column vector iterative expression meets the first iteration number threshold. The iteration result represents the second largest eigenvalue of the update matrix generated by the column vector iterative expression in the iteration. The first construction submodule is used to construct the optimal communication probability function based on the energy-aware random rumor algorithm, according to the trade-off factor, the second largest eigenvalue of the update matrix, and the average transmission energy consumption. The optimal communication probability function is shown in formula (2): Where P represents the node-pair communication probability matrix to be solved. Characterizing the pre-defined trade-off factors, Represents the second largest eigenvalue of the update matrix. The representation is taken as the second largest eigenvalue. Representation update matrix, The squared distance matrix representing the node pairs Characterizes average transmission energy consumption.
Citation Information
Patent Citations
Method for sound source direction estimation based on time frequency masking and deep neural network
CN109839612A
Voice Activity Detection based on Deep Neural Network Using EVS Codec Parameter and Voice Activity Detection Method thereof
KR101704925B1