Wireless acoustic sensor network speech enhancement method based on node selection

Through the sparse sensing model and triple norm concept, the complexity problem of frequency-dependent node selection in distributed WASN is solved, and efficient voice enhancement is achieved, suitable for a variety of application scenarios.

CN119943080APending Publication Date: 2025-05-06INNER MONGOLIA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510095309.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the frequency-dependent node selection method of distributed wireless acoustic sensor network (WASN) in the field of speech enhancement is complex, and the frequency invariant node selection (FI-SS) method is lacking in distributed topology, resulting in high solution complexity and low computing efficiency.

Method used

Using the sparse sensing model, the sparseness of the filter coefficients is minimized and the output noise power is limited, and the concept of triple norms is developed in a distributed environment is transformed into a convex semi-positive fixed planning problem to improve the solution efficiency.

Benefits of technology

Reduces the complexity of solution, improves computing efficiency, and improves voice quality. It is suitable for scenarios such as environmental monitoring, smart home and conference call.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943080A_ABST
    Figure CN119943080A_ABST
Patent Text Reader

Abstract

The invention discloses a wireless acoustic sensor network speech enhancement method based on node selection. Firstly, it is assumed that nodes are equipped with acoustic sensors, a signal model is established, signal frames are distinguished through a voice activity detector, and a correlation matrix is calculated. Then, filter coefficient group sparsity is improved, an FI-SS strategy is formulated, a triple norm concept is introduced and relaxed into a convex positive semidefinite programming problem to be solved, a weight matrix is used for replacement, and iterative calculation is carried out to promote a sparse solution; in the distributed node selection, graph correlation constraints are constructed, a cost function is inspired and constructed by a centralized weighting strategy, a weight updating standard is formulated, weights are updated by using two rules to ensure sparseness and connectivity of sub-networks, finally elements smaller than a threshold value are zeroed, and an MVDR beam former is operated by using residual nodes. According to the method, the solving complexity is reduced, the calculation efficiency is improved, the connectivity of the selected sub-network is ensured, the voice quality is improved, and the method has a good application prospect in multiple scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of wireless acoustic sensor network speech enhancement, in particular to a wireless acoustic sensor network speech enhancement method based on node selection. Background Art

[0002] With the rapid development of sensors, wireless communications, and integrated circuits, there are more and more smart terminals (such as smartphones, laptops, and tablets) with one or more microphones around us. Through network connection, these devices can form a large-aperture wireless acoustic sensor network (WASN), the so-called Internet of audio things (IoAuT). Compared with traditional microphone arrays, WASN can achieve higher spatial resolution and beamforming gain, so it has considerable potential in scenarios such as environmental monitoring, smart home, and teleconferencing.

[0003] As a next-generation audio / speech processing platform, WASN has shown a natural advantage in speech enhancement (SE), which aims to recover the target signal from noisy speech. A typical example is that some devices (nodes) in WASN may be close to the target sound source, so that high-quality audio signals can be collected, which is very beneficial for speech enhancement. Regarding the network topology of WASN, speech enhancement can be implemented in a centralized way or a distributed way. In the centralized way, the fusion center (FC) is responsible for physically connecting all nodes and performing computational processing. On the other hand, the distributed way allows all nodes to process data independently and in parallel without relying on the FC, and all nodes only communicate with their neighboring nodes. In contrast, distributed speech enhancement is more energy-efficient, but has higher requirements for data security. In both modes, in theory, activating the entire WASN can provide the best speech enhancement performance. However, in practical applications, considering factors such as power consumption, computational complexity, and sensor life, it is questionable whether to enable the entire WASN. Therefore, how to select a subset from WASN to find a good balance between performance and cost has become a research focus in the field of speech enhancement.

[0004] Recently, we designed a distributed sensor selection (SS) method to ensure the connectivity of the selected subnetwork by introducing more constraints and combining graph theory. Later, we also proposed a distributed spatiotemporal filter for speech enhancement in WASN and explored a near-optimal greedy SS algorithm by removing one node at each iteration until the subnetwork size reaches a preset value. However, the above methods are all designed in a frequency-dependent manner, which may cause conflicts in SS results between different frequencies, thus requiring additional switching processes. More seriously, some sensor nodes may not be equipped with computing units to support Fourier transform and switching operations. In contrast, frequency-invariant sensor selection (FI-SS) directly selects sensors from broadband audio signals, which avoids the complex switching operations required in frequency-dependent SS. In the latest FI-SS research, FI-SS is performed by constraining the broadband signal-to-noise ratio (SNR) while minimizing the total power consumption, which involves a non-convex Boolean programming problem. Then, the solution is obtained by relaxing the Boolean variables into continuous variables using the broadband optimization (Broadband optimization, BroadOpt) or narrowband voting (Narrowband voting, Navo) method. However, in this variable conversion stage, additional relaxation errors are inevitable. In order to improve the efficiency of FI-SS, some greedy search methods have been further proposed, including gradient removal (Gradient removal, GradR), weighted input signal-to-noise ratio removal (SNRremoval, SnrR) and broadband energy removal (Energy removal, EnergyR) methods. However, all these FI-SS methods are designed for centralized SE, and there is almost no research on distributed FI-SS.

[0005] Aiming at the challenges that the above non-greedy method contains NP-difficult Boolean optimization problems and the lack of FI-SS methods suitable for WASN under distributed topology, the present invention develops a novel sparse sensing model to avoid the Boolean planning problem in FI-SS and develops a distributed FI-SS method. Summary of the invention

[0006] 1. Technical issues to be solved

[0007] In view of the shortcomings of the prior art, the present invention provides a wireless acoustic sensor network speech enhancement method based on node selection, which has the advantages of reducing the solution complexity, improving the calculation efficiency, and improving the speech quality. It has good application prospects in scenarios such as environmental monitoring, smart home, and telephone conferencing, and solves the problem of complex solution caused by non-greedy algorithms and the lack of FI-SS methods for WASN under distributed topology.

[0008] (II) Technical solution

[0009] In order to achieve the above-mentioned purpose of reducing the complexity of solution, improving the calculation efficiency, improving the voice quality, and having good application prospects in scenarios such as environmental monitoring, smart home, and telephone conferencing, the present invention provides the following technical solutions:

[0010] A method for speech enhancement in a wireless acoustic sensor network based on node selection comprises the following steps:

[0011] Step 1:

[0012] Assume that each node in WASN is equipped with an acoustic sensor. In the discrete Fourier transform (DFT) domain, let k and t be the frequency index and frame index respectively, and the noise signal y recorded by node i is i (k,t), can be expressed as:

[0013] y i (k,t)=x i (k,t)+v i (k,t) (1)

[0014] where x i (k,t) and v i (k, t) represent the signal component and additive noise component at the target node i, respectively. For the case of a single target source, x i (k,t) can be written as:

[0015] x i (k,t)=a i (k)s(k,t) (2)

[0016] where a i (k) is the acoustic transfer function from the target source to node i, and s(k,t) is the corresponding target signal. In general, a i (k) can be derived from the locations of the sound source and the node. In addition, the relative transfer function a i(k) / a1(k) can be estimated by covariance subtraction / whitening (CS / CW) method. Let y(k,t)=[y1(k,t),y2(k,t),...,y I (k,t)] T , then the vector form of (1) is:

[0017] y(k,t)=x(k,t)+v(k,t)=a(k)s(k,t)+v(k,t) (3)

[0018] where x(k,t), v(k,t) and a(k) are vectors containing x i (k,t),v i (k,t) and a i (k). Assuming that the target speech and noise are uncorrelated, the correlation matrix of the received signal can be calculated as:

[0019]

[0020] in In practice, after applying the voice activity detector (VAD) technology (setting an energy threshold, when the energy of the current frame is greater than the threshold, it is judged as voice activity; when the frame energy is lower than the threshold, it is judged as a noise frame) to distinguish the received signal into noise-only frames and target plus noise frames, the noise correlation matrix R can be estimated from the noise-only frames. vv (k). Based on this, we can yy Subtract R from (k) vv (k) to calculate the speech correlation matrix R xx (k), i.e. R xx (k) = R yy (k)-R vv (k).

[0021] Step 2:

[0022] FI-SS is formulated by improving the group sparsity of the filter coefficients. First, the global vector w is defined as:

[0023]

[0024] where w i is a vector of filter coefficients for all frequencies of node i, w i The q-norm of is:

[0025]

[0026] If wi If the q-norm of [it] is equal to zero, we can put the [i]-th node into the sleep state. Therefore, we enhance the group sparsity of [w] by minimizing the l 1,q norm of [w] while restricting the output noise power, resulting in:

[0027]

[0028] where, ‖[w] i ‖ q The l1 norm is a convex approximation of the l0 norm, C1 is the constraint for the target signal without distortion, and C2 constrains the output noise (Output noise, oNoise) with a higher bound r(k):

[0029] oNoise min (k) ≤ r(k) ≤ oNoise max (k) (8)

[0030] where, oNoise min (k) and oNoise max (k) represent the minimum and maximum output noise powers at frequency k. Usually, oNoise max (k) can be determined by the input noise power, while oNoise min (k) can be obtained by activating all nodes. In the present invention, we introduce a scaling factor α to control r(k):

[0031]

[0032] where α min < α ≤ 1, The larger the value of α, the higher the expected noise reduction performance.

[0033] Step 3:

[0034] To improve the solution efficiency, the present invention develops the concept of the triple norm. Define the triple norm l 1,q[N],z of [w] as:

[0035]

[0036] where N is the frequency division factor that divides K frequencies into N parts. If N = 1, then the l 1,q[N],z norm reduces to the l 1,z norm; if N = K, then the l 1,q[N],z norm reduces to the l 1,q norm; 1 < N < K is a more general setting that can flexibly adjust the complexity of the subsequent FI-SS problem.

[0037] In this paper, we choose q = ∞ and z = 2. Then, by minimizing while maintaining constraints C1 and C2 A new cost function is established:

[0038]

[0039] Since constraint C2 is non-convex, we relax the above problem into a convex semi-definite programming (SDP) problem for efficient solution.

[0040] Step 4:

[0041] To further generalize the sparse solution, we use the weight matrix U I Replacement Matrix 1 I , and is calculated iteratively. Specifically, let l be the step index of the iteration, and the SDP problem in (23) can be modified as follows:

[0042]

[0043] Among them U (l) With all 1 matrix 1 I Initialization, that is, U (0) =1 I , and U (l) The (i,j)th element of It can be calculated iteratively using the unit rank penalty strategy:

[0044]

[0045] where ∈ is a small constant that prevents the denominator from being zero. Represents Γ (l) The (i,j)th element of is calculated as:

[0046] Γ (l) =γ (l) (γ (l) ) T (26)

[0047]

[0048] The operator Calculate the principal eigenvector of the matrix. It is worth noting that Γ (l) is a rank-one matrix consisting of each block diagonal matrix In addition, this reweighting method makes Small elements in can be penalized aggressively, encouraging them to go to zero as iterations proceed. As a result, a sparser solution can be obtained while ensuring performance constraints (since C3 is satisfied in all reweighting iterations).

[0049] Due to computational errors, it is difficult to obtain a strict sparse Therefore, we will Elements in η smaller than a very small threshold η0 are set to zero, which can achieve strict sparsity. But it will change the form of the solution obtained from (24), which may lead to a slight degradation in performance. To this end, we put a node to sleep if its filter coefficients on all frequencies are zero and rerun the minimum variance distortionless response (MVDR) method using the remaining nodes.

[0050] Step 5:

[0051] Rerun the MVDR beamformer using the remaining nodes. Output signal after speech enhancement It is expressed as:

[0052]

[0053] Where w(k)=[w1(k),w2(k),…,w I (k)] T is the filter coefficient of the kth frequency. In the classic MVDR method, w(k) is estimated by minimizing the output noise power under the constraint of maintaining the distortion-free target source. Specifically, it can be expressed as:

[0054]

[0055] By using the Lagrange multiplier method, the analytical solution for w can be expressed as:

[0056]

[0057] When you rerun the MVDR method using the remaining nodes, the vector a(k), the vector y(k,t), and the matrix R vv (k) should only contain the information of the remaining nodes, so the dimension of w(k) is the number of remaining selected nodes. Finally, use formula (28) for speech enhancement.

[0058] Preferably, in step three:

[0059] set up is another vector consisting of all filter coefficients, It can be expressed as:

[0060]

[0061] In addition, we introduce the rank 1 matrix As:

[0062]

[0063] in, Therefore, averaging C2 over frequency equals:

[0064]

[0065] where operation tr(·) is the trace of the computation matrix. Compared with C2, C3 specifies the overall denoising performance and expands the feasible domain of (11) because the number of constraints is reduced from K to 1. Based on the above, problem (11) can be equivalently rewritten as:

[0066]

[0067] where constraints C4 and C5 are used to replace (13), the symbol ≥ denotes a matrix inequality, i.e. X ≥ Y means that XY is positive semidefinite, and the operation rank(·) computes the rank of the matrix, is an auxiliary variable defined as follows:

[0068]

[0069] in is the index variable of k. The objective of (15) can be derived as:

[0070]

[0071] in It can be expressed as:

[0072]

[0073] Since the non-diagonal matrix does not appear in the constraints C1 and C3 in (15), we set them to zero according to the objective of (15), i.e. In order to remove these non-diagonal matrices from the SDP problem (15). Under this parameter setting, (18) can be further simplified as:

[0074]

[0075] in is a submatrix Note that in (19), the derivation from the second to the third line is due to It was established.

[0076] By defining an upper bound matrix Its (i,j)th element Always meet:

[0077]

[0078] After substituting (19) into (17), the objective in (17) can be further expressed as:

[0079]

[0080] In addition, since the non-diagonal matrix It has been deleted in (18) below, and constraint C4 can be simplified to:

[0081]

[0082] Therefore, by abandoning the non-convex constraint C5, problem (15) can be relaxed as:

[0083]

[0084] where C8 is an additional redundant constraint constructed from C1 and (13). The relaxed problem (21) is convex and can be solved efficiently using interior point methods or solvers such as the CVX toolbox.

[0085] Preferably, in steps 1 to 5, the distributed node selection includes the following steps:

[0086] Step 1: Construct graph-related constraints to ensure connectivity in the distributed topology:

[0087]

[0088] Among them I w is the number of non-zero groups in w, in is the adjacency matrix of the subnetwork obtained by sparse w. Since C9 involves nonlinear and nonconvex formulas, it is difficult to solve. Therefore, we invented a heuristic method to find a near-optimal solution.

[0089] Step 2: Inspired by the reweighting strategy in centralized node selection, we try to find appropriate weights to ensure the connectivity of the selected subnetwork. At this time, a cost function with the same form as (24) is constructed:

[0090]

[0091] in is the weighting matrix introduced at the lth iteration.

[0092] Step 3: We formulate a weight update criterion. For a given At the lth iteration, we can solve (32) to obtain At the same time, according to definitions (5) and (12), w (l) Can be obtained from Then, we choose the vector p (l) Can be constructed as:

[0093]

[0094] in Yes (l) The i-th element of Yes (l) The i-th group of vectors. (l) The sizes of adjacent sets in the selected subset can be calculated:

[0095]

[0096] where ξ (l) The i-th element of It represents the number of neighbors selected by node i in the lth iteration.

[0097] Step 4: To obtain sparse and connected sub-networks, we exploit the following two weight updating rules: 1) If the selected sub-network is connected in the current iteration, the leaf nodes in the sub-network should be penalized in the next iteration to improve sparsity.

[0098] 2) If the selected subnetwork is disconnected in the current iteration, the neighbors of the selected subnetwork should be encouraged to be selected in the next iteration to promote the connectivity of the network.

[0099] For the first case, The (i,j)th element of Can be updated to:

[0100]

[0101] Where β≥1 is a constant that controls the convergence speed, is the matrix Λ (l) The (i,j)th element of (l) Defined as:

[0102] Λ (l) =ξ (l) (ξ (l) ) T (36)

[0103] Three important features can be found from (35): (a) the weights corresponding to the unselected nodes will not be updated in the next iteration; (b) The corresponding weights of other nodes will be appropriately reduced; (c) The more neighbor nodes of the node are selected in the current iteration, the smaller the weight assigned to the node in the next iteration. On the one hand, based on features one and two, it can be concluded that the nodes that are not selected in the current iteration will not be selected in the next iteration (because the weights of the selected nodes will be further reduced). On the other hand, according to features two and three, since the number of neighbors of the leaf nodes is small, the weights of the leaf nodes in the sub-network selected in the current iteration are larger than the weights of other nodes. Therefore, in the next iteration, the filter coefficients of the leaf nodes will gradually decrease, thereby achieving a penalty for the leaf nodes. Therefore, under the premise of meeting the specified performance constraint C3, the leaf nodes in the sub-network are gradually closed.

[0104] For the second case, The (i,j)th element of Can be updated to:

[0105]

[0106] Only the weights corresponding to the neighbors of the selected subnetwork are appropriately reduced. This is because the first line in (37) makes the weights of the selected subnetwork not updated in the next iteration, and only the neighbors of the selected subnetwork are updated in the second line. Therefore, the selected nodes (in the current iteration) will continue to be activated (due to the sex constraint) and encourage their neighbors to be selected in the next iteration. The selected subnetwork can be gradually expanded through the update rule of (37), so that the subnetwork tends to be connected.

[0107] Step 5: The elements in that are smaller than a very small threshold η0 are set to zero, and the MVDR method is rerun for speech enhancement using the remaining nodes.

[0108] Preferably, in steps 1 to 5, the centralized node selection implementation method includes the following steps:

[0109] Step 1: Determine the number of microphones M, then perform frame division and windowing operations on the signals received by the microphones, and use VAD technology to obtain noise frames to obtain R vv (k) and R xx (k), and then find oNoise min (k), then input the frequency division factor N, scaling factor α, threshold δ and threshold η0, and β which controls the convergence speed

[0110] Step 2, perform iterative calculation, l from 1 to Lmax ,

[0111] Calculate formula (25) and solve formula (24), if Break out of the loop;

[0112] Step 3: Elements in the vector that are less than the threshold η0 are set to 0 and assigned to the vector w;

[0113] Step 4: Use formula (30) to solve again And perform voice enhancement.

[0114] Preferably, in the implementation of distributed node selection, steps one to five include the following steps:

[0115] Step 1: Determine the number of microphones M, then determine the connectivity radius r0 in the network, then perform frame windowing on the signal received by the microphone, and then estimate the global R by transmitting data between adjacent sensor nodes and recursively updating the inverse of the noise correlation matrix through a distributed consensus algorithm. vv (k) and R xx (k), and then find oNoise min (k), then input the frequency division factor N, scaling factor α, threshold δ and threshold η0, use the CVX toolbox in MATLAB to solve formula (32) and assign the solution to

[0116] Step 2, perform iterative calculation, l from 1 to L max , use formulas (33) and (34) to calculate the parameters

[0117] and (l) , if expression C9 is satisfied, then use formula (35) to calculate the weight matrix, otherwise use formula (37) to calculate the weight matrix, solve formula (32), if

[0118] Break out of the loop.

[0119] Step 3: Elements in the vector that are less than the threshold η0 are set to 0 and assigned to the vector w;

[0120] Step 4: Use formula (30) to solve again And perform voice enhancement.

[0121] Preferably, in the experiment, 25 microphones are placed in a room of (10×10×3)m, all nodes and sources are deployed at a height of 1m, there are two interference sources in the scene, respectively placed at (2.0, 3.0, 1)m and (7.5, 2.0, 1.0)m, and the target source is placed at (2.5, 6.0, 1.0)m. The fusion center is placed at the center of the room (5.0, 5.0, 1.0)m. Both the interference source and the target source are from the TIMIT database, and the room impulse response is generated by the IMAGE model. The speed of sound is 343m / s, the sampling rate is 16kHz, there are 256 sampling points per frame, and the overlap is 75%. Due to the conjugate symmetry of the speech signal in the discrete Fourier transform domain, we set K=128. The transmission power consumption (Power) is the sum of the squares of the distances from the selected node to the FC, and then divided by the total power consumption (the sum of the squares of the distances from all nodes in the network to the FC) to normalize it to between 0 and 1. Threshold δ = 1, threshold η0 = 0.1, L max =8, β=3.5.

[0122] Preferably, the experimental results under different segmentation factors N are studied in the C-FI-SS method, and the parameters are set as follows: the reverberation time is set to 200ms, the input signal-to-noise ratio is fixed to 10dB, and the scaling factor α=0.6. N is set to 1 to 128, where N=1 makes l 1,2[N],∞ The norm degenerates to l 1,2 norm, N = 128 allows l 1,2[N],∞ Norm conversion to l 1,∞ Norm. From Figure 2 It can be seen that the oNoise of the C-FI-SS method always satisfies the given oNoise constraint and increases with the increase of N. At the same time, the transmission power decreases with the increase of N. This is because as N increases, the solution of C-FI-SS becomes sparser, resulting in increased oNoise and reduced transmission power. Since the performance constraints can be flexibly adjusted, we believe that the method with lower transmission power will perform better as long as the pre-specified oNoise constraint is met. Therefore, the larger the value of N, the higher the energy efficiency of the C-FI-SS method. However, as shown in Table 1, the running time of each reweighting iteration increases with the increase of N, where the running time of N = 128 is more than 3 times that of N = 1. The main reason is that the number of constraints in C6 increases with the increase of N, which makes the SDP problem (33) more complicated. In addition, from Figure 3 It can be seen from Figure 1 that when the threshold is set to δ = 1, the reweighting strategy converges in the first three to four iterations. Moreover, the larger the value of N, the faster the reweighting process converges. In short, we should choose a suitable N according to the actual situation to achieve a trade-off between performance and complexity.

[0123] Preferably, in the comparison results of the C-FI-SS method and the other six centralized node selection methods, the parameters are set as follows: N = 8, the input signal-to-noise ratio is 15dB, and the reverberation time is 200ms. The six comparison methods are BroadOpt, NaVo, GradR, SnrR, EnergyR and UtilityR. The number of selected nodes I0 is obtained by adjusting the scaling factor α. Figure 4 As shown in Figure 2, the oSNR (Output SNR, oSNR) of all methods increases with the increase of network size I0. At the same time, as I0 increases, the growth rate becomes very slow, for example, when I0 ≥ 9, resulting in reduced computational efficiency. This phenomenon indicates that only a few nodes in WASN make significant contributions to the speech enhancement task. It can be seen that the C-FI-SS and UtilityR methods provide the highest. In most cases, when I0 ≤ 9, the oSNR of C-FI-SS is slightly higher than that of UtilityR. From Figure 4 We can also see that the transmission power of all methods increases approximately linearly with the increase of α, resulting in lower energy efficiency for subnetworks with larger sizes. The GradR method consumes the lowest power, while the transmission power of the utility and C-FI-SS methods is relatively high. This is because both the utility and C-FI-SS methods do not consider power consumption in their objective functions. Figure 4 We can conclude that the C-FI-SS approach performs well in finding information subsets from WASNs and is therefore computationally efficient for a particular performance constraint. However, the energy efficiency of C-FI-SS is lower than that of BroadOpt, NaVo, GradR, SnrR, and EnergyR approaches.

[0124] Preferably, for the performance evaluation of the MVDR method based on C-FI-SS, the parameters are set as follows: N=8. Figure 5 The reverberation time is 200ms, the scaling factor α changes from 0.2 to 0.9, the input signal-to-noise ratio changes from 0dB to 20dB, and the step size is 5dB. The input signal-to-noise ratio in Table 2 is 15dB, the scaling factor α changes from 0.3 to 0.9, the step size is 0.3, and the reverberation time changes from 100ms to 500ms, and the step size is 200ms. The evaluation criteria for speech quality include the perceptual evaluation of speech quality (PESQ), with a score range of -0.5 to 4.5, and the second is the short-time objective intelligibility (STOI), with a score range of 0-1. The larger the score of both, the better the speech quality. Figure 5As shown in Table 2, oSNR generally grows with increasing α. It is worth noting that oSNR sometimes remains unchanged with increasing α, because in these cases, the oNoise constraint can be satisfied without expanding the subnet. In addition, when α ≥ 0.5, the increasing trend of oSNRs gradually weakens. The main reason is that the nodes activated when the α value is small contribute more to the SE task, while the contributions of other nodes (activated later as the α value increases) are very limited. This reveals the necessity of sensor selection when performing speech enhancement in WASN. As can be seen from Table 2, both PESQ and STOI decrease with increasing reverberation time, which reveals the adverse effect of reverberation time on speech quality. In addition, when α increases from 0.6 to 0.9, both PESQ and STOI decrease slightly. This is because the subnet size becomes larger with increasing α, and some nodes with lower speech quality are added to the selected subset, resulting in reduced speech intelligibility.

[0125] Optimized node selection results of C-FI-SS method under different target sources. The parameters are set as follows: reverberation time is 200ms, input signal-to-noise ratio is 10dB. Scaling factor α=0.5, N=8. Different target sources are placed at (2.5, 6.0, 1.0) m, (7.5, 6.0, 1.0) m, (2.5, 1.5, 1.0) m, (7.5, 1.5, 1.0) m, and (5.0, 3.8, 1.0) m, respectively. Table 3 gives the indexes of selected nodes under different target locations. Obviously, C-FI-SS can always find information-rich nodes close to the target. Since the sound source at position (5.0, 3.8) m is close to FC, its power consumption is the lowest. In addition, when the target location is (7.5, 1.5) m, the subnet size is the largest, resulting in the highest power consumption. For the target at position (2.5, 1.5) m, the oSNR is relatively small, where only two nodes are selected. This is because oNoise in (17) max (k) is relatively small for target sources close to high SPL disturbances.

[0126] Preferably, the convergence speed of the D-FI-SS method in the reweighting process under different connectivity radii and the node selection results are evaluated. Table 5 shows the performance evaluation of the MVDR method based on D-FI-SS under different connectivity radii. The parameters are set as follows: input signal-to-noise ratio is 20 dB, reverberation time is 200 ms, scaling factor is α = 0.2, N = 8, and connectivity radius is r0 = [1.60, 1.75, 1.90, 2.05] m. Table 4 shows the changes in subnet size and subnet connectivity during the reweighting process. It can be seen that when r0 = 2.05 m, the reweighting method converges in one round. The reason is that in a dense network with more edges, the connectivity of the subnet is easily satisfied. For example, when r0 = 2.05 m, the best subset {3, 4} is connected, so after the first reweighting iteration, the subnet will not expand (37). In contrast, when r0 = 1.60m, r0 = 1.75m, and r0 = 1.90m, the reweighting strategy converges within 5 iterations. More specifically, we see from Table 44 that when l≤4, the size of the subnet increases, which triggers the update rule (37) to expand the subnet to induce connectivity. When l≤4, this process leads to Figure 7 Some fluctuations in . When l = 4, once the subnet expands to a connected subnet, it will trigger (35) to promote subnet sparsity (as shown in Table 4). In the sparsity induction phase, the subnetwork will no longer be disconnected because only leaf nodes can be deleted. This result is consistent with the discussion in Remark 2. In addition, the convergence speed of the reweighting rule (35) is very fast, and it converges in one round when l increases from 4 to 5. This experiment verifies the effectiveness of the reweighting strategy in D-FI-SS. Figure 7In Figure 5, we show the sensor selection results of D-FI-SS under different network topologies. We note that the subset {3,4} is selected in the first reweighting iteration for all communication radii. However, since nodes 3 and 4 are disconnected when r0 = 1.60m, r0 = 1.75m, and r0 = 1.90m, the subnetwork expands with the increase of reweighting iterations. Once the subnetwork is connected, the sparsity of the subnetwork is promoted by penalizing leaf nodes. Due to the differences in network topologies, the final subnetworks for r0 = 1.60m, r0 = 1.75m, and r0 = 1.90m are slightly different. For example, for r0 = 1.60m, the subset {6,7,8,11,22,23,24} is selected, while for r0 = 1.75m, the subset {6,7,10,11,22,23,24} is selected. The difference is that when r0 = 1.75m, node 10 has a higher priority than node 8 because in this case, node 10 has more selected neighbors than node 8, so the penalty in (35) is smaller. This phenomenon shows that the optimal subset may not be found in the D-FI-SS method. However, the connectivity of the subnet can always be satisfied, which is the basis of SS in distributed scenarios. For the performance evaluation of the D-FI-SS-based MVDR method, the more nodes are selected, the greater the oSNR in SE. For example, when r0 = 1.90m, 9 nodes are selected, providing the highest oSNR. In addition, neither PESQ nor STOI increases with the increase of subnet size. Specifically, the subset {3,4} provides the highest voice quality. This is because as the number of selected sensors increases, some nodes with lower voice quality are connected, resulting in a decrease in voice quality.

[0127] (III) Beneficial effects

[0128] Compared with the prior art, the present invention provides a wireless acoustic sensor network speech enhancement method based on node selection, which has the following beneficial effects:

[0129] 1. This node selection-based wireless acoustic sensor network speech enhancement method avoids the Boolean planning problem in FI-SS through a new sparse sensing model, reduces the solution complexity and improves the computational efficiency.

[0130] 2. This wireless acoustic sensor network speech enhancement method based on node selection can ensure the connectivity of the selected subnetwork through the developed distributed FI-SS method. It is suitable for WASN under distributed topology. It ensures the speech enhancement effect while saving energy and improves the speech quality. It has good application prospects in environmental monitoring, smart home, telephone conferencing and other scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0131] Figure 1This is a flow chart of speech enhancement based on node selection in the present invention, where “SS” stands for node selection and “SE” stands for speech enhancement.

[0132] Figure 2 Schematic diagram of the effect of the C-FI-SS method of the present invention under different frequency division factors N.

[0133] Figure 3 Schematic diagram of the convergence of the reweighting strategy of the C-FI-SS method under different N in the present invention.

[0134] Figure 4 Schematic diagram of the effects of different numbers of selected nodes in the present invention, (a) output signal-to-noise ratio, (b) transmission power consumption.

[0135] Figure 5 It is a schematic diagram of the noise reduction effect of the C-FI-SS-based MVDR method in the present invention under different input signal-to-noise ratios.

[0136] Figure 6 Schematic diagram of node selection results of the D-FI-SS method in different communication graphs in the present invention:

[0137] (a) r0 = 1.60m, (b) r0 = 1.75m, (c) r0 = 1.90m, (d) r0 = 2.05m.

[0138] Figure 7 Schematic diagram of the convergence speed of the reweighting strategy of the D-FI-SS method under different r0 in the present invention. DETAILED DESCRIPTION

[0139] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0140] A method for speech enhancement in a wireless acoustic sensor network based on node selection comprises the following steps:

[0141] Step 1: We assume that each node in WASN is equipped with an acoustic sensor. In the discrete Fourier transform (DFT) domain, let k and t be the frequency index and frame index respectively, and the noise signal y recorded by node i is i (k,t), can be expressed as:

[0142] y i (k,t)=x i(k,t)+v i (k,t) (1)

[0143] where x i (k,t) and v i (k, t) represent the signal component and additive noise component at the target node i, respectively. For the case of a single target source, x i (k,t) can be written as:

[0144] x i (k,t)=a i (k)s(k,t) (2)

[0145] where a i (k) is the acoustic transfer function from the target source to node i, and s(k,t) is the corresponding target signal. In general, a i (k) can be derived from the locations of the sound source and the node. In addition, the relative transfer function a i (k) / a1(k) can be estimated by covariance subtraction / whitening (CS / CW) method. Let y(k,t) =

[0146] [y1(k,t),y2(k,t),...,y I (k,t)] T , then the vector form of (1) is:

[0147] y(k,t)=x(k,t)+v(k,t)=a(k)s(k,t)+v(k,t) (3)

[0148] where x(k,t), v(k,t) and a(k) are vectors containing x i (k,t),v i (k,t) and a i (k). Assuming that the target speech and noise are uncorrelated, the correlation matrix of the received signal can be calculated as:

[0149]

[0150] in In practice, after applying a voice activity detector (VAD) (setting an energy threshold, when the energy of the current frame is greater than the threshold, it is judged as voice activity; when the frame energy is lower than the threshold, it is judged as a noise frame) to distinguish the received signal into noise-only frames and target plus noise frames, the noise correlation matrix R can be estimated from the noise-only frames. vv (k). Based on this, we can yySubtract R from (k) vv (k) to calculate the speech correlation matrix R xx (k), i.e. R xx (k) = R yy (k)-R vv (k).

[0151] Step 2: We formulate FI-SS by improving the group sparsity of the filter coefficients. First, the global vector w is defined as:

[0152]

[0153] where w i is a vector of filter coefficients for all frequencies of node i, w i The q-norm of is:

[0154]

[0155] If w i The q-norm of is equal to zero, and we can put the i-th node to sleep. Therefore, we minimize the l of w 1,q norm, while limiting the output noise power, to enhance the group sparsity of w, thus obtaining:

[0156]

[0157] Among them, ‖w i ‖ q The l1 norm of is a convex approximation of the l0 norm, C1 is the target signal distortion-free constraint, and C2 constrains the output noise (Output noise, oNoise) with a higher limit r(k):

[0158] oNoise min (k)≤r(k)≤oNoise max (k) (8)

[0159] Among them, oNoise min (k) and oNoise max (k) represents the minimum and maximum output noise power at frequency k. Usually, oNoise max (k) can be determined by the input noise power, and oNoise min (k) can be obtained by activating all nodes. In this invention, we introduce a scaling factor α to control r(k):

[0160]

[0161] where α min <α≤1, The larger the α value, the higher the expected noise reduction performance.

[0162] Step 3: To improve the solution efficiency, the present invention develops the concept of triple norm. Define the triple norm l of w 1,q[N],z as:

[0163]

[0164] where N is the frequency division factor that divides K frequencies into N parts. If N = 1, then the l 1,q[N],z norm is reduced to the l 1,z norm; if N = K, then the l 1,q[N],z norm is reduced to the l 1,q norm; 1 < N < K is a more general setting that can flexibly adjust the complexity of the subsequent FI-SS problem.

[0165] In this paper, we choose q = ∞ and z = 2. Then, by minimizing while maintaining the constraints C1 and C2 a new cost function is established:

[0166]

[0167] Since the constraint C2 is non-convex. Therefore, we relax the above problem to a convex semi-definite programming (SDP) problem for effective solution.

[0168] Step 3: Let be another vector composed of all filter coefficients, which can be expressed as:

[0169]

[0170] In addition, we introduce the rank-1 matrix as:

[0171]

[0172] where, Therefore, averaging C2 over frequencies is equal to:

[0173]

[0174] where the operation tr(·) is to calculate the trace of the matrix. Compared with C2, C3 specifies the overall noise reduction performance and expands the feasible region of (11) due to the reduction of the number of constraints from K to 1. Based on the above, problem (11) can be equivalently rewritten as:

[0175]

[0176] where constraints C4 and C5 are used to replace (13), the symbol ≥ denotes a matrix inequality, i.e. X ≥ Y means that XY is positive semidefinite, and the operation rank(·) computes the rank of the matrix, is an auxiliary variable defined as follows:

[0177]

[0178] in is the index variable of k. The objective of (15) can be derived as:

[0179]

[0180] in It can be expressed as:

[0181]

[0182] Since the non-diagonal matrix does not appear in the constraints C1 and C3 in (15), we set them to zero according to the objective of (15), i.e. In order to remove these non-diagonal matrices from the SDP problem (15). Under this parameter setting, (18) can be further simplified as:

[0183]

[0184] in is a submatrix Note that in (19), the derivation from the second to the third line is due to It was established.

[0185] By defining an upper bound matrix Its (i,j)th element Always meet:

[0186]

[0187] After substituting (19) into (17), the objective in (17) can be further expressed as:

[0188]

[0189] In addition, since the non-diagonal matrix It has been deleted in (18) below, and constraint C4 can be simplified to:

[0190]

[0191] Therefore, by abandoning the non-convex constraint C5, problem (15) can be relaxed as:

[0192]

[0193] where C8 is an additional redundant constraint constructed from C1 and (13). The relaxed problem (21) is convex and can be solved efficiently using interior point methods or solvers such as the CVX toolbox.

[0194] Step 4: To further generalize the sparse solution, we use the weight matrix U I Replacement Matrix 1 I , and is calculated iteratively. Specifically, let l be the step index of the iteration, and the SDP problem in (23) can be modified as follows:

[0195]

[0196] Among them U (l) With all 1 matrix 1 I Initialization, that is, U (0) =1 I , and U (l) The (i,j)th element of It can be calculated iteratively using the unit rank penalty strategy:

[0197]

[0198] where ∈ is a small constant that prevents the denominator from being zero. Represents Γ (l) The (i,j)th element of is calculated as:

[0199] Γ (l) =γ (l) (γ (l) ) T (26)

[0200]

[0201] The operator Calculate the principal eigenvector of the matrix. It is worth noting that Γ (l) is a rank-one matrix consisting of each block diagonal matrix In addition, this reweighting method makes Small elements in can be penalized aggressively, encouraging them to go to zero as iterations proceed. As a result, a sparser solution can be obtained while ensuring performance constraints (since C3 is satisfied in all reweighting iterations).

[0202] Due to computational errors, it is difficult to obtain a strict sparse Therefore, we will Elements in η smaller than a very small threshold η0 are set to zero, which can achieve strict sparsity. But it will change the form of the solution obtained from (24), which may lead to a slight degradation in performance. To this end, we put a node to sleep if its filter coefficients on all frequencies are zero and rerun the minimum variance distortionless response (MVDR) method using the remaining nodes.

[0203] Step 5: Rerun the MVDR beamformer using the remaining nodes. Output signal after speech enhancement It is expressed as:

[0204]

[0205] Where w(k)=[w1(k),w2(k),…,w I (k)] T is the filter coefficient of the kth frequency. In the classic MVDR method, w(k) is estimated by minimizing the output noise power under the constraint of maintaining the distortion-free target source. Specifically, it can be expressed as:

[0206]

[0207] By using the Lagrange multiplier method, the analytical solution for w can be expressed as:

[0208]

[0209] When you rerun the MVDR method using the remaining nodes, the vector a(k), the vector y(k,t), and the matrix R vv (k) should only contain the information of the remaining nodes, so the dimension of w(k) is the number of remaining selected nodes. Finally, use formula (28) for speech enhancement.

[0210] Distributed node selection:

[0211] Step 1: Construct graph-related constraints to ensure connectivity in the distributed topology:

[0212]

[0213] Among them I w is the number of non-zero groups in w, in is the adjacency matrix of the subnetwork obtained by sparse w. Since C9 involves nonlinear and nonconvex formulas, it is difficult to solve. Therefore, we invented a heuristic method to find a near-optimal solution.

[0214] Step 2: Inspired by the reweighting strategy in centralized node selection, we try to find appropriate weights to ensure the connectivity of the selected subnetwork. At this time, a cost function with the same form as (24) is constructed:

[0215]

[0216] in is the weighting matrix introduced at the lth iteration.

[0217] Step 3: We formulate a weight update criterion. For a given At the lth iteration, we can solve (32) to obtain At the same time, according to definitions (5) and (12), w (l) Can be obtained from Then, we choose the vector p (l) Can be constructed as:

[0218]

[0219] in Yes (l) The i-th element of Yes (l) The i-th group of vectors. (l) The sizes of adjacent sets in the selected subset can be calculated:

[0220] ξ (l) =Ap (l) (34)

[0221] where ξ (l) The i-th element of It represents the number of neighbors selected by node i in the lth iteration.

[0222] Step 4: To obtain sparse and connected sub-networks, we used the following two weight update rules:

[0223] 1) If the selected subnetwork is connected in the current iteration, the leaf nodes in the subnetwork should be penalized in the next iteration to improve sparsity.

[0224] 2) If the selected subnetwork is disconnected in the current iteration, the neighbors of the selected subnetwork should be encouraged to be selected in the next iteration to promote the connectivity of the network.

[0225] For the first case, The (i,j)th element of Can be updated to:

[0226]

[0227] Where β≥1 is a constant that controls the convergence speed, is the matrix Λ (l) The (i,j)th element of (l) Defined as:

[0228] Λ (l) =ξ (l) (ξ (l) ) T (36)

[0229] Three important features can be found from (35): (a) the weights corresponding to the unselected nodes will not be updated in the next iteration; (b) The corresponding weights of other nodes will be appropriately reduced; (c) The more neighbor nodes of the node are selected in the current iteration, the smaller the weight assigned to the node in the next iteration. On the one hand, based on features one and two, it can be concluded that the nodes that are not selected in the current iteration will not be selected in the next iteration (because the weights of the selected nodes will be further reduced). On the other hand, according to features two and three, since the number of neighbors of the leaf nodes is small, the weights of the leaf nodes in the sub-network selected in the current iteration are larger than the weights of other nodes. Therefore, in the next iteration, the filter coefficients of the leaf nodes will gradually decrease, thereby achieving a penalty for the leaf nodes. Therefore, under the premise of meeting the specified performance constraint C3, the leaf nodes in the sub-network are gradually closed.

[0230] For the second case, The (i,j)th element of Can be updated to:

[0231]

[0232] Only the weights corresponding to the neighbors of the selected subnetwork are appropriately reduced. This is because the first line in (37) makes the weights of the selected subnetwork not updated in the next iteration, and only the neighbors of the selected subnetwork are updated in the second line. Therefore, the selected nodes (in the current iteration) will continue to be activated (due to the sex constraint) and encourage their neighbors to be selected in the next iteration. The selected subnetwork can be gradually expanded through the update rule of (37), so that the subnetwork tends to be connected.

[0233] Step 5: The elements in that are smaller than a very small threshold η0 are set to zero, and the MVDR method is rerun for speech enhancement using the remaining nodes.

[0234] Centralized node selection:

[0235] First, determine the number of microphones M, then perform frame division and windowing operations on the signals received by the microphones, and use VAD technology to obtain noise frames to obtain R vv (k) and R xx (k), and then find oNoise min (k), then input the frequency division factor N, scaling factor α, threshold δ and threshold η0

[0236] 1. Use the CVX toolbox in MATLAB to solve equation (23) and assign the solution to

[0237] 2.for l=1,2,…,L max

[0238] 2.1. Calculation formula (25)

[0239] 2.2. Solve formula (24)

[0240]

[0241] Break out of the loop

[0242] End if

[0243] End for

[0244] 3. The obtained Elements in the vector that are less than the threshold η0 are set to 0 and assigned to the vector w

[0245] 4. Use formula (30) to solve again And perform speech enhancement

[0246] Distributed node selection:

[0247] First, the number of microphones M is determined, and then the connection radius r0 in the network is determined. Then, the signal received by the microphone is framed and windowed. Then, the global R is estimated by transmitting data between adjacent sensor nodes and recursively updating the inverse of the noise correlation matrix through a distributed consensus algorithm. vv (k) and R xx (k), and then find oNoise min (k), then input the frequency division factor N, scaling factor α, threshold δ and threshold η0, and β which controls the convergence speed

[0248] 1. Use the CVX toolbox in MATLAB to solve equation (32) and assign the solution to

[0249] 2.for l=1,2,…,L max

[0250] 2.1. Calculate parameters using formulas (33) and (34) and (l)

[0251] If expression C9 is satisfied

[0252] Then use formula (35) to calculate the weight matrix

[0253] else

[0254] The weight matrix is ​​calculated using formula (37)

[0255] End if

[0256] 2.2. Solve formula (32)

[0257]

[0258] Break out of the loop

[0259] End if

[0260] End for

[0261] 3. The obtained Elements in the vector that are less than the threshold η0 are set to 0 and assigned to the vector w

[0262] 4. Use formula (30) to solve again And perform speech enhancement

[0263] In this invention, we place 25 microphones in a room of (10×10×3) m, as Figure 6 As shown. All nodes and sources are deployed at a height of 1m. There are two interference sources in the scene, placed at (2.0, 3.0, 1)m and (7.5, 2.0, 1.0)m respectively, and the target source is placed at (2.5, 6.0, 1.0)m. The fusion center is placed in the center of the room (5.0, 5.0, 1.0)m. Both the interference source and the target source are from the TIMIT database, and the room impulse response is generated by the IMAGE model. The speed of sound is 343m / s, the sampling rate is 16kHz, there are 256 sampling points per frame, and the overlap is 75%. Due to the conjugate symmetry of the speech signal in the discrete Fourier transform domain, we set K=128. The transmission power consumption (Power) is the sum of the squares of the distances from the selected node to the FC, and then divided by the total power consumption (the sum of the squares of the distances from all nodes in the network to the FC) to normalize it to between 0-1. Threshold δ=1, threshold η0=0.1, L max =8, β=3.5.

[0264] Figure 2 , Table 1 and Figure 3In order to study the experimental results under different segmentation factors N in the C-FI-SS method, the parameters are set as follows: the reverberation time is set to 200ms, the input signal-to-noise ratio is fixed to 10dB, and the scaling factor α = 0.6.

[0265] Table 1 is a schematic diagram of the running time of the C-FI-SS method under different N in each re-weighted iteration in the present invention, as shown in the following figure:

[0266]

[0267] Figure 4 Comparison results of the C-FI-SS method and six other centralized node selection methods. The parameters are set as follows: N = 8, input signal-to-noise ratio is 15dB, and reverberation time is 200ms. The six comparison methods are BroadOpt, NaVo, GradR, SnrR, and UtilityR. The number of selected nodes I0 is obtained by adjusting the scaling factor α.

[0268] Figure 5 and Table 2 show the performance evaluation of the MVDR method based on C-FI-SS, with the parameters set as follows: N=8. Figure 5 The reverberation time is 200ms, the scaling factor α varies from 0.2 to 0.9, the input signal-to-noise ratio varies from 0dB to 20dB, and the step size is 5dB. The input signal-to-noise ratio in Table 2 is 15dB, the scaling factor α varies from 0.3 to 0.9, the step size is 0.3, and the reverberation time varies from 100ms to 500ms, and the step size is 200ms. The evaluation criteria for speech quality are the perceptual evaluation of speech quality (PESQ), with a score range of -0.5 to 4.5, and the second is the short-time objective intelligibility (STOI), with a score range of 0-1. The larger the score of both, the better the speech quality.

[0269] Table 2 is a schematic diagram of the speech quality of the MVDR method based on C-FI-SS in the present invention at different reverberation times, as shown in the following figure:

[0270]

[0271] Table 3 shows the node selection results of the C-FI-SS method under different target sources. The parameter settings are as follows: the reverberation time is 200ms, the input signal-to-noise ratio is 10dB, the scaling factor α=0.5, N=8. The different target sources are placed at (2.5, 6.0, 1.0)m, (7.5, 6.0, 1.0)m, (2.5, 1.5, 1.0)m, (7.5, 1.5, 1.0)m, and (5.0, 3.8, 1.0)m.

[0272] Table 3 is a schematic diagram of the node selection results of the C-FI-SS method under different target source positions in the present invention, as shown in the following figure:

[0273]

[0274] Figure 7 and Table 4 are the evaluations of the convergence speed during the reweighting process of the D-FI-SS method. Figure 6 is the node selection results of the D-FI-SS method under different connectivity radii. Table 5 shows the performance evaluation of the MVDR method based on D-FI-SS under different connectivity radii. The parameters are set as follows: the input signal-to-noise ratio is 20 dB, the reverberation time is 200 ms, the scaling factor is α = 0.2, N = 8, and the connectivity radius is r0 = [1.60, 1.75, 1.90, 2.05] m.

[0275] Table 4 is a schematic diagram of the size and connectivity of the network subsets during the iteration process of the D-FI-SS method under different r0 in the present invention, as shown in the following figure:

[0276]

[0277] Table 5 is a schematic diagram of the voice quality of the distributed MVDR method based on D-FI-SS in the present invention at different connection radii, as shown in the following figure:

[0278]

[0279] Evaluation of the convergence speed and node selection results of the D-FI-SS method in the reweighting process under different connectivity radii, Figure 6 Table 4 shows the node selection results of the D-FI-SS method under different connectivity radii, and Table 5 shows the performance evaluation of the MVDR method based on D-FI-SS under different connectivity radii. Figure 7 To evaluate the convergence speed of the D-FI-SS method under different connectivity radii, the parameters are set as follows: input signal-to-noise ratio is 20 dB, reverberation time is 200 ms, scaling factor is α = 0.2, N = 8, and connectivity radius is r0 = [1.60, 1.75, 1.90, 2.05] m.

[0280] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for speech enhancement in a wireless acoustic sensor network based on node selection, characterized in that: The following steps are involved: Step 1: Assume that each node in WASN is equipped with an acoustic sensor. In the discrete Fourier transform (DFT) domain, let k and t be the frequency index and frame index respectively. The noise signal y recorded by node i is i (k, t), can be expressed as: y i (k,t)=x i (k,t)+v i (k,t) (1) where x i (k, t) and v i (k, t) represent the signal component and additive noise component at the target node i, respectively. For a single target source, x i (k, t) can be written as: x i (k,t)=a i (k)s(k,t) (2) where a i (k) is the acoustic transfer function from the target source to node i, s(k, t) is the corresponding target signal, and in general, a i (k) can be derived from the locations of the sound source and the node. In addition, the relative transfer function a i (k) / a1(k) can be estimated by covariance subtraction / whitening (CS / CW) method, assuming y(k, t) = [y1(k, t), y2(k, t), ..., y I (k, t)] T , then the vector form of (1) is: y(k,t)=x(k,t)+v(k,t)=a(k)s(k,t)+v(k,t) (3) where x(k, t), v(k, t) and a(k) are vectors containing x i (k, t), v i (k, t) and a i (k), assuming that the target speech and noise are uncorrelated, the correlation matrix of the received signal can be calculated as: in Represents mathematical expectation. In practice, after applying a voice activity detector (VAD) (setting an energy threshold, when the energy of the current frame is greater than the threshold, it is judged as voice activity; when the frame energy is lower than the threshold, it is judged as a noise frame) to distinguish the received signal into noise-only frames and target plus noise frames, the noise correlation matrix R can be estimated from the noise-only frames. vv (k), based on this, we can yy Subtract R from (k) vv (k) to calculate the speech correlation matrix R xx (k), i.e. R xx (k) = R yy (k)-R vv (k); Step 2: FI-SS is formulated by improving the group sparsity of the filter coefficients. First, the global vector w is defined as: where w i is a vector of filter coefficients for all frequencies of node i, w i The q-norm of is: If w i The q-norm of is equal to zero, we can put the i-th node into sleep state, so we minimize the w norm, while limiting the output noise power, to enhance the group sparsity of w, thus obtaining: Among them, ||w i || q of The norm is The convex approximation of the norm, C1 is the target signal distortion-free constraint, and C2 constrains the output noise (Output noise, oNoise) with a higher limit r(k): oNoise min (k)≤r(k)≤oNoise max (k) (8) Among them, oNoise min (k) and oNoise max (k) represents the minimum and maximum output noise power at frequency k. Usually, oNoise max (k) can be determined by the input noise power, and oNoise min (k) can be obtained by activating all nodes. In this invention, we introduce a scaling factor α to control r(k): where α min <α≤1, The larger the α value, the higher the noise reduction performance is expected; Step 3: To improve the solution efficiency, the present invention develops the concept of triple norm and defines the triple norm of w for: Where N is the frequency division factor that divides K frequencies into N parts. If N = 1, then The norm is reduced to norm; if N = K, then The norm is reduced to norm; 1<N<K is a more general setting that can flexibly adjust the complexity of subsequent FI-SS problems; In this paper, we choose q = ∞ and z = 2, and then minimize A new cost function is established: Since constraint C2 is non-convex, we relax the above problem into a convex semi-definite programming (SDP) problem for efficient solution; Step 4: To further generalize the sparse solution, we use the weight matrix U I Replacement Matrix 1 I , and is calculated iteratively. Specifically, let l be the step index of the iteration. The SDP problem in (23) can be modified as follows: Among them U (l) With all 1 matrix 1 I Initialization, that is, U (0) =1 I , and U (l) The (i, j)th element of It can be calculated iteratively using the unit rank penalty strategy: where ∈ is a small constant that prevents the denominator from being zero. Represents Γ (l) The calculation formula for the (i, j)th element of is: C (l) =c (l) (c (l) ) T (26) The operator Calculate the main eigenvector of the matrix. It is worth noting that Γ (l) is a rank-one matrix consisting of each block diagonal matrix In addition, this reweighting method makes Small elements in can be actively penalized, encouraging them to go to zero as iterations proceed, thus leading to a sparser solution while ensuring performance constraints (since C3 is satisfied in all reweighting iterations); Due to computational errors, it is difficult to obtain a strict sparse Therefore, we will Elements in η smaller than a very small threshold η0 are set to zero, which can achieve strict sparsity. But it will change the form of the solution obtained from (24), which may lead to a slight degradation in performance. To this end, if the filter coefficients of a node on all frequencies are zero, we put it to sleep and rerun the minimum variance distortionless response (MVDR) method using the remaining nodes; Step 5: Rerun the MVDR beamformer using the remaining nodes, output signal after speech enhancement (SE) It is expressed as: Where, w(k)=[w1(k),w2(k),...,w I (k)] T is the filter coefficient of the kth frequency. In the classic MVDR method, w(k) is estimated by minimizing the output noise power under the constraint of maintaining the distortion-free target source. Specifically, it can be expressed as: By using the Lagrange multiplier method, the analytical solution for w can be expressed as: When you rerun the MVDR method using the remaining nodes, the vector a(k), the vector y(k, t), and the matrix R vv (k) should only contain the information of the remaining nodes, so the dimension of w(k) is the number of remaining selected nodes. Finally, formula (28) is used for speech enhancement.

2. According to the method for speech enhancement in a wireless acoustic sensor network based on node selection in claim 1, it is characterized in that: In the step three: set up is another vector consisting of all filter coefficients, It can be expressed as: In addition, we introduce the rank 1 matrix As: in, Therefore, averaging C2 over frequency equals: where operation tr(·) is the trace of the computation matrix. Compared with C2, C3 specifies the overall denoising performance and expands the feasible domain of (11) because the number of constraints is reduced from K to 1. Based on the above, problem (11) can be equivalently rewritten as: where constraints C4 and C5 are used to replace (13), and the symbol represents the matrix inequality, that is, Indicates that XY is semi-positive definite, and the operation rank(·) calculates the rank of the matrix. is an auxiliary variable defined as follows: in is the index variable of k, and the objective of (15) can be derived as: in It can be expressed as: Since the non-diagonal matrix does not appear in the constraints C1 and C3 in (15), we set them to zero according to the objective of (15), i.e. In order to remove these non-diagonal matrices from the SDP problem (15), under this parameter setting, (18) can be further simplified to: in is a submatrix The (i, j)th element of , note that in (19), the derivation from the second to the third line is due to established; By defining an upper bound matrix Its (i, j)th element Always meet: After substituting (19) into (17), the objective in (17) can be further expressed as: In addition, since the non-diagonal matrix It has been deleted in (18) below, and constraint C4 can be simplified to: Therefore, by abandoning the non-convex constraint C5, problem (15) can be relaxed as: where C8 is an additional redundant constraint constructed through C1 and (13). The relaxed problem (21) is convex and can be solved efficiently using interior point methods or solvers such as the CVX toolbox.

3. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: In steps 1 to 5, the distributed node selection includes the following steps: Step 1: Construct graph-related constraints to ensure connectivity in the distributed topology: Among them I w is the number of non-zero groups in w, in is the adjacency matrix of the subnetwork obtained by sparse w. Since C9 involves nonlinear and nonconvex formulas that are difficult to solve, we invented a heuristic method to find a near-optimal solution; Step 2: Inspired by the reweighting strategy in centralized node selection, we try to find appropriate weights to ensure the connectivity of the selected subnetwork. At this time, we construct a cost function with the same form as (24): in is the weighting matrix introduced at the lth iteration; Step 3: We formulate a weight update criterion for a given At the lth iteration, we can solve (32) to obtain At the same time, according to definitions (5) and (12), w (l) Can be obtained from It follows that, then, we choose the vector p (l) Can be constructed as: in Yes (l) The i-th element of Yes (l) The i-th group of vectors, through p (l) The sizes of adjacent sets in the selected subset can be calculated: ξ (l) =Ap (l) (34) where ξ (l) The i-th element of represents the number of neighbors selected by node i in the first iteration; Step 4: To obtain sparse and connected sub-networks, we used the following two weight update rules: 1) If the selected subnetwork is connected in the current iteration, the leaf nodes in the subnetwork should be penalized in the next iteration to improve sparsity; 2) If the selected subnetwork is disconnected in the current iteration, the neighbors of the selected subnetwork should be encouraged to be selected in the next iteration to promote the connectivity of the network; For the first case, The (i, j)th element of Can be updated to: Where β≥1 is a constant that controls the convergence speed, is the matrix Λ (l) The (i, j)th element of (l) Defined as: L (l) =ξ (l) (x) (l) ) T (36) Three important features can be found from (35): (a) the weights corresponding to the unselected nodes will not be updated in the next iteration; (b) The corresponding weights of other nodes will be appropriately reduced; (c) The more neighbor nodes of the node are selected in the current iteration, the smaller the weight assigned to the node in the next iteration. On the one hand, based on features 1 and 2, it can be concluded that the nodes that are not selected in the current iteration will not be selected in the next iteration (because the weight of the selected node will be further reduced). On the other hand, according to features 2 and 3, since the number of neighbors of the leaf node is small, the weight of the leaf node in the sub-network selected in the current iteration is larger than the weight of other nodes. Therefore, in the next iteration, the filter coefficient of the leaf node will gradually decrease, thereby realizing the penalty of the leaf node. Therefore, under the premise of meeting the specified performance constraint C3, the leaf nodes in the sub-network are gradually closed; For the second case, The (i, j)th element of Can be updated to: Only the weights corresponding to the neighbors of the selected subnetwork are appropriately reduced. This is because the first line in (37) makes the weight of the selected subnetwork not updated in the next iteration, and only the neighbors of the selected subnetwork are updated in the second line. Therefore, the selected nodes (in the current iteration) will continue to be activated (due to the sex constraint) and encourage their neighbors to be selected in the next iteration. The selected subnetwork can be gradually expanded through the update rule of (37), so that the subnetwork tends to be connected; Step 5: The elements in that are smaller than a very small threshold η0 are set to zero, and the MVDR method is rerun for speech enhancement using the remaining nodes.

4. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: In steps 1 to 5, the centralized node selection implementation method includes the following steps: Step 1: Determine the number of microphones M, then perform frame division and windowing operations on the signals received by the microphones, and use VAD technology to obtain noise frames to obtain R vv (k) and R xx (k), and then find oNoise min (k), then input the frequency division factor N, scaling factor α, threshold δ and threshold η0, and β which controls the convergence speed Step 2, perform iterative calculation, l from 1 to L max , Calculate formula (25) and solve formula (24), if Break out of the loop; Step 3: Elements in the vector that are less than the threshold η0 are set to 0 and assigned to the vector w; Step 4: Use formula (30) to solve again And perform voice enhancement.

5. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: In the steps 1 to 5, in the implementation of the distributed node selection, the following steps are included: Step 1: Determine the number of microphones M, then determine the connectivity radius r0 in the network, then perform frame windowing on the signal received by the microphone, and then estimate the global R by transmitting data between adjacent sensor nodes and recursively updating the inverse of the noise correlation matrix through a distributed consensus algorithm. vv (k) and R xx (k), and then find oNoise min (k), then input the frequency division factor N, scaling factor α, threshold δ and threshold η0, use the CVX toolbox in MATLAB to solve formula (32) and assign the solution to Step 2, perform iterative calculation, l from 1 to L max , use formulas (33) and (34) to calculate the parameters and (l) , if expression C9 is satisfied, then use formula (35) to calculate the weight matrix, otherwise use formula (37) to calculate the weight matrix, solve formula (32), if Break out of the loop; Step 3: Elements in the vector that are less than the threshold η0 are set to 0 and assigned to the vector w; Step 4: Use formula (30) to solve again And perform voice enhancement.

6. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: In the experiment, 25 microphones are placed in a room of (10×10×3)m. All nodes and sources are deployed at a height of 1m. There are two interference sources in the scene, which are placed at (2.0, 3.0, 1)m and (7.5, 2.0, 1.0)m respectively. The target source is placed at (2.5, 6.0, 1.0)m. The fusion center is placed at the center of the room (5.0, 5.0, 1.0)m. The interference source and the target source are both from the TIMIT database. The room impulse response is generated by the IMAGE model. The sound speed is 343m / s, the sampling rate is 16kHz, there are 256 sampling points per frame, and the overlap is 75%. Due to the conjugate symmetry of the speech signal in the discrete Fourier transform domain, we set K=128, the transmission power consumption (Power) is the sum of the squares of the distances from the selected node to the FC, and then divided by the total power consumption (the sum of the squares of the distances from all nodes in the network to the FC) to normalize it to between 0 and 1, the threshold δ=1, the threshold η0=0.1, and L max =8, β=3.

5.

7. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: The experimental results under different segmentation factors N are studied in the C-FI-SS method, and the parameters are set as follows: the reverberation time is set to 200ms, the input signal-to-noise ratio is fixed to 10dB, and the scaling factor α=0.

6.

8. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: In the comparison results between the C-FI-SS method and the other six centralized node selection methods, the parameters are set as follows: N = 8, the input signal-to-noise ratio is 15 dB, the reverberation time is 200 ms, the six compared methods are BroadOpt, NaVo, GradR, SnrR, and UtilityR, and the number of selected nodes I0 is obtained by adjusting the scaling factor α.

9. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: Performance evaluation of the MVDR method based on C-FI-SS, the parameters are set as follows: N = 8, reverberation time is 200ms, scaling factor α varies from 0.2 to 0.9, input signal-to-noise ratio varies from 0dB to 20dB, step size is 5dB, input signal-to-noise ratio is 15dB, scaling factor α varies from 0.3 to 0.9, step size is 0.3, reverberation time varies from 100ms to 500ms, step size is 200ms, the evaluation criteria of speech quality are perceptual evaluation of speech quality (PESQ), with a score range of -0.5 to 4.5, and the second is short-time objective intelligibility (STOI), with a score range of 0-1. The larger the score of both, the better the speech quality.

10. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: Node selection results of the C-FI-SS method under different target sources. The parameter settings are as follows: reverberation time is 200ms, input signal-to-noise ratio is 10dB, scaling factor α=0.5, N=8, and different target sources are placed at (2.5, 6.0, 1.0)m, (7.5, 6.0, 1.0)m, (2.5, 1.5, 1.0)m, (7.5, 1.5, 1.0)m, and (5.0, 3.8, 1.0)m respectively.

11. The method for speech enhancement in a wireless acoustic sensor network based on node selection according to claim 1, characterized in that: Evaluation of the convergence speed and node selection results of the D-FI-SS method in the reweighting process under different connectivity radii, node selection results of the D-FI-SS method under different connectivity radii, performance evaluation of the MVDR method based on D-FI-SS under different connectivity radii, evaluation of the convergence speed of the D-FI-SS method under different connectivity radii, the parameter settings are as follows: input signal-to-noise ratio is 20dB, reverberation time is 200ms, scaling factor is α=0.2, N=8, and connectivity radius is r0=[1.60, 1.75, 1.90, 2.05]m.