An unsupervised classification and supervised correction fusion method for speech separation related to spatial structural features
By integrating unsupervised classification and supervised correction methods in the speech separation algorithm, using time-delay cell neural network and dynamic growth self-organized mapping neural network to extract and classify speech features, and correcting model parameters through particle swarm optimization algorithm, the problem of insufficient generalization and accuracy of speech separation algorithm in the existing technology in complex acoustic environments and multi-speaker scenarios is solved, and a more efficient speech separation effect is achieved.
Patent Information
- Application Number
- CN202010966976.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-09-15
AI Technical Summary
Existing speech separation algorithms lack generalization and accuracy when dealing with mixed speech with complex acoustic environments and unknown number of speakers.
The speech separation method that integrates unsupervised classification and supervised correction is adopted for spatial structural features, and the speech fragment characteristics are extracted through time-delay cell neural networks, and the dynamic growth self-organized mapping neural network is used for unsupervised classification. The particle swarm optimization algorithm adapts the model parameters, and the speech reconstruction is performed using binary masking.
It improves the generalization and accuracy of the speech separation model, and enhances the ability to deal with complex acoustic environments and multi-talk scenes.
Smart Images

Figure CN112133323B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech signal processing, and particularly to a speech separation method that combines unsupervised classification related to spatial structural features and supervised correction. Background Art
[0002] In a complex acoustic environment, the speech signal of the target speaker is often interfered by various noises, which seriously affects the recognition performance of the target speech. Speech separation technology can effectively remove the noise interference in the actual environment and provide more accurate and reliable information for subsequent speech signal processing. The application scenarios of speech separation technology are very extensive. For example, in the field of national defense and military, in the background of war environment and conference monitoring, etc., simply using voiceprint recognition technology cannot accurately analyze whether there is a specific speaker in the intercepted conference recordings from the enemy, while speech separation technology can improve the accuracy of voiceprint recognition. In the field of public security, on the streets in a noisy multi-speaker scenario, using speech separation technology can more accurately find specific words and lock pedestrians with dangerous intentions; in the field of smart home, when controlling smart devices through voice commands, other family members are often having language exchanges, so speech separation technology is needed to accurately obtain the commands of the target speech and thus correctly execute its intention.
[0003] As an important foundation for technologies such as speech recognition and speech synthesis, speech separation is an important and key research topic and has received the key attention of researchers. From the perspective of application, single-channel systems have fewer restrictions on deployment and do not have potential problems such as the configuration stability of multi-microphone systems, and are easier to implement on devices. Therefore, single-channel speech separation is the most ideal research object. Single-channel speech separation technologies include spectral subtraction, Wiener filtering method, speech spectrum estimation method based on minimum mean square error, method based on auditory scene analysis, and method based on models. However, existing speech separation algorithms still have problems of insufficient generalization and accuracy in separating mixed speech with unknown number of speakers. Summary of the Invention
[0004] Aiming at the problems existing in speech segment feature extraction, speech segment classification, and speech separation model correction in the prior art, the present invention proposes a speech separation method that combines unsupervised classification related to spatial structural features and supervised correction.
[0005] The present invention is implemented by the following technical solutions: A speech separation method that combines unsupervised classification related to spatial structural features and supervised correction includes the following steps:
[0006] Step A: Extract the features of speech segments based on a time-delay cellular neural network;
[0007] Step A1: Calculate the modulation amplitude spectrum and phase spectrum based on envelope detection;
[0008] Step A2: Extract the features of the debugging amplitude spectrum based on the time-delay cellular neural network;
[0009] Step A3: Generate speech segment windows based on the mutation point detection method;
[0010] Step A4: Unify the feature dimensions of speech segments based on multi-scale spatial pyramid pooling.
[0011] Step B: Unsupervised adaptive classification of the speech segments obtained in Step A based on the dynamically growing self-organizing mapping neural network;
[0012] Step C: Adaptively correct the parameters of the speech separation model constructed in Steps A and B based on the particle swarm optimization algorithm;
[0013] Step D: Speech reconstruction of speech segments of the same class based on binary masking to obtain the target speech.
[0014] Further, when calculating the modulation amplitude spectrum by envelope detection in Step A1, the following specific method is adopted:
[0015] (1) Establish a filter bank with 128 channels using Gammatone filters;
[0016] (2) Perform envelope detection using Hilbert transform based on incoherent demodulation;
[0017] (3) Obtain the modulation amplitude spectrum through Fourier transform of 1024 points;
[0018] (4) Smooth the modulation amplitude spectrum through a low-pass filter.
[0019] Further, in Step A2, when extracting the features of the smoothed modulation amplitude spectrum, the following method is adopted: Construct a time-delay cellular neural network with a 128×1024 two-dimensional structure, and the output of the network is completely determined by the feedback template A, the control template B, the time-delay feedback template A τ , the time-delay control template B τ , the threshold I, and the time delay τ.
[0020] Further, in Step A3, generate speech segment windows according to the mutation point detection method:
[0021] (1) Calculate the first derivative of the smoothed modulation amplitude spectrum in each channel to obtain candidate onset / offset mutation points, where onset corresponds to the maximum point and offset corresponds to the minimum point; further screen the onset by setting a threshold; retain all the offsets with the minimum value between adjacent onsets and delete the remaining offsets;
[0022] (2) Take the average of the distances between all adjacent onsets in the current channel as the threshold; screen the onset set of adjacent channels whose distances from the onset of the current channel are less than the threshold, and select the onset with the smallest distance in the set for connection; use the same screening and connection method for offsets; cancel the connections with spans less than three adjacent channels.
[0023] (3) For the connection of onsets of consecutive channels, select the offset adjacent to the right of the onset, and construct an offset set of size Z; select an offset connection that passes through the largest number of points in the offset set as the offset connection that best matches the onset connection; end when all onsets of consecutive Z channels are successfully matched, otherwise repeat this process for the channels with failed matches; take the area between the matched onset and offset connections as the speech segment.
[0024] (4) Select the largest rectangular area in the speech segment as the speech segment window, and take the modulation amplitude spectrum features within the speech segment window as the speech segment features.
[0025] Further, in step A4, the multi-scale spatial pyramid pooling method is used to unify the dimensions of the speech segment features:
[0026] (1) Divide the speech window using 10 different scales of windows, and each scale represents one layer of the pyramid.
[0027] (2) Perform max-pooling operations on the speech segments within each pooling window, and expand the resulting R-dimensional feature vectors as the input to the speech segment classification neural network.
[0028] Further, in step B, an unsupervised adaptive classification of speech segments is performed based on a dynamically growing self-organizing map neural network, specifically as follows:
[0029] (1) The weight vector W of the root node (0) is assigned a random value in the interval [0, 1];
[0030] (2) Set the initial learning rate η(0) and growth factor α, and calculate the growth threshold G.
[0031] (3) Select a sample vector v from the sample set (i) , where i is the sample serial number, and find the nearest competitive layer node j according to the weighted distance to f (n) ; * ;
[0032] (4) Calculate the distance between v (i) and the competitive layer j *Error distance E between nodes: If E > G at this time, go to step e to perform the growth operation; otherwise, go to step 6 to perform the adjustment operation;
[0033] (5) Generate the child node of j * whose weight
[0034] (6) Adjust the weights of j * and its child nodes;
[0035] (7) Adopt a dynamic learning rate from large to small, and the learning rate can be adjusted according to the following formula:
[0036] η(t + 1) = λ.η(t)
[0037] (8) Repeat steps 3 - 7 until all samples are trained;
[0038] (9) Repeat step 8 to enter the next training cycle until no new nodes are generated in the network.
[0039] Furthermore, in step C, the parameters of the voice separation model are adaptively corrected based on the particle swarm optimization algorithm. Specifically as follows:
[0040] (1) Set the maximum number of generations G max for the particle swarm, the population size is S, and the acceleration factors are D 1 , D 2 , and the inertia weight ω;
[0041] (2) Use the parameter combination [A, A τ , B, B τ , I, τ, λ, α] as particles, randomly generate the parameter combinations of the population size as the initial positions of the particles, and randomly initialize the moving speeds of each particle;
[0042] (3) Select the mean square error of the classification error rate of the voice segment as the fitness function, calculate the value of the fitness function according to the positions of the particles, and use the simulated annealing algorithm to correct the fitness of the particles;
[0043] (4) Compare the magnitudes of the fitness values, and update the speeds and positions of the particles according to the individual local extreme values and the population global extreme values;
[0044] (5) Use the learning factor, inertia weight, and performance evaluation parameters as the inputs of the fuzzy controller, and use the percentages of the changes in the learning factor and inertia weight as the outputs, and use the fuzzy rules to adjust the inertia weight and learning factor simultaneously;
[0045] (6) Repeat steps 4 - 6. When the number of iterations reaches the maximum number of generations, obtain the global optimal particle position.
[0046] Further, in step D, based on the binary mask, the speech segments of the same class are speech-reconstructed to obtain the target speech. Specifically as follows:
[0047] (1) According to the result of the speech segment classification, the modulated amplitude spectrum of the nth speaker after separation is obtained by using the binary mask;
[0048] (2) Combining the modulated phase spectrum of the mixed speech, the speech envelope after separation is obtained by using the inverse Fourier transform;
[0049] (3) Combining the carriers of the mixed speech, the time-domain signals of different channels are obtained;
[0050] (4) The separated speech is synthesized by passing through an inverse Gammatone filter.
[0051] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0052] This solution uses the characteristics of the time-delay cellular neural network, such as parallel high speed and easy implementation on VLSI, to extract the characteristics of the timing signal, further improving the real-time performance of processing; at the same time, using the locally connected characteristic of the time-delay cellular neural network, fully considering the time-delay factor of information transmission between cells, highlighting the frequency band correlation and spatial structural information; proposing an unsupervised speech segment adaptive classification algorithm based on the dynamic growing self-organizing map neural network, simulating the load balancing process of neurons in the brain neural network, dynamically generating new sub-nodes in the output layer, and realizing the unsupervised adaptive classification of speech segments; proposing a speech separation algorithm that fuses the unsupervised speech segment classification related to spatial structural features and the supervised speech model adaptive correction, and through the organic combination of supervised learning and unsupervised learning, improving the generalization and accuracy of the speech separation model at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the principle block diagram of the speech separation described in the embodiment of the present invention;
[0054] Figure 2 is the schematic diagram of the calculation of the modulated amplitude spectrum and phase spectrum described in the embodiment of the present invention;
[0055] Figure 3 is the schematic diagram of the speech segment classification described in the embodiment of the present invention;
[0056] Figure 4 is the schematic diagram of the speech reconstruction described in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] To more clearly understand the above objects, features, and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0058] As Figure 1 shown, this embodiment proposes a voice separation method that fuses unsupervised classification and supervised correction related to spatial structural features, including the following steps:
[0059] Step 1: Extract voice segment features based on a time-delay cellular neural network;
[0060] Step 2: Adaptive classification of voice segments based on a dynamically growing self-organizing map neural network;
[0061] Step 3: Supervised language adaptive correction of the voice separation model based on the particle swarm optimization algorithm;
[0062] Step 4: Synthesize the target voice based on binary masking.
[0063] Step 1: Feature extraction of voice segments based on a time-delay cellular neural network
[0064] Voice feature extraction methods usually cannot fully utilize the correlation information between adjacent frequency bands and at the same time disrupt the spatial structure information inside the voice. To address this problem, using the feedback template and control template of the time-delay cellular neural network, through the nonlinear dynamic propagation effect between neighboring cells, the modulation amplitude spectrum features containing frequency band correlation and spatial structure information are obtained, and the accuracy and stability of feature extraction are improved based on the time-delay characteristics; furthermore, a mutation point detection method is used to generate a voice segment window to obtain the features of the voice segment; finally, the multi-scale spatial pyramid pooling method is used to unify the dimensions of the voice segment features to achieve effective extraction of the voice segment features. As Figure 2 shown, this embodiment uses the following method to extract features from voice segments:
[0065] 1. Modulation amplitude spectrum and phase spectrum calculation method
[0066] (1) Use a Gammatone filter to establish a filter bank with 128 channels. Specifically:
[0067] ① According to the empirical formula proposed by Moore and Glasberg, calculate the equivalent rectangular bandwidth (ERB) at a specific frequency:
[0068] ERB(f) = 24.7 * (0.00437f + 1)
[0069] ②Integrate the reciprocal of the product of the equivalent rectangular bandwidth and the fixed frequency interval to obtain a function that maps frequency to channel number, and further solve the inverse function that maps channel number to frequency;
[0070] ③Obtain the center frequencies of all channels by solving the fixed frequency interval, and then obtain the impulse response gt(t,k) of the Gammatone filter corresponding to the k-th channel;
[0071] ④The output of the k-th channel is obtained as:
[0072] y(t,k) = y(t) * gt(t,k)
[0073] where y(t) is the original input signal, and * represents the convolution operation.
[0074] (2) Use Hilbert transform based on non-coherent demodulation for envelope detection:
[0075]
[0076] where e(t,k) and c k (t) are the envelope and carrier of the k-th channel of the mixed speech signal respectively; Hilbert[.] is the Hilbert transform; LPF[.] is the low-pass filter; the cut-off frequency is 20 Hz.
[0077] Obtain the modulation amplitude spectrum Y a (i,k) through a 1024-point Fourier transform:
[0078]
[0079] (3) To eliminate the interference of weak fluctuations in the modulation spectrum, smooth the modulation amplitude spectrum through a low-pass filter. After smoothing the modulation amplitude spectrum of the k-th channel, we get:
[0080] Y s (i,k) = Y a (i,k) * g s (i)
[0081] where g s (i) is a low-pass FIR filter with a cut-off frequency of s Hz, and * represents the convolution operation. The value of s determines the degree of smoothing after filtering.
[0082] 2. Modulation Amplitude Spectrum Feature Extraction Method
[0083] Construct a time-delay cellular neural network with a 128×1024 two-dimensional structure. The output of the network is completely determined by the feedback template A, the control template B, the time-delay feedback template A τ , the time-delay control template Bτ , determined by the threshold I and the time delay τ. The network state and output are updated as follows:
[0084]
[0085] v yef (t) = tanh(v xef (t))
[0086] where v x (t), v u (t) and v y (t) are the state, input, and output of the cell respectively; C(h, l) is the cell in the h-th row and l-th column; N r (e, f) is the r-neighborhood of the cell in the e-th row and f-th column.
[0087] Through comparative experiments, in this embodiment, the neighborhood radius r is set to 2, the initial state v x (0) is set to 0, and the smooth modulation amplitude spectrum Y s (i, k) is used as the initial input v u (0) of the time-delay cellular neural network, and the output when the network converges is the feature of the modulation amplitude spectrum.
[0088] 3. Method for generating speech segment windows
[0089] Use the speech segment window generation method based on mutation point detection to obtain the mapping relationship between the modulation amplitude spectrum feature and the speech segment feature. Specifically:
[0090] (1) Calculate the first derivative of the smooth modulation amplitude spectrum in each channel to obtain candidate onset / offset mutation points, where onset corresponds to the maximum point and offset corresponds to the minimum point; further screen the onset by setting a threshold; retain all the offsets with the minimum value between adjacent onsets and delete the remaining offsets;
[0091] (2) Take the average value of the distances between all adjacent onsets in the current channel as the threshold; screen the set of onsets in adjacent channels whose distances from the onsets in the current channel are less than the threshold, and select the onset with the minimum distance in the set for connection; use the same screening and connection method for the offsets; cancel the connections with a span of less than three adjacent channels;
[0092] (3) For the connection line of the onset of a continuous channel, select the offset adjacent to the right side of the onset, and construct an offset set of size Z; select an offset connection line that passes through the largest number of points in the offset set as the offset connection line that best matches the onset connection line; end when the onsets of Z consecutive channels are all successfully matched, otherwise repeat this process for the channels with failed matches; use the area between the matched onset and offset connection lines as the speech segment;
[0093] (4) Select the largest rectangular area in the speech segment as the speech segment window, and use the modulation amplitude spectrum features within the speech segment window as the speech segment features, as Figure 1 shown.
[0094] 4. Feature Dimension Unification Method
[0095] The number of neurons in the input layer of the speech segment classification neural network is fixed, while the dimensions of the speech segment features are different. Further use multi-scale spatial pyramid pooling to unify the dimensions of the speech segment features. Specifically:
[0096] (1) Divide the speech window using windows of 10 different scales (30, 20, 15, 10, 8, 6, 4, 3, 2, 1). Each scale represents one layer of the pyramid. The window size of the pooling layer in the m-th layer is:
[0097]
[0098] where win_w and win_h are the width and height of the pooling window respectively, s m is the scale of the m-th layer, W in and H in are the width and height of the speech segment window respectively.
[0099] (2) Perform max-pooling operations on the speech segments within each pooling window, and expand the resulting R-dimensional feature vector as the input to the speech segment classification neural network.
[0100]
[0101] In this embodiment, R is 1755.
[0102] Step 2. Unsupervised Adaptive Classification of Speech Segments Based on Dynamically Growing Self-Organizing Map Neural Network
[0103] A voice separation model is usually a predefined static network structure before training. When the number of speakers in the mixed voice is unknown, a large number of attempts must be made to obtain a suitable network structure, resulting in a reduction in the generalization of the voice separation model. To address this problem, taking the features of voice segments as input, a flexible tree structure is adopted. Based on the root node in the initial state, the relationship between the growth threshold and the distance error is analyzed, simulating the load balancing process of neurons in the brain neural network, and dynamically generating new child nodes in the output layer to achieve unsupervised adaptive classification of voice segments. As Figure 3 shown, the following method is used in this embodiment to classify voice segments:
[0104] (1) The weight vector W (0) of the root node is assigned a random value in the interval [0, 1];
[0105] (2) Set the initial learning rate η(0) and the growth factor α to 0.5, and calculate the growth threshold G:
[0106]
[0107] where s is the total number of nodes in the growing self-organizing neural network;
[0108] (3) Select a sample vector v (i) from the sample set, where i is the sample serial number, and find the nearest competitive layer node j (n) to f according to the weighted distance * ; n ∈ {1, 2,..., N}
[0109] (4) Calculate the error distance E between v (i) and the competitive layer node j * :
[0110]
[0111] where R is the dimension of v (i) . If E > G at this time, go to step e to perform the growth operation, otherwise go to step 6 to perform the adjustment operation;
[0112] (5) Generate a child node of j * , and its weight
[0113] (6) Adjust the weights of j * and its child nodes:
[0114]
[0115] where, is the winning neighborhood of node j * ;
[0116] (7) A large learning rate can accelerate the learning speed but is not easy to converge. Therefore, a dynamic learning rate from large to small can be adopted, and the learning rate can be adjusted according to the following formula:
[0117] η(t + 1) = λ·η(t)
[0118] where λ is the adjustment factor, and its value range is (0, 1). In this embodiment, it is initialized to 0.5;
[0119] (8) Repeat steps 3 - 7 until all samples are trained;
[0120] (9) Repeat step 8 to enter the next training cycle until no new nodes are generated in the network.
[0121] Step 3: Adaptive correction of the supervised speech separation model based on the particle swarm optimization algorithm
[0122] The control template, feedback template, time delay, threshold, neighborhood radius of the time-delay cellular neural network and the growth factor, learning rate weight parameters of the dynamic growing self-organizing map neural network determine the accuracy of the speech separation model. However, the nodes in the output layer of the speech segment classification network represent different classification patterns, and it is difficult to adjust the parameters through the traditional backpropagation algorithm. To address this problem, using the particle swarm optimization algorithm, a combination of six model parameters, namely the control template, feedback template, time delay, threshold, growth factor, and learning rate weight, is selected as particles, and the error rate of speech segment classification is used as the fitness function. The inertia weight and learning factor of the particle swarm optimization algorithm are adaptively adjusted through a fuzzy controller, and the local extreme jump ability of simulated annealing is used to avoid falling into local minima, so as to globally optimize and update the parameters in the speech segment feature extraction and classification algorithm, and realize the adaptive correction of the speech separation model. In this embodiment, the following method is used to correct the speech model:
[0123] (1) Set the maximum number of generations G of the particle swarm max = 30, the population size is S = 30, and the acceleration factors D 1 = D 2 = 2, and the inertia weight ω = 1;
[0124] (2) Use the parameter combination [A, A τ , B, B τ , I, τ, λ, α] as particles, and randomly generate the parameter combinations of the population size number as the initial positions of the particles, and randomly initialize the moving speeds of each particle;
[0125] parameters combination as the initial positions of the particles, and randomly initialize the moving speeds of each particle;
[0126] (3) Select the mean square error of the speech segment classification error rate as the fitness function, and according to the position of the particle
[0127] Calculate the value of the fitness function, and use the simulated annealing algorithm to correct the fitness of the particles:
[0128]
[0129] where g is the generation number of evolution, p is the attenuation factor, T 0 is the initial temperature of simulated annealing, F and F new are the fitness before and after correction respectively;
[0130] (4) Compare the magnitudes of the fitness values, and update the velocity and position of the particles according to the local extreme value of the individual and the global extreme value of the population;
[0131] (5) Use the learning factor, inertia weight, and performance evaluation parameter as the inputs of the fuzzy controller, and use the percentages of the changes in the learning factor and inertia weight as the outputs, and use fuzzy rules to adjust the inertia weight and learning factor simultaneously;
[0132] (6) Repeat steps 4 - 6. When the number of iterations reaches the maximum generation number of evolution, obtain the global optimal particle position.
[0133] Step Four: Speech Reconstruction Based on Binary Masking
[0134] As Figure 4 shown, the following method is adopted in this embodiment to reconstruct the target speech:
[0135] (1) According to the result of speech segment classification, use binary masking to obtain the modulated amplitude spectrum of the nth speaker after separation
[0136]
[0137] (2) Combine the modulated phase spectrum of the mixed speech, and use the inverse Fourier transform to obtain the speech envelope after separation:
[0138]
[0139] (3) Combine the carrier wave of the mixed speech to obtain the time-domain signals of different channels:
[0140] y (n) (t,k) = e (n) (t,k).c(t,k)
[0141] (4) Synthesize the separated speech through the inverse Gammatone filter.
[0142] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An unsupervised classification and supervised correction fusion speech separation method related to spatial structural features, characterized in that, it includes the following steps: Step A: Extract the features of the speech segment based on the time-delay cellular neural network; Step A1: Calculate the modulation amplitude spectrum and phase spectrum based on envelope detection; Step A2: Extract the features of the debug amplitude spectrum based on the time-delay cellular neural network; Step A3: Generate the speech segment window based on the mutation point detection method; Step A4: Unify the feature dimensions of the speech segment based on the multi-scale spatial pyramid pooling; Step B: Perform unsupervised adaptive classification on the speech segments obtained in Step A based on the dynamic growing self-organizing map neural network; Step C: Adaptive correction of the parameters of the speech separation model constructed in Steps A and B based on the particle swarm optimization algorithm; Step D: Speech reconstruction of the speech segments of the same class based on binary masking to obtain the target speech.
2. The unsupervised classification and supervised correction fusion speech separation method related to spatial structural features according to claim 1, characterized in that: When calculating the modulation amplitude spectrum by envelope detection in Step A1, the following specific method is adopted: (1) Establish a filter bank with 128 channels using Gammatone filters; (2) Use Hilbert transform envelope detection based on incoherent demodulation; (3) Obtain the modulation amplitude spectrum through Fourier transform of 1024 points; (4) Smooth the modulation amplitude spectrum through a low-pass filter.
3. The unsupervised classification and supervised correction fusion speech separation method related to spatial structural features according to claim 1, characterized in that: In Step A2, the following method is adopted when extracting the features of the smoothed modulation amplitude spectrum: Construct a 128×1024 two-dimensional structure time-delay cellular neural network, and the output of the network is completely determined by the feedback template A, the control template B, the time-delay feedback template A τ , the time-delay control template B τ , the threshold I and the time-delay τ; the update methods of the network state and output are as follows: v yef (t) = tanh(v xef (t)) where v x (t), v u (t) and v y (t) are the state, input, and output of the cell, respectively; C(h, l) is the cell in the h-th row and l-th column; N r (e, f) is the r-neighborhood of the cell in the e-th row and f-th column.
4. The unsupervised classification and supervised correction fusion speech separation method related to spatial structural features according to claim 1, characterized in that: In Step A3, the speech segment window is generated according to the mutation point detection method: (1) Calculate the first derivative of the smoothed modulation amplitude spectrum in each channel to obtain candidate onset / offset mutation points, where onset corresponds to the maximum value point and offset corresponds to the minimum value point; further screen onset by setting a threshold; retain all offsets with the smallest value between adjacent onsets, and delete the remaining offsets; (2) Take the average value of the distances between all adjacent onsets in the current channel as the threshold; screen the set of onsets in adjacent channels whose distances from the onset of the current channel are less than the threshold, and select the onset with the smallest distance in the set for connection; use the same screening and connection method for offsets; cancel the connection with a span of less than three adjacent channels; (3) For the connection lines of onsets of individual consecutive channels, select the offset adjacent to the right of the onset to construct an offset set of size Z; select an offset connection line that passes through the largest number of points in the offset set as the offset connection line that best matches the onset connection line; end when the onsets of consecutive Z channels are all successfully matched, otherwise repeat this process for the channels with failed matches; take the area between the matched onset and offset connection lines as the speech segment; (4) Select the largest rectangular area in the speech segment as the speech segment window, and take the modulation amplitude spectrum features within the speech segment window as the speech segment features.
5. The unsupervised classification and supervised correction fusion speech separation method related to the spatial structural feature according to claim 1, characterized in that, in step A4, the multi-scale spatial pyramid pooling method is used to unify the dimensions of the speech segment features: (1) Use windows of 10 different scales (30, 20, 15, 10, 8, 6, 4, 3, 2, 1) to divide the speech window. Each scale represents one layer of the pyramid. The window size of the pooling layer of the m-th layer is: where win_w and win_h are the width and height of the pooling window respectively, and s m is the scale of the m-th layer, W in and H in are the width and height of the speech segment window respectively; (2) Perform a max pooling operation on the speech segments within each pooling window, and expand the obtained R-dimensional feature vector as the input of the speech segment classification neural network.
6. The unsupervised classification and supervised correction fusion speech separation method related to the spatial structural feature according to claim 1, characterized in that, in step B, the speech segments are classified unsupervised and adaptively based on the dynamic growing self-organizing map neural network, specifically as follows: (1) The weight vector W of the root node (0) is assigned a random value in the interval [0, 1]; (2) Set the initial learning rate η(0) and the growth factor α to 0.5, and calculate the growth threshold G: where s is the total number of nodes in the growing self-organizing neural network; (3) Select the sample vector v from the sample set (i) , where i is the sample serial number, and find the competitive layer node j closest to f according to the weighted distance (n) ; n ∈ {1, 2,..., N} * (4) Calculate v (i) with the competitive layer j * The error distance E between nodes: where R is the dimension of v (i) ; if E > G at this time, go to step e to perform the growth operation, otherwise go to step 6 to perform the adjustment operation; (5) Generate the child nodes of j * whose weights (6) Adjust j * and the weights of its child nodes: Among them, is the winning neighborhood of node j * ; (7) A large learning rate can accelerate the learning speed but is not easy to converge. Therefore, a dynamic learning rate from large to small can be adopted, and the learning rate can be adjusted according to the following formula: η(t + 1) = λ.η(t) where λ is the adjustment factor, and the value range is (0, 1); (8) Repeat steps 3 - 7 until all samples are trained; (9) Repeat step 8 to enter the next training cycle until no new nodes are generated in the network.
7. The unsupervised classification and supervised correction fusion speech separation method related to the spatial structural feature according to claim 1, characterized in that, in step C, the parameters of the speech separation model are adaptively corrected based on the particle swarm optimization algorithm, specifically as follows: (1) Set the maximum number of generations G for the particle swarm optimization max = 30, the population size is S = 30, and the acceleration factor D 1 = D 2 = 2, and the inertia weight ω = 1; (2) Using the parameter combination [A, A τ , B, B τ , I, τ, λ, α] as particles, randomly generate the parameter combinations of the population size as the initial positions of the particles, and randomly initialize the moving speeds of each particle; (3) Select the mean square error of the speech segment classification error rate as the fitness function, calculate the value of the fitness function according to the position of the particle, and use the simulated annealing algorithm to correct the fitness of the particle: Among them, g is the number of evolutionary generations, p is the attenuation factor, and T 0 is the initial temperature of simulated annealing, and F and F new are the fitness before and after correction, respectively; (4) Compare the fitness values, and update the velocity and position of the particle according to the individual local extreme value and the population global extreme value; (5) Use the learning factor, inertia weight, and performance evaluation parameter as the inputs of the fuzzy controller, and take the percentages of the changes in the learning factor and inertia weight as the outputs, and use the fuzzy rules to adjust the inertia weight and learning factor simultaneously; (6) Repeat steps 4 - 6. When the number of iterations reaches the maximum number of evolutionary generations, obtain the globally optimal particle position.
8. The unsupervised classification and supervised correction fusion speech separation method related to the spatial structural features according to claim 1, characterized in that, in the said step D, based on binary masking, speech reconstruction is performed on speech segments of the same class to obtain the target speech, specifically as follows: (1) According to the result of speech segment classification, use binary masking to obtain the modulation amplitude spectrum of the nth speaker after separation (2) Combine the modulation phase spectrum of the mixed speech and use the inverse Fourier transform to obtain the speech envelope after separation: (3) Combine the carriers of the mixed speech to obtain the time-domain signals of different channels: y (n) (t, k) = e (n) (t, k).c(t, k) (4) Synthesize the separated speech through the inverse Gammatone filter.
Citation Information
Patent Citations
Voice data processing method and equipment
CN101136199A
Cellular nerve network with genetic algorithm (GACNN)-based multisource image fusion method
CN103971329A