Multi-channel speech separation method and device based on multi-scale feature channel fusion

By introducing multi-scale feature channel fusion and convolutional separation modules into the multi-channel speech separation algorithm, the shortcomings of existing algorithms in capturing context information at different scales are solved, and the performance and separation effect of speech separation are significantly improved.

CN119252272BActive Publication Date: 2025-05-13SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411765064.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-13
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The separation capability of the existing time domain multi-channel speech separation algorithm still needs to be improved, and it is impossible to effectively capture context information at different scales, resulting in poor separation effect.

Method used

The multi-channel speech separation method based on multi-scale feature channel fusion is adopted, and the multi-scale feature separation is fused through the high-dimensional feature extraction module, the spatial feature extraction module and the multi-scale feature extraction module, and the speaker mask is calculated using a convolutional separation network based on feature channel fusion, and finally the speech separation is achieved.

Benefits of technology

By introducing multi-scale feature extraction module and convolutional separation module, the model can better capture context information at different scales, improve the performance and generalization ability of speech separation, and significantly improve the separation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119252272B_ABST
    Figure CN119252272B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-channel speech separation method and device based on multi-scale feature channel fusion, the method comprising: obtaining a plurality of multi-channel mixed speech signals with different noises, reverberations and speakers to form a training data set; constructing a multi-channel speech separation network based on multi-scale feature channel fusion, specifically comprising a high-dimensional feature extraction module, a spatial feature extraction module, a multi-scale feature extraction module, a convolutional separation network based on feature channel fusion, and a speech reconstruction module; inputting the training data set into the multi-channel speech separation network to perform network training; inputting the mixed multi-channel speech signal containing noise, reverberation and multiple speakers to be separated into the trained multi-channel speech separation network to obtain the single-channel speech signal of each speaker. The present invention has stronger separation ability and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to speech separation technology, and in particular to a multi-channel speech separation method and device based on multi-scale feature channel fusion. Background Art

[0002] The purpose of speech separation technology is to separate the target speech signal from the mixed speech signal. This technology is the core task in the field of speech signal processing. Realizing the separation of the target speaker's speech can improve the intelligibility and perceptual quality of the separated speech, thereby greatly improving the performance of systems such as speech recognition, speech emotion recognition, and speech translation. In addition to time domain and frequency domain features, multi-channel speech separation technology can also use the spatial features of inter-channel signals to separate mixed signals, which has better generalization and robustness than single-channel speech separation.

[0003] In the speech separation algorithm based on deep learning, relying on the powerful nonlinear modeling ability of deep neural networks, the speech separation model can separate the speech signal of the target speaker without any statistical assumptions after a large amount of data training. Multi-channel speech separation algorithms based on deep learning are usually divided into two categories: time-frequency domain and time domain. Due to the shortcomings of time-frequency domain methods such as reconstructed phase and high computational complexity of time-frequency domain features, multi-channel speech separation algorithms based on time domain have received more and more attention. However, the separation ability of multi-channel speech separation algorithms in time domain still needs to be improved. Summary of the invention

[0004] In view of the problems existing in the prior art, the purpose of the present invention is to provide a multi-channel speech separation method, device and storage medium based on multi-scale feature channel fusion with better separation capability.

[0005] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0006] A multi-channel speech separation method based on multi-scale feature channel fusion, comprising:

[0007] Step 1: obtain a number of multi-channel mixed speech signals with different noises, reverberations and speakers, and use the corresponding pure single-speaker single-channel speech signals as labels to form a training data set;

[0008] Step 2: Construct a multi-channel speech separation network based on multi-scale feature channel fusion, which specifically includes:

[0009] A high-dimensional feature extraction module is used to extract high-dimensional features of a reference channel mixed speech signal, wherein the reference channel mixed speech signal is a mixed speech signal on any channel of a multi-channel mixed speech signal;

[0010] A spatial feature extraction module, used to extract inter-channel convolution difference parameters of multi-channel mixed speech signals;

[0011] A multi-scale feature extraction module, used for fusing the high-dimensional features and the inter-channel convolution difference parameters at multiple scales to obtain multi-scale fusion features;

[0012] A convolutional separation network based on feature channel fusion is used to calculate the mask of each speaker in the multi-channel mixed speech signal based on the multi-scale fusion features;

[0013] A speech reconstruction module, used to reconstruct the pure single-channel speech signal of each speaker according to the mask of each speaker and the high-dimensional features, so as to achieve speech separation;

[0014] Step 3: input the training data set into the multi-channel speech separation network to perform network training;

[0015] Step 4: Input the mixed multi-channel speech signal containing noise, reverberation and multiple speakers to be separated into the trained multi-channel speech separation network to obtain the single-channel speech signal of each speaker.

[0016] Furthermore, the high-dimensional feature extraction module includes a connected one-dimensional convolutional layer and a PReLU activation function.

[0017] Furthermore, the spatial feature extraction module includes two parallel feature extraction branches, each of which includes a connected two-dimensional convolutional layer and a PReLU activation function.

[0018] Furthermore, the multi-scale feature extraction module includes a first scale branch, a second scale branch, a third scale branch and a fusion branch, the input of the fusion branch is the feature of the output splicing of the first scale branch, the second scale branch and the third scale branch, the first scale branch includes a first scale one-dimensional convolution layer, a group normalization layer and a PReLU activation function connected in sequence, the second scale branch includes a second scale one-dimensional convolution layer, a group normalization layer and a PReLU activation function connected in sequence, the third scale branch includes a third scale one-dimensional convolution layer, a group normalization layer and a PReLU activation function connected in sequence, and the fusion branch includes a two-dimensional convolution layer, a group normalization layer and a PReLU activation function.

[0019] Furthermore, the convolutional separation network based on feature channel fusion includes a global normalization layer, a one-dimensional convolutional layer, B stacked convolutional separation modules, a PReLU activation function, a one-dimensional convolutional layer and a sigmoid activation function connected in sequence, wherein the B convolutional separation modules are specifically dilated by factors of 2 b-1 Convolution separation module, b=1,…,B, B is a positive integer greater than 1.

[0020] Furthermore, the convolution separation module specifically includes a first one-dimensional convolution layer, a PReLU activation function, a first global normalization layer, a global average pooling layer, a second one-dimensional convolution layer, a fully connected layer, a sigmoid activation function, a dot product operation, a first splicing operation, a depth convolution layer, a second PReLU activation function, a second global normalization layer, a third one-dimensional convolution layer and a second splicing operation connected in sequence, wherein the output of the first global normalization layer and the output of the sigmoid activation function perform the dot product operation, the output of the first global normalization layer also performs the first splicing operation with the output of the dot product operation, the input of the first one-dimensional convolution layer also performs the second splicing operation with the output of the third one-dimensional convolution layer, and the output of the second splicing operation is the output of the convolution separation module.

[0021] Furthermore, the speech reconstruction module is specifically a one-dimensional transposed convolutional layer.

[0022] Furthermore, the loss function used in the training of the multi-channel speech separation network based on multi-scale feature channel fusion is:

[0023]

[0024] in, represents the loss function, is the scale-invariant signal-to-noise ratio of the network output to the label.

[0025] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above method.

[0026] A computer-readable storage medium having a computer program / instruction stored thereon, wherein the computer program / instruction implements the above method when executed by a processor.

[0027] Compared with the prior art, the present invention has the following beneficial effects: by introducing a multi-scale feature extraction module, the present invention overcomes the limitation that the traditional convolutional separation network cannot well capture context information of different scales, provides more comprehensive and richer context information, and improves the separation performance and generalization ability of the model. In addition, a convolutional separation module is added to the convolutional separation network based on feature channel fusion. The convolutional separation module takes into account the correlation between feature channels, realizes the cross-channel interaction of feature information, can capture the relationship between feature channels, strengthens key feature information, and improves the accuracy of the separation module in estimating the speaker mask. The present invention has achieved significant improvements in various evaluation indicators of speech separation and achieved better separation effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flow chart of a multi-channel speech separation method based on multi-scale feature channel fusion provided by the present invention;

[0029] Figure 2 It is a structural schematic diagram of a multi-channel speech separation network based on multi-scale feature channel fusion provided by the present invention;

[0030] Figure 3 It is a structural schematic diagram of the high-dimensional feature extraction module provided by the present invention;

[0031] Figure 4 It is a structural schematic diagram of the spatial feature extraction module provided by the present invention;

[0032] Figure 5 It is a structural schematic diagram of the multi-scale feature extraction module provided by the present invention;

[0033] Figure 6 It is a structural schematic diagram of a convolutional separation network based on feature channel fusion provided by the present invention;

[0034] Figure 7 It is a structural schematic diagram of the convolution separation module provided by the present invention. DETAILED DESCRIPTION

[0035] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0036] This embodiment 1 provides a multi-channel speech separation method based on multi-scale feature channel fusion, such as Figure 1 As shown, the following steps are included:

[0037] Step 1: Obtain several multi-channel mixed speech signals with different noises, reverberations and speakers, and use the corresponding pure single-speaker single-channel speech signals as labels to form a training data set.

[0038] In the specific implementation, firstly, a number of pure, noise-free and reverberation-free single-speaker single-channel speech signals are obtained, and an open source speech database can be used. In this embodiment, the single-speaker single-channel speech signal comes from the Librispeech dataset. The data in the Librispeech dataset are all pure speaker speech collected in a quiet environment by a single channel, and a total of 100 hours of corpus are included.

[0039] For each pure single-speaker single-channel speech signal, the impulse response of the specified direction is generated by the mirror method, and convolved with the pure single-speaker single-channel speech signal to obtain a single-speaker multi-channel speech signal of the specified direction, and the single-speaker multi-channel speech signals of different directions are added together to obtain a multi-channel mixed speech signal containing multiple speakers. When the speech of different speakers is added, the signal energy ratio is adjusted to a random value between 0 and 5dB.

[0040] For multi-channel mixed speech signals, noise with different signal-to-noise ratios and different types of reverberation are added to obtain several multi-channel mixed speech signals with different noise, reverberation and speakers. The noise signal comes from the 100 Nonspeech corpus, which contains 100 noises of different lengths and in various life scenarios. The relative signal-to-noise ratio between the pure speech signal and the noise is randomly sampled between 10 and 20 dB, and the reverberation time RT60 is uniformly sampled between 0.1 and 0.5 seconds.

[0041] All multi-channel mixed speech signals with different noises, reverberations and speakers are used as samples, and the corresponding pure single-speaker single-channel speech signals are used as labels to form a training data set.

[0042] Step 2: Construct a multi-channel speech separation network based on multi-scale feature channel fusion.

[0043] Among them, the structure of the multi-channel speech separation network is as follows Figure 2 As shown, it specifically includes a high-dimensional feature extraction module, a spatial feature extraction module, a multi-scale feature extraction module, a convolutional separation network based on feature channel fusion, and a speech reconstruction module.

[0044] The high-dimensional feature extraction module is used to extract high-dimensional features of the reference channel mixed speech signal, wherein the reference channel mixed speech signal is a mixed speech signal on any channel of the multi-channel mixed speech signal. Figure 3 As shown in the figure, the high-dimensional feature extraction module includes a connected one-dimensional convolution layer (1-D Conv) and a PReLU activation function. The convolution kernel size of the one-dimensional convolution layer is 32, the step size is 16, and the number of output channels is 512.

[0045] The spatial feature extraction module is used to extract the inter-channel convolution difference parameters of the multi-channel mixed speech signal, such as Figure 4As shown, it specifically includes two parallel feature extraction branches, each of which includes a connected two-dimensional convolution layer (2-D Conv) and a PReLU activation function. For example, for a multi-channel speech signal containing 6 channels, the channel convolution difference parameters between channel pairs with channel indices (1,2), (3,4), (5,6), (1,4), (2,5), (3,6) are selected as the inter-channel convolution difference parameters to be extracted. The convolution kernel size of the first two-dimensional convolution layer is (2,32), the expansion factor is (1,1), and the step size is (2,16), which is used to extract the channel convolution difference parameters of the three channel pairs (1,2), (3,4), and (5,6); the convolution kernel size of the second two-dimensional convolution layer is (2,32), the expansion factor is (3,1), and the step size is (1,16), which is used to extract the channel convolution difference parameters of the three channel pairs (1,4), (2,5), and (3,6).

[0046] The multi-scale feature extraction module is used to perform multi-scale fusion of the high-dimensional features and the inter-channel convolution difference parameters to obtain multi-scale fusion features. Figure 5 As shown, the multi-scale feature extraction module includes a first scale branch, a second scale branch, a third scale branch and a fusion branch. The input of the fusion branch is the output spliced ​​features of the first scale branch, the second scale branch and the third scale branch. The first scale branch includes a first scale one-dimensional convolution layer (1-D Conv with a convolution kernel size of 1), a group normalization layer (GN) and a PReLU activation function connected in sequence. The second scale branch includes a second scale one-dimensional convolution layer (1-D Conv with a convolution kernel size of 3), a group normalization layer (GN) and a PReLU activation function connected in sequence. The third scale branch includes a third scale one-dimensional convolution layer (1-D Conv with a convolution kernel size of 5), a group normalization layer (GN) and a PReLU activation function connected in sequence. The fusion branch includes a two-dimensional convolution layer (2-D Conv with a convolution kernel size of (1,1)), a group normalization layer (GN) and a PReLU activation function.

[0047] The convolutional separation network based on feature channel fusion is used to calculate the mask of each speaker in the multi-channel mixed speech signal based on the multi-scale fusion features. Figure 6 As shown in Figure 1, it includes a global normalization layer (gLN), a one-dimensional convolution layer (1-D Conv with a convolution kernel size of 1), B stacked convolution separation modules, a PReLU activation function, a one-dimensional convolution layer (1-D Conv with a convolution kernel size of 1) and a sigmoid activation function, wherein the B convolution separation modules are specifically connected with a dilation factor d of 2. b-1The convolution separation module is b=1,…,B, where B is a positive integer greater than 1. The exponentially growing expansion factor can exponentially expand the receptive field without increasing the amount of computation, fully capture the contextual information of long sequence inputs and retain the long-term dependencies of speech signals, and better calculate the mask of each speaker in the mixed speech signal.

[0048] like Figure 7 As shown, the convolution separation module specifically includes a first one-dimensional convolution layer (1-D Conv with a convolution kernel size of 1), a PReLU activation function, a first global normalization layer (gLN), a global average pooling layer (GAP), a second one-dimensional convolution layer (1-D Conv with a convolution kernel size of 3), a fully connected layer (FC), a sigmoid activation function, a dot product operation, a first splicing operation, a depthwise convolution layer (Depthwise Conv), a second PReLU activation function, a second global normalization layer (gLN), a third one-dimensional convolution layer (1-D Conv with a convolution kernel size of 1) and a second splicing operation, wherein the output of the first global normalization layer performs the dot product operation with the output of the sigmoid activation function, the output of the first global normalization layer also performs the first splicing operation with the output of the dot product operation, the input of the first one-dimensional convolution layer also performs the second splicing operation with the output of the third one-dimensional convolution layer, and the output of the second splicing operation is the output of the convolution separation module. The convolutional separation module takes into account the correlation between feature channels and uses convolution operations to achieve cross-channel interaction of feature information. It can capture the relationship between feature channels, strengthen key feature information, and improve the effectiveness of features. By adding the convolutional separation module, the subsequent separation operation of the convolutional separation network is facilitated.

[0049] The speech reconstruction module is used to reconstruct the pure single-channel speech signal of each speaker according to the mask of each speaker and the high-dimensional features to achieve speech separation. The speech reconstruction module is specifically a one-dimensional transposed convolution layer.

[0050] Step 3: Input the training data set into the multi-channel speech separation network to perform network training.

[0051] Specifically, the multi-channel mixed speech signal is input into the multi-channel speech separation network, and the reference channel mixed speech signal is input into the feature extraction module to obtain the single-channel speech signal of each speaker after separation as the network output. Then, the network is trained based on the label using the forward propagation and back propagation algorithms. The loss function used in the training is: ,in, represents the loss function, is the scale-invariant signal-to-noise ratio of the network output to the label.

[0052] Step 4: Input the mixed multi-channel speech signal containing noise, reverberation and multiple speakers to be separated into the trained multi-channel speech separation network to obtain the single-channel speech signal of each speaker.

[0053] In order to verify the effect of the present invention, the method of this embodiment is simulated and verified, and the improved signal distortion ratio (Source to Distortion Ratio improvement, SDRi) and the improved scale-invariant signal-to-noise ratio (Scale Invariant Source-to-Noise Ratio improvement, SISNRi) are used to evaluate the separation effect, PESQ is used to evaluate the speech quality, and STOI is used to evaluate the speech intelligibility. The perceptual evaluation of speech quality (PESQ) score is based on the ITU-T P.862 standard. It is an objective speech quality evaluation method that uses the original signal as a reference to measure the quality of the degraded signal and returns a score in the range of -0.5~4.5. The short-time objective intelligibility (STOI) score is an objective evaluation method for measuring the human auditory perception system for speech intelligibility. It is expressed as a percentage and uses the original signal as a reference to measure the intelligibility of the degraded signal. The present invention compares the performance of a multi-channel speech separation network based on multi-scale features and feature channel fusion and a traditional time-domain convolution multi-channel separation algorithm, thereby analyzing the impact of multi-scale feature parameters and feature channel fusion on the multi-channel speech separation algorithm. The final performance evaluation is as follows:

[0054] A. The evaluation of speech separation algorithm is shown in Table 1:

[0055] Table 1 Evaluation of speech separation algorithms

[0056]

[0057] B. The influence of reverberation time on speech separation algorithm is shown in Table 2:

[0058] Table 2 Effect of reverberation time on speech separation algorithm

[0059]

[0060] Table 2 shows the experimental results of various speech separation algorithms when the reverberation time is set to different ranges. It can be seen that compared with the traditional time domain convolution multi-channel separation algorithm, the present invention has improved various evaluation indicators and achieved better separation effect.

[0061] Embodiment 2 of the present invention provides a computer device, and the embodiment of the present invention provides services for implementing the method of embodiment 1 of the present invention. The device may include: a memory storing a computer executable program; a processor coupled to the memory; the processor calls the computer executable program stored in the memory to execute the steps in the method described in embodiment 1.

[0062] The memory may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the memory may be used to read and write non-removable, non-volatile magnetic media (commonly referred to as a "hard drive"). A program / utility having a set (at least one) of program modules may be stored in, for example, the memory, such program modules including but not limited to an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The computer executable program of the program module typically performs the functions and / or methods in the embodiments described in the present invention.

[0063] The processor executes various functional applications and data processing by running the program stored in the memory, such as implementing the method provided in the first embodiment of the present invention.

[0064] The code of the computer executable program can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages.

[0065] Embodiment 3 of the present invention provides a storage medium including a computer executable program. When the computer executable program is executed by a computer processor, it is used to execute the method of embodiment 1.

[0066] The storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device.

[0067] Of course, the storage medium containing a computer executable program provided by an embodiment of the present invention, the computer executable program of which is not limited to the above method operations, can also execute related operations in the method provided by any embodiment of the present invention.

[0068] It should be noted that, in this document, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0069] It should be understood that the above embodiments and descriptions only describe the principles, main features and advantages of the present invention. Without departing from the spirit and scope of the present invention, the present invention may be subject to various changes and improvements, and these changes and improvements all fall within the scope of protection of the present invention.

Claims

1. A multi-channel speech separation method based on multi-scale feature channel fusion, characterized in that: include: Step 1: obtain a number of multi-channel mixed speech signals with different noises, reverberations and speakers, and use the corresponding pure single-speaker single-channel speech signals as labels to form a training data set; Step 2: Construct a multi-channel speech separation network based on multi-scale feature channel fusion, which specifically includes: A high-dimensional feature extraction module is used to extract high-dimensional features of a reference channel mixed speech signal, wherein the reference channel mixed speech signal is a mixed speech signal on any channel of a multi-channel mixed speech signal; A spatial feature extraction module, used to extract inter-channel convolution difference parameters of multi-channel mixed speech signals; A multi-scale feature extraction module, used for fusing the high-dimensional features and the inter-channel convolution difference parameters at multiple scales to obtain multi-scale fusion features; A convolutional separation network based on feature channel fusion is used to calculate the mask of each speaker in the multi-channel mixed speech signal based on the multi-scale fusion features; A speech reconstruction module, used to reconstruct the pure single-channel speech signal of each speaker according to the mask of each speaker and the high-dimensional features, so as to achieve speech separation; Step 3: input the training data set into the multi-channel speech separation network to perform network training; Step 4: Input the mixed multi-channel speech signal containing noise, reverberation and multiple speakers to be separated into the trained multi-channel speech separation network to obtain the single-channel speech signal of each speaker; The multi-scale feature extraction module includes a first scale branch, a second scale branch, a third scale branch and a fusion branch, the input of the fusion branch is the output spliced ​​features of the first scale branch, the second scale branch and the third scale branch, the first scale branch includes a first scale one-dimensional convolution layer, a group normalization layer and a PReLU activation function connected in sequence, the second scale branch includes a second scale one-dimensional convolution layer, a group normalization layer and a PReLU activation function connected in sequence, the third scale branch includes a third scale one-dimensional convolution layer, a group normalization layer and a PReLU activation function connected in sequence, and the fusion branch includes a two-dimensional convolution layer, a group normalization layer and a PReLU activation function; The convolutional separation network based on feature channel fusion includes a global normalization layer, a one-dimensional convolutional layer, B stacked convolutional separation modules, a PReLU activation function, a one-dimensional convolutional layer and a sigmoid activation function connected in sequence, wherein the B convolutional separation modules are specifically dilated by factors of 2 b-1 Convolution separation module, b=1,…,B, B is a positive integer greater than 1; The convolution separation module specifically includes a first one-dimensional convolution layer, a PReLU activation function, a first global normalization layer, a global average pooling layer, a second one-dimensional convolution layer, a fully connected layer, a sigmoid activation function, a dot product operation, a first splicing operation, a depth convolution layer, a second PReLU activation function, a second global normalization layer, a third one-dimensional convolution layer and a second splicing operation connected in sequence, wherein the output of the first global normalization layer and the output of the sigmoid activation function perform the dot product operation, the output of the first global normalization layer also performs the first splicing operation with the output of the dot product operation, the input of the first one-dimensional convolution layer also performs the second splicing operation with the output of the third one-dimensional convolution layer, and the output of the second splicing operation is the output of the convolution separation module.

2. The multi-channel speech separation method based on multi-scale feature channel fusion according to claim 1, characterized in that: The high-dimensional feature extraction module includes a connected one-dimensional convolutional layer and a PReLU activation function.

3. The multi-channel speech separation method based on multi-scale feature channel fusion according to claim 1, characterized in that: The spatial feature extraction module includes two parallel feature extraction branches, each of which includes a connected two-dimensional convolutional layer and a PReLU activation function.

4. The multi-channel speech separation method based on multi-scale feature channel fusion according to claim 1 is characterized in that: The speech reconstruction module is specifically a one-dimensional transposed convolutional layer.

5. The multi-channel speech separation method based on multi-scale feature channel fusion according to claim 1 is characterized in that: The loss function used in the training of the multi-channel speech separation network based on multi-scale feature channel fusion is: , in, represents the loss function, is the scale-invariant signal-to-noise ratio of the network output to the label.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor executes the computer program to implement the method according to any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: The computer program / instructions, when executed by a processor, implement the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Voice data separation method and device, equipment and storage medium

    CN113470688A

  • Multi-modal time domain voice separation method based on multiple scales

    CN115881156A