METHOD AND SPEECH SEPARATION DEVICE

DE102025101110A1Pending Publication Date: 2025-07-31HYUNDAI MOTOR CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025101110
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-01
Filing Date
2025-01-14
Publication Date
2025-07-31

Smart Images

  • Figure 00000019_0000
    Figure 00000019_0000
  • Figure 00000019_0001
    Figure 00000019_0001
  • Figure 00000020_0000
    Figure 00000020_0000
Patent Text Reader

Abstract

A method and apparatus for separating speech based on a local and global modeling network (LAGNet) are provided. A speech separation apparatus comprises an encoder configured to progressively compress a sequence of mixed speech using a one-dimensional convolution-based local block to generate local information, a bottleneck configured to generate global information using a multiple-self-attention-based global block, and a decoder configured to progressively reconstruct the sequence using the local block. A speech separation apparatus also uses a skip connection comprising gates configured to filter the local information using the global information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELDThe present disclosure relates to a method and apparatus for separating speech contributions based on a local and global modeling network (LAGNet).BACKGROUNDThe statements in this section merely provide information related to the present disclosure and do not necessarily represent the prior art.A speech separation extracts the individual speech of each speaker from mixed speech contributions produced by two or more simultaneously speaking speakers. The speech separation model does not use a separate short-time Fourier transform (STFT) for processing mixed speech contributions, but uses an end-to-end processing method to handle mixed speech contributions directly in the time domain. To learn a representation replacing the STFT domain, the speech separation model uses a one-dimensional convolutional audio encoder and decoder replacing the STFT and inverse STFT (iSTFT) operations. The conv-TasNet (Time-domain Audio Separation Network) model is used as a speech separation model that processes mixed speech contributions in the time domain based on one-dimensional convolution. By accumulating a one-dimensional convolution-based temporal convolutional network (TCN) based on dilation of powers of two, the conv-TasNet implements local modeling and ensures a sufficient receptive field. Compared to STFT-based models, the conv-TasNet (time-domain audio separation network) model exhibits excellent performance in terms of scale-invariant signal-to-noise ratio (SI-SNRi) improvement.However, the problem of TasNet is that it does not have global modeling despite the need to process a very long sequence in the time domain. In order to effectively process a long sequence, a bi-LSTM (Bi-LSTM)-based dual path recurrent neural network (DPRNN) or a transformer-based sepformer is proposed as a speech separation model based on a dual path sequence model. The dual path model-based speech separation model divides a very long sequence into chunks of predetermined length and uses a path for processing intra-chunk sequences and a path for processing external inter-chunk sequences. Intra-chunk sequences and external inter-chunk sequences correspond to local modeling and global modeling, respectively, for long sequences. While the dual path model-based speech separation model significantly improves the performance of TasNet, it still has the problem of requiring significant computational resources to create chunks and process inter-chunk sequences.In order to reduce the excessive computational load associated with the dual path model, the SuDoRM RF (Successive Down-sampling and Re-sampling of Multi-Resolution Features) model is proposed as a speech separation model using a multi-scale sequence model. The SuDoRM RF model uses a multiscale U-net based sequence model to expand the receptive field, thereby reducing computational effort. However, since all layers of the multiscale sequence model are implemented as one-dimensional convolutions, the SuDoRM RF model still has the problem of not effectively modelling the global information of the sequence.SUMMARYAccordingly, a speech separation model capable of reducing the computational burden while improving the performance of speech separation by also processing global information in addition to local information of the sequence should be considered.Embodiments of the present disclosure provide a method and apparatus for separating speech contributions using the Local And Global Modeling Network (LAGNet), which is a deep learning-based speech separation model that includes a multi-scale sequence model.Moreover, embodiments of the present disclosure provide a method and apparatus for separating speech contributions using a multi-scale sequence model. The multi-scale sequence model includes an encoder that progressively compresses a sequence of mixed phrases using a one-dimensional convolution-based local modeling block, a Bottleneck that uses a multiple self attention-based global modeling block, and a decoder that progressively reconstructs the sequence using the local modeling block.According to at least one aspect of the present disclosure, a voice separation device is provided. The speech separation apparatus comprises a local encoder having local encoding levels which allows each local encoding level to generate local information by progressively compressing latent representations of mixed speech contributions using the local encoding levels and generating a compressed output based on an output of the last local encoding level. The speech separation device also includes a Bottleck having a global encoding level and generates global information from the compressed output of the local encoder using the global level. The speech separation apparatus also includes a local decoder having local decoding stages that enables each local decoding stage to generate reconstructed local information by progressively expanding the global information using the local decoding stages and providing an output of the last local decoding stage as a reconstructed latent output. The speech separation device also includes a skip connection having global gates and local gates that filters local information of the local encoding stages using the global gates and the global information, and filters an output of the global gates using the local gates and the reconstructed local information of the local decoding stages.According to another aspect of the present disclosure, a method of separating speech contributions performed by a speech separation device is provided. The method includes enabling each local encoding stage to generate local information by progressively compressing latent representations of mixed speech contributions using local encoding stages within a local encoder. The method also includes generating a compressed output based on an output of a last local encoding stage. The method also includes generating global information from the compressed output of the local encoder using a global stage. The method includes enabling each local decoding stage to generate reconstructed local information by progressively expanding the global information using local decoding stages within a local decoder. The method also includes providing an output of a last local decode stage as a reconstructed latent output. The method also includes filtering local information of the local encoding stages using global gates and the global information. The method also includes filtering an output of the global gates using local gates and reconstructed local information of the local decoding stages.According to still another aspect of the present disclosure, there is provided a non-transitory computer-readable recording medium storing instructions. The instructions, when executed by a computer, enable the computer to perform operations. The operations include enabling each local encoding stage to generate local information by progressively compressing latent representations of mixed speech contributions using local encoding stages within a local encoder. The operations also include generating a compressed output based on an output of a last local encoding stage. Moreover, the operations include generating global information from the compressed output of the local encoder using a global level. The operations further include enabling each local decoding stage to generate reconstructed local information by progressively expanding the global information using local decoding stages within a local decoder. The operations further include providing an output of a last local decode stage as a reconstructed latent output. The operations also include filtering local information of the local encoding stages using global gates and the global information. Moreover, the operations include filtering an output of the global gates using local gates and reconstructed local information of the local decoding stages.As described above, the present disclosure provides a method and apparatus for separating speech contributions using LAGNet, a deep learning-based speech separation model that includes a multi-scale sequence model. Thus, the method and apparatus for separating speech contributions improve the performance of speech separation and at the same time reduce the computational cost of the speech separation model.Moreover, as described above, the present disclosure provides a method and apparatus for separating speech contributions using a multi-scale sequence model. The multi-scale sequence model includes an encoder that progressively compresses a sequence of mixed phrases using a one-dimensional convolution-based local modeling block, a Bottleneck that uses a multiple self-attention-based global modeling block, and a decoder that progressively restores the sequence using the local modeling block. Thus, the method and apparatus for separating speech contributions effectively process local and global information of the sequence.BRIEF DESCRIPTION OF THE DRAWINGSFIG. 1 conceptually illustrates a speech separation apparatus according to an embodiment of the present disclosure. FIG. 2 illustrates a global-local multiscale sequence model according to an embodiment of the present disclosure. Figure 3 illustrates a dual path sequence model. FIG. 4 illustrates a multi-scale sequence model. FIG. 5 illustrates a speech separation model using a global-local multiscale sequence model, according to an embodiment of the present disclosure. FIG. 6 illustrates a global block according to an embodiment of the present disclosure. FIG. 7 illustrates a local block according to an embodiment of the present disclosure. FIG. 8 illustrates a global gate according to an embodiment of the present disclosure. FIG. 9 illustrates a local gate according to an embodiment of the present disclosure. FIG. 10 is a flow diagram illustrating a method for separating speech contributions by a speech separation device according to an embodiment of the present disclosure. FIG. 11 is a block diagram illustrating an example of a computing device according to an embodiment of the present disclosure.DETAILED DESCRIPTIONHereinafter, some embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the accompanying drawings, like reference numerals designate like elements, although the elements are shown in different drawings. Further, in the following description, detailed descriptions of related known components and functions have been omitted for clarity and brevity, provided that they obscure the subject matter of the present disclosure.Moreover, various terms such as "first / first / first", "second / second / second", "A", "B", "a", "b", etc. are used only to distinguish one component from the other, but not to imply or suggest the substances, order or sequence of the components. In this specification, when a part is referred to as having, having, comprising, or the like, a component, it is intended that the part further comprises one or more other components and not that other components are excluded unless expressly stated to the contrary. The terms such as "unit", "module", and the like refer to one or more units for processing at least one function or operation, which may be realized by hardware, software, or a combination thereof.When a component, device, unit, element, or the like of the present disclosure is described as dedicated or for performing an operation, function, or the like, the component, device, or element should be considered herein to be / is configured to perform this purpose or perform this operation or function.The following detailed description, taken in conjunction with the accompanying drawings, is intended to illustrate embodiments of the present disclosure. The detailed description is not intended to represent the only embodiments in which the present disclosure may be practiced.The present disclosure relates, in some embodiments, to speech separation for separating speech contributions of individual speakers from mixed speech contributions. More specifically, the present disclosure provides a method and apparatus for speaker verification for separating speech contributions using the Local and Global Modeling Network (LAGNet), which is a deep learning-based speech separation model that includes a multi-scale sequence model.FIG. 1 conceptually illustrates a speech separation apparatus according to an embodiment of the present disclosure.The speech separation apparatus according to embodiments of the present disclosure encodes mixed phrases to generate (mixed) latent representations, processes the latent representations locally and globally to generate a separation output and individual masks, separates masked latent representations by applying individual masks to the latent representations, and reconstructs the individual phrases based on the masked latent representations. Here, mixed speech contributions refer to speech contributions generated by simultaneous utterances from two or more speakers. The speech separation device comprises, in whole or in part, an audio encoder 110, a speech separation model 120 and an audio decoder 130. The constituent elements included in the voice separation device according to the embodiments of the present disclosure are not necessarily limited to the above specific example. Moreover, the speech separation apparatus may include a trainer (not shown) for training the audio encoder 110, the speech separation model 120, and the audio decoder 130, or may be realized in a form connected to an external trainer.In one embodiment, in the example of Figure 1, two speakers are assumed. The speech contributions resulting from utterances of the two speakers are represented by s 1 and s 2 and the mixed speech x represents the mixture of s 1 and s 2. In general, mixed speech may include utterances from J speakers (where J is a natural number of 2 or greater).In the following description, the speech separation model 120 and the separator are used interchangeably.The audio encoder 110 replaces the STFT and performs a one-dimensional convolution in the time domain to generate a latent representation X from the mixed language.The speech separation model 120 processes the latent representation X locally and globally to generate the separator output Y(Y 1, Y 2). The speech separator generates masks M 1, M 2 by applying an activation function, for example a sigmoid function, to the separator output. The speech separation device separates the masked latent representations S 1rec, S 2_rec by applying an element-by-element product between the latent representation X and the masks.The audio decoder 130 replaces the iSTFT based on the masked latent representations S 1rec, S 2_rec and performs a one-dimensional convolution in the time domain to finally reconstruct the individual speech contributions s 1_rec, s 2_rec.When the audio encoder 110, the speech separation model 120, and the audio decoder 130 are continuously trained, the speech according to the utterance of each speaker may be used as ground truth (GT).Since the audio encoder 110 and the audio decoder 130 perform transformation and inverse transformation between latent representations and speech contributions in the time domain, the speech separation model 120 will be described generally below with respect to speech separation.In order to effectively separate mixed-speech contributions in the time domain, the speech separation model 120 must effectively use and obtain detailed information obtained from the input speech, similar to speech enhancement. To process the detailed information, the speech separation model 120 must effectively process the local information. In addition, unlike speech enhancement, the speech separation model 120 must effectively process the global information inherent to each speaker to solve the permutation on the time axis for each speech source. The speech separation model 120 according to embodiments of the present disclosure effectively separates speech contributions in the time domain by processing an extended sequence by simultaneously combining local modeling and global modeling. Additionally, speech separation model 120 according to embodiments of the present disclosure may require a low computational effort in processing both local and global information.In the following description, the terms "sequence" and "representation" may be used interchangeably. In other words, the representation can comprise temporal information.The speech separation model 120 according to embodiments of the present disclosure uses a global-local multiscale sequence modelling technique as shown in FIG. 2. By using a global-local multiscale sequence modelling technique, the speech separation model 120 according to embodiments of the present disclosure may maintain significantly less computational effort compared to models such as DPRNN or Sepformer that use the dual-path sequence model shown in FIG. 3. In addition, the speech separation model 120 according to embodiments of the present disclosure may effectively utilize global information compared to a model such as SuDoRM-RF using a multi-scale sequence model as illustrated in FIG. 4.As shown in FIG. 2, in one embodiment, the speech separation model 120 includes a local encoder 210 that progressively compresses a sequence using a one-dimensional convolution-based local modeling block, a Bottleneck (e.g., a Bottleneck module) 220 that uses a multiple self attention-based global modeling block, a local decoder 230 that progressively reconstructs the sequence using the local modeling block, and a skip connection. In the example of FIG. 2, F represents the number of reduced features generated from the latent representation X. F features represent a frame. T represents the number of frames (i.e., F-dimensional vector) and represents the time length of the sequence of frames. The local encoder 210 and the local decoder 230 each include R stages, and the bottleck 220 includes a central stage. A detailed structure and an operation of the local encoder 210, the bottle corner 220, and the local decoder 230 according to embodiments will be described later.In the following description, local modeling blocks are used interchangeably with local blocks. In addition, global modeling blocks are used interchangeably with global blocks.Further, in the following description, the terms "sequence" and "frame sequence" may be used interchangeably.As shown in FIG. 3, a separator using a dual path sequence model includes segmentation, intra-chunk, inter-chunk and overlap-add steps to model a long sequence. The step of segmenting divides a long sequence into chunks of fixed length, allowing overlap between the chunks. In the example of FIG. 3, T L represents the size of the chunk obtained by dividing frames, and T G represents the number of chunks. If 50% overlap is allowed, the number of frames may be defined as T=T G T L / 2. The intra-chunk step corresponds to local modeling for the long sequence and the inter-chunk step corresponds to global modeling for the long sequence. The overlap add step generates a separator output based on the sequence processed by the intra-chunk and inter-chunk steps.As illustrated in FIG. 4, a separator using a multi-scale sequence model includes an encoder, a decoder, and a skip connection based on the U-net to reduce the computational cost. Compared to the speech separation model 120 according to embodiments of the present disclosure, the separator using the multi-scale sequence model does not include the bottleck 220 and thus does not effectively model the global information of the sequence.FIG. 5 illustrates a speech separation model using a global-local multiscale sequence model, according to an embodiment of the present disclosure.As described above, the speech separation model 120 according to embodiments of the present disclosure includes a local encoder 210, a bottleck 220, and a local decoder 230. The speech separation model 120 additionally comprises three linear layers 510, 520, 530. As shown in FIG. 5, the speech separation model 120 additionally includes the skip connection, which includes global and local gates.The LAGNet according to embodiments of the present disclosure represents a voice separation device as shown in FIG. 1. Alternatively, the LAGNet may specifically represent the speech separation model 120, as shown in FIG. 5.The latent representation X is generated by the audio encoder 110 and has the dimensions F 0×TF 0. F 0 represents the number of features in the latent representation and T represents the number of frames. In the latent representation X, a frame F contains 0 features. In general, the latent representation X can be generated from the mixed speech based on utterances from J speakers.The linear layer 510 generates a reduced feature F.F features from the input latent representation form a frame. To normalize the output of linear layer 510, linear layer 510 is followed by a layer normalization (LayerNorm) layer. The output of the linear layer 510 has the dimensions F×T. The output of the linear layer 510 is provided to the local encoder 210. The LayerNorm normalizes samples using the mean and variance of the output samples of the layer.The local encoder 210 is a time-contracting encoder that is composed of R (where R is a natural number of 1 or more) local encoding stages (local E stages). Each local coding stage consists of repeated local blocks. Each local coding stage is followed by a down sampling layer. The local encoder 210 progressively compresses the sequence using down sampling. Before the down sampling is performed, each local coding stage captures local context to generate local information.As shown in the example of FIG. 5, the output of the r-th (r=1, 2,..., R) local encoding stage is represented by L r and has the dimensions F×T / 2 r-1. The dimension of the output of the r-th local coding stage is changed to F×T / 2 r by the r-th down sampling layer. The output of the r-th (r=1, 2,..., R-1) down-sampling layer is provided to the (r+1)-th local encoding stage. The output of the last Rth down sampling layer and the output of the rth (r=1, 2,..., R) local encoder stage are provided to the bottleck 220.The bottle corner 220 includes a global level consisting of B repeated global blocks (where B is a natural number greater than or equal to 1). The bottle corner 220 processes the global information based on a sequence compressed by the local encoder 210. The output of the global stage is denoted G R and has the dimensions F×T / 2 r. The output of the global stage is provided to the local decoder 230.The local decoder 230 is a time-expanding decoder and is composed of R local decoding stages (local D stages) corresponding to the local encoding stages. Each local decoding stage consists of repeated local blocks. At the front of each local decoding stage, a local gate with an upsampling layer is arranged. The local decoder 230 reconstructs the sequence stepwise based on repeated local gates and local decoding stages.As shown in the example of FIG. 5, the input of the r-th (r=R, R -1, 1) local gate is represented by G r and has the dimensions F×T / 2 r. The output of the r-th local gate is changed to the dimensions F×T / 2 r-1 by the upsampling layer. The output of the rth local gate is provided to the rth local decode stage. The output of the last local decoding stage, i.e. the reconstructed output of the local decoder 230, is provided to the linear layer 520. The reconstructed output of local decoder 230 has the same dimensions F×T as the input of local encoder 210.A skip connection may be used to progressively expand / restore sequences. The skip connection connects corresponding encoding and decoder stages using global and local gates. The global gate filters local information and the local gate filters the output of the global gate. Like the local gate, the global gate also includes the upsampling layer and helps in the stepwise extension of the sequence. Similar to the local gate, R global gates are used. The global gate is included in the bottleck 220, while the local gate is included in the local decoder 230, as described above.As shown in the example of FIG. 5, the output of the r-th (r=1, 2,..., R) local encoding stage and the output of the global stage are provided to the r-th global gate. The output of the rth global gate is denoted L r_tilde. The output of the rth global gate is provided to the corresponding rth local gate within the local decoder 230.As shown in the example of FIG. 5, the local decoding stage corresponding to the local encoding stage is symmetrically arranged around the bottle corner. A global gate and a local gate (i.e., the corresponding local gate) may form a skip connection for the local encoding stage and the local decoding stage, which correspond to each other. Alternatively, a local gate and a global gate (i.e., the corresponding global gate) may also form a skip connection.The linear layer 520 and the transform layer transform the reconstructed output of the local decoder 230 to generate a distinct representation for each speaker. The output of the transform layer has the dimensions J×F×T, where, as described above, J represents the number of speakers included in the mixed language. The output of the forming layer is provided to the linear layer 530.The linear layer 530 extends the characteristics of the separate representations generated by the previous linear layer 520 and ultimately generates a separator output Y. The separator output Y has the dimensions J×F 0×T. The output of the linear layer 530 is provided to the audio decoder 130.The linear layers 510, 520, 530 may each be layers completely connected to one another.FIG. 6 illustrates a global block according to an embodiment of the present disclosure.As described above, the global level within the bottle corner 220 includes a plurality of global blocks. Global blocks generate global information GR from a sequence compressed by the local encoder 210. A global block includes the transformer's multi-head self-attention (MHSA) module and the feed-forward network (FFN). As shown in FIG. 6, the FFN includes linear layers supporting hidden features of dimension 4F, and further includes a Gaussian Error Linear Unit (GELU) activation function and a dropout layer. Both the MHSA module and the FFN use a residual link. By positioning the LayerScale layer and the LayerNorm layer before and after the MHSA module and the FFN, the training speed and the global level performance can be improved. Element-by-element multiplication is applied between the layer scale layer and the MHSA module and between the layer scale layer and the FFN. The dropout layer has the task of deleting a part of the network to prevent overfitting.FIG. 7 illustrates a local block according to an embodiment of the present disclosure.As described above, the local encoding stage in the local encoder 210 and the local decoding stage in the local decoder 210 include a plurality of local blocks. Local blocks in the r-th local encoding stage generate local information L r, and local blocks in the r-th local decoding stage generate local information G r-1. A local block includes a convolutional local attenuation (CLA) module and the FFN. In other words, the local block uses the CLA module and thus replaces the MHSA module of the global block.The CLA module carefully captures the local context based on a gated linear unit (GLU) that performs dynamic gating of input features. As shown in FIG. 7, the GLU comprises two parallel pointwise convolution (P-Conv) blocks. The GLU applies an activation function (e.g., a sigmoid function) to a P-Conv output and multiplies the output of the activation function by the output of the remaining P-Conv blocks element by element.The CLA module comprises a GLU, a 1-dimensional depthwise convolution (D-Conv1D) block and a P-Conv block. The kernel size of the D-Conv1D block is K. In other words, the D-Conv1D block processes information corresponding to K frames. The D-Conv1D block is followed by a batch normalization (BatchNormal) layer and a GELU. In addition, the CLA module comprises a P-Conv block and a dropout layer on the back side of the GELU. The BatchNorm layer normalizes the samples using the mean and variance of the batch output samples.Both the CLA module and the FFN use a residual compound. By positioning the layer scale layer and the layer norm layer before and after the CLA module and the FFN, the training speed and performance of the local encoder / decoder stages can be improved. Element-by-element multiplication is applied between the layer scale layer and the CLA module and between the layer scale layer and the FFN.The first local encoding stage and the first local decoding stage comprise C 1 local blocks. The remaining local encoding and decoding stages include C 2(< C 1) local blocks.FIG. 8 illustrates a global gate according to an embodiment of the present disclosure.As described above, the global gate exists on a skip connection and is included in the bottle corner 220. The global gate filters the output of the rth local encoding stage on the skip connection using G R, which are the output sequence of the global stage, and then provides the filtered sequence to the corresponding local gate.In the example of FIG. 8, the operation of the r-th global gate will be described. To adapt dimensionality to the output of the corresponding local coding stage, the r-th global gate performs upsampling of G R, of the global stage output sequence using the upsampling layer. For example, the upsampling layer may upsampling the input sequence using nearest neighbor interpolation. The output of the upsampling layer has the dimensions F×T / 2 r-1. The rth global gate generates gate values by applying an activation function, for example a sigmoid function, to the output of the upsampling layer. The rth global gate generates the output of the rth global gate by applying an element-by-element product between the output of the corresponding local encoding stage and gate values.FIG. 9 illustrates a local gate according to an embodiment of the present disclosure.As described above, the local gate is located at the front of the local decode stage within the local decoder 230. The local gate connects the previous local decode stage and the current local decode stage and establishes a skip connection with the corresponding local encode stage. The local gate generates a filtered sequence by filtering the global gate output using the output of the previous local decode stage and completing the skip connection. The local gate generates an output by adding the output of the previous local decode stage and the filtered sequence, and provides the generated output to the next local decode stage. The first local gate uses the global level output. By using the local gate based on the output of the global gate, the local decoder 230 can efficiently reconstruct locally missing information.In the example of FIG. 9, the operation of the r-th local gate will be described. To adapt dimensionality to the output of the corresponding global gate, the r-th local gate performs an upsampling of the input sequence G R. For example, the upsampling layer may perform upsampling of the input sequence using nearest neighbor interpolation. The output of the upsampling layer has the dimension F×T / 2 r-1. The output of the upsampling layer at the rth local gate is applied to two separate D-Conv1D blocks. The r-th local gate generates gate values by applying an activation function to the output from a D-Conv1D layer. The rth local gate completes the skip connection by generating a filtered sequence through an element-by-element product between the output of the corresponding global gate and gate values. Finally, the rth local gate generates the output of the rth local gate by adding the output of the remaining D-Conv1D blocks and the filtered sequence on the skip connection.In the following description, the operations of the Conv1D layer, the P-Conv block, and the D-Conv1D block according to embodiments will be described. The Conv1D layer generates an output in the dimension F 0 × T from input samples. It is assumed that the P-Conv block and the D-Conv1D block have outputs in the dimensions F × T. As described above, F represents the number of features (i.e., channels) and T represents the length of a frame sequence. The direction in which the channel index increases is defined as the channel direction, and the direction in which the index of the frame within the sequence increases is defined as the time direction. Since the Conv1D layer, the P-Conv block, and the D-Conv1D block all perform 1D convolution, their kernels are all defined in the time direction.The Conv1D layer generates features in the time direction by applying the kernel in the time direction and adding the generated results in the channel direction. The Conv1D layer uses F 0 kernels to generate an output in the dimension F 0 ×T. Since channel-by-channel operations are performed, the Conv1D layer may be used to adjust the number of channels between input and output. For example, the Conv1D layer in the audio encoder 110 sets the number of channels between input and output from 1 to F 0.The P-Conv block generates features along the channel direction by applying a weighted sum to the output generated using a kernel of size 1 in the channel direction. The kernel of the P-Conv block has a size of 1 in the time direction and has, for example, weights corresponding to the number of input channels in the channel direction. The P-Conv block uses F kernels to generate an output in the dimension F×T. Because the P-Conv block performs operations in the channel direction without including operations in the time direction, features may be generated that reflect context within the sample. As channel-direction operations are performed, the P-Conv block may be used to adjust the number of channels between input and output. In the present disclosure, the P-Conv block is used to maintain, increase, or decrease the number of channels between input and output.The D-Conv1D block generates features in the time direction by applying a large-sized kernel in the time direction. The D-Conv1D block uses F kernels to generate an output in the dimension F×T. Because the D-Conv1D block performs operations in the time direction without including operations in the channel direction, features may be generated that reflect context between frames. Since no operations are performed in the channel direction, the D-Conv1D block maintains the number of channels between input and output.Referring to the example of Fig. 10, a method of separating speech contributions by a speech separation device will be described below.FIG. 10 is a flow diagram illustrating a method for separating speech contributions by a speech separation device according to an embodiment of the present disclosure.The speech separation device uses local encoding levels within the local encoder to gradually compress the latent representations of the mixed speech, thereby allowing each local encoding level to generate local information in an operation S1000.The speech separator generates latent representations from the mixed speech based on a one-dimensional convolution. The speech separation device may reduce the number of features of the latent representations using the linear layer 510 and the LayerNorm layer. The mixed speech is generated by mixing utterances from multiple speakers.Each local coding stage comprises local blocks. Local blocks generate local information. Each local block includes a CLA (Convolutional Local Attention) module and a feed-forward network (FFN). The CLA module includes a gated linear unit (GLU) and a D-Conv1D block. The CLA module captures local context by dynamically gating input features based on the GLU.Moreover, the speech separation device also comprises down sampling layers. Each down-sampling layer performs down-sampling of the output of each local encoding stage and provides the down-calculated output to the next local encoding stage.In an operation S1002, the speech separation device generates a compressed output based on the output of the last local encoding stage. The last down-sampling layer performs down-sampling of the output of the last local encoding stage to generate a compressed output, and provides the compressed output to the global stage.In an operation S1004, the speech separation device generates global information from the compressed output of the local encoder using a global level.The global level includes global blocks. Global blocks generate global information. Each global block includes a multi-head self attenuation (MHSA) module and a feedforward network.In an operation S1006, the speech separator uses local decoding stages within the local decoder to step-expand the global information, thereby allowing each local decoding stage to generate reconstructed local information.Each local decoding stage comprises local blocks. Local blocks generate reconstructed local information. Each local block comprises a CLA module and a feedforward network.In an operation S1008, the speech separator provides the output to the last local decoding stage as reconstructed latent output.In an operation S1010, the speech separator filters local information from local encoding levels using global gates and global information.The speech separator performs global information upsampling using an upsampling layer within each global gate. The speech separator filters the local information of each local encoding stage using the up-calculated global information, thereby generating a filtered output. The speech separator provides the output of each global gate to the corresponding local gate.In an operation S 1012, the speech separator filters the output of the global gates using the reconstructed local information of the local gates and the local decoding stages.The speech separator uses the upsampling layer in each gate to upsampling the reconstructed local information of the previous local decoding stage of each local decoding stage. The speech separator generates a filtered sequence by filtering the output of the corresponding global gate using the up-converted reconstructed local information. The speech separator generates an output by adding the up-converted reconstructed local information and the filtered sequence, and provides the generated output to each local decoding stage. Each global gate and corresponding local gate connect each local encoding stage and corresponding local decoding stage.The speech separator transforms the reconstructed latent output using the linear layer 520 and the reform layer, thus generating separate representations for each speaker.The speech separation device may increase the number of features of the separated representations using the linear layer 530.The speech separation device reconstructs the speech of each speaker from separated representations and latent representations for each speaker based on a one-dimensional convolution.In the following description, experimental results demonstrating the performance of the speech separation model LAGNet are described.The performance of the speech separation model may be measured based on scale-invariant signal-to-noise ratio (SI-SNRi) enhancement, signal-to-distortion ratio (SDRi), size of the model, computational load of the model, and computational speed of the model. The present disclosure presents simulation results that measure the performance of the above-described elements.As described above, the speech separation apparatus according to the present disclosure may enable trade-off between the performance and the computation speed by flexibly adjusting the number of stages and blocks included in the local encoder 210, the bottle corner 220, and the local decoder 230 in the speech separation model 120. In other words, the speech separation apparatus according to the present disclosure may use a multi-scale sequence model that requires a small amount of computation while maintaining the same level of performance as compared to the latest speech separation models such as DPRNN and Sepformer.As described above, both the local encoder 210 and the local decoder 230 may repeatedly include R local stages. C 1 or C 2 local blocks may be included in each local stage. The global level of the bottleck 220 may repeatedly include B global blocks.In the simulation, a database for separating speech (e.g., WSJ0-2Mix or WHAM!) is used for training and verification. The mixed speech contributions comprising database WSJ0-2Mix are created by mixing utterances of different speakers at a relative signal-to-noise ratio (SNR) between -5 and 5 dB. The mixed speech contributions have a sample value of 8 kHz and 4-second segments are used for training. WHAM! is a database created by adding noise to the mixed phrases of WSJ0-2Mix.In the LAGNet according to the embodiments of the present disclosure, the kernel size and the step size of the audio encoder 110 and the audio decoder 130 for the one-dimensional convolution are set to 4 ms and 1 ms, respectively. In other words, for overlapping strides, each frame includes features corresponding to a length of 4 ms. The number F 0 of filters (i.e., features) used by the audio encoder 110 is set to 256. The number F of features reduced by the linear layer 510 is set to 96. Thus, a frame F includes features corresponding to a length of 4 ms. The number R of local stages within the local encoder 210 and local stages within the local decoder 230 is set to 4. The number of MHSA heads in the global block is set to 8. In the CLA module of the local block, the kernel size K of the D-Conv1D block is set to 65. In other words, the D-Conv1D block in the CLA module processes information of 65 frames.Table 1 shows the LAGNet with various configurations and the corresponding performance metrics. [Table 1] Table 1] [Table 1] Table 1]18(4, 2)(4, 2)3.47.2119.42216(2, 1)(2, 1)3.44.9918.35328(0, 0)(0, 0)3.53.4714.814(6, 3)(6, 3)3.69.5215.4658(8, 4)(0, 0)3.47.2118.2168(0, 0)(8, 4)3.47.2116.7878(6, 3)(2, 1)3.47.2118.7588(2, 1)(6, 3)3.47.2119.37In Table 1, the system represents the LAGNet having various configurations depending on the combination of B and (C 1, C 2). The number of parameters indicates the size of the model, and the multiply and accumulate (MACs) indicates the computational effort required for the model. As shown in Table 1, by adjusting B and (C 1, C 2) different LAGetzes can be constructed to have similar sizes or similar computational amounts. In Table 1, the performance of the speech separation is measured from the SI SNRi.According to Table 1, the configuration for the LAGNet can be derived with the optimal performance, the optimal size, and the optimal computational cost. As compared to the reference configuration, System 1, Systems 2 to 4 comprise a different number of global blocks, local blocks in the encoding stage and local blocks in the decoding stage. The performance degradation observed in systems 2 to 4 indicates the need for the individual components. The systems 5 to 8 comprise a different number of local blocks in the encoding stage and local blocks in the decoding stage, while maintaining the same number of global blocks. Since system 8 has better performance than system 7, it can be determined that the decoder's temporal reconstruction capability is of greater importance in the multiscale sequence model.Tables 2 and 3 show simulation results for the LAGNet and the latest speech separation models. [Table 2] [Table 2]Conv-TasNet (2019)5.110.515.315.612.7-DPRNN (2020)2.688.518.819.013.714.1SuDoRM-RF (2020)6.410.118.9-13.714.1Sepformer (2021)26.086.920.420.514.415.0LAGNet3.47.219.619.815.215.4LAGNet-Large3.813.420.120.315.515.7[Table 3][Table 3]Conv-TasNet (2019)5.110.214.10.43DPRNN (2020)2.685.559.93.52SuDoRM-RF (2020)6.410.043.31.05Sepformer (2021)26.7115.537.14.17LAGNet3.47.216.70.49LAGNet-Large3.813.420.61.02In Tables 2 and 3, the performance of speech separation is evaluated using SI SNRi and SDRi. Higher SI SNRi and SDRi values indicate better performance in speech separation. The real time factor (RTF) representing the computational speed of the model is calculated by dividing the processing time of the model by the duration of the input. The duration and processing time in Tables 2 and 3 are measured in seconds.As shown in Table 2, the LAGNet achieves comparable performance to the latest speech separation models (DPRNN, Sepformer) with respect to SI SNRi, while using less than 10% of the computational resources with respect to MACs. Moreover, the LAGNet may be trained and evaluated in noisy environments (WHAM!) to simultaneously perform noise removal and voice separation.As shown in Table 3, the LAGNet improves the RTF by 72% and 55% in a GPU (Graphic Processing Unit) environment compared to the latest speech separation models (DPRNN, Sepformer). Compared to the latest speech separation models (DPRNN, Sepformer), the LAGNet improves the RTF by more than 86 to 90% in a CPU (Central Processing Unit) environment.As shown in the simulation results of Tables 2 and 3, the LAGNet according to the present disclosure has similar separation performance as compared with the latest speech separation models, while greatly improving performance in terms of reducing the model size, the computational resources required for the model, and the computational speed of the model.In the simulations shown in Tables 2 and 3, the LAGNet-large is a model that increases the computational capacity to improve the performance.The LAGNet according to the present disclosure may be used as a preprocessing module for a voice recognition device using the voice separation.Tables 4 and 5 show the speech recognition performance of the LAGNet based on the pre-processing of the speech separation. Table 4 shows the speech recognition performance based on the pre-processing of speech separation of speech contributions with echoes. Table 5 shows the speech recognition performance based on the preprocessing of speech separation of noisy speech contributions. [Table 4] [Table 4]Input11.811.718.827.235.643.3DPRNN-fast10.610.412.716.620.823.5LAGNet-fast11.110.811.914.517.419.5LAGNet10.910.912.114,417.519.5[Table 5][Table 5]Input1010.51115.518.622.927.4DPRNN-fast8.08.98.710.411.911.9LAGNet-fast6.47.17.48.79.18.9LAGNet6.37.37.17.98.78.7Input59.710.617.322.32832.6DPRNN-fast7.98.29. 910.411.612.4LAGNetz-fast6.37.37.58.19.29.4LAGNet6.56.67.28.08.48.6Input09.610.517.925.231.737.3DPRNN-fast7.78.59.711.612.813.0LAGNet-fast6.37.47.89.110.010.1LAGNet6.47.37.48.28.78.8In Tables 4 and 5, the word error rate (WER) is used as a performance scale representing the speech recognition rate. The overlap rate is given in %. Both 0S and 0L have an overlap rate of 0%, but differ based on whether the mute portion is short or long. The LibriCSS database, which is used for continuous speech separation, allows overlaps between utterances from different speakers.As shown in Tables 4 and 5, the LAGNet achieves a higher recognition rate using less computational resources compared to DPRNN. In other words, for non-overlapping and partially overlapping talk contributions, the LAGNet may serve as a robust preprocessing module for voice separation. In addition, the LAGNet may contribute to improving speech recognition rate by effectively separating the speech sources in the presence of reflections or noise.In the simulations shown in Tables 4 and 5, LAGNet-fast and DPRNN-fast correspond to the models requiring less computational resources to improve the processing speed.FIG. 11 is a block diagram illustrating an example of a computing device according to an embodiment of the present disclosure.A method for separating speech contributions by a deep learning model according to embodiments of the present disclosure may be realized by the computing device 3000 shown in FIG. 11.As shown in FIG. 11, computing device 3000 may include at least a processor 3010, a memory 3020, a network interface 3030, and an input / output interface 3040. A bus 3050 provides a mechanism that allows the components of computing device 3000 to communicate with each other as intended. Although bus 3050 is schematically depicted as a single bus, multiple buses could be used in alternative implementations.Processor 3010 may be configured to process instructions of a computer program by performing basic arithmetic, logical, and input / output operations. Instructions may be provided to the processor 3010 via the memory 300 or the network interface 3030. For example, processor 3010 may be configured to execute instructions stored in accordance with program codes stored in a device such as memory 3020.The memory 3020 is a computer readable storage medium and may include a nonvolatile mass storage device such as a random-access memory (RAM), a read-only memory (ROM), and a hard disk. Here, permanent mass recording devices such as ROM and hard disks may be included in the computing device 3000 as an individual permanent storage device separate from the memory 3020.In addition, the memory 3020 may store an operating system and at least one program code. These software components are separate from the memory 3020 and may be loaded into the memory 3020 from a computer readable recording medium. These separate recording media may include computer readable recording media such as floppy disk drives, hard disks, tapes, DVD / CD-ROM drives, and memory cards. In another embodiment, software components may be loaded into memory 3020 via network interface 3030 rather than via a computer readable recording medium. For example, software components based on a computer program installed through files received via network interface 3030 may be loaded into memory 3020 of computing device 3000.The network interface 3030 may provide functions to allow the computing device 3000 to communicate with other external devices (e.g., servers or terminals external to the computing device 3000) via a wired or wireless communication network. For example, requests, commands, data, or files generated by the processor 3010 of the computing device 3000 according to program codes stored in a recording device such as the memory 3020 may be transmitted to other external devices via a wired or wireless communication network according to the control of the network interface 3030. Conversely, signals, commands, data, or files from other external devices may be received from computing device 3000 via network interface 3030 of computing device 3000 via a wired or wireless communication network. Signals, commands, or data received via the network interface 3030 may be transmitted to the processor 3010 or the memory 3020, while files may be stored in a storage medium that may further include the computing device 3000 (e.g., the separate recording medium described above).The input / output interface 3040 may be a means for interfacing with an input / output device. For example, input devices may include devices such as a microphone, keyboard, or mouse, while output devices may include devices such as displays or speakers. In another example, the input / output interface 3040 may be a means for interfacing with a device that integrates input and output functions, such as a touch screen. The input / output device may be integrated into a single device along with the computing device 3000.In other embodiments, computing device 3000 may also include fewer or more components than those shown in FIG. 11. For example, computing device 3000 may be implemented to include at least a portion of the input / output devices, or may further include other components such as a transceiver and a database.Each component of the apparatus or method according to embodiments of the present disclosure may be realized as hardware or software or a combination of hardware and software. Further, a function of each component may be realized as software, and a microprocessor may also be realized to execute the function of the software corresponding to each component.Although the steps in the respective flowcharts are described as being performed sequentially, the steps merely illustrate the technical idea of some embodiments of the present disclosure. Therefore, a person skilled in the art to which the present disclosure pertains could perform the steps by changing the sequences described in the respective drawings or performing two or more of the steps in parallel. Therefore, the steps in the respective flowcharts are not limited to the illustrated timings.It should be understood that the above description represents illustrative embodiments that may be implemented in various other ways. The functions described in some embodiments may be realized by hardware, software, firmware, and / or a combination thereof. It should also be understood that the functional components described in the present disclosure are labeled "..Unit" to clearly emphasize the possibility of independently realizing them.Meanwhile, various methods or functions described in some embodiments may be realized as instructions stored in a non-transitory recording medium that may be read and executed by one or more processors. The non-transitory recording medium may include, for example, various types of recording devices in which data is stored in a form readable by a computer system. For example, the non-volatile recording medium may include storage media such as an erasable programmable read-only memory (EPROM), a flash drive, an optical drive, a magnetic hard drive, and a solid state drive (SSD), just to name a few.Although embodiments of the present disclosure have been described for illustrative purposes, those skilled in the art to which this disclosure pertains should understand that various changes, additions and substitutions are possible, without departing from the spirit and scope of the present disclosure. Embodiments of the present disclosure have therefore been described for brevity and clarity. The scope of the technical idea of the embodiments of the present disclosure is not limited by the illustrations. Accordingly, it should be understood by those skilled in the art to which the present disclosure pertains that the scope of the present disclosure is not limited by the embodiments expressly described above, but by the claims and their equivalents.

Claims

A speech separation apparatus comprising: a local encoder comprising local encoding stages, each local encoding stage being configured to generate local information by progressively compressing latent representations of mixed speech using the local encoding stages and generating a compressed output based on an output of the last local encoding stage; a bottle corner comprising a global stage, the bottle corner being configured to generate global information from the compressed output of the local encoder using the global stage; a local decoder comprising local decoding stages, each local decoding stage being configured to generate reconstructed local information by incrementally expanding the global information using the local decoding stages and providing an output of the last local decoding stage as the reconstructed latent output; and a skip connection comprising global gates and local gates, wherein the skip connection is configured to filter local information of the local encoding stages using the global gates and the global information, and to filter an output of the global gates using the local gates and the reconstructed local information of the local decoding stages.The speech separation apparatus of claim 1, wherein the local encoder additionally comprises down-sampling layers, each down-sampling layer being configured to perform down-sampling of an output of each of the local encoder stages and provide the down-calculated output to the next local encoder stage.The speech separation device of claim 1, wherein each global gate is configured to: perform an upsampling of the global information using an upsampling layer; filter local information of each of the local encoding stages using the up-calculated global information to generate a filtered output; and provide the filtered output to the corresponding local gate.The speech separation apparatus of claim 3, wherein each local gate is configured to: perform upsampling reconstructed local information of a previous local decoding stage of each of the local decoding stages using an upsampling layer; and generate a filtered sequence by using an output of a corresponding global gate using the up-calculated reconstructed local information.The speech separation apparatus of claim 4, wherein each local gate is configured to: generate an output by adding the up-converted reconstructed local information and the filtered sequence; and provide the generated output to each of the local decoding stages.The speech separation apparatus of claim 5, wherein each global gate and the corresponding local gate connect each of the local encoding stages and the corresponding local decoding stage to each other.The speech separation apparatus of claim 1, wherein: each of the local encoding stages comprises local blocks; the local blocks are configured to generate the local information; and each local block comprises a convolutional local attenuation (CLA) module and a feed forward network (FFN).The speech separation device of claim 7, wherein the CLA module comprises a gated linear unit (GLU) and a D-Conv1D (1-dimensional depthwise convolution) block, and wherein the CLA module is configured to detect local context by dynamically gating input features based on the GLU.The speech separation apparatus of claim 1, wherein: the global level comprises global blocks; the global blocks generate the global information; and each global block comprises an MHSA (Multi-head Self Attention) module and a feed-forward network (FFN).The speech separation apparatus of claim 1, wherein: each of the local decoding stages comprises local blocks; the local blocks are configured to generate the reconstructed local information; and each local block comprises a convolutional local attenuation (CLA) module and a feed-forward network (FFN).The speech separation apparatus of claim 1, further comprising: an audio encoder configured to generate the latent representations from the mixed speech based on 1-D convolution, wherein the mixed speech is generated by mixing utterances from a plurality of speakers; a linear layer configured to transform a reconstructed latent output of the local decoder to generate separate representations for each speaker; and an audio decoder configured to reconstruct the speech from each of the speakers based on the separate representations for each of the speakers and the latent representations based on the 1-D convolution.A method for separating speech contributions performed by a speech separation apparatus, the method comprising: enabling each local encoding stage to generate local information by incrementally compressing latent representations of mixed speech using local encoding stages within a local encoder; generating a compressed output based on an output of the last local encoding stage; generating global information from the compressed output of the local encoder using a global stage; enabling each local decoding stage to generate reconstructed local information by incrementally expanding the global information using local decoding stages within a local decoder; providing an output of the last local decoding stage as the reconstructed latent output; filtering local information of the local encoding stages using global gates and the global information; filtering an output of the global gates using local gates and reconstructed local information of the local decoding stages.The method of claim 12, wherein generating the local information by each local encoding stage comprises: down sampling an output of each local encoding stage using a down sampling layer; and providing the down-calculated output to the next local encoding stage.The method of claim 12, wherein filtering the local information comprises: upsampling the global information using an upsampling layer within each global gate; generating a filtered output by filtering local information of the individual local encoding stages using the up-calculated global information; and providing the filtered output of each of the global gates to the corresponding local gate.The method of claim 14, wherein filtering the output of the global gates comprises: upsampling reconstructed local information of a previous local decoding stage of each local decoding stage using an upsampling layer within each local gate; and generating a filtered sequence by filtering an output of the corresponding global gate using the up-calculated reconstructed local information.The method of claim 15, wherein filtering the output of the global gates comprises: generating an output by adding the up-calculated reconstructed local information and the filtered sequence; and providing the generated output to each local decoding stage.The method of claim 12, further comprising: generating the latent representations from the mixed speech based on 1-D convolution, wherein the mixed speech is generated by mixing utterances from a plurality of speakers; generating separate representations for each speaker by reshaping reconstructed latent output of the local decoder; and reconstructing the speech of each of the speakers based on the separate representations for each of the speakers and the latent representations based on 1-D convolution.A non-transitory computer readable recording medium storing instructions, the instructions, when executed by a processor, cause the processor to: enable each local encoding stage to generate local information by incrementally compressing latent representations of mixed speech using local encoding stages within a local encoder; generate a compressed output based on an output of the last local encoding stage; generate global information from the compressed output of the local encoder using a global stage; enable each local decoding stage to generate reconstructed local information by incrementally expanding the global information using local decoding stages within a local decoder; provide an output of a last local decoding stage as the reconstructed latent output; filter local information of the local encoding stages using global gates and the global information; filtering an output of the global gates using local gates and reconstructed local information of the local decoding stages.