A method, apparatus, device, and storage medium for processing voice data

By using in-block and inter-block processing units in the speech separation method to extract and fuse the speech blocking features to generate speech prediction features, the problem of low signal-to-noise ratio of speech separation in the prior art is solved, and a better speech separation effect is achieved.

CN114913870BActive Publication Date: 2025-06-10AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210511095.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2025-06-10
Estimated Expiration
2042-05-11

AI Technical Summary

Technical Problem

The existing time domain-based speech separation method cannot effectively model the source speaker's voice, resulting in a low signal-to-noise ratio and fail to achieve good speech separation effect.

Method used

The first dimension features and the second dimension features of the speech blocking features are extracted by the in-block processing unit and the inter-block processing unit of the separator, and fuse these features to generate the speech prediction features. Then, the speech separation result is determined based on the predicted characteristics and the speech characteristics to be separated by the waveform reconstruction.

Benefits of technology

The speech separation signal-to-noise ratio is improved, the model parameters are reduced, and the speech separation effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913870B_ABST
    Figure CN114913870B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device, and storage medium for processing voice data. The method includes: obtaining voice data to be separated, and performing feature extraction on the voice data to be separated to obtain voice features to be separated; segmenting the voice features to be separated according to a preset block length and a preset step length to obtain at least two voice block features; extracting first-dimensional features of each voice block feature through an intra-block processing unit; and extracting second-dimensional features of each voice block feature through an inter-block processing unit; fusing the first-dimensional features and the second-dimensional features of each voice block feature to obtain at least two voice prediction features; and determining each voice separation result according to each voice prediction feature and the voice features to be separated. This technical solution solves the problem of low signal-to-noise ratio in voice separation of time-domain-based separation methods, can improve the signal-to-noise ratio while reducing model parameters, and thus achieve a good voice separation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular, to a method, apparatus, device, and storage medium for processing speech data. Background Art

[0002] With the rise of deep learning, speech separation technology based on deep models, such as Convolutional Neural Network (CNN), has gradually become a trend due to its good speech separation effect.

[0003] Currently, multi-speaker speech separation technology is mainly divided into two categories: frequency-domain based separation methods and time-domain based separation methods. Among them, the time-domain based separation method has more prominent performance. For example, the Conv-TasNet (Convolutional time-domain audio separation network) method, the DPRNN (Dual-path recurrent neural

[0004] network, Dual-path recurrent neural network) method, etc.

[0005] However, in the prior art, the time-domain based separation method cannot effectively model the source speaker's speech, has a low signal-to-noise ratio, and cannot achieve a good speech separation effect. Summary of the Invention

[0006] The present invention provides a method, apparatus, device, and storage medium for processing speech data to solve the problem of low signal-to-noise ratio in speech separation of the time-domain based separation method, and can reduce the model parameters while improving the signal-to-noise ratio, thereby achieving a good speech separation effect.

[0007] According to one aspect of the present invention, there is provided a method for processing speech data, the method comprising:

[0008] Obtaining the speech data to be separated, and extracting features of the speech data to be separated through a feature extractor to obtain speech features to be separated;

[0009] Segmenting the speech features to be separated by a separator according to a preset block length and a preset step length to obtain at least two speech block features;

[0010] Extracting first-dimensional features of each speech block feature through an intra-block processing unit of the separator; and extracting second-dimensional features of each speech block feature through an inter-block processing unit of the separator;

[0011] Fusing the first - dimension feature and the second - dimension feature of each voice segment feature through the separator to obtain at least two voice prediction features;

[0012] Through a waveform reconstructor, determining each voice separation result according to each voice prediction feature and the voice feature to be separated.

[0013] According to another aspect of the present invention, there is provided a processing device for voice data, the device comprising:

[0014] A voice feature to be separated generation module, configured to obtain voice data to be separated, and extract features from the voice data to be separated through a feature extractor to obtain a voice feature to be separated;

[0015] A voice segment feature generation module, configured to segment the voice feature to be separated through a separator according to a preset block length and a preset step length to obtain at least two voice segment features;

[0016] A two - dimension feature extraction module, configured to extract the first - dimension feature of each voice segment feature through the intra - block processing unit of the separator; and extract the second - dimension feature of each voice segment feature through the inter - block processing unit of the separator;

[0017] A voice prediction feature generation module, configured to fuse the first - dimension feature and the second - dimension feature of each voice segment feature through the separator to obtain at least two voice prediction features;

[0018] A voice separation result determination module, configured to determine each voice separation result through a waveform reconstructor according to each voice prediction feature and the voice feature to be separated.

[0019] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0020] At least one processor; and

[0021] A memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the voice data processing method according to any embodiment of the present invention.

[0023] According to another aspect of the present invention, there is provided a computer - readable storage medium, the computer - readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the voice data processing method according to any embodiment of the present invention when executed.

[0024] In the technical solution of the embodiment of the present invention, the in-block processing unit and the inter-block processing unit of the separator respectively extract the first-dimensional features and the second-dimensional features of each voice block feature, and fuse the first-dimensional features and the second-dimensional features of each voice block feature to obtain at least two voice prediction features. Through the waveform reconstructor, according to each voice prediction feature and the voice feature to be separated, each voice separation result is determined. This solution solves the problem of low signal-to-noise ratio in voice separation of the time-domain-based separation method, and can improve the signal-to-noise ratio while reducing the model parameters, thereby achieving a good voice separation effect.

[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1A is a flowchart of a method for processing voice data according to Embodiment 1 of the present invention;

[0028] Figure 1B is a schematic structural diagram of a separation model according to an embodiment of the present invention;

[0029] Figure 2A is a flowchart of a method for processing voice data according to Embodiment 2 of the present invention;

[0030] Figure 2B is a schematic structural diagram of a separator according to Embodiment 2 of the present invention;

[0031] Figure 2C is a schematic structural diagram of an in-block processing unit according to Embodiment 2 of the present invention;

[0032] Figure 2D is a schematic structural diagram of a temporal convolutional network according to Embodiment 2 of the present invention;

[0033] Figure 2E is a schematic structural diagram of an inter-block processing unit according to Embodiment 2 of the present invention;

[0034] Figure 2F is a schematic structural diagram of a forward adaptive sub-unit according to Embodiment 2 of the present invention;

[0035] Figure 2G It is a schematic structural diagram of a separation enhancement unit provided according to Embodiment 2 of the present invention;

[0036] Figure 3 It is a schematic structural diagram of a voice data processing device provided according to Embodiment 3 of the present invention;

[0037] Figure 4 It is a schematic structural diagram of an electronic device for implementing the voice data processing method of the embodiments of the present invention. Detailed implementation manners

[0038] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0039] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solutions of this application all comply with the relevant regulations of national laws and regulations.

[0040] Embodiment 1

[0041] Figure 1A This is a flowchart of a voice data processing method provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation of voice data processing, especially the separation scenario of mixed voice data. This method can be executed by a voice data processing device, and the device can be implemented in the form of hardware and / or software, and the device can be configured in an electronic device. As Figure 1A shown, the method includes:

[0042] S110. Obtain the voice data to be separated, and perform feature extraction on the voice data to be separated through a feature extractor to obtain the voice features to be separated.

[0043] This solution can be executed by electronic devices such as computers and servers. The voice data to be separated can be a mixed voice signal. The voice data to be separated can be pre-input into the electronic device or obtained by the electronic device through a data acquisition terminal. A separation model for processing the voice data to be separated can be deployed in the electronic device, and the separation model can be obtained through pre-training. Figure 1B It is a schematic diagram of the separation model structure provided by an embodiment of the present invention, as Figure 1B shown, the separation model can include a feature extractor, a separator, and a waveform reconstructor. Among them, the feature extractor can be used to extract the features of the mixed voice signal. The separator can separate the mixed voice features into component voice signal features, and then reconstruct the component voice signal features into each voice separation signal through the waveform reconstructor.

[0044] After obtaining the voice data to be separated, the electronic device can extract the features of the voice data to be separated through the feature extractor of the separation model to obtain the voice features to be separated. Among them, the feature extractor can include structures such as a convolutional layer, an activation function layer, and a pooling layer. It should be noted that voice data is usually observed from the perspective of the relationship between vibration amplitude and time. Voice data can be a one-dimensional signal of amplitude with respect to time. Therefore, the convolutional layer in the feature extractor can use one-dimensional convolution to extract features from the voice data to be separated. The activation function of the feature extractor can be any one of activation functions such as Sigmoid, Tanh, Relu, and Softplus. The pooling layer can be a pooling method such as max pooling or average pooling.

[0045] S120: Segment the voice features to be separated by the separator according to a preset block length and a preset step length to obtain at least two voice block features.

[0046] After obtaining the voice features to be separated, the separator can perform processing such as normalization and dimensionality reduction on the voice features to be separated to accelerate the separation process. The voice features to be separated can be one-dimensional voice features with a certain length. The separator can segment the one-dimensional voice features according to a preset block length and a preset step length to obtain multiple voice block features, that is, transform the voice features to be separated from one-dimensional to two-dimensional. For example, assume that the length of the voice features to be separated is L, and L is divided into S voice block features with a block length of K and a step length of K / 2. It should be noted that when at the end of the segmentation process, if the remaining length of the voice features to be separated is less than the block length, the separator can fill the voice block features in a default filling manner, such as filling the insufficient part with 0.

[0047] S130. Extract the first - dimensional features of each speech segment feature through the intra - block processing unit of the separator; and extract the second - dimensional features of each speech segment feature through the inter - block processing unit of the separator.

[0048] Taking the segmentation method described in S120 as an example, the speech features to be separated can be understood as being transformed from a one - dimensional feature of length L to a feature of length K and width S. The intra - block processing unit of the separator can extract the intra - block features of each speech segment feature, that is, the K - dimensional features of the speech features to be separated, and the inter - block processing unit of the separator can extract the correlation features between the speech segment features, that is, the S - dimensional features of the speech features to be separated. Specifically, each speech segment feature can correspond to the speech features of a time period. The first - dimensional features can represent the dependency relationship between the key time points within each time period, and the second - dimensional features can represent the dependency relationship between the speech features of each time period.

[0049] S140. Through the separator, fuse the first - dimensional features and the second - dimensional features of each speech segment feature to obtain at least two speech prediction features.

[0050] After obtaining the features of the two dimensions, the separator can fuse the information of the two aspects of each speech segment feature to obtain the multi - dimensional features of each speech segment feature. After the separator extracts the multi - dimensional features of each speech segment, it can merge each speech segment to obtain the multi - dimensional speech features to be separated with the length of the speech features to be separated, and then further separate according to the multi - dimensional speech features to be separated. For example, split each speaker feature through an additive gating mechanism to obtain multiple speech prediction features.

[0051] S150. Through the waveform reconstructor, determine each speech separation result according to each speech prediction feature and the speech features to be separated.

[0052] It is easy to understand that the waveform reconstructor can have the opposite function to the feature extractor, and is used to restore each speech prediction feature to the speech signals of each speaker, that is, each speech separation result. The waveform reconstructor can include structures such as de - convolution and de - pooling. It should be noted that, in order to ensure the consistency of restoration, the parameter settings in the waveform reconstructor are usually the same as those in the feature extractor, such as the convolution kernel size, convolution step, etc.

[0053] In this technical solution, the in-block processing unit and the inter-block processing unit of the separator respectively extract the first-dimensional features and the second-dimensional features of each speech segment feature, and fuse the first-dimensional features and the second-dimensional features of each speech segment feature to obtain at least two speech prediction features. Through the waveform reconstructor, according to each speech prediction feature and the speech feature to be separated, each speech separation result is determined. This solution solves the problem of low signal-to-noise ratio in speech separation of time-domain-based separation methods, and can improve the signal-to-noise ratio while reducing the model parameters, thereby achieving a good speech separation effect.

[0054] Embodiment 2

[0055] Figure 2A FIG. is a flowchart of a method for processing speech data provided in Embodiment 2 of the present invention. This embodiment is refined based on the above embodiment. As Figure 2A shown, the method includes:

[0056] S210. Obtain the speech data to be separated, and extract features from the speech data to be separated through a feature extractor to obtain the speech feature to be separated.

[0057] This embodiment is illustrated with a specific separation model. As Figure 1B shown, the speech data to be separated can be represented by x. As a one-dimensional speech signal, the value range of x can be represented as Taking x as the input of the feature extractor, the number of channels of x is increased through one-dimensional convolution, and the input sequence is mapped to a higher-dimensional feature space. Then, the ReLU activation function can be used to activate the up-dimensioned x to obtain the speech feature e to be separated.

[0058] S220. Through the separator, segment the speech feature to be separated according to a preset block length and a preset step length to obtain at least two speech segment features.

[0059] Figure 2B FIG. is a schematic structural diagram of a separator provided in Embodiment 2 of the present invention. As Figure 2B shown, the output e after feature extraction can first be normalized to perform a linear transformation on the original data to scale its value to the range of 0-1, accelerate the gradient descent and the speed of finding the optimal solution, and at the same time, a one-dimensional convolution (Conv1d) can be used to reduce the channel dimension. Then, perform segmentation processing to obtain the feature e'. Assume that the length of the speech feature to be separated is L, and divide L into S speech segment features with a block length of K and a step length of K / 2. When the final length is less than K, padding is used for filling.

[0060] The above solution can effectively reduce the computational amount of the attention matrix in the self-attention mechanism, so as to reduce the computational amount from the original O(L 2) reduced to O(K 2 ) + O(S 2 ).

[0061] S230. Input each voice chunk feature into the intra-block processing unit, and extract the intra-block dependency feature of each voice chunk feature based on a preset block length through the intra-block transformation subunit.

[0062] As Figure 2B shown, each voice chunk feature can be represented by e′, and e′ is input into the processing unit of the separator. Among them, the processing unit includes an intra-block processing unit and an inter-block processing unit, and the connection relationship between the intra-block processing unit and the inter-block processing unit can be a dual-path architecture as Figure 2B shown. It is easy to understand that the voice features to be separated are segmented according to a preset block length and a preset step length. Therefore, the intra-block dependency feature of each voice chunk feature is extracted based on the preset block length. The intra-block processing unit processes the K dimension of e′ and extracts the dependency relationship inside each voice chunk feature, that is, the intra-block dependency feature z.

[0063] The intra-block dependency feature of the voice chunk feature is mainly extracted through the intra-block transformation subunit in the intra-block processing unit. Figure 2C is a schematic structural diagram of an intra-block processing unit provided according to Embodiment 2 of the present invention. As Figure 2C shown, the intra-block transformation subunit may include one or more Transformer structures. The Transformer structure is a model that uses an attention mechanism to improve the model training speed. Compared with transformation models such as recurrent neural networks, the Transformer structure is more suitable for parallel computing, and the complexity of the Transformer model also makes it far superior to recurrent neural networks in terms of accuracy and performance. It should be noted that the Transformer structure in the intra-block transformation subunit can be called an intra-block Transformer. As Figure 2C shown, its feed-forward network layer may include a Temporal Convolutional Networks (TCN). Figure 2D is a schematic structural diagram of a temporal convolutional network provided according to Embodiment 2 of the present invention. The temporal convolutional network may include a depth convolution and a dilated convolution structure, such as Depth wise and Dilated Conv1d, for widening the receptive field and facilitating the construction of long-term memory.

[0064] S240. Add each voice chunk feature and the corresponding intra-block dependency feature as input data and input it into the inter-block processing unit, and extract the inter-block dependency feature of each voice chunk feature based on a preset block length and a preset step length.

[0065] AsFigure 2B For the dual-path architecture shown, after obtaining the intra-block dependent features, the separator needs to add the speech chunk feature e' and its corresponding intra-block dependent feature z, and use the result as input data to the inter-block processing unit to facilitate the fusion of intra-block dependent features and inter-block dependent features. The inter-block dependent features of each speech chunk feature are extracted based on a preset chunk length and a preset stride. The inter-block processing unit processes the S dimension of e', and extracts the dependency relationship between speech chunk features, that is, the inter-block dependent feature z".

[0066] Similarly, the intra-block dependent features of speech chunk features are mainly extracted by the inter-block transformation sub-unit in the inter-block processing unit. Figure 2E It is a schematic structural diagram of an inter-block processing unit provided according to Embodiment 2 of the present invention. As Figure 2E shown, the inter-block transformation sub-unit may include a bidirectional long short-term memory network. The bidirectional long short-term memory network can add to and forget previous input information through internal gate structures, such as a forget gate, an update gate, and an output gate, etc., so as to achieve the purpose of using both past information and future information to solve current problems. The gate structures in the bidirectional long short-term memory network can be regarded as a way of extracting useful information.

[0067] In this solution, the intra-block processing unit further includes a forward adaptation sub-unit; the forward adaptation sub-unit is arranged before the intra-block transformation sub-unit;

[0068] Before extracting the first dimension features of each speech chunk feature through the intra-block processing unit, the method further includes:

[0069] Inputting each speech chunk feature into the forward adaptation sub-unit, and optimizing each speech chunk feature through the forward adaptation sub-unit.

[0070] Figure 2F It is a schematic structural diagram of a forward adaptation sub-unit provided according to Embodiment 2 of the present invention. As Figure 2F shown, the forward adaptation sub-unit may include structures such as a multilayer perceptron (MLP), depthwise convolution (Dw-Conv2d), and depthwise group convolution (Dw-G-Conv2d). Referring to Figure 2F , the optimization process of the forward adaptation sub-unit for each speech chunk feature can be expressed by the following formula:

[0071] e1' = MLP(e') * e';

[0072] e2′ = Norm(DwConv2d(PReLU(Norm(DwConv2d(e1′)))))*e1′;

[0073] e′_out = (MLP(Norm(DwConv2d(e2′)) N ;

[0074] where * represents element-wise multiplication, and Norm is the abbreviation of the normalization function (Normalization).

[0075] The forward adaptive sub-unit in this scheme can further enhance the perception ability of the speech chunk features in the channel and spatial dimensions, thereby improving the performance of the speech separation network.

[0076] It should be noted that the inter-block processing unit may also include a Figure 2F forward adaptive sub-unit with the same structure as shown. Similarly, the forward adaptive sub-unit can be set before the inter-block transformation sub-unit to simultaneously increase the channel adaptability and spatial adaptability of the separation network.

[0077] Taking the output e′_out of the forward network as the input, the intra-block Transformer structure is reused M (M>0) times to deeply extract the dependency information inside the speech chunk features, and finally layer normalization is applied to obtain the local feature z. The following formula:

[0078] z = Norm(IntraT(Forward Adaptability(e′)+e′) M )

[0079] where the function IntraT represents the transformation operation of the intra-block Transformer structure, and the Forward Adaptability function represents the optimization operation of the forward adaptive sub-unit.

[0080] The specific calculation formula of IntraT is as follows:

[0081] IntraT = Dropout(FFNintra(Norm(MHA(Norm(·))+·)));

[0082] where (·) represents the input information of each function, MHA represents the multi-head attention mechanism (Multi-Head Attention); the FFNintra function represents the extraction operation of the intra-block feed-forward network layer (Feed Forward Network, FFN); the Dropout function represents the operation of temporarily discarding features from the network with a certain probability.

[0083] The specific calculation formula of FFNintra is as follows:

[0084] FFNintra = TCN(ReLU(Conv1d(·))).

[0085] The input of the inter-block processing unit can be represented by z′, and it can be understood that z′ = z + e′.

[0086] The output of the inter-block processing unit can be expressed by the following formula:

[0087] z″ = Norm(InterT(Forward Adaptability(z′) + z′) M );

[0088] where InterT = Dropout(FFNinter(Norm(MHA(Norm(·)) + ·)));

[0089] FFNinter = BiLSTM(ReLU(Conv1d(·))); The FFNinter function represents the extraction operation of the inter-block feed-forward network layer.

[0090] Take e″ = z + z″ as the output of the entire processing module and input it into the subsequent network structure of the processing unit as shown in Figure 2B the following.

[0091] S250. Add the intra-block dependent features and the matching inter-block dependent features of each block to obtain the fused features of each block.

[0092] It can be understood that in order to fuse the two aspects of the speech chunk features, the separator needs to perform an addition operation on the intra-block dependent features and the matching inter-block dependent features of each block to obtain the fused features. Specifically, as shown in Figure 2B the following, the calculation formula for obtaining the fused features through the processing unit can be expressed as:

[0093] e″ = finter(P(fintra(e′) + e′)) + z;

[0094] where the processing process of the intra-block processing unit can be represented by the function finter(·), and the processing process of the inter-block processing unit is represented by the function finter(·). P is used to represent the dimension conversion function permute, which is used to exchange the two dimensions of the speech chunk features.

[0095] S260. Input each fused feature into the superposition unit to obtain the combined feature.

[0096] Based on the above solution, after the e″ output by the processing unit undergoes PReLU activation processing, e″′ is obtained through a two-dimensional convolution, and then an overlap-and-add operation is performed to merge the S voice chunk features with a length of K back to the feature length before segmentation, thereby obtaining the merged feature e″″. The specific calculation process is as follows:

[0097] e″′ = Conv2d(PReLU(e″));

[0098] e″″ = OverlapAdd(e″′).

[0099] S270. Input the merged feature into the gating unit to obtain at least two voice prediction features.

[0100] Specifically, the gating unit can adopt an additive gating mechanism, as shown in the network structure of the additive gating mechanism in Figure 2B . The specific calculation method is as follows:

[0101] e″″ = Conv1d(Tanh(Conv1d(e″″)) + Sigmoid(Conv1d(e″″)));

[0102] The additive gating mechanism can split e″″ into at least two voice prediction features, such as N speaker features spk1, spk2,..., spkN. As shown in the following formula:

[0103] (spk1, spk2,..., spkN) = (e″″[0], e″″[1],..., e″″[N]).

[0104] S280. Concatenate the at least two voice prediction features, input the concatenated feature into each branch of the separation enhancement unit, and enhance each voice prediction feature based on the multiplication masks output by each branch.

[0105] Figure 2G FIG.

[0106] is a schematic structural diagram of a separation enhancement unit according to Embodiment 2 of the present invention. The separation enhancement unit includes at least two branches, where the number of branches in the separation enhancement unit is the same as the number of voice prediction features. In fact, each branch of the separation enhancement unit can correspond to each voice prediction feature, such as each speaker feature. N (N = 1, 2,...).

[0107] Taking the separation of the features of two speakers as an example, first, the features of the two speakers to be separated are concatenated. Then, the concatenated features are input into a two-dimensional depth convolution to obtain a multiplicative mask Mask, which is used to enhance the speech features belonging to the same speaker and suppress the speech features not belonging to the same speaker. As shown in the following formula:

[0108] (spk1, spk2) = (spk1 * Mask(spk1, spk2), spk2 * Mask(spk1, spk2));

[0109] Among them, Mask(·) represents a set of operations such as concatenation, two-dimensional depth convolution (Dw-Conv2d), normalization, and activation (Sigmoid). * represents element-wise multiplication.

[0110] Stack spk1 and spk2 together to form the final output e N , N represents the Nth speaker, that is, as shown in the following formula:

[0111] (e 1 , e 2 ) = (spk1, spk2).

[0112] S290. Through the waveform reconstructor, according to each speech prediction feature and the speech features to be separated, determine each speech separation result.

[0113] The waveform reconstructor can reconstruct the waveform from the speech prediction features using a one-dimensional transposed convolution with learnable parameters. The convolution stride and convolution kernel size in the waveform reconstructor can be consistent with the feature extractor, and the length of the speech data to be separated before inputting into the feature extractor is restored. The input of the waveform reconstructor is the output e of the separator N Multiplied by the e output by the feature extractor, as shown in the following formula:

[0114] Among them, can represent the speech signal of the Nth separated speaker.

[0115] This solution effectively models the channel features of speech data using a forward adaptive sub-unit, effectively models the spatial dimension of the speech sequence block using depth convolution and depth group convolution, and fully utilizes the speech correlation obtained from the local and global features of the speech data to be separated. At the same time, by designing a separation enhancement unit, this solution can further improve the separation performance of the separation model. It provides an effective solution for speech pickup and speech recognition in multi-speaker scenarios.

[0116] Example 3

[0117] Figure 3 This is a schematic structural diagram of a voice data processing device provided in Embodiment 3 of the present invention. As Figure 3 shown, the device includes:

[0118] A voice feature to be separated generation module 310, configured to obtain voice data to be separated, and perform feature extraction on the voice data to be separated through a feature extractor to obtain voice features to be separated;

[0119] A voice block feature generation module 320, configured to segment the voice features to be separated through a separator according to a preset block length and a preset step length to obtain at least two voice block features;

[0120] A two-dimensional feature extraction module 330, configured to extract first-dimensional features of each voice block feature through an intra-block processing unit of the separator; and, extract second-dimensional features of each voice block feature through an inter-block processing unit of the separator;

[0121] A voice prediction feature generation module 340, configured to fuse the first-dimensional features and the second-dimensional features of each voice block feature through the separator to obtain at least two voice prediction features;

[0122] A voice separation result determination module 350, configured to determine each voice separation result through a waveform reconstructor according to each voice prediction feature and the voice features to be separated.

[0123] In this solution, optionally, the intra-block processing unit includes an intra-block transformation subunit; the first-dimensional feature is an intra-block dependence feature;

[0124] Correspondingly, the two-dimensional feature extraction module 330 is specifically configured to:

[0125] Input each voice block feature into the intra-block processing unit, and extract the intra-block dependence feature of each voice block feature through the intra-block transformation subunit based on the preset block length.

[0126] On the basis of the above solution, optionally, the inter-block processing unit includes an inter-block transformation subunit; the second-dimensional feature is an inter-block dependence feature;

[0127] Correspondingly, the two-dimensional feature extraction module 330 is specifically configured to:

[0128] Add each voice block feature and the corresponding intra-block dependence feature as input data and input the same into the inter-block processing unit, and extract the inter-block dependence feature of each voice block feature through the inter-block transformation subunit based on the preset block length and the preset step length.

[0129] In a feasible solution, optionally, the separator further includes a stacking unit and a gating unit;

[0130] Correspondingly, the voice prediction feature generation module 340 is specifically configured to:

[0131] Add the intra-block dependent features and the matching inter-block dependent features to obtain respective fusion features;

[0132] Input the respective fusion features into the stacking unit to obtain combined features;

[0133] Input the combined features into the gating unit to obtain at least two voice prediction features.

[0134] Based on the above solution, optionally, the separator further includes a separation enhancement unit; the separation enhancement unit includes at least two branches, and the number of branches in the separation enhancement unit is the same as the number of voice prediction features;

[0135] The device further includes:

[0136] A voice prediction feature enhancement module, configured to splice the at least two voice prediction features, input the spliced features into each branch of the separation enhancement unit, and enhance each voice prediction feature based on the multiplication masks output by each branch.

[0137] In a preferred solution, the intra-block transformation sub-unit includes a temporal convolutional network; the inter-block transformation sub-unit includes a bidirectional long short-term memory network.

[0138] In this solution, optionally, the intra-block processing unit further includes a forward adaptation sub-unit; the forward adaptation sub-unit is arranged before the intra-block transformation sub-unit;

[0139] The device further includes:

[0140] A voice block feature optimization module, configured to input each voice block feature into the forward adaptation sub-unit to optimize each voice block feature through the forward adaptation sub-unit.

[0141] The voice data processing device provided by the embodiments of the present invention can execute the voice data processing method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0142] Embodiment Four

[0143] Figure 4FIG. 0 shows a schematic structural diagram of an electronic device 410 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0144] As Figure 4 shown, the electronic device 410 includes at least one processor 411, and a memory communicatively connected to the at least one processor 411, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc. The memory stores a computer program executable by the at least one processor. The processor 411 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 412 or the computer program loaded from the storage unit 418 into the random access memory (RAM) 413. In the RAM 413, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 411, the ROM 412, and the RAM 413 are connected to each other via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.

[0145] Multiple components in the electronic device 410 are connected to the I / O interface 415, including: an input unit 416, such as a keyboard, a mouse, etc.; an output unit 417, such as various types of displays, speakers, etc.; a storage unit 418, such as a magnetic disk, an optical disc, etc.; and a communication unit 419, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 419 allows the electronic device 410 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0146] The processor 411 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 411 executes the various methods and processes described above, such as the method for processing voice data.

[0147] In some embodiments, the method for processing voice data may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 410 via the ROM 412 and / or the communication unit 419. When the computer program is loaded into the RAM 413 and executed by the processor 411, one or more steps of the method for processing voice data described above may be performed. Alternatively, in other embodiments, the processor 411 may be configured to execute the method for processing voice data by any other suitable means (e.g., by means of firmware).

[0148] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0149] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on a remote machine or server.

[0150] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0152] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0153] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0154] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0155] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for processing voice data, characterized in that, the method includes: obtaining voice data to be separated, and extracting features of the voice data to be separated through a feature extractor to obtain voice features to be separated; segmenting the voice features to be separated by a separator according to a preset block length and a preset step length to obtain at least two voice block features; extracting first-dimensional features of each voice block feature through an intra-block processing unit of the separator; and extracting second-dimensional features of each voice block feature through an inter-block processing unit of the separator; fusing the first-dimensional features and the second-dimensional features of each voice block feature through the separator to obtain at least two voice prediction features; determining each voice separation result through a waveform reconstructor according to each voice prediction feature and the voice features to be separated; wherein, the intra-block processing unit includes an intra-block transformation subunit; the first-dimensional feature is an intra-block dependence feature; correspondingly, the extracting first-dimensional features of each voice block feature through the intra-block processing unit of the separator includes: inputting each voice block feature into the intra-block processing unit, and extracting the intra-block dependence feature of each voice block feature based on the preset block length through the intra-block transformation subunit; the inter-block processing unit includes an inter-block transformation subunit; the second-dimensional feature is an inter-block dependence feature; correspondingly, the extracting second-dimensional features of each voice block feature through the inter-block processing unit of the separator includes: adding each voice block feature and the corresponding intra-block dependence feature as input data and inputting the input data into the inter-block processing unit, and extracting the inter-block dependence feature of each voice block feature based on the preset block length and the preset step length through the inter-block transformation subunit; the separator further includes a superposition unit and a gating unit; correspondingly, the fusing the first-dimensional features and the second-dimensional features of each voice block feature through the separator to obtain at least two voice prediction features includes: adding each intra-block dependence feature and the matched inter-block dependence feature to obtain each fused feature; inputting each fused feature into the superposition unit to obtain a combined feature; inputting the combined feature into the gating unit to obtain at least two voice prediction features.

2. The method according to claim 1, characterized in that, the separator further includes a separation enhancement unit; the separation enhancement unit includes at least two branches, and the number of branches in the separation enhancement unit is the same as the number of voice prediction features; after obtaining at least two voice prediction features, the method further includes: concatenating the at least two voice prediction features, inputting the concatenated features into each branch of the separation enhancement unit, and enhancing each voice prediction feature based on the multiplication masks output by each branch.

3. The method according to claim 1, characterized in that, the intra-block transformation subunit includes a temporal convolutional network; the inter-block transformation subunit includes a bidirectional long short-term memory network.

4. The method according to claim 1, characterized in that, the intra-block processing unit further includes a forward adaptation subunit; the forward adaptation subunit is arranged before the intra-block transformation subunit; Before extracting the first - dimensional features of each speech segment feature through the in - block processing unit of the separator, the method further includes: Inputting each speech segment feature into the forward - adaptive subunit, and optimizing each speech segment feature through the forward - adaptive subunit.

5. A processing device for speech data, characterized in that, it includes: A speech feature to be separated generation module, configured to obtain speech data to be separated, and perform feature extraction on the speech data to be separated through a feature extractor to obtain speech features to be separated; A speech segment feature generation module, configured to segment the speech features to be separated through a separator according to a preset block length and a preset step length to obtain at least two speech segment features; A two - dimensional feature extraction module, configured to extract the first - dimensional features of each speech segment feature through the in - block processing unit of the separator; and extract the second - dimensional features of each speech segment feature through the inter - block processing unit of the separator; A speech prediction feature generation module, configured to fuse the first - dimensional features and the second - dimensional features of each speech segment feature through the separator to obtain at least two speech prediction features; A speech separation result determination module, configured to determine each speech separation result through a waveform reconstructor according to each speech prediction feature and the speech features to be separated; wherein, the in - block processing unit includes an in - block transformation subunit; the first - dimensional feature is an in - block dependent feature; Correspondingly, the two - dimensional feature extraction module is specifically configured to: Input each speech segment feature into the in - block processing unit, and extract the in - block dependent features of each speech segment feature based on a preset block length through the in - block transformation subunit; The inter - block processing unit includes an inter - block transformation subunit; the second - dimensional feature is an inter - block dependent feature; Correspondingly, the two - dimensional feature extraction module is specifically configured to: Add each speech segment feature and the corresponding in - block dependent feature as input data and input it into the inter - block processing unit, and extract the inter - block dependent features of each speech segment feature based on a preset block length and a preset step length through the inter - block transformation subunit; The separator further includes a superposition unit and a gating unit; Correspondingly, the speech prediction feature generation module is specifically configured to: Add each in - block dependent feature and the matched inter - block dependent feature to obtain each fused feature; Input each fused feature into the superposition unit to obtain a combined feature; Input the combined feature into the gating unit to obtain at least two speech prediction features.

6. An electronic device, characterized in that, the electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for processing speech data according to any one of claims 1 - 4.

7. A computer - readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing the method for processing voice data according to any one of claims 1-4 when executed by a processor.

Citation Information

Patent Citations

  • Single-channel voice separation method and device and electronic equipment

    CN111429938A

  • Speech processing method, device and equipment, and storage medium

    CN111899758A

  • Voice deep neural network training method and device, storage medium and electronic device

    CN114067785A