Voice extraction method, device and electronic equipment
By jointly using a speaker extraction network and a speaker suppression network, and iteratively processing the audio domain sequence representation, the problem of speech information loss in existing technologies is solved, achieving higher speech extraction accuracy and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING CHANGAN AUTOMOBILE CO LTD
- Filing Date
- 2023-05-15
- Publication Date
- 2026-05-12
AI Technical Summary
In existing target speaker extraction and speaker suppression technology frameworks, the use of a single network framework leads to loss of speech information, causing distortion and reducing user experience.
采用联合使用说话人抽取网络和说话人抑制网络的方法,通过多个网络模块迭代处理语音频域序列表征,共享信息以弥补丢失的语音信息,提高提取准确性。
This greatly improves the accuracy of speech extraction and enhances the user experience.
Smart Images

Figure CN116434767B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech separation technology, specifically to a speech extraction method, apparatus, and electronic device. Background Technology
[0002] Speech extraction has become a research hotspot in speech signal processing in recent years. Speech extraction includes two branches of speech separation: target speaker extraction and target speaker suppression. The purpose of target speaker extraction is to extract the speech belonging to the target speaker from a mixed speech by utilizing the characteristics of the target speaker. The purpose of target speaker suppression is to suppress the speaker's voice information from a mixed speech by utilizing the speaker's information features.
[0003] Existing target speaker extraction and speaker suppression are implemented through a main network framework of "encoder-speaker extraction network / speaker suppression network-decoder". The limitation of this network framework is mainly due to the fact that the performance of the target speaker extraction and suppression network frameworks both borrow from a single separation network. When using a single speaker extraction network or speaker suppression network, it will cause certain damage to other sound source information and lose speech information, thereby causing distortion, making speech extraction inaccurate and reducing the user experience. Summary of the Invention
[0004] One of the objectives of this invention is to provide a speech extraction method, apparatus, and electronic device that can jointly use a speaker extraction network and a speaker suppression network in the process of target speaker extraction and speaker suppression, effectively avoiding the distortion problem caused by the loss of speech information due to the use of a single network, greatly improving the accuracy of speech extraction and enhancing the user experience.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A speech extraction method, wherein the method includes:
[0007] The process involves obtaining the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the target speaker's registered speech data; wherein the multi-speaker aliased speech data contains the target speaker's speech.
[0008] The audio domain sequence representation and voiceprint features are simultaneously input into the trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker. The audio domain sequence representation and voiceprint features are simultaneously input into the trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker.
[0009] The target speaker's latent audio domain sequence representation and the first target speaker's audio domain sequence representation, along with the audio domain sequence representation of the first non-target speaker, are simultaneously input into the second speaker extraction network module to obtain the second target speaker's audio domain sequence representation. Additionally, the non-target speaker's latent audio domain sequence representation and the first non-target speaker's audio domain sequence representation are simultaneously input into the second speaker suppression network module to obtain the second non-target speaker's audio domain sequence representation.
[0010] The second target speaker's audio domain sequence representation is input into the trained speech decoding and transformation network to reconstruct the target speaker's time-domain speech signal. Similarly, the second non-target speaker's audio domain sequence representation is input into the trained speech decoding and transformation network to reconstruct the non-target speaker's time-domain speech signal.
[0011] Furthermore, the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the target speaker's registered speech data are obtained, including:
[0012] Collect multi-speaker aliased speech data to be extracted and voiceprint registration speech data of the target speaker;
[0013] The multi-speaker aliased speech data to be extracted is input into the trained short-time Fourier transform network to obtain multi-speaker aliased speech domain data, and the voiceprint registration speech data of the target speaker is input into the trained short-time Fourier transform network to obtain voiceprint registration speech domain data.
[0014] Input multi-speaker aliased speech domain data into a trained speech coding network to obtain speech domain sequence representations;
[0015] Voiceprint registration audio domain data is input into a trained speaker assistance network to obtain voiceprint features.
[0016] Furthermore, the network structure of the speech coding network consists of an input layer, a first 2D convolutional layer, a first LN normalization layer, a first Prelu activation function layer, a DenseNet model layer, a second 2D convolutional layer, a second LN normalization layer, and a second Prelu activation function layer connected in sequence.
[0017] The DenseNet model consists of four dilated convolutional layers with dilation factors of 1, 2, 4, and 8, respectively.
[0018] Furthermore, the second speaker extraction network module includes at least one pre-trained second speaker extraction network;
[0019] The target speaker's latent audio domain sequence representation and the first target speaker's audio domain sequence representation are simultaneously input into the second speaker extraction network module to obtain the second target speaker's audio domain sequence representation, including:
[0020] When the second speaker extraction network module includes a pre-trained second speaker extraction network, the target speaker's latent audio domain sequence representation and the first target speaker's audio domain sequence representation are simultaneously input into the second speaker extraction network along with the first non-target speaker's audio domain sequence representation to obtain the second target speaker's audio domain sequence representation; or...
[0021] In the case where the second speaker extraction network module includes multiple trained second speaker extraction networks, the target speaker latent speech domain sequence representation and the first target speaker speech domain sequence representation, along with the speech domain sequence representation of the first non-target speaker, are simultaneously input into the second speaker extraction network connected to the first speaker extraction network in the second speaker extraction network module to obtain the intermediate target speaker speech domain sequence representation.
[0022] For second speaker extraction networks other than those connected to the first speaker extraction network, the target speaker latent audio domain sequence representation and the target speaker latent audio domain sequence representation of the first non-target speaker audio domain sequence representation, as well as the intermediate target speaker audio domain sequence representation output by the previous second speaker extraction network, are simultaneously input into the second speaker extraction network to obtain the intermediate target speaker audio domain sequence representation output by the second speaker extraction network.
[0023] The intermediate target speaker audio domain sequence representation of the last second speaker extracted from the network output is used as the second target speaker audio domain sequence representation.
[0024] Furthermore, the second speaker suppression network module includes at least one trained second speaker suppression network;
[0025] The latent speech domain sequence representation of the non-target speaker, along with the speech domain sequence representation of the first target speaker and the speech domain sequence representation of the first non-target speaker, are simultaneously input into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation, including:
[0026] When the second speaker suppression network module includes a pre-trained second speaker suppression network, the speech domain sequence representation, the non-target speaker latent speech domain sequence representation, and the first non-target speaker speech domain sequence representation are simultaneously input into the second speaker suppression network to obtain the second non-target speaker speech domain sequence representation; or...
[0027] In the case where the second speaker suppression network module includes multiple trained second speaker suppression networks, the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation are simultaneously input into the second speaker suppression network connected to the first speaker suppression network in the second speaker suppression network module to obtain the intermediate non-target speaker speech domain sequence representation.
[0028] For second speaker suppression networks other than those connected to the first speaker suppression network, the speech domain sequence representation, the non-target speaker latent speech domain sequence representation of the first target speaker speech domain sequence representation, and the intermediate non-target speaker speech domain sequence representation output by the previous second speaker suppression network are simultaneously input into the second speaker suppression network to obtain the intermediate non-target speaker speech domain sequence representation output by the second speaker suppression network.
[0029] The intermediate non-target speaker audio domain sequence representation of the last second speaker suppression network output is used as the second non-target speaker audio domain sequence representation.
[0030] Furthermore, the first speaker extraction network, the second speaker extraction network, the first speaker suppression network, and the second speaker suppression network are all masking networks;
[0031] The masking network structure includes an input layer, multiple stacked blocks, a ReLU activation function layer, a first sigmoid activation function layer, and a first dot product output layer connected in sequence. The input layer is also connected to the first dot product output layer.
[0032] Each stacked block contains multiple recurrent neural networks connected in sequence.
[0033] Furthermore, the second target speaker's audio domain sequence representation is input into the trained speech decoding and transformation network to reconstruct the target speaker's time-domain speech signal; and the second non-target speaker's audio domain sequence representation is input into the trained speech decoding and transformation network to reconstruct the non-target speaker's time-domain speech signal, including:
[0034] The second target speaker's audio domain sequence representation is input into the trained speech decoding network to obtain the decoded second target speaker's audio domain sequence representation.
[0035] The decoded audio domain sequence representation of the second target speaker is input into a trained inverse Fourier transform network to reconstruct the time-domain speech signal of the target speaker; and,
[0036] The second non-target speaker audio domain sequence representation is input into the trained speech decoding network to obtain the decoded second non-target speaker audio domain sequence representation.
[0037] The decoded audio domain sequence representation of the second non-target speaker is input into a trained inverse Fourier transform network to reconstruct the time-domain speech signal of the non-target speaker.
[0038] Furthermore, the network structure of the speech decoding network includes a DenseNet model layer, a sub-pixel2D convolutional layer, a third LN normalization layer, a third 2-D convolutional layer, and a dual-path model layer connected in sequence.
[0039] The dual-path model layer includes a fourth 2D convolutional layer, a fifth 2D convolutional layer, a tanh activation function layer, a second sigmoid activation function layer, a second dot product output layer, a sixth 2D convolutional layer, and a third sigmoid activation function layer.
[0040] A speech extraction device, wherein the device comprises:
[0041] The acquisition module is used to acquire the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the target speaker's voiceprint registration speech data; wherein, the multi-speaker aliased speech data contains the target speaker's speech;
[0042] The first input module is used to simultaneously input the audio domain sequence representation and the voiceprint features into the trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker, and to simultaneously input the audio domain sequence representation and the voiceprint features into the trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker.
[0043] The second input module is used to simultaneously input the target speaker latent speech domain sequence representation and the first target speaker speech domain sequence representation, along with the speech domain sequence representation of the first non-target speaker, into the second speaker extraction network module to obtain the second target speaker speech domain sequence representation; and to simultaneously input the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation.
[0044] The speech extraction and reconstruction module is used to input the audio domain sequence representation of the second target speaker into the trained speech decoding and transformation network to reconstruct the time domain speech signal of the target speaker, and to input the audio domain sequence representation of the second non-target speaker into the trained speech decoding and transformation network to reconstruct the time domain speech signal of the non-target speaker.
[0045] An electronic device includes a processor and a memory, the processor being configured to execute a speech extraction program stored in the memory to implement the speech extraction method described above.
[0046] A readable storage medium, characterized in that the readable storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described speech extraction method.
[0047] The beneficial effects of this invention are:
[0048] The speech extraction method, apparatus, and electronic device provided in this invention include: acquiring the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the voiceprint registration speech data of the target speaker; simultaneously inputting the audio domain sequence representation and the voiceprint features into a trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker; and simultaneously inputting the audio domain sequence representation and the voiceprint features into a trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker; and combining the audio domain sequence representation with the audio domain sequence representation of the first non-target speaker and the audio domain sequence representation of the target speaker's latent audio domain, and the audio domain sequence representation of the first target speaker. The list of features is simultaneously input into the second speaker extraction network module to obtain the second target speaker audio domain sequence representation. Then, the audio domain sequence representation, along with the non-target speaker latent audio domain sequence representation and the first non-target speaker audio domain sequence representation, are simultaneously input into the second speaker suppression network module to obtain the second non-target speaker audio domain sequence representation. The second target speaker audio domain sequence representation is then input into a trained speech decoding and transformation network to reconstruct the target speaker's temporal speech signal. Finally, the second non-target speaker audio domain sequence representation is input into the trained speech decoding and transformation network to reconstruct the non-target speaker's temporal speech signal. This invention combines a speaker extraction network and a speaker suppression network in the process of target speaker extraction and speaker suppression. Since the inputs of the speaker extraction network and the speaker suppression network are the same and their outputs are complementary, the outputs of the speaker extraction network and the speaker suppression network each contain speech information missing from the other network. Information can be shared between the target speaker latent speech domain sequence representation of the speaker extraction network and the non-target speaker latent speech domain sequence representation of the speaker suppression network to compensate for the lost speech information, thereby greatly improving the accuracy of speech extraction and enhancing the user experience. Attached Figure Description
[0049] Figure 1 A schematic diagram of the network architecture extracted for the speaker;
[0050] Figure 2 A schematic diagram of a network architecture for speaker suppression;
[0051] Figure 3 A flowchart illustrating an embodiment of a speech extraction method provided by this invention;
[0052] Figure 4 This is a schematic diagram of a speech coding network structure provided in an embodiment of the present invention;
[0053] Figure 5A schematic diagram of a masking network structure provided in an embodiment of the present invention;
[0054] Figure 6 This is a schematic diagram of a network structure for a voice decoding network provided in an embodiment of the present invention;
[0055] Figure 7 A schematic diagram of a speech extraction framework provided in an embodiment of the present invention;
[0056] Figure 8 This is a schematic diagram of the structure of a speech extraction device provided in an embodiment of the present invention;
[0057] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0059] Currently, the technical frameworks commonly used in academia and industry for target speaker extraction and speaker suppression are as follows: Figure 1 and Figure 2As shown in the figure, the overall framework of speaker extraction and speaker suppression is similar, both consisting of two parts. One part is the main network architecture composed of "encoder - speaker extraction network / speaker suppression network - decoder," and the other part is an encoder and auxiliary network specifically providing target speaker features. The encoder's role is to transform the time-domain signals of the input mixed speech and target speaker speech into another latent space, forming feature vector representations of the mixed speech and target speaker speech in this latent space. The speaker extraction network / speaker suppression network's role is to extract or suppress the target speech in the latent space. The decoder's role is to transform the extracted target speaker features or the features of other sound sources after suppressing the target speaker back into the time-domain signal. The auxiliary network's role is to provide the main network architecture with information related to the target speaker, that is, to tell the main network architecture which person's voice should be extracted or suppressed. This is the existing framework for target speaker extraction technology. Each module is a neural network (or a traditional signal processing module, such as a Fourier transform). For example, the encoder / decoder can use a 1D convolutional network (1d-CNN), while the speaker extraction and suppression networks can use LSTM (Long Short-Term Memory), TCN (Temporal Convolutional Network), RNN (Recurrent Neural Network), etc. The auxiliary network can use an output layer from a speaker recognition network. The training criteria can be MSE (Mean Square Error) or SI-SDR (Scale Invariant Signal-to-Distortion Ratio). Thus, the different network structures, objective functions, and training methods of different modules can be combined to create numerous algorithms, which represents the current state of target speaker extraction and suppression technology.
[0060] The limitations of existing target speaker extraction and suppression technology frameworks are mainly due to the fact that the main network architectures for these techniques rely on single, separate networks. After using these networks, the success depends entirely on the quality of the target speaker's voiceprint features transmitted by the auxiliary network. The closer the voiceprint features provided by the auxiliary network are to the speaker's features in the mixed speech, the better the effect; conversely, the effect deteriorates. Furthermore, the extraction or suppression network can damage other sound source information, causing loss of speech information and distortion, resulting in inaccurate speech extraction and a reduced user experience. Therefore, the speech extraction method, apparatus, and electronic device provided in this invention can alleviate the above problems.
[0061] To facilitate understanding of the embodiments of the present invention, further explanations and descriptions will be provided below with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.
[0062] This embodiment proposes a speech extraction method, see [link to relevant documentation]. Figure 3 The above is a flowchart of an embodiment of a speech extraction method provided by the present invention. Figure 3 The process shown may include the following steps:
[0063] Step 301: Obtain the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the target speaker's registered speech data;
[0064] Among them, the multi-speaker aliased speech data includes the speech of the target speaker; the audio domain sequence is represented as the high-dimensional feature of the multi-speaker aliased speech data in the frequency domain; and the voiceprint registration speech data refers to the clean speech of the target speaker used for voiceprint registration.
[0065] Step 302: Input the audio domain sequence representation and the voiceprint features into the trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker; and input the audio domain sequence representation and the voiceprint features into the trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker.
[0066] In this embodiment, the first speaker extraction network extracts the representation of the target speaker from the multi-speaker aliased speech data to obtain the first target speaker audio domain sequence representation. Furthermore, the first speaker suppression network extracts the representation of the non-target speaker from the multi-speaker aliased speech data to obtain the first non-target speaker audio domain sequence representation.
[0067] Step 303: Simultaneously input the target speaker latent speech domain sequence representation and the first target speaker speech domain sequence representation, along with the speech domain sequence representation of the first non-target speaker, into the second speaker extraction network module to obtain the second target speaker speech domain sequence representation; and simultaneously input the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation.
[0068] The latent audio domain sequence representation of the target speaker is obtained by subtracting the audio domain sequence representation of the first non-target speaker from the audio domain sequence representation of the target speaker. The latent audio domain sequence representation of the target speaker includes the speech information lost by the first speaker extraction network and the audio domain sequence representation of the first target speaker. In this embodiment, the latent audio domain sequence representation of the target speaker and the audio domain sequence representation of the first target speaker are used as inputs to the second speaker extraction network module to compensate for the speech information lost by the first speaker extraction network. This enables the second speaker extraction network model to further accurately extract the representation belonging to the target speaker from the multi-speaker aliased speech data, greatly improving the accuracy of the target speaker's speech.
[0069] The latent audio domain sequence representation of the non-target speaker is obtained by subtracting the audio domain sequence representation of the first target speaker from the audio domain sequence representation. This latent audio domain sequence representation of the non-target speaker includes the speech information lost by the first speaker suppression network and the audio domain sequence representation of the first non-target speaker. In this embodiment, the latent audio domain sequence representation of the non-target speaker and the audio domain sequence representation of the first non-target speaker are used as inputs to the second speaker suppression network module to compensate for the speech information lost by the first speaker suppression network. This enables the second speaker suppression network model to further accurately extract representations that do not belong to the target speaker from multi-speaker aliased speech data, greatly improving the accuracy of non-target speaker speech.
[0070] Since the speaker extraction network output only contains the speech features of the target speaker, while the speaker suppression network output contains the speech features of other speakers besides the target speaker, in this embodiment, the joint speaker extraction and suppression process essentially involves subtracting the output representation from the audio domain sequence representation obtained through the first speaker extraction and suppression networks. The results of this subtraction include the missing speech information from the first speaker extraction and suppression networks. The subtracted results are then iterated again with the outputs of the first speaker extraction and suppression networks using their respective networks, namely the second speaker extraction and suppression network modules, to achieve the lowest possible distortion in the final result and thus accurately extract the speech of both the target and non-target speakers.
[0071] Step 304: Input the second target speaker's audio domain sequence representation into the trained speech decoding and transformation network to restore the target speaker's time-domain speech signal; and input the second non-target speaker's audio domain sequence representation into the trained speech decoding and transformation network to restore the non-target speaker's time-domain speech signal.
[0072] The speech extraction method provided in this invention uses multiple speaker extraction networks and multiple speaker suppression networks in conjunction during the target speaker extraction and speaker suppression processes. Since the inputs of the first speaker extraction network and the first speaker suppression network are the same and their outputs are complementary, the outputs of the first speaker extraction network and the first speaker suppression network each contain speech information missing from another network. The intermediate information of the two networks, namely the latent speech domain sequence representation of the target speaker and the latent speech domain sequence representation of the non-target speaker, can be shared to compensate for the lost speech information, thereby greatly improving the accuracy of speech extraction and enhancing the user experience.
[0073] In one implementation, step 301 can be achieved through steps A1 to A4:
[0074] Step A1: Collect the multi-speaker aliased speech data to be extracted and the voiceprint registration speech data of the target speaker;
[0075] Specifically, taking a sampling rate of 16kHz as an example, the speech segment to be extracted of arbitrary length and the voiceprint registration speech segment of the specified target speaker are collected, thus obtaining the multi-speaker aliased speech data to be extracted and the voiceprint registration speech data of the target speaker.
[0076] Step A2: Input the multi-speaker aliased speech data to be extracted into the trained short-time Fourier transform network to obtain multi-speaker aliased speech domain data, and input the voiceprint registration speech data of the target speaker into the trained short-time Fourier transform network to obtain voiceprint registration speech domain data.
[0077] The time-domain multi-speaker aliased speech data to be extracted and the voiceprint registration speech data of the target speaker are transformed to the frequency domain by a short-time Fourier transform network. Since the frequency domain data has better separability, it is easier for the speaker extraction network to extract the target speaker representation better, and for the speaker suppression network to extract the non-target speaker representation better.
[0078] Step A3: Input the multi-speaker aliased audio domain data into the trained speech coding network to obtain audio domain sequence representations;
[0079] The limitations of existing speaker extraction and speaker suppression technology frameworks also include: since the encoder is generally composed of a 1-dimensional convolutional neural network, the speech signal is mapped to a new high-dimensional space after passing through the encoder, but the distinguishability of this space is poor, resulting in poor speech extraction performance of speaker extraction networks or speaker suppression networks.
[0080] In this embodiment, instead of using a one-dimensional convolutional neural network as the encoder, a speech coding network composed of multiple networks is used as the encoder. The speech coding network receives multi-speaker aliased audio domain data as input, and maps the frequency domain of the multi-speaker aliased audio domain data to higher-dimensional features through the speech coding network, outputting audio domain sequence representations. Since the audio domain sequence representations have higher separability and more abundant frequency domain data, the speaker extraction network and speaker suppression network can better distinguish between target speaker representations and non-target speaker representations, ensuring that the output results have less distortion.
[0081] The aforementioned audio domain sequence represents the deep, high-dimensional features of multiple speakers in the multi-speaker aliased speech data to be extracted. The network structure of the speech coding network is as follows: Figure 4 As shown, it includes an input layer, a first 2D convolutional layer, a first LN normalization layer, a first Prelu activation function layer, a DenseNet model layer, a second 2D convolutional layer, a second LN normalization layer, and a second Prelu activation function layer connected in sequence; wherein, the DenseNet model layer includes four dilated convolutional layers, and the dilation factors of the four dilated convolutional layers are 1, 2, 4, and 8, respectively.
[0082] Step A4: Input the voiceprint registration audio domain data into the trained speaker assistance network to obtain voiceprint features.
[0083] Voiceprint features are high-dimensional features that map voiceprint registration audio domain data to a higher dimension.
[0084] In this embodiment, the speaker-assisted network is the same as... Figure 1 and Figure 2 The auxiliary network in the process of obtaining voiceprint features is as follows:
[0085] Step B1: Model the frequency domain dependency of the voiceprint registration audio domain data using convolutional or recurrent neural networks;
[0086] Specifically, the frequency domain dependencies of voiceprint registration speech audio domain data can be modeled by stacking multiple layers of convolutional networks with residual connections or bidirectional long short-term memory networks. On one hand, a convolutional network with n ≥ 5 layers can be used for modeling, where the first layer has (D, O) input and output channels, the middle layers have (O, O), the third-to-last layer has (O, P), the second-to-last layer has (P, P), and the last layer has (P, H). Furthermore, all convolutional networks except the first and last layers use residual connections, and layer normalization is added before the first layer. On the other hand, a bidirectional long short-term memory network with n layers, an input dimension of D, and a hidden layer dimension of H can also be used for modeling, followed by processing using a ReLU activation function and a fully connected layer with an input dimension of H.
[0087] Step B2 involves using a pooling layer based on a self-attention mechanism to extract the voiceprint features of the target speaker from the voiceprint registration audio domain data after modeling.
[0088] Specifically, the pooling layer based on the self-attention mechanism consists of a feedforward network and a pooling network. The feedforward network comprises two fully connected layers with input and output channels of (H,h) and (h,1), respectively. The pooling network first calculates the attention coefficients through masking, then uses the softmax function to obtain the probability weights at each time point, and performs pooling operation by weighted averaging. Finally, after processing by the fully connected layer and the tanh activation function, the voiceprint features of the target speaker are obtained.
[0089] The aforementioned second speaker extraction network module includes at least one trained second speaker extraction network. In practical applications, the more second speaker extraction networks there are, the more accurate the target speaker representation will be in the iterative process, and the more accurately the target speaker's speech can be extracted from the multi-speaker aliased speech data to be extracted.
[0090] When the second speaker extraction network module includes a trained second speaker extraction network, the process in step 303 above of simultaneously inputting the target speaker's latent audio domain sequence representation and the first target speaker's audio domain sequence representation, along with the first non-target speaker's audio domain sequence representation, into the second speaker extraction network module to obtain the second target speaker's audio domain sequence representation includes: simultaneously inputting the target speaker's latent audio domain sequence representation and the first target speaker's audio domain sequence representation into the second speaker extraction network to obtain the second target speaker's audio domain sequence representation.
[0091] Since the latent speech domain sequence representation of the target speaker includes speech information lost in the first speaker extraction network, the latent speech domain sequence representation of the target speaker and the first target speaker speech domain sequence representation output by the first speaker extraction network can be concatenated using the sum or cat method to compensate for the speech information lost in the first speaker extraction network, thereby reducing speech distortion. Then, the concatenated representation is input into the second speaker extraction network, enabling the second speaker extraction network to more accurately extract the second target speaker speech domain sequence representation of the target speaker's speech from the multi-speaker aliased speech data to be extracted.
[0092] When the second speaker extraction network module includes multiple trained second speaker extraction networks, the process in step 303 above, which involves simultaneously inputting the target speaker's latent audio domain sequence representation and the first target speaker's audio domain sequence representation, along with the first non-target speaker's audio domain sequence representation, into the second speaker extraction network module to obtain the second target speaker's audio domain sequence representation, includes:
[0093] Step C1: The target speaker latent speech domain sequence representation and the first target speaker speech domain sequence representation, along with the speech domain sequence representation of the first non-target speaker, are simultaneously input into the second speaker extraction network module, which is connected to the first speaker extraction network, to obtain the intermediate target speaker speech domain sequence representation.
[0094] Step C2: For second speaker extraction networks other than those connected to the first speaker extraction network, the target speaker latent audio domain sequence representation and the target speaker latent audio domain sequence representation of the first non-target speaker audio domain sequence representation, as well as the intermediate target speaker audio domain sequence representation output by the previous second speaker extraction network, are simultaneously input into the second speaker extraction network to obtain the intermediate target speaker audio domain sequence representation output by the second speaker extraction network.
[0095] For all second speaker extraction networks except those connected to the first speaker extraction network, the network input is the intermediate target speaker audio domain sequence representation and the target speaker latent audio domain sequence representation output by the previous second speaker extraction network. This compensates for the speech information lost by the first speaker extraction network. Through multiple iterations of multiple second speaker extraction networks, the more accurate the output target speaker representation is, the more accurately the speech representation of the target speaker can be extracted from the multi-speaker aliased speech data to be extracted.
[0096] Step C3: The intermediate target speaker audio domain sequence representation extracted from the network output of the last second speaker is used as the second target speaker audio domain sequence representation.
[0097] In this embodiment, the intermediate target speaker audio domain sequence representation output by the last second speaker extraction network among the multiple second speaker extraction networks connected in sequence in the second speaker extraction network module is regarded as the second target speaker audio domain sequence representation of the target speaker speech finally extracted from the multi-speaker aliased speech data to be extracted.
[0098] The aforementioned second speaker suppression network module includes at least one trained second speaker suppression network. In practical applications, the more second speaker suppression networks there are, the more accurate the representation of the non-target speaker generated in the iteration will be, and the more accurately the speech of the non-target speaker can be extracted from the multi-speaker aliased speech data to be extracted.
[0099] When the second speaker suppression network module includes a trained second speaker suppression network, the process in step 303 above of simultaneously inputting the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation includes: simultaneously inputting the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation into the second speaker suppression network to obtain the second non-target speaker speech domain sequence representation.
[0100] Since the latent speech domain sequence representation of the non-target speaker includes speech information lost in the first speaker suppression network, the latent speech domain sequence representation of the non-target speaker and the first non-target speaker speech domain sequence representation output by the first speaker suppression network can be concatenated using the sum or cat method to compensate for the speech information lost by the first speaker suppression network, thereby reducing speech distortion. Then, the concatenated representation is input into the second speaker suppression network, enabling the second speaker suppression network to more accurately extract the second non-target speaker speech domain sequence representation of the non-target speaker speech from the multi-speaker aliased speech data to be extracted.
[0101] When the second speaker suppression network module includes multiple trained second speaker suppression networks, the process in step 303 above, which involves simultaneously inputting the speech domain sequence representation, the non-target speaker latent speech domain sequence representation, and the first non-target speaker speech domain sequence representation into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation, includes:
[0102] Step D1: The speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation are simultaneously input into the second speaker suppression network module, which is connected to the first speaker suppression network, to obtain the intermediate non-target speaker speech domain sequence representation.
[0103] Step D2: For any second speaker suppression network other than the one connected to the first speaker suppression network, the speech domain sequence representation, the non-target speaker latent speech domain sequence representation of the first target speaker speech domain sequence representation, and the intermediate non-target speaker speech domain sequence representation output by the previous second speaker suppression network are simultaneously input into the second speaker suppression network to obtain the intermediate non-target speaker speech domain sequence representation output by the second speaker suppression network.
[0104] For other second speaker suppression networks besides the one connected to the first speaker suppression network, the network input is the intermediate non-target speaker audio domain sequence representation and the non-target speaker latent audio domain sequence representation output by the previous second speaker suppression network. This compensates for the speech information lost by the first speaker suppression network. Through multiple iterations of multiple second speaker suppression networks, the more accurate the output non-target speaker representation is, the more accurately the speech representation of the non-target speaker can be extracted from the multi-speaker aliased speech data to be extracted.
[0105] Step D3: The intermediate non-target speaker audio domain sequence representation of the last second speaker suppression network output is used as the second non-target speaker audio domain sequence representation.
[0106] In this embodiment, the intermediate non-target speaker audio domain sequence representation output by the last of the multiple second speaker suppression networks connected in sequence in the second speaker suppression network module is regarded as the second non-target speaker audio domain sequence representation of the non-target speaker speech finally extracted from the multi-speaker aliased speech data to be extracted.
[0107] In practical applications, the first speaker extraction network, the second speaker extraction network, the first speaker suppression network, and the second speaker suppression network are all masking networks; the network structure of a masking network is as follows: Figure 5 As shown, it includes an input layer, multiple stacked blocks, a ReLU activation function layer, a first sigmoid activation function layer, and a first dot product output layer connected in sequence. The input layer is also connected to the first dot product output layer. Each stacked block includes multiple recurrent neural networks (RNNs) connected in sequence.
[0108] In this embodiment, a recurrent neural network (RNN) can be selected as the base network for the masking network. Multiple RNNs are connected in series to form a stack, and then the stacks are combined in series to form a complete network. The number of RNNs and the number of stacks can be adjusted according to performance and algorithm requirements when combining RNNs into stacks and stacks into a network, and are not limited here. The output representation of the first sigmoid activation function layer in the masking network is multiplied with the speech domain sequence representation in the first dot product output layer to obtain the latent representation, namely, the speech domain sequence representation of the first target speaker, the speech domain sequence representation of the first non-target speaker, the speech domain sequence representation of the second target speaker, the speech domain sequence representation of the second non-target speaker, the speech domain sequence representation of the intermediate target speaker, or the speech domain sequence representation of the intermediate non-target speaker.
[0109] In one implementation, step 304 can be achieved through steps E1 to E4:
[0110] Step E1: Input the second target speaker audio domain sequence representation into the trained speech decoding network to obtain the decoded second target speaker audio domain sequence representation.
[0111] The network structure of the above-mentioned speech decoding network is as follows: Figure 6 As shown, it includes: a DenseNet model layer, a sub-pixel2D convolutional layer, a third LN normalization layer, a third 2D convolutional layer, and a dual-path model layer connected in sequence; wherein, the dual-path model layer includes a fourth 2D convolutional layer, a fifth 2D convolutional layer, a tanh activation function layer, a second sigmoid activation function layer, a second dot product output layer, a sixth 2D convolutional layer, and a third sigmoid activation function layer.
[0112] In this embodiment, the speech decoding network is used as a decoder. The speech decoding network can be used to decode the high-dimensional second target speaker audio domain sequence representation to a low dimension, so as to obtain the decoded second target speaker audio domain sequence representation.
[0113] Step E2: Input the decoded audio domain sequence representation of the second target speaker into the trained inverse Fourier transform network to reconstruct the time domain speech signal of the target speaker;
[0114] After obtaining the audio domain sequence representation of the second target speaker, an inverse Fourier transform network is used to convert the frequency domain signal into a time domain signal. Experiments have shown that the speech signal obtained by frequency domain conversion plus the mapping of encoder and decoder has a lower distortion rate.
[0115] Step E3: Input the second non-target speaker audio domain sequence representation into the trained speech decoding network to obtain the decoded second non-target speaker audio domain sequence representation.
[0116] Step E4: Input the decoded audio domain sequence representation of the second non-target speaker into the trained inverse Fourier transform network to reconstruct the time domain speech signal of the non-target speaker.
[0117] The process of restoring the time-domain speech signal of the non-target speaker in steps E3 to E4 is similar to the process of restoring the time-domain speech signal of the target speaker in steps E1 to E2, and will not be described in detail here. Furthermore, the execution order of steps E3-E4 and steps E1-E2 can be performed simultaneously or sequentially, and is not limited here.
[0118] To facilitate understanding of the entire speech extraction process, Figure 7 The speech extraction framework includes multiple second speaker extraction networks and multiple second speaker suppression networks as an example to illustrate speech extraction. Figure 7 As shown, the multi-speaker aliased speech data to be extracted and the voiceprint registration speech data of the target speaker are respectively input into a short-time Fourier transform network 700. The multi-speaker aliased speech domain data output from the short-time Fourier transform network 700 is input into a speech coding network 701 to obtain a speech domain sequence representation. The speech domain sequence representation output from the short-time Fourier transform network 700 is input into a speaker auxiliary network 702 to obtain voiceprint features. Then, the speech domain sequence representation and voiceprint features are input into a first speaker extraction network 703 to obtain a first target speaker speech domain sequence representation. At the same time, the speech domain sequence representation and voiceprint features are input into a first speaker suppression network 704 to obtain a first non-target speaker speech domain sequence representation. Then, for the target speaker speech extraction, the first target speaker speech domain sequence representation and the speechprint registration speech data of the target speaker are first processed by... Figure 7The minus sign in the diagram represents the subtraction of the audio domain sequence representation from the first non-target speaker's audio domain sequence representation. The resulting latent audio domain sequence representation of the target speaker is then input into the second speaker extraction network 705, which is ranked first. This yields the intermediate target speaker audio domain sequence representation output by this second speaker extraction network. The intermediate target speaker audio domain sequence representation and the latent target speaker audio domain sequence representation are then input into the second speaker extraction network 705, which is ranked second. This yields the intermediate target speaker audio domain sequence representation output by this second speaker extraction network. Finally, the intermediate target speaker audio domain sequence representation and the latent target speaker audio domain sequence representation are input into the third-ranked second speaker extraction network 705. The intermediate target speaker audio domain sequence representation output by the network is taken, and the above process is repeated until the second speaker extraction network 705, which is ranked last, is reached. The intermediate target speaker audio domain sequence representation output by the second speaker extraction network 705, which is ranked last, is regarded as the second target speaker audio domain sequence representation. This is to accurately extract the speech representation of the target speaker from the multi-speaker aliased speech data to be extracted. Finally, the second target speaker audio domain sequence representation is first input into the speech decoding network 706 to obtain the decoded second target speaker audio domain sequence representation. Then, the decoded second target speaker audio domain sequence representation is input into the inverse Fourier transform network 707 to restore the time-domain speech signal of the target speaker. The restored time-domain speech signal of the target speaker is the speech signal of the target speaker extracted from the multi-speaker aliased speech data to be extracted.
[0119] For non-target speaker speech extraction, after obtaining the first target speaker's audio domain sequence representation, the first non-target speaker's audio domain sequence representation, and the first non-target speaker's audio domain sequence representation, the first non-target speaker's audio domain sequence representation is first processed, and then... Figure 7The minus sign in the diagram represents the subtraction of the audio domain sequence representation from the first target speaker's audio domain sequence representation. The resulting non-target speaker latent audio domain sequence representation is then input into the second speaker suppression network 708, which is ranked first. This yields the intermediate non-target speaker audio domain sequence representation output by the second speaker suppression network. The intermediate non-target speaker audio domain sequence representation and the non-target speaker latent audio domain sequence representation output by the second speaker suppression network are then input into the second speaker suppression network 708, which is ranked second. This yields the intermediate non-target speaker audio domain sequence representation output by the second speaker suppression network. Finally, the intermediate target speaker audio domain sequence representation and the target speaker latent audio domain sequence representation output by the second speaker suppression network are input into the third speaker suppression network 708. This yields the second speaker suppression network... The output intermediate non-target speaker audio domain sequence representation is repeated in the above process until the second speaker suppression network 705, which is ranked last, is reached. The intermediate non-target speaker audio domain sequence representation output by the second speaker suppression network 708, which is ranked last, is regarded as the second non-target speaker audio domain sequence representation. This is to accurately extract the speech representation of the non-target speaker from the multi-speaker aliased speech data to be extracted. Finally, the second non-target speaker audio domain sequence representation is first input into the speech decoding network 706 to obtain the decoded second non-target speaker audio domain sequence representation. Then, the decoded second non-target speaker audio domain sequence representation is input into the inverse Fourier transform network 707 to restore the time-domain speech signal of the non-target speaker. The restored time-domain speech signal of the non-target speaker is the speech signal of the non-target speaker extracted from the multi-speaker aliased speech data to be extracted.
[0120] For the speech extraction process of a speech extraction framework that includes only one second speaker extraction network and one second speaker suppression network, the extraction of the target speaker's speech and... Figure 7 The difference in the speech extraction process lies only in that the intermediate target speaker audio domain sequence representation output by the second speaker extraction network (ranked first) is treated as the second target speaker audio domain sequence representation. Then, a speech decoding network and an inverse Fourier transform network are used to obtain the target speaker's time-domain speech signal. For non-target speaker speech extraction... Figure 7 The only difference in the speech extraction process is that the intermediate non-target speaker audio domain sequence representation output by the second speaker suppression network, which is ranked first above, is regarded as the second non-target speaker audio domain sequence representation. Then, the speech decoding network and the inverse Fourier transform network are used to obtain the time-domain speech signal of the non-target speaker, so as to realize the speech extraction of both the target speaker and the non-target speaker.
[0121] Corresponding to the above method embodiments, this embodiment provides a speech extraction device, see [link to relevant documentation]. Figure 8 The diagram shown illustrates the structure of a speech extraction device, which includes:
[0122] The acquisition module 801 is used to acquire the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the voiceprint registration speech data of the target speaker; wherein, the multi-speaker aliased speech data contains the speech of the target speaker;
[0123] The first input module 802 is used to simultaneously input the audio domain sequence representation and the voiceprint features into the trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker, and to simultaneously input the audio domain sequence representation and the voiceprint features into the trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker.
[0124] The second input module 803 is used to simultaneously input the target speaker latent speech domain sequence representation and the first target speaker speech domain sequence representation, along with the speech domain sequence representation of the first non-target speaker, into the second speaker extraction network module to obtain the second target speaker speech domain sequence representation; and to simultaneously input the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation.
[0125] The speech extraction and reconstruction module 804 is used to input the audio domain sequence representation of the second target speaker into the trained speech decoding and transformation network to reconstruct the time domain speech signal of the target speaker, and to input the audio domain sequence representation of the second non-target speaker into the trained speech decoding and transformation network to reconstruct the time domain speech signal of the non-target speaker.
[0126] The speech extraction apparatus provided in this embodiment of the invention includes: acquiring the audio domain sequence representation of multi-speaker aliased speech data to be extracted and the voiceprint features of the voiceprint registration speech data of the target speaker; simultaneously inputting the audio domain sequence representation and the voiceprint features into a trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker; and simultaneously inputting the audio domain sequence representation and the voiceprint features into a trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker; and combining the audio domain sequence representation with the target speaker latent audio domain sequence representation and the first target speaker audio domain sequence representation. The second target speaker's audio domain sequence representation is input into the second speaker extraction network module to obtain the second target speaker's audio domain sequence representation. Simultaneously, the audio domain sequence representation, along with the non-target speaker's latent audio domain sequence representation and the first non-target speaker's audio domain sequence representation, are input into the second speaker suppression network module to obtain the second non-target speaker's audio domain sequence representation. The second target speaker's audio domain sequence representation is then input into a trained speech decoding and transformation network to reconstruct the target speaker's temporal speech signal. Finally, the second non-target speaker's audio domain sequence representation is input into the trained speech decoding and transformation network to reconstruct the non-target speaker's temporal speech signal. This invention combines a speaker extraction network and a speaker suppression network in the process of target speaker extraction and speaker suppression. Since the inputs of the speaker extraction network and the speaker suppression network are the same and their outputs are complementary, the outputs of the speaker extraction network and the speaker suppression network each contain speech information missing from the other network. Information can be shared between the target speaker latent speech domain sequence representation of the speaker extraction network and the non-target speaker latent speech domain sequence representation of the speaker suppression network to compensate for the lost speech information, thereby greatly improving the accuracy of speech extraction and enhancing the user experience.
[0127] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 9 The illustrated electronic device 900 includes at least one processor 901, a memory 902, at least one network interface 904, and other user interfaces 903. The various components in the electronic device 900 are coupled together via a bus system 905. It is understood that the bus system 905 is used to implement communication between these components. In addition to a data bus, the bus system 905 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 9 The general labeled all buses as Bus System 905.
[0128] The user interface 903 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).
[0129] It is understood that the memory 902 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 902 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0130] In some implementations, memory 902 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 9021 and application program 9022.
[0131] The operating system 9021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 9022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 9022.
[0132] In this embodiment of the invention, the processor 901 executes the method steps provided in each method embodiment by calling the program or instructions stored in the memory 902, specifically the program or instructions stored in the application program 9022.
[0133] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 901. Processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in processor 901. The processor 901 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 902. Processor 901 reads the information in memory 902 and, in conjunction with its hardware, completes the steps of the above method.
[0134] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.
[0135] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0136] The electronic device provided in this embodiment may be as follows: Figure 9 The electronic device shown can perform the following: Figure 3 All steps of the Chinese speech extraction method are then implemented to achieve... Figure 3 For details on the technical effectiveness of the speech extraction method shown, please refer to [link / reference]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.
[0137] This invention also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; the memory may also include combinations of the above types of memory.
[0138] When one or more programs in the storage medium can be executed by one or more processors to implement the above data fusion method.
[0139] The processor is used to execute the data fusion program stored in the memory to implement the steps of the speech extraction method.
[0140] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0141] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0142] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A speech extraction method, characterized in that, The method includes: The process involves acquiring the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the target speaker's registered speech data; wherein the multi-speaker aliased speech data includes the target speaker's speech. The audio domain sequence representation and the voiceprint features are simultaneously input into a trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker. The audio domain sequence representation and the voiceprint features are simultaneously input into a trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker. The speech domain sequence representation, along with the target speaker latent speech domain sequence representation and the first target speaker speech domain sequence representation, are simultaneously input into the second speaker extraction network module to obtain the second target speaker speech domain sequence representation. Additionally, the speech domain sequence representation, along with the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation, are simultaneously input into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation. The second target speaker's audio domain sequence representation is input into the trained speech decoding and transformation network to restore the target speaker's time-domain speech signal; and the second non-target speaker's audio domain sequence representation is input into the trained speech decoding and transformation network to restore the non-target speaker's time-domain speech signal. The target speaker latent audio domain sequence representation is obtained by subtracting the first non-target speaker audio domain sequence representation from the audio domain sequence representation. The target speaker latent audio domain sequence representation includes the speech information lost by the network extracted by the first speaker and the first target speaker audio domain sequence representation. The non-target speaker latent audio domain sequence representation is obtained by subtracting the first target speaker audio domain sequence representation from the audio domain sequence representation. The non-target speaker latent audio domain sequence representation includes the speech information lost by the first speaker suppression network and the first non-target speaker audio domain sequence representation.
2. The method according to claim 1, characterized in that, The acquisition of the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the target speaker's registered speech data includes: Collect the multi-speaker aliased speech data to be extracted and the voiceprint registration speech data of the target speaker; The multi-speaker aliased speech data to be extracted is input into the trained short-time Fourier transform network to obtain multi-speaker aliased speech domain data, and the voiceprint registration speech data of the target speaker is input into the trained short-time Fourier transform network to obtain voiceprint registration speech domain data. The multi-speaker aliased audio domain data is input into a trained speech coding network to obtain audio domain sequence representations. The voiceprint registration audio domain data is input into the trained speaker assistance network to obtain voiceprint features.
3. The method according to claim 2, characterized in that, The network structure of the speech coding network is as follows: an input layer, a first 2D convolutional layer, a first LN normalization layer, a first Prelu activation function layer, a DenseNet model layer, a second 2D convolutional layer, a second LN normalization layer, and a second Prelu activation function layer connected in sequence. The DenseNet model layer includes four dilated convolutional layers with dilation factors of 1, 2, 4, and 8, respectively.
4. The method according to claim 1, characterized in that, The second speaker extraction network module includes at least one trained second speaker extraction network; The step of simultaneously inputting the target speaker latent audio domain sequence representation and the first target speaker audio domain sequence representation, along with the first non-target speaker audio domain sequence representation, into the second speaker extraction network module to obtain the second target speaker audio domain sequence representation includes: When the second speaker extraction network module includes a trained second speaker extraction network, the speech domain sequence representation, the target speaker latent speech domain sequence representation, and the first target speaker speech domain sequence representation are simultaneously input into the second speaker extraction network to obtain the second target speaker speech domain sequence representation; or... In the case where the second speaker extraction network module includes multiple trained second speaker extraction networks, the target speaker latent audio domain sequence representation and the first target speaker audio domain sequence representation, along with the first non-target speaker audio domain sequence representation, are simultaneously input into the second speaker extraction network connected to the first speaker extraction network in the second speaker extraction network module to obtain the intermediate target speaker audio domain sequence representation. For any second speaker extraction network other than the second speaker extraction network connected to the first speaker extraction network, the speech domain sequence representation, the target speaker latent speech domain sequence representation of the first non-target speaker speech domain sequence representation, and the intermediate target speaker speech domain sequence representation output by the previous second speaker extraction network are simultaneously input into the second speaker extraction network to obtain the intermediate target speaker speech domain sequence representation output by the second speaker extraction network. The intermediate target speaker audio domain sequence representation of the last second speaker extracted from the network output is used as the second speaker audio domain sequence representation.
5. The method according to claim 4, characterized in that, The second speaker suppression network module includes at least one trained second speaker suppression network; The step of simultaneously inputting the speech domain sequence representation, the non-target speaker latent speech domain sequence representation, and the first non-target speaker speech domain sequence representation into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation includes: When the second speaker suppression network module includes a trained second speaker suppression network, the speech domain sequence representation, the non-target speaker latent speech domain sequence representation, and the first non-target speaker speech domain sequence representation are simultaneously input into the second speaker suppression network to obtain the second non-target speaker speech domain sequence representation; or... In the case where the second speaker suppression network module includes multiple trained second speaker suppression networks, the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation are simultaneously input into the second speaker suppression network connected to the first speaker suppression network in the second speaker suppression network module to obtain the intermediate non-target speaker speech domain sequence representation. For any second speaker suppression network other than the second speaker suppression network connected to the first speaker suppression network, the speech domain sequence representation, the non-target speaker latent speech domain sequence representation of the first target speaker speech domain sequence representation, and the intermediate non-target speaker speech domain sequence representation output by the previous second speaker suppression network are simultaneously input into the second speaker suppression network to obtain the intermediate non-target speaker speech domain sequence representation output by the second speaker suppression network. The intermediate non-target speaker audio domain sequence representation of the last second speaker suppression network output is used as the second non-target speaker audio domain sequence representation.
6. The method according to claim 5, characterized in that, The first speaker extraction network, the second speaker extraction network, the first speaker suppression network, and the second speaker suppression network are all masking networks; The network structure of the masking network includes an input layer, multiple stacked blocks, a ReLU activation function layer, a first sigmoid activation function layer, and a first dot product output layer connected in sequence. The input layer is also connected to the first dot product output layer. Each of the stacked blocks includes multiple recurrent neural networks connected in sequence.
7. The method according to claim 3, characterized in that, The step of inputting the second target speaker's audio domain sequence representation into the trained speech decoding and transformation network to reconstruct the target speaker's time-domain speech signal, and inputting the second non-target speaker's audio domain sequence representation into the trained speech decoding and transformation network to reconstruct the non-target speaker's time-domain speech signal, includes: The second target speaker audio domain sequence representation is input into the trained speech decoding network to obtain the decoded second target speaker audio domain sequence representation; The decoded second target speaker's audio domain sequence representation is input into a trained inverse Fourier transform network to reconstruct the target speaker's time-domain speech signal; and... The second non-target speaker audio domain sequence representation is input into the trained speech decoding network to obtain the decoded second non-target speaker audio domain sequence representation. The decoded second non-target speaker audio domain sequence representation is input into the trained inverse Fourier transform network to reconstruct the non-target speaker's time-domain speech signal.
8. The method according to claim 7, characterized in that, The network structure of the speech decoding network includes the DenseNet model layer, the sub-pixel2D convolutional layer, the third LN normalization layer, the third 2-D convolutional layer, and the dual-path model layer connected in sequence. The dual-path model layer includes a fourth 2D convolutional layer, a fifth 2D convolutional layer, a tanh activation function layer, a second sigmoid activation function layer, a second dot product output layer, a sixth 2D convolutional layer, and a third sigmoid activation function layer.
9. A speech extraction device, characterized in that, The device includes: The acquisition module is used to acquire the audio domain sequence representation of the multi-speaker aliased speech data to be extracted and the voiceprint features of the voiceprint registration speech data of the target speaker; wherein, the multi-speaker aliased speech data includes the speech of the target speaker; The first input module is used to simultaneously input the audio domain sequence representation and the voiceprint features into a trained first speaker extraction network to obtain the audio domain sequence representation of the first target speaker, and to simultaneously input the audio domain sequence representation and the voiceprint features into a trained first speaker suppression network to obtain the audio domain sequence representation of the first non-target speaker. The second input module is used to simultaneously input the speech domain sequence representation and the target speaker latent speech domain sequence representation and the first target speaker speech domain sequence representation of the first non-target speaker speech domain sequence representation into the second speaker extraction network module to obtain the second target speaker speech domain sequence representation; and to simultaneously input the speech domain sequence representation and the non-target speaker latent speech domain sequence representation and the first non-target speaker speech domain sequence representation into the second speaker suppression network module to obtain the second non-target speaker speech domain sequence representation. The speech extraction and reconstruction module is used to input the speech domain sequence representation of the second target speaker into the trained speech decoding and transformation network to reconstruct the time domain speech signal of the target speaker, and to input the speech domain sequence representation of the second non-target speaker into the trained speech decoding and transformation network to reconstruct the time domain speech signal of the non-target speaker. The target speaker latent audio domain sequence representation is obtained by subtracting the first non-target speaker audio domain sequence representation from the audio domain sequence representation. The target speaker latent audio domain sequence representation includes the speech information lost by the network extracted by the first speaker and the first target speaker audio domain sequence representation. The non-target speaker latent audio domain sequence representation is obtained by subtracting the first target speaker audio domain sequence representation from the audio domain sequence representation. The non-target speaker latent audio domain sequence representation includes the speech information lost by the first speaker suppression network and the first non-target speaker audio domain sequence representation.
10. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute a speech extraction program stored in the memory to implement the speech extraction method according to any one of claims 1 to 8.