Speech processing method and apparatus, storage medium, and computer device
By employing a speech separation method based on feature extraction and neural network models, the problem of low consistency in speech separation systems under real-world conditions is solved, achieving higher accuracy in speech processing.
Patent Information
- Application Number
- CN202411415239.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-10-11
AI Technical Summary
In real-world environments, existing speech separation systems suffer from low consistency in speech separation due to missing labeled data, making them unable to effectively separate blind sources.
By acquiring fused speech data for feature extraction, and using trained feature perceptrons and characteristic perceptrons to determine speech feature masks and features, speech separation is performed in conjunction with a decoder, reducing the difference between the separated speech and the original speech.
It improves the speech consistency of the speech separation system in real-world scenarios and enhances the accuracy of speech processing.
Smart Images

Figure CN119479683B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a speech processing method, apparatus, storage medium and computer equipment. Background Technology
[0002] In modern life, people are often surrounded by multiple sounds, such as mixed speech from multiple speakers, background noise, and music. Selective attention mechanisms enable people to focus on the voice of a target speaker in complex noisy environments. With the rapid advancement and widespread application of artificial intelligence (AI) technology, the Blind Source Separation (BSS) algorithm has been proposed. It aims to separate different sound sources into independent sound sources from mixed speech from multiple speakers. High-quality speech separation systems can play an important role in teleconferencing systems, voice recording systems, hearing aids and hearing enhancement devices, smart homes, and security monitoring.
[0003] In related technologies, although AI technology has significantly improved the performance of BSS tasks, the lack of labeled data in real-world environments makes it impossible to train deep learning. Consequently, when speech separation systems perform blind source separation in real-world scenarios, the consistency between the separated speech and the original speech is low. Summary of the Invention
[0004] The main objective of this application is to provide a speech processing method, apparatus, storage medium, and computer device that can reduce the difference between the separated speech and the original speech, improve consistency, and thus enhance the accuracy of speech processing.
[0005] In a first aspect, embodiments of this application provide a voice processing method, including:
[0006] Acquire fused speech data, extract features from the fused speech data to obtain fused speech features, wherein the fused speech data includes speech data generated after mixing at least two sound sources;
[0007] The fused speech features are input into the trained feature perceptron to determine the speech feature mask for each speech source corresponding to the speech data.
[0008] Based on each of the aforementioned speech feature masks, the speech feature representation of the speech data corresponding to each sound source is determined from the fused speech features;
[0009] Each of the aforementioned speech feature representations is input into the trained feature perceptron to determine the speech feature mask for the speech data corresponding to each sound source.
[0010] Based on each of the speech feature masks, the speech features corresponding to each sound source are determined from the fused speech features;
[0011] Each of the aforementioned speech features is input into the trained decoder to decode the speech data corresponding to each of the aforementioned sound sources.
[0012] Secondly, embodiments of this application provide a voice processing device, including:
[0013] A feature extraction unit is used to acquire fused speech data, perform feature extraction on the fused speech data to obtain fused speech features, wherein the fused speech data includes speech data generated after mixing at least two sound sources;
[0014] The feature perception unit is used to input the fused speech features into the trained feature perceptron to determine the speech feature mask of the speech data corresponding to each sound source.
[0015] The first determining unit is used to determine the speech feature representation of each speech source corresponding to the speech data from the fused speech features based on each of the speech feature masks.
[0016] The feature perception unit is used to input each of the speech feature representations into the trained feature perceptron to determine the speech feature mask of the speech data corresponding to each sound source.
[0017] The second determining unit is used to determine the speech features of each sound source corresponding to the speech data from the fused speech features based on each of the speech feature masks;
[0018] The decoding unit is used to input each of the speech features into the trained decoder and decode the speech data corresponding to each of the sound sources.
[0019] In some embodiments, the apparatus further includes:
[0020] Obtain the first sample fused speech features of the first sample fused speech data, and the first sample real speech data of each first sound source included in the first sample fused speech data;
[0021] The first sample fused speech features are input into the feature perceptron to be trained to determine the first sample speech feature mask of the speech data corresponding to each first sound source.
[0022] The first predicted feature of the speech data corresponding to each first sound source is determined by using the speech feature mask of each first sample and the predictor to be trained.
[0023] Each of the first sample's real speech data is input into the trained feature extraction model to extract the first real feature of the speech data corresponding to each first sound source;
[0024] The first loss value is determined based on the first predicted feature and the corresponding first true feature of each first sound source;
[0025] When the first loss value is greater than the first preset loss value, the network parameters of the feature perceptron to be trained and the predictor to be trained are adjusted according to the first loss value, and the process of inputting the first sample fused speech features into the feature perceptron to be trained and determining the first sample speech feature mask of the speech data corresponding to each first sound source is repeated until the first loss value is less than or equal to the first preset loss value, so as to obtain the feature perceptron and the predictor after preliminary training.
[0026] In some embodiments, the apparatus includes:
[0027] Based on each first sample speech feature mask, the first sample speech feature representation of each sound source corresponding to the speech data is determined from the first sample fused speech features;
[0028] Each of the first sample speech features is input into the predictor to be trained to obtain the first predicted feature of the speech data corresponding to each first sound source.
[0029] In some embodiments, the apparatus includes:
[0030] Determine the similarity between the first predicted feature and the corresponding first true feature of each first sound source;
[0031] Select the target first sound source with the highest similarity from each first sound source;
[0032] The similarity between the first predicted feature of the target first sound source and the first predicted feature of each other first sound source is obtained to obtain multiple first predicted similarities, wherein the other first sound sources are first sound sources other than the target first sound source among the first sound sources;
[0033] A first loss value is determined based on the target similarity between the first predicted feature of the target first sound source and the corresponding first real feature, as well as multiple first predicted similarities.
[0034] In some embodiments, the apparatus further includes:
[0035] The second sample fused speech features of the second sample fused speech data are obtained, as well as the second sample real speech data of each second sound source included in the second sample fused speech data;
[0036] The second sample fused speech features are input into the pre-trained feature perceptron to determine the second sample speech feature mask for each second sound source corresponding to the speech data.
[0037] Based on each second sample speech feature mask, the second sample speech feature representation of the speech data corresponding to each second sound source is determined from the second sample fused speech features;
[0038] Each second sample speech feature representation is input into the feature perceptron to be trained to determine the sample speech feature mask of each second sound source corresponding to the speech data.
[0039] Based on each sample speech feature mask, the sample speech features of each second sound source corresponding to the speech data are determined from the second sample fused speech features;
[0040] Each of the sample speech features is input into the decoder to be trained to decode the second sample predicted speech data corresponding to each second sound source;
[0041] The second loss value is determined based on the second sample predicted speech data and the corresponding second sample real speech data of each second sound source.
[0042] When the second loss value is greater than the second preset loss value, the network parameters of the feature perceptron to be trained are adjusted according to the second loss value, and the process returns to the step of inputting the second sample fused speech features into the pre-trained feature perceptron and determining the second sample speech feature mask for each second sound source speech data, until the second loss value is less than or equal to the second preset loss value, thus obtaining the trained feature perceptron, the trained feature perceptron and the trained decoder.
[0043] In some embodiments, the apparatus includes:
[0044] Based on the predicted speech data of the second sample of each second sound source and the corresponding real speech data of the second sample, determine the speech distortion degree of each second sound source.
[0045] The target speech distortion with the smallest speech distortion is selected from each of the stated speech distortion values, and the target speech distortion value is determined as the second loss value.
[0046] In some embodiments, the apparatus includes:
[0047] For each second sound source, calculate the similarity energy value of the similar parts in the corresponding second sample predicted speech data and the corresponding second sample real speech data;
[0048] Calculate the difference energy value of the difference portion between the corresponding second sample predicted speech data and the corresponding second sample real speech data;
[0049] The speech distortion degree corresponding to each second sound source is determined based on the energy ratio of the similar energy value to the difference energy value.
[0050] Thirdly, embodiments of this application provide a storage medium that stores multiple instructions adapted for loading by a processor to execute any of the above-mentioned speech processing methods.
[0051] Fourthly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the voice processing method as described above.
[0052] In this embodiment, fused speech data is acquired, and features are extracted from the fused speech data to obtain fused speech features. The fused speech data includes speech data generated by mixing at least two sound sources. The fused speech features are input to a trained feature perceptron to determine a speech feature mask for each sound source's corresponding speech data. Based on each speech feature mask, a speech feature representation for each sound source's corresponding speech data is determined from the fused speech features. Each speech feature representation is input to a trained feature perceptron to determine a speech feature mask for each sound source's corresponding speech data. Based on each speech feature mask, a speech feature for each sound source's corresponding speech data is determined from the fused speech features. Each speech feature is input to a trained decoder to decode the speech data corresponding to each sound source. Compared to related technologies where blind source separation is impossible or the separated speech has low consistency due to missing labeled data during training, this embodiment uses a feature perceptron to perceive the speech features of speech data corresponding to different sound sources, thereby performing blind source separation based on speech features. This reduces the difference between the separated speech and the original speech, improves consistency, and ultimately enhances the accuracy of speech processing.
[0053] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1This is a schematic diagram of a scenario for the voice processing system provided in an embodiment of this application.
[0056] Figure 2 This is a flowchart illustrating the speech processing method provided in an embodiment of this application.
[0057] Figure 3 This is a schematic diagram of the architecture of the speech processing method provided in the embodiments of this application.
[0058] Figure 4 This is a schematic diagram of the training architecture for the feature perceptron in this application.
[0059] Figure 5 This is a schematic diagram of the training architecture for the feature perceptron in this application.
[0060] Figure 6 This is a schematic diagram of the structure of the voice processing device provided in the embodiments of this application.
[0061] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0062] To enable those skilled in the art to better understand the solutions of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0063] It should be noted that while some processes described in the specification, claims, and accompanying drawings contain multiple steps that appear in a specific order, it should be clearly understood that these steps may not be performed in the order they appear herein, or may be performed in parallel. The step numbers are merely used to distinguish different steps and do not represent any particular order of execution. Furthermore, descriptions such as "first," "second," or "objective" in this document are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0064] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:
[0065] Artificial Intelligence (AI) is an important component of the discipline of intelligence, attempting to understand the nature of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI is a very broad science, encompassing robotics, speech recognition, image recognition, natural language processing, expert systems, machine learning, computer vision, and more.
[0066] An encoder is a component or algorithm that transforms input data into a specific encoded or feature representation. In computer science and machine learning, the role of an encoder is to compress input information, extract key features, or perform some form of transformation for subsequent processing, analysis, or transmission. The design and functionality of an encoder depend on the specific application scenario and task requirements, and its purpose is to make data easier to process, classify, generate, or compare with other data in a new representation.
[0067] For example, in deep learning, an encoder, typically part of a neural network, maps input images, text, or other data into a low-dimensional feature space, thus achieving an efficient representation of the data. These encoded features usually contain key information from the input data while removing redundant or unimportant parts.
[0068] A decoder is a neural network module that corresponds to an encoder. Its main function is to transform the encoded low-dimensional representation (usually generated by the encoder) back into the original data space or generate new data samples. Its main functions and characteristics include: 1. Data Reconstruction: A primary function of the decoder is to reconstruct the original data from the encoded low-dimensional representation. For example, in an autoencoder, the encoder compresses the input data into a low-dimensional code, while the decoder attempts to recover an output from this code that is as close as possible to the original input. By training the autoencoder, the decoder learns how to map the encoded representation back to the original data space, thus achieving data reconstruction. 2. Generating New Data: The decoder can also be used to generate new data samples. In models such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs), the decoder is often combined with random noise or latent variables to generate new data samples similar to the training data. For example, in a VAE, the encoder maps the input data to a latent space, and then the decoder samples from this latent space and generates new data samples. By adjusting the values of the latent variables, the features of the generated data can be controlled. 3. Working in Collaboration with the Encoder: The decoder typically works closely with the encoder. The encoder transforms the input data into a low-dimensional representation, while the decoder transforms this low-dimensional representation back into the original data space or generates new data. This encoder-decoder architecture is highly effective in many tasks, such as image generation, machine translation, and speech synthesis. During training, the encoder and decoder are typically optimized together to minimize reconstruction errors or the discrepancy between generated and real data.
[0069] Currently, in real-world environments, people are often surrounded by various sounds, such as mixed speech from multiple speakers, background noise, and music. Selective attention mechanisms enable people to focus on the voice of the target speaker in complex noisy environments. In some scenarios, such as remote conferencing, to ensure that participants can clearly hear the speech of the target participant, the Blind Source Separation (BSS) algorithm is used to separate the speech of the target participant in a noisy environment from the noisy mixed speech, allowing participants to focus on the voice of the target participant.
[0070] However, due to the lack of labeled data during the training process, the methods in related technologies result in low consistency between the separated speech and the original speech when performing blind source separation in real-world scenarios.
[0071] To address the aforementioned problems, this application embodiment acquires fused speech data, extracts features from the fused speech data to obtain fused speech features, wherein the fused speech data includes speech data generated by mixing at least two sound sources; the fused speech features are input into a trained feature perceptron to determine a speech feature mask for each sound source's corresponding speech data; based on each speech feature mask, a speech feature representation for each sound source's corresponding speech data is determined from the fused speech features; each speech feature representation is input into a trained feature perceptron to determine a speech feature mask for each sound source's corresponding speech data; based on each speech feature mask, a speech feature for each sound source's corresponding speech data is determined from the fused speech features; and each speech feature is input into a trained decoder to decode the speech data corresponding to each sound source. Compared to related technologies where blind source separation is impossible or the separated speech has low consistency due to missing labeled data during training, this application embodiment uses a feature perceptron to perceive the speech features of speech data corresponding to different sound sources, thereby performing blind source separation based on speech features, reducing the difference between the separated speech and the original speech, and improving consistency. Please refer to the following specific embodiments for details.
[0072] Please see Figure 1 , Figure 1 This is a schematic diagram of a voice processing system provided in an embodiment of this application. It includes a terminal 140, an Internet 130, a gateway 120, a computer device 110, etc.
[0073] Terminal 140 includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, and aircraft. Furthermore, it can be a single device or a collection of multiple devices. Terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data.
[0074] Computer equipment refers to a computer system that can provide certain services to terminal 140. Compared to ordinary terminal 140, computer equipment 110 has higher requirements in terms of stability, security, and performance. Computer equipment 110 can be a single high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a single high-performance computer (e.g., a virtual machine), or a combination of portions of multiple high-performance computers (e.g., virtual machines).
[0075] Gateway 120, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from terminal 140 to computer device 110 are routed through gateway 120 to the corresponding computer device 110. Messages sent from computer device 110 to terminal 140 are also routed through gateway 120 to the corresponding terminal 140.
[0076] The data transmission method of this disclosure embodiment can be implemented in computer device 110.
[0077] It should be noted that, Figure 1 The schematic diagram of the speech processing system shown is merely an example. The speech processing system and scenario described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of image processing technology and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0078] In this embodiment, the description will be from the perspective of a voice processing device, which can be integrated into a computer device that has a storage unit and a microprocessor and thus computing power.
[0079] Please see Figure 2 , Figure 2 This is a flowchart illustrating a speech processing method provided in an embodiment of this application. The speech processing method includes:
[0080] In step 201, fused speech data is acquired, and feature extraction is performed on the fused speech data to obtain fused speech features. The fused speech data includes speech data generated after mixing at least two sound sources.
[0081] Fusion speech data is speech data generated by mixing at least two sound sources. These sound sources can be the voices of different people or other audio recordings involving at least two people speaking. For example, in a multi-person dialogue scenario, fusion speech data is an audio signal where the voices of multiple speakers are mixed together. Fusion speech data can be acquired from the actual environment using audio acquisition devices such as microphones, or it can be read from existing audio files.
[0082] Specifically, the feature extraction process for the fused speech data is performed by a trained encoder, which is mainly used to convert the original fused speech data into more abstract and representative fused speech features.
[0083] Specifically, the feature extraction process of the encoder can be described as follows: a convolutional kernel slides across the input data, performing weighted summation on local regions to obtain a new feature map. Different convolutional kernels can extract different features, such as components of different frequencies or different speech patterns. As the convolutional layers deepen, more advanced and abstract features can be extracted. The feature map output by the convolutional layer is downsampled to reduce the feature dimensionality while retaining the main features.
[0084] In step 202, the fused speech features are input into the trained feature perceptron to determine the speech feature mask for each sound source corresponding to the speech data.
[0085] The feature perceptron is a trained neural network model whose function is to extract a speech feature mask for each sound source from the fused speech features. Features include syllables, prosody, vocabulary, and syntax information in the audio data. Different people have different speaking characteristics, so the feature perceptron perceives the speech features of each sound source to form a speech feature mask.
[0086] Specifically, the input to the feature perceptron is the fused speech features, and the output is a speech feature mask with the same number of sound sources. Each speech feature mask is a vector or matrix with the same dimension as the fused speech features, representing the part of the fused speech features that is related to the speech features of a certain sound source.
[0087] The process of determining the speech feature mask is as follows: After the fused speech features are input into the feature perceptron, the feature perceptron performs a series of nonlinear transformations and calculations on the fused speech features. These transformations and calculations typically include convolution operations, fully connected layers, activation functions, etc. By learning the speech feature patterns in the training data, the feature perceptron gradually adjusts its parameters so that the output speech feature mask can accurately represent the speech features of each sound source.
[0088] For example, if the speech feature of sound source A is reduplication, and the speech feature of sound source B is rhyme, then for sound source A, the corresponding speech feature mask output by the feature perceptron will have a larger value in the reduplication part, and a smaller value in other prosodic parts; similarly, for sound source B, the corresponding speech feature mask output by the feature perceptron will have a larger value in the rhyme part, and a smaller value in other prosodic parts. In this way, the trained feature perceptron can quickly determine the speech feature mask for each sound source based on the input fused speech features.
[0089] In step 203, based on each of the speech feature masks, the speech feature representation of the speech data corresponding to each sound source is determined from the fused speech features.
[0090] In step S202, after determining the speech feature mask for each sound source, the speech feature representation of the corresponding speech data for each sound source can be determined from the fused speech features. The speech feature representation is an abstract representation of the speech data of each sound source in terms of specific speech features. It contains important feature information of the speech data of that sound source, such as syllables, prosody, vocabulary, and syntax. This information is crucial for subsequent speech separation and recognition.
[0091] Specifically, the method for determining the speech feature representation of each sound source's corresponding speech data is as follows: Since the speech feature mask is a vector or matrix with the same dimension as the fused speech features, and each element in the speech feature mask represents the weight of the corresponding element in the fused speech features, the result tensor can be obtained by calculating the tensor product of the speech feature mask of each sound source and the fused speech features. This result tensor is the speech feature representation of the speech data corresponding to each sound source. It highlights the parts of the fused speech features that are related to the speech features of that sound source, while suppressing other irrelevant parts.
[0092] In step 204, each of the speech feature representations is input into the trained feature perceptron to determine the speech feature mask for each speech source corresponding to the speech data.
[0093] The feature perceptron is also a trained neural network model. Its role is to further extract the speech feature mask of each sound source from the speech feature representation, so as to ensure that the speech features determined by the speech feature mask are not distorted after being decoded into speech data.
[0094] The input to a feature perceptron is a speech feature representation for each sound source, and the output is a speech feature mask equal to the number of sound sources. Each speech feature mask is a vector or matrix with the same dimension as the speech feature representation, representing the portion of the speech feature representation related to a specific speech feature of a particular sound source. Specifically, similar to the feature perceptron, the feature perceptron also performs a series of nonlinear transformations and calculations on the input speech feature representation. The feature perceptron enables the output speech feature mask to more accurately represent the speech features of each sound source.
[0095] In step 205, the speech features corresponding to each sound source are determined from the fused speech features based on each of the speech feature masks.
[0096] In step S204, after determining the speech feature mask for each sound source, the speech features corresponding to the speech data of each sound source are determined again from the fused speech features. The method is similar to determining the speech feature representation from the fused speech features; it involves performing a tensor product calculation between each speech feature mask and the fused speech features, implementing element-wise multiplication between each speech feature mask and the fused speech features. The result is the speech feature of the speech data corresponding to each sound source, which more accurately represents the information of each sound source's speech data regarding specific speech features.
[0097] Specifically, speech features are a more detailed and specific representation of the speech data of each sound source. They include various characteristic information of the speech data from that source, such as frequency, amplitude, and phase. This information is crucial for the final speech separation and decoding.
[0098] In step 206, each of the speech features is input to the trained decoder to decode the speech data corresponding to each sound source.
[0099] The decoder is a trained neural network model whose function is to decode the speech features of each sound source into corresponding speech data. The input of the decoder is the speech features of each sound source, and the output is the separated speech data corresponding to each sound source.
[0100] Specifically, the decoding process of the decoder includes: the decoder performs a series of inverse transforms and calculations on the input speech features to recover the original speech signal. These inverse transforms and calculations typically include deconvolution operations, upsampling, activation functions, etc. The decoder learns the relationship between the speech features in the training data and the original speech data, gradually adjusting its parameters so that the output speech data is as close as possible to the original sound source speech data. After processing by the decoder, the separated speech data corresponding to each sound source is finally obtained, achieving the goal of speech separation.
[0101] Therefore, by acquiring fused speech data, feature extraction is performed on the fused speech data to obtain fused speech features. The fused speech data includes speech data generated after mixing at least two sound sources. The fused speech features are input into a trained feature perceptron to determine the speech feature mask corresponding to each sound source. Based on each speech feature mask, a speech feature representation of each sound source corresponding to the speech data is determined from the fused speech features. Each speech feature representation is input into a trained feature perceptron to determine the speech feature mask corresponding to each sound source. Based on each speech feature mask, a speech feature of each sound source corresponding to the speech data is determined from the fused speech features. Each speech feature is input into a trained decoder to decode the speech data corresponding to each sound source. By using a feature perceptron to perceive the speech features of speech data corresponding to different sound sources, blind source separation is performed based on the speech features, reducing the difference between the separated speech and the original speech, improving consistency, and thus improving the accuracy of speech processing.
[0102] For details, please refer to Figure 3 , Figure 3 This is a schematic diagram of the architecture of the speech processing method provided in this application embodiment. When fused speech data is acquired, a trained encoder extracts features from the fused speech data to obtain fused speech features. A trained feature perceptron analyzes the fused speech features to determine the speech feature mask for each sound source. Based on these speech feature masks, the speech feature representation of each sound source can be determined from the fused speech features. Then, each speech feature representation is input to the trained feature perceptron, which determines the speech feature mask for each sound source. Based on these speech feature masks, the speech features for each sound source are then determined from the fused speech features. Finally, the speech features of each sound source are input to the trained decoder, which decodes these speech features to separate the speech data corresponding to each sound source. This process gradually extracts the independent speech information of each sound source from the fused speech, achieving effective separation of speech data.
[0103] In some implementations, before inputting the fused speech features into a trained feature perceptron to determine the speech feature mask for each sound source's corresponding speech data, the method further includes:
[0104] (1) Obtain the first sample fused speech features of the first sample fused speech data, and the first sample real speech data of each first sound source included in the first sample fused speech data;
[0105] (2) Input the first sample fused speech features into the feature perceptron to be trained, and determine the first sample speech feature mask of each first sound source corresponding to the speech data;
[0106] (3) Determine the first prediction feature of the speech data corresponding to each first sound source by using the speech feature mask of each first sample and the predictor to be trained;
[0107] (4) Input each of the first sample real speech data into the trained feature extraction model to extract the first real feature of each first sound source corresponding to the speech data;
[0108] (5) Determine the first loss value based on the first predicted feature and the corresponding first true feature of each first sound source;
[0109] (6) When the first loss value is greater than the first preset loss value, adjust the network parameters of the feature perceptron to be trained and the predictor to be trained according to the first loss value, and return to execute the step of inputting the first sample fused speech features into the feature perceptron to be trained and determining the first sample speech feature mask of each first sound source corresponding to the speech data until the first loss value is less than or equal to the first preset loss value, and obtain the feature perceptron and the predictor after preliminary training.
[0110] Please refer to Figure 4 , Figure 4 This is a schematic diagram of the training architecture for the feature perceptron in this application. The first sample fused speech data is speech data mixed from multiple first sound sources. The encoder to be trained obtains the first sample fused speech features, which are an abstract representation of the fused speech data. At the same time, it is also necessary to obtain the individual first sample real speech data for each first sound source as the true value for subsequent comparison and loss calculation.
[0111] The feature perceptron to be trained is a neural network model whose purpose is to identify speech features associated with each first sound source from the fused speech features. After inputting the first sample of fused speech features, the feature perceptron undergoes a series of calculations and learning processes, outputting a mask of the first sample speech features corresponding to each first sound source. This mask can filter out the parts of the fused speech features that are related to a specific first sound source.
[0112] By using the speech feature mask of each first sample and the predictor to be trained, the first predicted feature of the speech data corresponding to each first sound source is determined. The obtained first sample speech feature mask is then used in conjunction with the predictor to be trained for further processing. Based on the mask and other information, the predictor predicts the first predicted feature of the speech data corresponding to each first sound source. This predicted feature is an estimate of the true features of the first sound source.
[0113] Each first sample of real speech data is input into the trained feature extraction model to extract the first real feature of the speech data corresponding to each first sound source.
[0114] For the first sample of real speech data from each first sound source, a pre-trained feature extraction model is used to extract its true features. These true features serve as the standard for subsequent loss calculations, measuring the accuracy of the predicted features.
[0115] A first loss value is determined based on the first predicted feature and the corresponding first true feature of each first sound source. The first loss value can be calculated using a loss function (such as mean squared error) by comparing the first predicted feature and the first true feature of each first sound source. The loss value reflects the degree of difference between the predicted feature and the true feature.
[0116] When the first loss value is greater than the first preset loss value, the network parameters of the feature perceptron and the predictor to be trained are adjusted according to the first loss value, and the process returns to the step of inputting the first sample fused speech features into the feature perceptron to be trained and determining the first sample speech feature mask of the speech data corresponding to each first sound source, until the first loss value is less than or equal to the first preset loss value, and the feature perceptron and the predictor after preliminary training are obtained.
[0117] If the initial loss value is greater than the preset loss value, it indicates that the performance of the current feature perceptron and predictor is not good enough. At this point, the network parameters of the encoder, feature perceptron, and predictor are adjusted using algorithms such as backpropagation based on the loss value, enabling them to learn and predict better. This process is then iteratively optimized until the initial loss value is less than or equal to the preset loss value. At this point, the encoder, feature perceptron, and predictor are preliminarily trained and can extract the features of each first sound source from the fused speech data relatively accurately.
[0118] Therefore, by inputting the first sample of real speech data into the trained feature extraction model, the first real feature is extracted and used as a standard to measure prediction accuracy. By comparing the first predicted feature with the first real feature, a first loss value is calculated using a loss function, which reflects the degree of difference between the prediction and the reality. When the first loss value is greater than a first preset loss value, the network parameters of the feature perceptron and predictor are adjusted through backpropagation, iteratively optimizing until the loss value is less than or equal to the preset value. The resulting pre-trained feature perceptron and trained predictor can extract the features of each first sound source from the fused speech data relatively accurately, which helps improve the feature perception capability of the feature perceptron.
[0119] In some implementations, a first predicted feature of the speech data corresponding to each first sound source is determined using a mask of each first sample speech feature and a predictor to be trained, including:
[0120] (3.1) Based on each first sample speech feature mask, determine the first sample speech feature representation of each sound source corresponding to the speech data from the first sample fused speech features;
[0121] (3.2) Input the speech feature representation of each first sample into the predictor to be trained to obtain the first predicted feature of the speech data corresponding to each first sound source.
[0122] The first predicted feature of the speech data corresponding to each sound source is obtained by first calculating the tensor product of the speech feature mask of each first sample and the fused speech feature of the first sample, and then inputting the speech feature representation of the first sample to the predictor to be trained, so as to obtain the first predicted feature of the speech data corresponding to each first sound source predicted by the predictor.
[0123] In some implementations, determining the first loss value based on the first predicted feature and the corresponding first true feature of each first sound source includes:
[0124] (5.1) Determine the similarity between the first predicted feature and the corresponding first true feature of each first sound source;
[0125] (5.2) Select the target first sound source with the highest similarity from each first sound source;
[0126] (5.3) Obtain the similarity between the first predicted feature of the target first sound source and the first predicted feature of each other first sound source to obtain multiple first predicted similarities, wherein the other first sound sources are first sound sources other than the target first sound source among the first sound sources;
[0127] (5.4) Determine a first loss value based on the target similarity between the first predicted feature of the target first sound source and the corresponding first real feature, as well as multiple first predicted similarities.
[0128] This involves measuring the similarity between the first predicted feature of each first sound source obtained through the predictor to be trained and the first real feature extracted from actual real speech data by the trained feature extraction model. Various similarity metrics can be used, such as cosine similarity and Pearson correlation coefficient. This similarity value reflects how close the prediction is to the reality; the higher the similarity, the more accurate the prediction.
[0129] After calculating the similarity between the first predicted feature and the corresponding first true feature of each first sound source, the sound source with the highest similarity is selected as the target first sound source from all first sound sources. The purpose of this step is to determine which sound source's prediction result is closest to the true situation in the current state, so as to facilitate further analysis of information related to that sound source.
[0130] After identifying the target primary sound source, the similarity between the first predicted feature of the target primary sound source and the first predicted features of all other non-target primary sound sources is calculated. This yields multiple first predicted similarities. These similarity values reflect the degree of difference in predicted features between the target primary sound source and other sound sources, helping to assess the uniqueness of the target primary sound source and its distinguishability from other sound sources.
[0131] Finally, the first loss value is determined by combining the similarity between the predicted features of the target first sound source and the true features (target similarity) and the predicted similarity between the target first sound source and other sound sources (multiple first predicted similarities). One possible approach is to design a loss function that considers the relationship between target similarity and multiple first predicted similarities, maximizing target similarity while minimizing the predicted similarity between the target first sound source and other sound sources, thus ensuring the accuracy and uniqueness of the target first sound source's predictions. By adjusting the parameters of the feature perceptron and predictor to be trained, this loss value is minimized, thereby improving the model's performance.
[0132] Specifically, the calculation method for the first loss value can be referred to the following formula:
[0133]
[0134] The loss function employed is Noise Contrastive Estimation (NCE). Its aim is to learn effective feature representations by maximizing the similarity between positive sample pairs while minimizing the similarity between negative sample pairs. i The primary and most authentic characteristic of the target sound source. For the first predictive characteristic of the target sound source, ω is the temperature coefficient. It is any similarity function. Q represents the target similarity between the first predicted feature of the first sound source and the corresponding first true feature. U The first predicted feature for each first sound source, To calculate the similarity between the first predicted feature of the target first sound source and the first predicted feature of each first sound source, the value of the similarity function is mapped to the positive real number domain using the exponential function exp(), so that sample pairs with high similarity have larger exponential values. This is the sum of the similarities between the first predicted feature of the target first sound source and the first predicted features of every other first sound source, after an exponential transformation. It represents the value of the entire loss function. By taking the logarithm, division can be transformed into subtraction, facilitating calculation and optimization. The goal of the loss function is to minimize its value, which means that the similarity between positive sample pairs should be as high as possible, while the similarity with negative sample pairs should be as low as possible.
[0135] In some implementations, before inputting each of the speech feature representations into a trained feature perceptron to determine the speech feature mask for each speech source, the method further includes:
[0136] (7) Obtain the second sample fused speech features of the second sample fused speech data, and the second sample real speech data of each second sound source included in the second sample fused speech data;
[0137] (8) Input the second sample fused speech features into the feature perceptron after preliminary training to determine the second sample speech feature mask of each second sound source corresponding to the speech data;
[0138] (9) Based on each second sample speech feature mask, determine the second sample speech feature representation of the speech data corresponding to each second sound source from the second sample fused speech features;
[0139] (10) Input each second sample speech feature representation into the feature perceptron to be trained to determine the sample speech feature mask of each second sound source corresponding to the speech data;
[0140] (11) Based on each sample speech feature mask, determine the sample speech features of each second sound source corresponding to the speech data from the second sample fused speech features;
[0141] (12) Input each of the sample speech features into the decoder to be trained, and decode the second sample predicted speech data corresponding to each second sound source;
[0142] (13) Determine the second loss value based on the second sample predicted speech data and the corresponding second sample real speech data of each second sound source;
[0143] (14) When the second loss value is greater than the second preset loss value, adjust the network parameters of the feature perceptron to be trained according to the second loss value, and return to execute the step of inputting the second sample fused speech features into the pre-trained feature perceptron and determining the second sample speech feature mask of each second sound source speech data until the second loss value is less than or equal to the second preset loss value, and obtain the trained feature perceptron, the trained feature perceptron and the trained decoder.
[0144] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the training architecture for the feature perceptron in this application. During the training process of the feature perceptron, a new set of sample data is prepared. The second sample fused speech data is speech data resulting from the mixing of multiple second sound sources. Its fused speech features are extracted using a specific method, while simultaneously acquiring the real speech data of each individual second sound source to provide a foundation for subsequent training and evaluation. The feature perceptron, initially trained in the previous steps, is used to process the new fused speech features. Through learning and computation, the feature perceptron generates a corresponding speech feature mask for each second sound source, used to filter out specific parts of the fused speech features related to that sound source.
[0145] Using the obtained speech feature mask, speech feature representations for each second sound source are extracted from the fused speech features. This representation highlights the features associated with that sound source, providing more specific information for further analysis and processing.
[0146] The feature perceptron to be trained takes speech feature representations as input and, through learning and computation, determines a sample speech feature mask for each second sound source. This mask is more precise for the specific speech features of each sound source than a speech feature mask.
[0147] Sample speech features are extracted from the fused speech features using sample speech feature masks, and these features more accurately reflect the unique characteristics of each speech source.
[0148] The decoder to be trained takes the sample speech features as input and attempts to recover the predicted speech data of each second sound source through inverse transformation and calculation, thus realizing the conversion from features to speech.
[0149] A second loss value is calculated by comparing the predicted speech data and the real speech data for each second sound source using an appropriate loss function. This loss value reflects the performance of the decoder and feature perceptron, i.e., the degree of difference between the predicted speech and the real speech.
[0150] If the second loss value is greater than the preset value, it indicates that the current performance of the feature perceptron and decoder is not good enough. Based on the loss value, the network parameters of the feature perceptron to be trained are adjusted using algorithms such as backpropagation. Then, subsequent steps are repeated, iteratively optimizing until the loss value is less than or equal to the preset value. At this point, the trained encoder, feature perceptron, and decoder are obtained, which can more accurately separate the speech data of each second sound source from the fused speech data, ensuring that the separated speech data is not distorted or that the possibility of distortion is reduced.
[0151] In some implementations, a second loss value is determined based on the predicted speech sample data of each second sound source and the corresponding second sample real speech data, including:
[0152] (13.1) Determine the speech distortion degree corresponding to each second sound source based on the second sample predicted speech data and the corresponding second sample real speech data of each second sound source;
[0153] (13.2) Select the target speech distortion with the smallest speech distortion from each of the speech distortion values, and determine the target speech distortion value as the second loss value.
[0154] In this training method, speech distortion is used as the loss function for the feature perceptron. Therefore, based on the predicted speech data of the second sample for each second sound source and the corresponding real speech data of the second sample, the speech distortion for each second sound source is determined. After obtaining the speech distortion for each second sound source, the minimum speech distortion is selected as the target speech distortion. This is because during training, the goal is to find parameter settings that best approximate the overall prediction results to reality. This minimum speech distortion is determined as the second loss value, used to subsequently evaluate model performance and adjust model parameters. If the second loss value is greater than a second preset loss value, it indicates that the model's prediction results differ significantly from the real data, requiring further adjustment of model parameters to reduce distortion. If the second loss value is less than or equal to the second preset loss value, it indicates that the model's performance is good, and training can be stopped.
[0155] In some implementations, the speech distortion degree corresponding to each second sound source is determined based on the second sample predicted speech data and the corresponding second sample real speech data of each second sound source, including:
[0156] (13.1.1) For each second sound source, calculate the similarity energy value of the similar part between the corresponding second sample predicted speech data and the corresponding second sample real speech data;
[0157] (13.1.2) Calculate the difference energy value of the difference part between the corresponding second sample predicted speech data and the corresponding second sample real speech data;
[0158] (13.1.3) Determine the speech distortion degree corresponding to each second sound source based on the energy ratio of the similar energy value and the difference energy value.
[0159] The speech distortion for each second sound source can be calculated using the following formula:
[0160]
[0161] Among them, L SI-SDRIt is the scale-invariant signal-to-distortion ratio loss function, which is used to measure the predicted speech data from the second sample. The degree of difference between the second sample real speech data s and the numerator, where the numerator is... It is the energy of the normalized projection of the predicted speech data of the second sample onto the direction of the real speech data of the second sample, that is, the similarity energy value of the similar part, and the denominator is... This is the energy difference between the normalized projection and the predicted signal, i.e., the difference energy value of the difference portion. Using this formula, the speech distortion corresponding to each second sound source can be calculated.
[0162] For details on the implementation of each of the above steps, please refer to the previous examples, which will not be repeated here.
[0163] To facilitate better implementation of the speech processing method provided in the embodiments of this application, the embodiments of this application also provide an apparatus based on the above speech processing method. The meanings of the terms used are the same as in the above speech processing method, and specific implementation details can be found in the descriptions in the method embodiments.
[0164] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a speech processing device provided in an embodiment of this application. The speech processing device is applied to a computer device. The speech processing device may include a feature extraction unit 601, a feature perception unit 602, a first determination unit 603, a feature perception unit 604, a second determination unit 605, and a decoding unit 606, etc.
[0165] The feature extraction unit 601 is used to acquire fused speech data, perform feature extraction on the fused speech data, and obtain fused speech features. The fused speech data includes speech data generated after mixing at least two sound sources.
[0166] The feature perception unit 602 is used to input the fused speech features into the trained feature perceptron to determine the speech feature mask of the speech data corresponding to each sound source.
[0167] The first determining unit 603 is used to determine the speech feature representation of each speech source corresponding to the speech data from the fused speech features based on each of the speech feature masks.
[0168] The feature perception unit 604 is used to input each of the speech feature representations into the trained feature perceptron to determine the speech feature mask of the speech data corresponding to each sound source.
[0169] The second determining unit 605 is used to determine the speech features of each sound source corresponding to the speech data from the fused speech features based on each of the speech feature masks.
[0170] The decoding unit 606 is used to input each of the speech features into the trained decoder to decode the speech data corresponding to each of the sound sources.
[0171] In some embodiments, the apparatus further includes:
[0172] Obtain the first sample fused speech features of the first sample fused speech data, and the first sample real speech data of each first sound source included in the first sample fused speech data;
[0173] The first sample fused speech features are input into the feature perceptron to be trained to determine the first sample speech feature mask of the speech data corresponding to each first sound source.
[0174] The first predicted feature of the speech data corresponding to each first sound source is determined by using the speech feature mask of each first sample and the predictor to be trained.
[0175] Each of the first sample's real speech data is input into the trained feature extraction model to extract the first real feature of the speech data corresponding to each first sound source;
[0176] The first loss value is determined based on the first predicted feature and the corresponding first true feature of each first sound source;
[0177] When the first loss value is greater than the first preset loss value, the network parameters of the feature perceptron to be trained and the predictor to be trained are adjusted according to the first loss value, and the process of inputting the first sample fused speech features into the feature perceptron to be trained and determining the first sample speech feature mask of the speech data corresponding to each first sound source is repeated until the first loss value is less than or equal to the first preset loss value, so as to obtain the feature perceptron and the predictor after preliminary training.
[0178] In some embodiments, the apparatus includes:
[0179] Based on each first sample speech feature mask, the first sample speech feature representation of each sound source corresponding to the speech data is determined from the first sample fused speech features;
[0180] Each of the first sample speech features is input into the predictor to be trained to obtain the first predicted feature of the speech data corresponding to each first sound source.
[0181] In some embodiments, the apparatus includes:
[0182] Determine the similarity between the first predicted feature and the corresponding first true feature of each first sound source;
[0183] Select the target first sound source with the highest similarity from each first sound source;
[0184] The similarity between the first predicted feature of the target first sound source and the first predicted feature of each other first sound source is obtained to obtain multiple first predicted similarities, wherein the other first sound sources are first sound sources other than the target first sound source among the first sound sources;
[0185] A first loss value is determined based on the target similarity between the first predicted feature of the target first sound source and the corresponding first real feature, as well as multiple first predicted similarities.
[0186] In some embodiments, the apparatus further includes:
[0187] The second sample fused speech features of the second sample fused speech data are obtained, as well as the second sample real speech data of each second sound source included in the second sample fused speech data;
[0188] The second sample fused speech features are input into the pre-trained feature perceptron to determine the second sample speech feature mask for each second sound source corresponding to the speech data.
[0189] Based on each second sample speech feature mask, the second sample speech feature representation of the speech data corresponding to each second sound source is determined from the second sample fused speech features;
[0190] Each second sample speech feature representation is input into the feature perceptron to be trained to determine the sample speech feature mask of each second sound source corresponding to the speech data.
[0191] Based on each sample speech feature mask, the sample speech features of each second sound source corresponding to the speech data are determined from the second sample fused speech features;
[0192] Each of the sample speech features is input into the decoder to be trained to decode the second sample predicted speech data corresponding to each second sound source;
[0193] The second loss value is determined based on the second sample predicted speech data and the corresponding second sample real speech data of each second sound source.
[0194] When the second loss value is greater than the second preset loss value, the network parameters of the feature perceptron to be trained are adjusted according to the second loss value, and the process returns to the step of inputting the second sample fused speech features into the pre-trained feature perceptron and determining the second sample speech feature mask for each second sound source speech data, until the second loss value is less than or equal to the second preset loss value, thus obtaining the trained feature perceptron, the trained feature perceptron and the trained decoder.
[0195] In some embodiments, the apparatus includes:
[0196] Based on the predicted speech data of the second sample of each second sound source and the corresponding real speech data of the second sample, determine the speech distortion degree of each second sound source.
[0197] The target speech distortion with the smallest speech distortion is selected from each of the stated speech distortion values, and the target speech distortion value is determined as the second loss value.
[0198] In some embodiments, the apparatus includes:
[0199] For each second sound source, calculate the similarity energy value of the similar parts in the corresponding second sample predicted speech data and the corresponding second sample real speech data;
[0200] Calculate the difference energy value of the difference portion between the corresponding second sample predicted speech data and the corresponding second sample real speech data;
[0201] The speech distortion degree corresponding to each second sound source is determined based on the energy ratio of the similar energy value to the difference energy value.
[0202] The specific implementation of each of the above units can be found in the previous embodiments, and will not be repeated here.
[0203] As described above, in this embodiment, the feature extraction unit 601 acquires fused speech data, performs feature extraction on the fused speech data to obtain fused speech features, and the fused speech data includes speech data generated after mixing at least two sound sources; the feature perception unit 602 inputs the fused speech features into a trained feature perceptron to determine the speech feature mask corresponding to each sound source; the first determination unit 603 determines the speech feature representation of each sound source corresponding to the speech data from the fused speech features based on each of the speech feature masks; the feature perception unit 604 inputs each of the speech feature representations into a trained feature perceptron to determine the speech feature mask corresponding to each sound source; the second determination unit 605 determines the speech features of each sound source corresponding to the speech data from the fused speech features based on each of the speech feature masks; and the decoding unit 606 inputs each of the speech features into a trained decoder to decode the speech data corresponding to each sound source.
[0204] Therefore, by acquiring fused speech data, feature extraction is performed on the fused speech data to obtain fused speech features. The fused speech data includes speech data generated after mixing at least two sound sources. The fused speech features are input into a trained feature perceptron to determine the speech feature mask corresponding to each sound source. Based on each speech feature mask, a speech feature representation of each sound source corresponding to the speech data is determined from the fused speech features. Each speech feature representation is input into a trained feature perceptron to determine the speech feature mask corresponding to each sound source. Based on each speech feature mask, a speech feature of each sound source corresponding to the speech data is determined from the fused speech features. Each speech feature is input into a trained decoder to decode the speech data corresponding to each sound source. By using a feature perceptron to perceive the speech features of speech data corresponding to different sound sources, blind source separation is performed based on the speech features, reducing the difference between the separated speech and the original speech, improving consistency, and thus improving the accuracy of speech processing.
[0205] The specific implementation of each of the above units can be found in the previous embodiments, and will not be repeated here.
[0206] Reference Figure 7 , Figure 7This is a partial structural block diagram of a computer device 110 implementing an embodiment of the present disclosure. The computer device 110 can vary significantly due to different configurations or performance characteristics, and may include one or more central processing units (CPUs) 622 (e.g., one or more processors) and a memory 632, and one or more storage media 630 (e.g., one or more mass storage devices) storing application programs 642 or data 644. The memory 632 and storage media 630 may be temporary or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server 600. Furthermore, the CPU 622 may be configured to communicate with the storage media 630 and execute the series of instruction operations in the storage media 630 on the server 600.
[0207] Computer device 110 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0208] The central processing unit 622 in the computer device 110 can be used to execute the speech processing method of the embodiments of this disclosure, for example:
[0209] Acquire fused speech data, extract features from the fused speech data to obtain fused speech features, wherein the fused speech data includes speech data generated after mixing at least two sound sources;
[0210] The fused speech features are input into the trained feature perceptron to determine the speech feature mask for each speech source corresponding to the speech data.
[0211] Based on each of the aforementioned speech feature masks, the speech feature representation of the speech data corresponding to each sound source is determined from the fused speech features;
[0212] Each of the aforementioned speech feature representations is input into the trained feature perceptron to determine the speech feature mask for the speech data corresponding to each sound source.
[0213] Based on each of the speech feature masks, the speech features corresponding to each sound source are determined from the fused speech features;
[0214] Each of the aforementioned speech features is input into the trained decoder to decode the speech data corresponding to each of the aforementioned sound sources.
[0215] This disclosure also provides a computer-readable storage medium for storing program code for executing the speech processing methods of the foregoing embodiments.
[0216] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the aforementioned voice processing method. For example:
[0217] Acquire fused speech data, extract features from the fused speech data to obtain fused speech features, wherein the fused speech data includes speech data generated after mixing at least two sound sources;
[0218] The fused speech features are input into the trained feature perceptron to determine the speech feature mask for each speech source corresponding to the speech data.
[0219] Based on each of the aforementioned speech feature masks, the speech feature representation of the speech data corresponding to each sound source is determined from the fused speech features;
[0220] Each of the aforementioned speech feature representations is input into the trained feature perceptron to determine the speech feature mask for the speech data corresponding to each sound source.
[0221] Based on each of the speech feature masks, the speech features corresponding to each sound source are determined from the fused speech features;
[0222] Each of the aforementioned speech features is input into the trained decoder to decode the speech data corresponding to each of the aforementioned sound sources.
[0223] Furthermore, the terms “comprising” and “including”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.
[0224] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0225] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.
[0226] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0227] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0228] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0229] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0230] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.
[0231] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0232] The above is a detailed description of the embodiments of this application. However, this application is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A speech processing method, characterized in that, include: Acquire fused speech data, extract features from the fused speech data to obtain fused speech features, wherein the fused speech data includes speech data generated after mixing at least two sound sources; Obtain the first sample fused speech features of the first sample fused speech data, and the first sample real speech data of each first sound source included in the first sample fused speech data; The first sample fused speech features are input into the feature perceptron to be trained to determine the first sample speech feature mask of the speech data corresponding to each first sound source. The first predicted feature of the speech data corresponding to each first sound source is determined by using the speech feature mask of each first sample and the predictor to be trained. Each of the first sample's real speech data is input into the trained feature extraction model to extract the first real feature of the speech data corresponding to each first sound source; The first loss value is determined based on the first predicted feature and the corresponding first true feature of each first sound source; When the first loss value is greater than the first preset loss value, the network parameters of the feature perceptron to be trained and the predictor to be trained are adjusted according to the first loss value, and the process of inputting the first sample fused speech features into the feature perceptron to be trained and determining the first sample speech feature mask of the speech data corresponding to each first sound source is repeated until the first loss value is less than or equal to the first preset loss value, so as to obtain the feature perceptron and the predictor after preliminary training. The fused speech features are input into the trained feature perceptron to determine the speech feature mask for each speech source corresponding to the speech data. Based on each of the aforementioned speech feature masks, the speech feature representation of the speech data corresponding to each sound source is determined from the fused speech features; The second sample fused speech features of the second sample fused speech data are obtained, as well as the second sample real speech data of each second sound source included in the second sample fused speech data; The second sample fused speech features are input into the pre-trained feature perceptron to determine the second sample speech feature mask for each second sound source corresponding to the speech data. Based on each second sample speech feature mask, the second sample speech feature representation of the speech data corresponding to each second sound source is determined from the second sample fused speech features; Each second sample speech feature representation is input into the feature perceptron to be trained to determine the sample speech feature mask of each second sound source corresponding to the speech data. Based on each sample speech feature mask, the sample speech features of each second sound source corresponding to the speech data are determined from the second sample fused speech features; Each of the sample speech features is input into the decoder to be trained to decode the second sample predicted speech data corresponding to each second sound source; The second loss value is determined based on the second sample predicted speech data and the corresponding second sample real speech data of each second sound source. When the second loss value is greater than the second preset loss value, the network parameters of the feature perceptron to be trained are adjusted according to the second loss value, and the process of inputting the second sample fused speech features into the pre-trained feature perceptron and determining the second sample speech feature mask for each second sound source speech data is repeated until the second loss value is less than or equal to the second preset loss value, thus obtaining the trained feature perceptron, the trained feature perceptron and the trained decoder. Each of the aforementioned speech feature representations is input into the trained feature perceptron to determine the speech feature mask for the speech data corresponding to each sound source. Based on each of the speech feature masks, the speech features corresponding to each sound source are determined from the fused speech features; Each of the aforementioned speech features is input into the trained decoder to decode the speech data corresponding to each of the aforementioned sound sources.
2. The speech processing method according to claim 1, characterized in that, The step of determining the first predicted feature of the speech data corresponding to each first sound source by using the speech feature mask of each first sample and the predictor to be trained includes: Based on each first sample speech feature mask, the first sample speech feature representation of each sound source corresponding to the speech data is determined from the first sample fused speech features; Each of the first sample speech features is input into the predictor to be trained to obtain the first predicted feature of the speech data corresponding to each first sound source.
3. The speech processing method according to claim 1, characterized in that, The step of determining the first loss value based on the first predicted feature and the corresponding first true feature of each first sound source includes: Determine the similarity between the first predicted feature and the corresponding first true feature of each first sound source; Select the target first sound source with the highest similarity from each first sound source; The similarity between the first predicted feature of the target first sound source and the first predicted feature of each other first sound source is obtained to obtain multiple first predicted similarities, wherein the other first sound sources are first sound sources other than the target first sound source among the first sound sources; A first loss value is determined based on the target similarity between the first predicted feature of the target first sound source and the corresponding first real feature, as well as multiple first predicted similarities.
4. The speech processing method according to claim 1, characterized in that, The determination of the second loss value based on the predicted speech sample data of each second sound source and the corresponding real speech data of the second sample includes: Based on the predicted speech data of the second sample of each second sound source and the corresponding real speech data of the second sample, determine the speech distortion degree of each second sound source. The target speech distortion with the smallest speech distortion is selected from each of the stated speech distortion values, and the target speech distortion value is determined as the second loss value.
5. The speech processing method according to claim 4, characterized in that, The step of determining the speech distortion degree corresponding to each second sound source based on the predicted speech data of the second sample of each second sound source and the corresponding real speech data of the second sample includes: For each second sound source, calculate the similarity energy value of the similar parts in the corresponding second sample predicted speech data and the corresponding second sample real speech data; Calculate the difference energy value of the difference portion between the corresponding second sample predicted speech data and the corresponding second sample real speech data; The speech distortion degree corresponding to each second sound source is determined based on the energy ratio of the similar energy value to the difference energy value.
6. A voice processing device, characterized in that, include: A feature extraction unit is used to acquire fused speech data, extract features from the fused speech data, and obtain fused speech features. The fused speech data includes speech data generated after mixing at least two sound sources. The device also includes: Obtain the first sample fused speech features of the first sample fused speech data, and the first sample real speech data of each first sound source included in the first sample fused speech data; The first sample fused speech features are input into the feature perceptron to be trained to determine the first sample speech feature mask of the speech data corresponding to each first sound source. The first predicted feature of the speech data corresponding to each first sound source is determined by using the speech feature mask of each first sample and the predictor to be trained. Each of the first sample's real speech data is input into the trained feature extraction model to extract the first real feature of the speech data corresponding to each first sound source; The first loss value is determined based on the first predicted feature and the corresponding first true feature of each first sound source; When the first loss value is greater than the first preset loss value, the network parameters of the feature perceptron to be trained and the predictor to be trained are adjusted according to the first loss value, and the process of inputting the first sample fused speech features into the feature perceptron to be trained and determining the first sample speech feature mask of the speech data corresponding to each first sound source is repeated until the first loss value is less than or equal to the first preset loss value, so as to obtain the feature perceptron and the predictor after preliminary training. The feature perception unit is used to input the fused speech features into the trained feature perceptron to determine the speech feature mask of the speech data corresponding to each sound source. The first determining unit is used to determine the speech feature representation of each speech source corresponding to the speech data from the fused speech features based on each of the speech feature masks. The device also includes: The second sample fused speech features of the second sample fused speech data are obtained, as well as the second sample real speech data of each second sound source included in the second sample fused speech data; The second sample fused speech features are input into the pre-trained feature perceptron to determine the second sample speech feature mask for each second sound source corresponding to the speech data. Based on each second sample speech feature mask, the second sample speech feature representation of the speech data corresponding to each second sound source is determined from the second sample fused speech features; Each second sample speech feature representation is input into the feature perceptron to be trained to determine the sample speech feature mask of each second sound source corresponding to the speech data. Based on each sample speech feature mask, the sample speech features of each second sound source corresponding to the speech data are determined from the second sample fused speech features; Each of the sample speech features is input into the decoder to be trained to decode the second sample predicted speech data corresponding to each second sound source; The second loss value is determined based on the second sample predicted speech data and the corresponding second sample real speech data of each second sound source. When the second loss value is greater than the second preset loss value, the network parameters of the feature perceptron to be trained are adjusted according to the second loss value, and the process of inputting the second sample fused speech features into the pre-trained feature perceptron and determining the second sample speech feature mask for each second sound source speech data is repeated until the second loss value is less than or equal to the second preset loss value, thus obtaining the trained feature perceptron, the trained feature perceptron and the trained decoder. The feature perception unit is used to input each of the speech feature representations into the trained feature perceptron to determine the speech feature mask of the speech data corresponding to each sound source. The second determining unit is used to determine the speech features of each sound source corresponding to the speech data from the fused speech features based on each of the speech feature masks; The decoding unit is used to input each of the speech features into the trained decoder and decode the speech data corresponding to each of the sound sources.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to execute the speech processing method according to any one of claims 1 to 5.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the speech processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Speech enhancement method, training method of speech enhancement network and electronic equipment
CN116959471A
Voice wake-up method and device based on generative adversarial network and storage medium
CN117690432A