Task processing method, task processing model training method, and conference speech separation method
By performing feature extraction and time-dimensional processing of mixed speech data, voice separation information is generated and separated, the problem of poor speech separation performance in the prior art is solved, and a high-performance speech separation effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/124573
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-10-12
- Publication Date
- 2025-05-08
AI Technical Summary
The prior art has large positioning errors in speech separation, resulting in poor speech separation performance, and a high-performance speech separation solution is urgently needed.
By obtaining mixed speech data, performing feature extraction, processing the mixed speech feature sequence based on the time dimension, generating speech separation information, and separating the speech feature sequence based on this information to obtain task processing results.
By fully considering the cyclic mode and complex time dependencies of the speech signal in the time dimension, the speech separation performance is improved, and the accuracy and efficiency of speech separation are enhanced.
Smart Images

Figure CN2024124573_08052025_PF_FP_ABST
Abstract
Description
Task processing, task processing model training, and conference speech separation methods
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 3, 2023, with application number 202311465109.8 and application name “Task processing, task processing model training and conference speech separation method”, the entire content of which is incorporated by reference in this disclosure. Technical Field
[0002] The embodiments of this specification relate to the field of computer technology, and in particular to task processing, task processing model training, and conference speech separation methods. Background Art
[0003] With the development of computer technology, scenarios involving multiple people communicating simultaneously are becoming more and more common. If speech separation is not performed after capturing the voices of multiple people simultaneously through microphones, it will directly affect the speech recognition system, auditory perception, and comprehension. Therefore, speech separation technology has gradually become a research focus.
[0004] Currently, the target speech is usually separated from the speech of multiple speakers through the sound source localization method. However, there are often large positioning errors in the sound source localization process, resulting in poor speech separation performance. Therefore, a high-performance speech separation solution is urgently needed.
[0005] Summary of the Invention
[0006] In view of this, embodiments of this specification provide a task processing method. One or more embodiments of this specification also relate to a task processing model training method, a conference speech separation method, a task processing device, a task processing model training device, a conference speech separation device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0007] According to a first aspect of an embodiment of this specification, a task processing method is provided, including:
[0008] Acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects;
[0009] Extracting features from mixed speech data to obtain a mixed speech feature sequence;
[0010] For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension;
[0011] The mixed speech feature sequence is separated according to the speech separation information to obtain the task processing result.
[0012] According to a second aspect of an embodiment of this specification, a task processing model training method is provided, comprising:
[0013] Acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects;
[0014] Extracting features from the sample mixed speech data to obtain a sample mixed speech feature sequence;
[0015] Inputting the sample mixed speech feature sequence into a loop processing unit in the task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence;
[0016] The model parameters of the task processing model are adjusted according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.
[0017] According to a third aspect of the embodiments of this specification, a conference speech separation method is provided, including:
[0018] Acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects;
[0019] Extracting features from mixed speech data to obtain a mixed speech feature sequence;
[0020] For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension;
[0021] The mixed speech feature sequence is separated according to the speech separation information to obtain the conference speech separation result.
[0022] According to a fourth aspect of the embodiments of this specification, there is provided a task processing device, including:
[0023] A first acquisition module is configured to acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects;
[0024] A first extraction module is configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;
[0025] A first processing module is configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing results corresponding to each feature dimension;
[0026] The first separation module is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a task processing result.
[0027] According to a fifth aspect of the embodiments of this specification, a task processing model training device is provided, comprising:
[0028] A second acquisition module is configured to acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects;
[0029] A second extraction module is configured to perform feature extraction on the sample mixed speech data to obtain a sample mixed speech feature sequence;
[0030] An input module is configured to input the sample mixed speech feature sequence into a recurrent processing unit in the task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the recurrent processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence;
[0031] The adjustment module is configured to adjust the model parameters of the task processing model according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.
[0032] According to a sixth aspect of the embodiments of this specification, a conference voice separation device is provided, including:
[0033] A third acquisition module is configured to acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects;
[0034] a third extraction module, configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;
[0035] a second processing module configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing results corresponding to each feature dimension;
[0036] The second separation module is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a conference speech separation result.
[0037] According to a seventh aspect of the embodiments of this specification, a computing device is provided, including:
[0038] memory and processor;
[0039] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.
[0040] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.
[0041] According to a ninth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect, the second aspect, or the third aspect above.
[0042] One embodiment of this specification provides a task processing method that obtains mixed speech data, where the mixed speech data includes speech data of multiple objects; performs feature extraction on the mixed speech data to obtain a mixed speech feature sequence; processes the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generates speech separation information corresponding to each object based on the processing results corresponding to each feature dimension; and separates the mixed speech feature sequence based on the speech separation information to obtain a task processing result. By processing the mixed speech feature sequence based on the time dimension, the cyclical patterns exhibited by the speech signal in the time dimension are fully considered, and the complex temporal dependencies within the speech signal are captured, thereby improving speech separation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] FIG1 is an architecture diagram of a task processing system provided by one embodiment of this specification;
[0044] FIG2 is an architecture diagram of another task processing system provided by one embodiment of this specification;
[0045] FIG3 is a flowchart of a task processing method provided by one embodiment of this specification;
[0046] FIG4 is a schematic diagram of a processing method of an attention processing unit in a task processing method provided by one embodiment of this specification;
[0047] FIG5 is a schematic diagram of a processing method of a loop processing unit in a task processing method provided by an embodiment of this specification;
[0048] FIG6 is an architecture diagram of a loop processing unit in a task processing method provided by one embodiment of this specification;
[0049] FIG7 is an architecture diagram of a convolutional layer in a task processing method provided by one embodiment of this specification;
[0050] FIG8 is an architecture diagram of an expanded feedforward layer in a task processing method provided by one embodiment of this specification;
[0051] FIG9 is an architecture diagram of a two-dimensional dilated convolutional layer in a task processing method provided by one embodiment of this specification;
[0052] FIG10 is a flowchart of a task processing model training method provided by one embodiment of this specification;
[0053] FIG11 is a flowchart of a conference speech separation method provided by one embodiment of this specification;
[0054] FIG12 is a flowchart of a task processing method according to an embodiment of the present specification;
[0055] FIG13 is a schematic diagram of the structure of a task processing device provided by one embodiment of this specification;
[0056] FIG14 is a schematic diagram of the structure of a task processing model training device provided by one embodiment of this specification;
[0057] FIG15 is a schematic diagram of the structure of a conference voice separation device provided by one embodiment of this specification;
[0058] FIG16 is a structural block diagram of a computing device provided in one embodiment of this specification. DETAILED DESCRIPTION
[0059] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0060] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0061] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0062] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0063] First, the terms involved in one or more embodiments of this specification are explained.
[0064] Speech separation technology: Speech separation refers to separating the mixed speech of multiple speakers and obtaining the individual speech of all speakers.
[0065] Self-attention: Self-attention is a neural network mechanism for processing sequential data. It allows the model to consider every element in the sequence simultaneously, not just its neighbors. This mechanism can be used in many natural language processing tasks, such as text classification, machine translation, and question-answering systems.
[0066] Deep learning algorithms: Deep learning algorithms are computational models that mimic neurons in the human brain and are used to handle complex machine learning tasks. They build multi-layer neural networks to learn the characteristics and patterns of data, enabling tasks such as data classification, prediction, and generation. Deep learning algorithms have been widely used in many fields, including computer vision, natural language processing, and speech recognition.
[0067] The "cocktail party problem" refers to the challenge of effectively identifying and understanding the target speaker's speech amidst a cacophony of sounds in a noisy environment. This issue is crucial in the field of computer speech recognition, as it requires models to accurately identify the target speech amidst complex background noise. In scenarios where multiple people are communicating simultaneously, the failure to perform speech separation after capturing their voices through microphones can directly impact the speech recognition system, auditory perception, and comprehension. Consequently, speech separation technology has become a research focus. The goal of speech separation is to separate the mixed speech of multiple speakers. The separation results can be used as input for speech recognition or directly played back to the listener, thereby improving recognition results and auditory perception.
[0068] Currently, speech separation can be performed using the following approaches: First, additional speaker labels are used during model training, allowing the trained model to be used for speech separation; second, speech separation is performed solely based on convolutional networks; and third, long speech sequences are truncated into shorter ones, followed by intra- and inter-sequence attention. However, the addition of speaker labels in the first approach increases model training costs; the second approach lacks advanced attention mechanisms and cannot process global information within the sequence; and the third approach still processes global information through implicit, indirect interactions, resulting in some performance degradation and a lack of local modeling capabilities. Furthermore, the third approach's use of two channels introduces significant processing overhead due to overlapping blocks, and the cross-processing is inefficient in modeling global information. Overall, these approaches prioritize long-range, coarse-grained dependencies, resulting in suboptimal speech separation results.
[0069] In order to solve the above problems, the embodiments of this specification propose a speech separation solution that focuses on fine-grained cyclic patterns in the speech separation process, which can be applied to the recognition or playback of recorded speech in multi-object speaking scenarios such as conference scenarios and cocktail parties. Specifically, the embodiments of this specification propose a task processing method, which obtains mixed speech data, wherein the mixed speech data includes speech data of multiple objects; performs feature extraction on the mixed speech data to obtain a mixed speech feature sequence; for any feature dimension of the mixed speech feature sequence, processes the mixed speech feature sequence based on the time dimension, and generates speech separation information corresponding to each object according to the processing results corresponding to each feature dimension; separates the mixed speech feature sequence according to the speech separation information to obtain the task processing result. By processing the mixed speech feature sequence based on the time dimension, the cyclic pattern of the speech signal in the time dimension is fully taken into account, and the complex time dependency within the speech signal is captured, thereby improving the speech separation performance.
[0070] In this specification, a task processing method is provided. This specification also involves a task processing model training method, a conference speech separation method, a task processing device, a task processing model training device, a conference speech separation device, a computing device, a computer-readable storage medium and a computer program, which are described in detail one by one in the following embodiments.
[0071] Referring to FIG1 , FIG1 shows an architecture diagram of a task processing system provided by an embodiment of this specification. The task processing system may include a client 100 and a server 200;
[0072] The client 100 is configured to send mixed voice data to the server 200, wherein the mixed voice data includes voice data of multiple objects;
[0073] The server 200 is configured to extract features from the mixed speech data to obtain a mixed speech feature sequence; process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object based on the processing results corresponding to each feature dimension; separate the mixed speech feature sequence based on the speech separation information to obtain a task processing result; and send the task processing result to the client 100;
[0074] The client 100 is also used to receive the task processing result sent by the server 200.
[0075] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.
[0076] Referring to FIG. 2 , FIG. 2 shows an architecture diagram of another task processing system provided in accordance with one embodiment of the present disclosure. The task processing system may include multiple clients 100 and a server 200. The clients 100 may be end-side devices, and the server 200 may be a cloud-side device. Multiple clients 100 may establish communication connections via the server 200. In a speech separation task processing scenario, the server 200 is used to provide speech separation services between multiple clients 100. Multiple clients 100 may serve as either senders or receivers, communicating via the server 200.
[0077] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100. In the speech separation task processing scenario, users can publish data streams to the server 200 through the client 100. The server 200 generates task processing results based on the data stream and pushes the task processing results to other clients that have established communication.
[0078] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.
[0079] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0080] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that provide background training to support models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that is integrated with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0081] It is worth noting that the task processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server and thus execute the task processing methods provided in the embodiments of this specification. In other embodiments, the task processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.
[0082] Referring to FIG3 , FIG3 shows a flowchart of a task processing method provided by an embodiment of this specification, which specifically includes the following steps:
[0083] Step 302: Acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects.
[0084] In one or more embodiments of the present specification, during task processing, mixed voice data corresponding to a target task may be obtained, and the mixed voice data may be separated to obtain separate voice data for each separated object.
[0085] Specifically, mixed voice data can be data from different scenarios, such as mixed voice data from a meeting or a cocktail party. Mixed voice data refers to voice data that simultaneously contains two or more voices and can be referred to as a mixed voice waveform. The voice data of each object in mixed voice data may be intertwined. An object refers to the source of the voice data and can be a virtual character, a recorder, or other similar device. Voice data can be referred to as a voice signal, including but not limited to spoken language data and music data.
[0086] In practical applications, there are multiple ways to obtain mixed voice data, and the specific method is selected according to the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, mixed voice data can be received from a user through a client. In another possible implementation of this specification, mixed voice data can be read from other data acquisition devices or databases.
[0087] Step 304: extract features from the mixed speech data to obtain a mixed speech feature sequence.
[0088] In one or more embodiments of the present specification, after obtaining the mixed speech data, further feature extraction may be performed on the mixed speech data to obtain a mixed speech feature sequence.
[0089] Specifically, the mixed speech feature sequence includes a set of feature vectors, each of which represents a small part of the speech signal. The mixed speech feature sequence is used to represent the attribute characteristics of the mixed speech data, such as speech frequency, energy, tone, etc.
[0090] In practical applications, there are many ways to extract features from mixed speech data and obtain a mixed speech feature sequence. The specific method to be selected depends on the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, traditional feature extraction methods such as filters, statistical feature extraction, and wavelet transforms can be used. In another possible implementation of this specification, a deep learning model can be used to automatically learn and generate a mixed speech feature sequence from mixed speech data, such as inputting the mixed speech data into the encoding unit in the task processing model to obtain a mixed speech feature sequence.
[0091] In an optional embodiment of the present specification, the encoding unit can use a one-dimensional (1-D) convolutional layer (Conv1D) combined with a rectified linear unit (ReLU) to encode the input mixed speech waveform. During the encoding process, the mixed speech waveform can be converted into a non-negative speech embedding sequence, that is, a mixed speech feature sequence. Assuming that the convolution kernel size of the encoding unit is K1 and the stride is K1 / 2, the length of the encoded mixed speech feature sequence can be calculated by the following formula (1): S = 2T / K1-1 (1)
[0092] Where T represents the length of the mixed speech waveform, N represents the embedding dimension, and S represents the length of the mixed speech feature sequence.
[0093] Step 306: for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension.
[0094] In one or more embodiments of the present specification, after obtaining mixed speech data and performing feature extraction on the mixed speech data to obtain a mixed speech feature sequence, the mixed speech feature sequence is further processed based on the time dimension for any feature dimension of the mixed speech feature sequence, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension.
[0095] Specifically, the dimensions of speech data include frequency and time, where the frequency dimension can be understood as a feature dimension. Speech separation information can be understood as masking information, which corresponds one-to-one with an object and is used to separate the individual speech data of each object from the mixed speech data.
[0096] It should be noted that, for any feature dimension of a mixed speech feature sequence, processing the mixed speech feature sequence based on the time dimension can be understood as simultaneously performing fine-grained, cyclic processing on the mixed speech feature sequence from both the frequency and time dimensions. After obtaining the processing results corresponding to each feature dimension, the processing results corresponding to each feature dimension can be directly used as the speech separation information corresponding to each object. Furthermore, the data identifier corresponding to each processing result can be obtained from a preset masking identifier library and used as the speech separation information.
[0097] In practical applications, there are multiple ways to process a mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence. The specific method to be selected depends on the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, a loop filtering process can be performed on the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence until a preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.
[0098] In another possible implementation of the present specification, the mixed speech feature sequence may be updated using a joint attention processing mechanism, and then the updated mixed speech feature sequence may be subjected to a cyclic filtering process. That is, the mixed speech feature sequence may be processed based on the time dimension for any feature dimension of the mixed speech feature sequence, and the following steps may be included:
[0099] Based on the time dimension, the mixed speech feature sequence is subjected to joint attention processing to obtain an updated mixed speech feature sequence;
[0100] For any feature dimension of the updated mixed speech feature sequence, the mixed speech feature sequence is subjected to cyclic filtering processing based on the time dimension until a preset cyclic stopping condition is reached, thereby obtaining processing results corresponding to each feature dimension.
[0101] It should be noted that in order to achieve efficient self-attention on a wide range of mixed speech feature sequences, the embodiments of this specification apply self-attention to the entire mixed speech feature sequence through joint attention processing, wherein the joint attention processing process adopts a joint local-global self-attention strategy.
[0102] In practical applications, when performing joint attention processing on mixed speech feature sequences based on the time dimension, full computational self-attention can be performed on non-overlapping local segments at the same time, and a linearized and resource-efficient self-attention mechanism can be used on the entire mixed speech feature sequence to obtain an updated mixed speech feature sequence.
[0103] Furthermore, since speech data inherently has cyclic patterns manifested in speech structure, prosody, and semantic associations, these patterns play an important role in speech separation. Therefore, the complex temporal dependencies within the speech signal can be modeled. Assuming that different embedding layers retain different cyclic patterns, cyclic filtering learning can be performed on each feature dimension. All feature dimensions can be learned in parallel until a preset cyclic stop condition is reached, and the processing results corresponding to each feature dimension can be obtained. The preset cyclic stop condition includes but is not limited to the number of cycles reaching a preset number of cycles, wherein the preset number of cycles, such as 20 times, is selected based on actual conditions and is not limited in this embodiment of the present specification.
[0104] Applying the solutions of the embodiments of this specification, joint attention processing is performed on the mixed speech feature sequence based on the time dimension to obtain an updated mixed speech feature sequence. For each feature dimension of the updated mixed speech feature sequence, cyclic filtering is performed on the mixed speech feature sequence based on the time dimension until a preset loop stop condition is met, obtaining the processing results corresponding to each feature dimension. By leveraging the combined advantages of self-attention and cyclic modeling, it helps capture the extensive dependencies and local cyclic patterns of speech data, improving speech separation performance.
[0105] In an optional embodiment of the present specification, a task processing model can be used to perform joint attention processing and recurrent filtering processing. That is, the above-mentioned joint attention processing is performed on the mixed speech feature sequence based on the time dimension to obtain an updated mixed speech feature sequence, which may include the following steps:
[0106] Inputting the mixed speech feature sequence into the attention processing unit in the task processing model to obtain an updated mixed speech feature sequence based on the time dimension;
[0107] For any feature dimension of the updated mixed speech feature sequence, performing cyclic filtering on the mixed speech feature sequence based on the time dimension until a preset loop stop condition is reached to obtain a processing result corresponding to each feature dimension may include the following steps:
[0108] For any feature dimension of the updated mixed speech feature sequence, the updated mixed speech feature sequence is input into the loop processing unit in the task processing model to obtain the initial processing result corresponding to the feature dimension, and the initial processing result and the mixed speech feature sequence are returned to the attention processing unit in the task processing model until the preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.
[0109] Specifically, the attention processing unit considers each element in the mixed speech feature sequence of the current layer to learn the embedding of each subsequent layer, independent of recurrence. Self-attention primarily captures long-range, coarse-grained dependencies. The recurrence processing unit is used to perform recurring learning on each feature dimension, enabling parallel learning across feature dimensions.
[0110] For example, referring to FIG4 , FIG4 shows a schematic diagram of the processing method of the attention processing unit in a task processing method provided by an embodiment of the present specification. As shown in FIG4 , when the attention processing unit determines the embedding of the target local segment (the block with black slashes in FIG4 ), it does so in the time dimension based on the global information in the layer that has completed learning. Referring to FIG5 , FIG5 shows a schematic diagram of the processing method of the loop processing unit in a task processing method provided by an embodiment of the present specification. As shown in FIG5 , for the target feature dimension (the dimension where the block with black slashes in FIG5 is located), the loop processing unit can perform loop filtering processing on the mixed speech feature sequence in the time dimension until a preset loop stop condition is reached, thereby obtaining a processing result corresponding to the target feature dimension.
[0111] Using the solution of the embodiments of this specification, a mixed speech feature sequence is input into the attention processing unit of the task processing model to obtain an updated mixed speech feature sequence based on the time dimension; for any feature dimension of the updated mixed speech feature sequence, the updated mixed speech feature sequence is input into the loop processing unit of the task processing model to obtain the initial processing result corresponding to the feature dimension, and the initial processing result and the mixed speech feature sequence are input back into the attention processing unit of the task processing model until a preset loop stop condition is met to obtain the processing result corresponding to each feature dimension. By leveraging the combined advantages of self-attention and loop modeling, it helps to capture the extensive dependencies and local loop patterns of speech data, thereby improving speech separation performance.
[0112] In an optional embodiment of the present specification, a recurrent processing unit can be constructed based on a long short-term memory network (LSTM) or a gated recurrent unit (GRU), but since both of the above methods are sequential processing, the amount of calculation is large and the processing speed is slow, resulting in obvious shortcomings in the processing process. Therefore, a recurrent neural network (RNN, Recurrent Neural Network) recurrent processing unit based on a feedforward sequential memory network (FSMN, Feedforward Sequential Memory Networks) can be used. Specifically, the recurrent processing unit includes a bottleneck layer, a gated convolutional layer and an output layer; the above-mentioned input of the updated mixed speech feature sequence into the recurrent processing unit in the task processing model to obtain the initial processing result corresponding to the feature dimension can include the following steps:
[0113] Inputting the updated mixed speech feature sequence into the bottleneck layer to obtain a first speech feature sequence, wherein the bottleneck layer is used to reduce the feature dimension of the updated mixed speech feature sequence;
[0114] Inputting the first speech feature sequence into a gated convolutional layer to obtain a second speech feature sequence, wherein the gated convolutional layer is used to perform a cyclic filtering process on the first speech feature sequence based on a time dimension;
[0115] The second speech feature sequence is input into the output layer to obtain an initial processing result.
[0116] Refer to Figure 6, which shows an architectural diagram of a cyclic processing unit in a task processing method provided by an embodiment of this specification. The cyclic processing unit includes a bottleneck layer (Bottleneck layer), a gated convolution layer (GCU layer, Gate Control Unit layer) and an output layer (Output layer); the bottleneck layer includes a 1×1 convolution kernel, a third activation layer (PReLU activation, Parametric Rectified Linear Unit, parameterized rectified linear unit) and a second normalization layer (LayerNorm layer); the gated convolution layer includes a dilated feedforward layer (Diliated FSMN), two convolution layers (Convolution-U) and a processing layer, in which addition and tensor product operations can be performed; the output layer includes a third normalization layer and a 1×1 convolution kernel. In the cyclic processing unit, the updated mixed speech feature sequence is input, and the outputs of the bottleneck layer, the gated convolution layer and the output layer are all different expressions of speech features. The bottleneck layer is used to reduce the feature dimension (from N to N'), and the output layer is used to restore the feature dimension (from N' to N).
[0117] It should be noted that in the bottleneck layer, the updated mixed speech feature sequence is sequentially input into a 1×1 convolution kernel, the third activation layer, and the second normalization layer to obtain the first speech feature sequence. In the output layer, the second speech feature sequence is sequentially input into the third normalization layer and the 1×1 convolution kernel to output the initial processing result.
[0118] By applying the solution of the embodiments of this specification, the updated mixed speech feature sequence is input into the cyclic processing unit, the feature dimension is reduced in the bottleneck layer while retaining the key features, the first speech feature sequence is subjected to cyclic filtering processing based on the time dimension in the gated convolution layer, and finally the initial processing result is obtained through the output layer, which captures the extensive dependencies and local cyclic patterns of the speech data and ensures the accuracy of the initial processing result.
[0119] In an optional embodiment of the present specification, the gated convolution layer includes a dilated feedforward layer, a convolution layer, and a processing layer; the step of inputting the first speech feature sequence into the gated convolution layer to obtain the second speech feature sequence may include the following steps:
[0120] Inputting the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence;
[0121] Input the convolution speech feature sequence into the dilated feedforward layer to obtain the dilated speech feature sequence;
[0122] The first speech feature sequence, the convolution speech feature sequence and the expansion speech feature sequence are input into the processing layer to obtain the second speech feature sequence.
[0123] Specifically, the design of the gated convolution layer is inspired by the effectiveness of the gating mechanism in the gated linear unit (GLU). The gated convolution layer is the main processing part of the recurrent processing unit. The gated convolution layer performs modeling of the recurrent pattern and is obtained based on the gated convolution unit. In the embodiment of this specification, a dilated feedforward layer is combined with the gated convolution layer. The dilated feedforward layer refers to a network structure that adds a "dilation" mechanism to the traditional FSMN. This mechanism can expand the receptive field of the convolution kernel. The dilated feedforward layer is used to enable the FSMN to cover a wider receptive field through the dilation operation while reducing the required memory resources. In order to further improve the dilated feedforward layer and allow for a reduction in the embedding dimension, the embodiment of this specification uses a convolution unit instead of a linear unit.
[0124] It should be noted that the gated convolutional layer maps the first speech feature sequence to the output. As shown in Figure 6, the gated convolutional layer includes two convolutional layers. To facilitate model training, the gated convolutional layer includes a skip connection that connects the input of the gated convolutional layer to its output. This connection allows the model to learn only the residual portion of the input, facilitating gradient backpropagation and accelerating model training. First, the convolution speech feature sequence can be obtained by the following formula (2) and formula (3); specifically, the first speech feature sequence can be input into two convolution layers respectively, and the first speech feature sequence can be calculated by the convolution operation (ConvU) to generate two convolution speech feature sequences (U and V); secondly, the expanded speech feature sequence can be obtained by the following formula (4); specifically, a convolution speech feature sequence (V) can be input into the expanded feedforward layer, and the convolution speech feature sequence (V) can be calculated by the DliatedFSMN operation in the expanded feedforward layer to obtain an expanded speech feature sequence (Y) with the same shape as the convolution speech feature sequence; finally, the second speech feature sequence can be obtained by the following formula (5); specifically, the first speech feature sequence (X1), the convolution speech feature sequence (U) and the expanded speech feature sequence (Y) can be calculated by element-by-element multiplication and element-level addition to obtain a second speech feature sequence (O) with the same shape as the first speech feature sequence, where N' is the number of rows of the matrix and S is the number of columns of the matrix. U = Conv U (X1)∈R N'×S (2) V=Conv U (X1)∈R N'×S (3) Y=Dilated FSMN (V)∈R N'×S (4)
[0125] By applying the solution of the embodiments of this specification, the first speech feature sequence is input into the convolution layer to obtain a convolution speech feature sequence; the convolution speech feature sequence is input into the expanded feedforward layer to obtain an expanded speech feature sequence; the first speech feature sequence, the convolution speech feature sequence and the expanded speech feature sequence are input into the processing layer to obtain a second speech feature sequence. The expansion feedforward layer is used to achieve coverage of a wider receptive field while reducing the required memory resources.
[0126] In an optional embodiment of the present specification, the convolution layer includes a first normalization layer, a first linear layer, a second activation layer, and a depthwise convolution layer; the step of inputting the first speech feature sequence into the convolution layer to obtain the convolution speech feature sequence may include the following steps:
[0127] Inputting the first speech feature sequence into the first normalization layer, the first linear layer, and the second activation layer in sequence to obtain an initial convolution speech feature sequence;
[0128] Input the initial convolution speech feature sequence into the deep convolution layer to obtain the reference convolution speech feature sequence;
[0129] A convolution speech feature sequence is determined according to an initial convolution speech feature sequence and a reference convolution speech feature sequence.
[0130] Referring to Figure 7, Figure 7 shows an architecture diagram of a convolutional layer in a task processing method provided by an embodiment of this specification. The convolutional layer is used to help the gated convolutional layer capture local patterns at a position. The convolutional layer includes a first normalization layer, a first linear layer (linear), a second activation layer (SiLU layer, Sigmoid-Weighted Linear Unit, S-shaped weighted linear unit), and a deep convolution layer (D-Convolution). Within the convolutional layer, the first normalization layer is used to normalize each batch of input features, the first linear layer is used to perform a linear transformation operation on the features, and the second activation layer uses a nonlinear activation function to process the features. The deep convolution layer is equivalent to performing a one-dimensional filtering operation on the sequence by sliding a small window, thereby identifying and extracting important local features and outputting a convolution speech feature sequence. It should be noted that after the second activation layer, a jump connection is used to promote model training, and the initial convolution speech feature sequence and the reference convolution speech feature sequence are added to determine the convolution speech feature sequence.
[0131] Applying the solution of the embodiments of this specification, the first speech feature sequence is sequentially input into the first normalization layer, the first linear layer, and the second activation layer to obtain an initial convolutional speech feature sequence; the initial convolutional speech feature sequence is input into the deep convolution layer to obtain a reference convolutional speech feature sequence; and the convolutional speech feature sequence is determined based on the initial convolutional speech feature sequence and the reference convolutional speech feature sequence. This improves speech separation performance.
[0132] In an optional embodiment of the present specification, the dilated feedforward layer includes a feedforward layer and a plurality of interconnected two-dimensional dilated convolutional layers; the step of inputting the convolutional speech feature sequence into the dilated feedforward layer to obtain the dilated speech feature sequence may include the following steps:
[0133] Input the convolution speech feature sequence into the feedforward layer to obtain a feedforward speech feature sequence;
[0134] Input the feedforward speech feature sequence into multiple two-dimensional dilated convolutional layers to obtain a reference speech feature sequence;
[0135] An expanded speech feature sequence is determined according to the feedforward speech feature sequence and the reference speech feature sequence.
[0136] It should be noted that multiple interconnected two-dimensional dilated convolutional layers can be called memory layers. The two-dimensional dilated convolutional layer includes a dilation factor, which is used to increase the receptive field and perform context aggregation at different resolutions. The two-dimensional dilated convolutional layer is easy to implement a grouping mechanism. The speech embedding from the feedforward layer is initially two-dimensional and then reshaped into a three-dimensional sequence. The embedding dimension is further converted into input channels. These input channels are then divided into different groups, and each group uses a dedicated filter for convolution. Among them, the dilation factor is the parameter setting of the dilated convolution. The purpose of the dilated convolution is to expand the receptive field of the convolution. The larger the dilation factor, the larger the receptive field obtained. Because different dilation factors obtain different receptive fields, that is, contexts of different resolutions, different dilation factors are used to obtain different resolutions. The ultimate effect is to perform context aggregation at different resolutions.
[0137] Referring to FIG8 , FIG8 shows an architecture diagram of an expanded feedforward layer in a task processing method provided by an embodiment of the present specification. The expanded feedforward layer includes a feedforward layer and a memory layer. The memory layer includes a plurality of two-dimensional expanded convolutional layers connected to each other. The two-dimensional expanded convolutional layer includes expansion factors (d=1, d=2, ..., d=2 L-1 ), the feedforward layer includes a second linear layer, a fourth activation layer, and a fifth linear layer. First, the convolutional speech feature sequence is sequentially input into the second linear layer, the fourth activation layer, and the fifth linear layer to obtain a feedforward speech feature sequence. Second, the feedforward speech feature sequence is further input into multiple two-dimensional dilated convolutional layers to obtain a reference speech feature sequence. Finally, the feedforward speech feature sequence and the reference speech feature sequence are added together to output a dilated speech feature sequence.
[0138] By applying the solution of the embodiments of this specification, the convolution speech feature sequence is input into the feedforward layer to obtain the feedforward speech feature sequence; the feedforward speech feature sequence is input into multiple two-dimensional dilated convolution layers to obtain a reference speech feature sequence; and the dilated speech feature sequence is determined based on the feedforward speech feature sequence and the reference speech feature sequence, thereby improving the speech separation performance.
[0139] In an optional embodiment of the present specification, the above-mentioned inputting the feedforward speech feature sequence into multiple two-dimensional dilated convolutional layers to obtain the reference speech feature sequence may include the following steps:
[0140] Obtaining a candidate speech feature sequence corresponding to a target two-dimensional dilated convolutional layer, wherein the target two-dimensional dilated convolutional layer is the last of a plurality of two-dimensional dilated convolutional layers, and the candidate speech feature sequence is obtained based on speech feature sequences inputted by each two-dimensional dilated convolutional layer before the target two-dimensional dilated convolutional layer;
[0141] The candidate speech feature sequence is input into the target two-dimensional dilated convolutional layer to obtain the reference speech feature sequence.
[0142] It should be noted that the use of feedforward dense connections can be extended in the memory layer to establish interconnections between each 2D dilated convolution layer and all other 2D dilated convolution layers to enhance information flow and promote gradient propagation. Specifically, the input of the L-th 2D dilated convolution layer is the speech embedding of all previous 2D dilated convolution layers, that is, the output of the L-th 2D dilated convolution layer can be determined by the following formula (6): X L =H L ([X0,...,X L-1 ]) (6)
[0143] Among them, [X0,...,X L-1 ] refers to the concatenation of the inputs of the two-dimensional dilated convolutional layer preceding the L-th two-dimensional dilated convolutional layer, H L Represents a 2D dilated convolution operation.
[0144] By applying the solution of the embodiment of this specification, a candidate speech feature sequence corresponding to the target two-dimensional dilated convolution layer is obtained, and the candidate speech feature sequence is input into the target two-dimensional dilated convolution layer to obtain a reference speech feature sequence, thereby improving the speech separation efficiency.
[0145] In an optional embodiment of the present specification, the target two-dimensional dilated convolution layer includes a constant padding layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer, and a connection layer; the above-mentioned inputting the candidate speech feature sequence into the target two-dimensional dilated convolution layer to obtain the reference speech feature sequence may include the following steps:
[0146] The candidate speech feature sequence is sequentially input into the constant padding layer, the two-dimensional convolution layer, the instance normalization layer, and the first activation layer to obtain the initial reference speech feature sequence;
[0147] The initial reference speech feature sequence and the candidate speech feature sequence are input into the connection layer to obtain the reference speech feature sequence.
[0148] It's important to note that the 2D dilated convolutional layer performs convolution operations on features to model local features. The output of the 2D dilated convolutional layer is concatenated with the input to promote dense connections. Therefore, the dilated feedforward layer only uses forward connections and convolutional memory layers to capture recurrent patterns.
[0149] Referring to FIG9 , FIG9 shows an architecture diagram of a two-dimensional dilated convolution layer in a task processing method provided in one embodiment of this specification. As shown in FIG9 , the two-dimensional dilated convolution layer includes a constant padding layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer (PReLU activation layer), and a connection layer. The constant padding layer is used to perform a zero-adding operation in the feature sequence so that the sequence length remains unchanged after convolution; the two-dimensional convolution layer is used to perform a two-dimensional convolution operation; the instance normalization layer is used to normalize the features on each channel; and the first activation layer is used to perform nonlinear processing on the features. The feature dimension does not change in the two-dimensional dilated convolution layer.
[0150] Step 308: Separate the mixed speech feature sequence according to the speech separation information to obtain a task processing result.
[0151] In one or more embodiments of the present specification, mixed speech data is obtained, features are extracted from the mixed speech data to obtain a mixed speech feature sequence, and for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension. Further, the mixed speech feature sequence can be separated according to the speech separation information to obtain a task processing result, that is, to obtain separate speech data for each object.
[0152] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.
[0153] In an optional embodiment of the present specification, the separation of the mixed speech feature sequence according to the speech separation information to obtain the task processing result may include the following steps:
[0154] Filtering the target speech feature sequence corresponding to each object from the mixed speech feature sequence according to the speech separation information;
[0155] The target speech feature sequence is decoded to generate target speech data corresponding to each object.
[0156] Specifically, the target voice data refers to voice data of an individual object, that is, the target voice data only includes voice data of one object.
[0157] It should be noted that when the target speech feature sequence corresponding to each object is screened out from the mixed speech feature sequence based on the speech separation information, the tensor product of the speech separation information and the mixed speech feature sequence can be calculated to determine the target speech feature sequence corresponding to each object. There are many ways to decode the target speech feature sequence and generate the target speech data corresponding to each object. The specific selection is based on the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, traditional feature decoding methods such as filters, inverse spectral transforms, and wavelet transforms can be used. In another possible implementation of this specification, a deep learning model can be used to automatically learn the representation of the speech signal from the target speech feature sequence, and then a backpropagation algorithm can be used to restore the original speech signal. For example, the target speech feature sequence corresponding to each object is input into the decoding unit in the task processing model to obtain the target speech data corresponding to each object, wherein the decoding unit can use a transposed one-dimensional convolution layer with the same convolution kernel size and stride as the encoding unit to reconstruct the original speech signal of length T.
[0158] By applying the solution of the embodiment of this specification, the target speech feature sequence corresponding to each object is screened out from the mixed speech feature sequence according to the speech separation information; the target speech feature sequence is decoded to generate the target speech data corresponding to each object, thereby improving the speech separation performance.
[0159] Referring to FIG10 , FIG10 shows a flowchart of a task processing model training method provided by an embodiment of this specification, which specifically includes the following steps:
[0160] Step 1002: Acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects.
[0161] Step 1004: extract features from the sample mixed speech data to obtain a sample mixed speech feature sequence.
[0162] Step 1006: Input the sample mixed speech feature sequence into the loop processing unit in the task processing model to obtain the predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence.
[0163] Step 1008: Adjust the model parameters of the task processing model according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.
[0164] It should be noted that the implementation of step 1002, step 1004 and step 1006 is the same as the processing method of the above-mentioned task processing method, and will not be described in detail in the embodiment of this specification.
[0165] In practical applications, the speech separation label of the sample object is the speech separation target. The speech separation label can be an independent sample speech label of each sample object, or a sample separation information label corresponding to each sample object. The sample separation information label can be obtained based on the sample speech label encoding.
[0166] By applying the solution of the embodiments of this specification, since the predicted speech separation information is obtained by the cyclic processing unit based on the time dimension and the various feature dimensions of the sample mixed speech feature sequence, it fully takes into account the cyclic pattern of the speech signal in the time dimension and captures the complex time dependency within the speech signal, thereby improving the speech separation performance of the task processing model.
[0167] The following further illustrates the task processing method provided in this specification using the application of the task processing method in a smart conference scenario as an example, in conjunction with FIG11. FIG11 shows a flowchart of a conference speech separation method provided in one embodiment of this specification, which specifically includes the following steps:
[0168] Step 1102: Acquire mixed voice data, where the mixed voice data includes conference voice data of multiple objects.
[0169] Step 1104: Extract features from the mixed speech data to obtain a mixed speech feature sequence.
[0170] Step 1106: for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension.
[0171] Step 1108: Separate the mixed speech feature sequence according to the speech separation information to obtain a conference speech separation result.
[0172] It should be noted that the implementation of steps 1102 to 1108 is the same as the implementation of steps 302 to 308 described above, and this embodiment of the specification does not impose any limitation on this.
[0173] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.
[0174] Referring to Figure 12, Figure 12 shows a flowchart of the processing process of a task processing method provided by one embodiment of this specification. The goal of the task processing method is to use a task processing model to perform speech separation on given mixed speech data to obtain independent speech data for each object. The task processing model adopts a time-domain masking network framework, including an encoding unit, a masking unit, and a decoding unit. The masking unit includes an attention processing unit and a recurrent processing unit. The encoding unit and the decoding unit are responsible for feature extraction and waveform reconstruction of the speech signal. The masking unit is used to map the output of the encoding unit into a set of speech separation information, specifically including:
[0175] Coding unit: input the mixed speech data into the coding unit to obtain a mixed speech feature sequence;
[0176] Masking unit: inputs the mixed speech feature sequence into the attention processing unit, and obtains an updated mixed speech feature sequence based on the time dimension; for any feature dimension of the updated mixed speech feature sequence, inputs the updated mixed speech feature sequence into the loop processing unit in the task processing model to obtain the initial processing result corresponding to the feature dimension, and returns the initial processing result and the mixed speech feature sequence to the attention processing unit in the task processing model until a preset loop stop condition is reached, and obtains the processing result corresponding to each feature dimension, and generates speech separation information corresponding to each object according to the processing result corresponding to each feature dimension; calculates the separated tensor product based on the mixed speech feature sequence and the speech separation information to obtain the target speech feature sequence;
[0177] Decoding unit: The target speech feature sequence is input into the decoding unit to generate the target speech data corresponding to each object.
[0178] It should be noted that the dimension of the mixed speech data is Bx1xT, the dimension of the mixed speech feature sequence is BxNxS, the dimension of the updated mixed speech feature sequence is BxNxS, and the dimension of the target speech feature sequence is CxBxNxS, where C is the number of objects, B is the batch size, N is the feature dimension, and S is the number of feature sequence frames.
[0179] By applying the solution of the embodiments of this specification, the mixed speech feature sequence is subjected to joint attention processing by the attention processing unit, which is able to fully consider the local feature information of each object and its global feature information in the entire speech, making the initial speech feature sequence more accurate. In addition, the initial speech feature sequence is subjected to cyclic filtering processing by the cyclic processing unit, which fully considers the cyclic pattern of the speech signal in terms of speech structure, rhythm and semantic association, captures the complex time dependency within the speech signal, and improves the speech separation performance. In addition, the cyclic processing unit uses linear and convolution operations to process the entire sequence through forward connections without the need for block division, thereby ensuring parallel computing, thereby more effectively solving the problems faced by the speech separation task, such as unsatisfactory separation effect and susceptibility to noise reverberation interference.
[0180] Corresponding to the above-mentioned task processing method embodiment, this specification also provides a task processing device embodiment. FIG13 shows a schematic diagram of the structure of a task processing device provided in one embodiment of this specification. As shown in FIG13, the device includes:
[0181] A first acquisition module 1302 is configured to acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects;
[0182] The first extraction module 1304 is configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;
[0183] The first processing module 1306 is configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing results corresponding to each feature dimension;
[0184] The first separation module 1308 is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a task processing result.
[0185] Optionally, the first processing module 1306 is further configured to perform joint attention processing on the mixed speech feature sequence based on the time dimension to obtain an updated mixed speech feature sequence; for any feature dimension of the updated mixed speech feature sequence, the mixed speech feature sequence is subjected to cyclic filtering processing based on the time dimension until a preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.
[0186] Optionally, the first processing module 1306 is further configured to input the mixed speech feature sequence into the attention processing unit in the task processing model, and obtain an updated mixed speech feature sequence based on the time dimension; for any feature dimension of the updated mixed speech feature sequence, input the updated mixed speech feature sequence into the loop processing unit in the task processing model, obtain the initial processing result corresponding to the feature dimension, and return the initial processing result and the mixed speech feature sequence to the attention processing unit in the task processing model until the preset loop stop condition is reached, and the processing results corresponding to each feature dimension are obtained.
[0187] Optionally, the cyclic processing unit includes a bottleneck layer, a gated convolution layer and an output layer; the first processing module 1306 is further configured to input the updated mixed speech feature sequence into the bottleneck layer to obtain a first speech feature sequence, wherein the bottleneck layer is used to reduce the feature dimension of the updated mixed speech feature sequence; input the first speech feature sequence into the gated convolution layer to obtain a second speech feature sequence, wherein the gated convolution layer is used to perform cyclic filtering processing on the first speech feature sequence based on the time dimension; input the second speech feature sequence into the output layer to obtain an initial processing result.
[0188] Optionally, the gated convolution layer includes an expanded feedforward layer, a convolution layer and a processing layer; the first processing module 1306 is further configured to input the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence; input the convolution speech feature sequence into the expanded feedforward layer to obtain an expanded speech feature sequence; input the first speech feature sequence, the convolution speech feature sequence and the expanded speech feature sequence into the processing layer to obtain a second speech feature sequence.
[0189] Optionally, the expanded feedforward layer includes a feedforward layer and multiple two-dimensional expanded convolutional layers connected to each other; the first processing module 1306 is further configured to input the convolution speech feature sequence into the feedforward layer to obtain a feedforward speech feature sequence; input the feedforward speech feature sequence into multiple two-dimensional expanded convolutional layers to obtain a reference speech feature sequence; and determine the expanded speech feature sequence based on the feedforward speech feature sequence and the reference speech feature sequence.
[0190] Optionally, the first processing module 1306 is further configured to obtain a candidate speech feature sequence corresponding to the target two-dimensional dilated convolution layer, wherein the target two-dimensional dilated convolution layer is the last of multiple two-dimensional dilated convolution layers, and the candidate speech feature sequence is obtained based on the speech feature sequences input by each two-dimensional dilated convolution layer before the target two-dimensional dilated convolution layer; the candidate speech feature sequence is input into the target two-dimensional dilated convolution layer to obtain a reference speech feature sequence.
[0191] Optionally, the target two-dimensional dilated convolution layer includes a constant padding layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer and a connection layer; the first processing module 1306 is further configured to input the candidate speech feature sequence into the constant padding layer, the two-dimensional convolution layer, the instance normalization layer and the first activation layer in sequence to obtain an initial reference speech feature sequence; and input the initial reference speech feature sequence and the candidate speech feature sequence into the connection layer to obtain a reference speech feature sequence.
[0192] Optionally, the convolution layer includes a first normalization layer, a first linear layer, a second activation layer, and a deep convolution layer; the first processing module 1306 is further configured to input the first speech feature sequence into the first normalization layer, the first linear layer, and the second activation layer in sequence to obtain an initial convolution speech feature sequence; input the initial convolution speech feature sequence into the deep convolution layer to obtain a reference convolution speech feature sequence; and determine the convolution speech feature sequence based on the initial convolution speech feature sequence and the reference convolution speech feature sequence.
[0193] Optionally, the first separation module 1308 is further configured to filter out a target speech feature sequence corresponding to each object from the mixed speech feature sequence according to the speech separation information; and decode the target speech feature sequence to generate target speech data corresponding to each object.
[0194] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.
[0195] The above is a schematic scheme of a task processing device of this embodiment. It should be noted that the technical scheme of the task processing device and the technical scheme of the task processing method described above are of the same concept. For details not described in detail in the technical scheme of the task processing device, please refer to the description of the technical scheme of the task processing method described above.
[0196] Corresponding to the above-mentioned task processing model training method embodiment, this specification also provides a task processing model training device embodiment. Figure 14 shows a schematic diagram of the structure of a task processing model training device provided by one embodiment of this specification. As shown in Figure 14, the device includes:
[0197] The second acquisition module 1402 is configured to acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects;
[0198] The second extraction module 1404 is configured to perform feature extraction on the sample mixed speech data to obtain a sample mixed speech feature sequence;
[0199] Input module 1406 is configured to input the sample mixed speech feature sequence into the recurrent processing unit in the task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the recurrent processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence;
[0200] The adjustment module 1408 is configured to adjust the model parameters of the task processing model according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.
[0201] By applying the solution of the embodiments of this specification, since the predicted speech separation information is obtained by the cyclic processing unit based on the time dimension and the various feature dimensions of the sample mixed speech feature sequence, it fully takes into account the cyclic pattern of the speech signal in the time dimension and captures the complex time dependency within the speech signal, thereby improving the speech separation performance of the task processing model.
[0202] The above is a schematic diagram of a task processing model training device according to this embodiment. It should be noted that the technical solution of the task processing model training device and the technical solution of the task processing model training method described above are based on the same concept. For details not described in detail in the technical solution of the task processing model training device, please refer to the description of the technical solution of the task processing model training method described above.
[0203] Corresponding to the above-mentioned conference speech separation method embodiment, this specification also provides a conference speech separation device embodiment. FIG15 shows a schematic diagram of the structure of a conference speech separation device provided in one embodiment of this specification. As shown in FIG15 , the device includes:
[0204] The third acquisition module 1502 is configured to acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple subjects;
[0205] The third extraction module 1504 is configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;
[0206] The second processing module 1506 is configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing results corresponding to each feature dimension;
[0207] The second separation module 1508 is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a conference speech separation result.
[0208] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.
[0209] The above is a schematic diagram of a conference voice separation device according to this embodiment. It should be noted that the technical solution of the conference voice separation device and the technical solution of the conference voice separation method described above are based on the same concept. For details not described in detail in the technical solution of the conference voice separation device, please refer to the description of the technical solution of the conference voice separation method described above.
[0210] Figure 16 shows a block diagram of a computing device according to one embodiment of the present disclosure. Components of the computing device 1600 include, but are not limited to, a memory 1610 and a processor 1620. The processor 1620 is connected to the memory 1610 via a bus 1630, and a database 1650 is used to store data.
[0211] Computing device 1600 also includes an access device 1640 that enables computing device 1600 to communicate via one or more networks 1660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 1640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0212] In one embodiment of the present specification, the aforementioned components of computing device 1600 and other components not shown in FIG16 may also be connected to each other, for example, via a bus. It should be understood that the block diagram of the computing device structure shown in FIG16 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0213] Computing device 1600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). Computing device 1600 may also be a mobile or stationary server.
[0214] Among them, the processor 1620 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned task processing method or task processing model training method or conference speech separation method.
[0215] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of this computing device is based on the same concept as the technical schemes of the task processing method, task processing model training method, and conference speech separation method described above. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical schemes of the task processing method, task processing model training method, or conference speech separation method described above.
[0216] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned task processing method or task processing model training method or conference speech separation method.
[0217] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium is based on the same concept as the technical schemes of the task processing method, task processing model training method, and conference speech separation method described above. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical schemes of the task processing method, task processing model training method, or conference speech separation method described above.
[0218] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned task processing method or task processing model training method or conference speech separation method.
[0219] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of this computer program is based on the same concept as the technical schemes of the task processing method, task processing model training method, and conference speech separation method described above. For details not described in detail in the technical scheme of the computer program, please refer to the description of the technical schemes of the task processing method, task processing model training method, or conference speech separation method described above.
[0220] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0221] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0222] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0223] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0224] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A task processing method, comprising: Acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects; Extracting features from the mixed speech data to obtain a mixed speech feature sequence; For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing result corresponding to each feature dimension; The mixed speech feature sequence is separated according to the speech separation information to obtain a task processing result.
2. The method according to claim 1, wherein for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, comprising: Based on the time dimension, performing joint attention processing on the mixed speech feature sequence to obtain an updated mixed speech feature sequence; For any feature dimension of the updated mixed speech feature sequence, a loop filtering process is performed on the mixed speech feature sequence based on the time dimension until a preset loop stop condition is reached, thereby obtaining a processing result corresponding to each feature dimension.
3. The method according to claim 2, wherein the step of performing joint attention processing on the mixed speech feature sequence based on the time dimension to obtain an updated mixed speech feature sequence comprises: Inputting the mixed speech feature sequence into the attention processing unit in the task processing model to obtain an updated mixed speech feature sequence based on the time dimension; The step of performing loop filtering on any feature dimension of the updated mixed speech feature sequence based on the time dimension until a preset loop stop condition is reached to obtain a processing result corresponding to each feature dimension includes: For any feature dimension of the updated mixed speech feature sequence, the updated mixed speech feature sequence is input into the loop processing unit in the task processing model to obtain the initial processing result corresponding to the feature dimension, and the initial processing result and the mixed speech feature sequence are returned to the attention processing unit in the task processing model until the preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.
4. The method according to claim 3, wherein the loop processing unit comprises a bottleneck layer, a gated convolution layer and an output layer; The step of inputting the updated mixed speech feature sequence into a loop processing unit in the task processing model to obtain an initial processing result corresponding to the feature dimension includes: Inputting the updated mixed speech feature sequence into the bottleneck layer to obtain a first speech feature sequence, wherein the bottleneck layer is used to reduce the feature dimension of the updated mixed speech feature sequence; Inputting the first speech feature sequence into the gated convolution layer to obtain a second speech feature sequence, wherein the gated convolution layer is used to perform a loop filtering process on the first speech feature sequence based on the time dimension; The second speech feature sequence is input into the output layer to obtain an initial processing result.
5. The method according to claim 4, wherein the gated convolutional layer comprises a dilated feed-forward layer, a convolutional layer, and a processing layer; The step of inputting the first speech feature sequence into the gated convolutional layer to obtain a second speech feature sequence comprises: Inputting the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence; Inputting the convolution speech feature sequence into the dilated feedforward layer to obtain a dilated speech feature sequence; The first speech feature sequence, the convolution speech feature sequence and the expansion speech feature sequence are input into the processing layer to obtain a second speech feature sequence.
6. The method according to claim 5, wherein the dilated feed-forward layer comprises a feed-forward layer and a plurality of two-dimensional dilated convolutional layers connected to each other; The step of inputting the convolution speech feature sequence into the dilated feedforward layer to obtain the dilated speech feature sequence comprises: Inputting the convolution speech feature sequence into the feedforward layer to obtain a feedforward speech feature sequence; Inputting the feedforward speech feature sequence into the multiple two-dimensional dilated convolutional layers to obtain a reference speech feature sequence; An expanded speech feature sequence is determined according to the feedforward speech feature sequence and the reference speech feature sequence.
7. The method according to claim 6, wherein the step of inputting the feedforward speech feature sequence into the plurality of two-dimensional dilated convolutional layers to obtain a reference speech feature sequence comprises: Obtaining a candidate speech feature sequence corresponding to a target two-dimensional dilated convolutional layer, wherein the target two-dimensional dilated convolutional layer is the last one of the multiple two-dimensional dilated convolutional layers, and the candidate speech feature sequence is obtained based on speech feature sequences input by each two-dimensional dilated convolutional layer before the target two-dimensional dilated convolutional layer; The candidate speech feature sequence is input into the target two-dimensional dilated convolutional layer to obtain a reference speech feature sequence.
8. The method according to claim 7, wherein the target two-dimensional dilated convolution layer comprises a constant padding layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer, and a connection layer; The step of inputting the candidate speech feature sequence into the target two-dimensional dilated convolutional layer to obtain a reference speech feature sequence comprises: Inputting the candidate speech feature sequence into the constant padding layer, the two-dimensional convolution layer, the instance normalization layer and the first activation layer in sequence to obtain an initial reference speech feature sequence; The initial reference speech feature sequence and the candidate speech feature sequence are input into the connection layer to obtain a reference speech feature sequence.
9. The method according to claim 5, wherein the convolution layer comprises a first normalization layer, a first linear layer, a second activation layer, and a depth convolution layer; The step of inputting the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence comprises: Inputting the first speech feature sequence into the first normalization layer, the first linear layer and the second activation layer in sequence to obtain an initial convolution speech feature sequence; Inputting the initial convolution speech feature sequence into the deep convolution layer to obtain a reference convolution speech feature sequence; According to the initial convolution speech feature sequence and the reference convolution speech feature sequence, a convolution speech feature sequence is determined.
10. The method according to any one of claims 1 to 9, wherein the step of separating the mixed speech feature sequence according to the speech separation information to obtain a task processing result comprises: Filtering the target speech feature sequence corresponding to each object from the mixed speech feature sequence according to the speech separation information; The target speech feature sequence is decoded to generate target speech data corresponding to each object.
11. A task processing model training method, comprising: Acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects; Extracting features from the sample mixed speech data to obtain a sample mixed speech feature sequence; Inputting the sample mixed speech feature sequence into a loop processing unit in a task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence; The model parameters of the task processing model are adjusted according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.
12. A conference voice separation method, comprising: Acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects; Extracting features from the mixed speech data to obtain a mixed speech feature sequence; For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing result corresponding to each feature dimension; The mixed speech feature sequence is separated according to the speech separation information to obtain a conference speech separation result.
13. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 10 or claim 11 or claim 12 are implemented.
14. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1 to 10 or claim 11 or claim 12.
15. A computer program, wherein: When the computer program is executed in a computer, the computer is caused to execute the steps of implementing the method described in any one of claims 1 to 10 or claim 11 or claim 12.
Citation Information
Patent Citations
Single-channel voice separation method and device for multiple speakers
CN112331218A
Voice separation method and device, computer equipment and storage medium
CN114724579A
Voice extraction method and device, neural network model training method and device and storage medium
CN115116448A
Speech separation method
CN116168717A
Auditory selection method and device based on memory and attention model
US20200227064A1
Cited By
Confirmation method and device of call key point, equipment, storage medium and program product
CN120932635A