Task processing method, task processing model training method and conference voice separation method

By performing feature extraction and time-dimensional processing of mixed speech data, voice separation information is generated and separated, the problem of large positioning errors in speech separation of multiple speakers is solved, and the speech separation performance is improved.

CN119943082APending Publication Date: 2025-05-06HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311465109.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has large positioning errors in the speech separation of multiple speakers, resulting in poor speech separation performance.

Method used

By obtaining mixed speech data, feature extraction and time dimension processing are performed, voice separation information is generated, and the speech feature sequence is separated based on this information to achieve task processing results.

Benefits of technology

The performance and effect of speech separation are improved by considering the cyclic mode and complex time dependencies of speech signals in the time dimension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943082A_ABST
    Figure CN119943082A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task processing method, a task processing model training method and a conference voice separation method, and the task processing method comprises the steps: obtaining mixed voice data which comprises voice data of a plurality of objects; performing feature extraction on the mixed voice data to obtain a mixed voice feature sequence; for any feature dimension of the mixed voice feature sequence, processing the mixed voice feature sequence based on a time dimension, and generating voice separation information corresponding to each object according to a processing result corresponding to each feature dimension; and separating the mixed voice feature sequence according to the voice separation information to obtain a task processing result. According to the method, the mixed voice feature sequence is processed based on the time dimension, the circulation mode of the voice signal in the time dimension is fully considered, and the complex time dependency relationship in the voice signal is captured, so that the voice separation performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to task processing, task processing model training, and conference speech separation methods. Background Art

[0002] With the development of computer technology, there are more and more scenarios where multiple people communicate at the same time. After collecting the voices of multiple people communicating at the same time through microphones, if the voice separation is not performed, it will directly affect the speech recognition system or auditory perception and understanding. Therefore, speech separation technology has gradually become a research focus.

[0003] At present, the target speech is usually separated from the speech of multiple speakers through the sound source localization method. However, there is often a large positioning error in the sound source localization process, resulting in poor speech separation performance. Therefore, a high-performance speech separation solution is urgently needed. Summary of the invention

[0004] In view of this, an embodiment of this specification provides a task processing method. One or more embodiments of this specification also involve a task processing model training method, a conference voice separation method, a task processing device, a task processing model training device, a conference voice separation device, a computing device, a computer-readable storage medium and a computer program to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a task processing method is provided, including:

[0006] Acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects;

[0007] Extracting features from mixed speech data to obtain a mixed speech feature sequence;

[0008] For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension;

[0009] The mixed speech feature sequence is separated according to the speech separation information to obtain the task processing result.

[0010] According to a second aspect of an embodiment of this specification, a task processing model training method is provided, including:

[0011] Acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects;

[0012] Extracting features from the sample mixed speech data to obtain a sample mixed speech feature sequence;

[0013] Inputting the sample mixed speech feature sequence into the loop processing unit in the task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence;

[0014] The model parameters of the task processing model are adjusted according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.

[0015] According to a third aspect of an embodiment of this specification, a conference voice separation method is provided, including:

[0016] Acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects;

[0017] Extracting features from mixed speech data to obtain a mixed speech feature sequence;

[0018] For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension;

[0019] The mixed speech feature sequence is separated according to the speech separation information to obtain the conference speech separation result.

[0020] According to a fourth aspect of the embodiments of this specification, there is provided a task processing device, including:

[0021] A first acquisition module is configured to acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects;

[0022] A first extraction module is configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;

[0023] A first processing module is configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing result corresponding to each feature dimension;

[0024] The first separation module is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a task processing result.

[0025] According to a fifth aspect of an embodiment of this specification, a task processing model training device is provided, comprising:

[0026] A second acquisition module is configured to acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects;

[0027] A second extraction module is configured to perform feature extraction on the sample mixed speech data to obtain a sample mixed speech feature sequence;

[0028] An input module is configured to input the sample mixed speech feature sequence into a loop processing unit in the task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence;

[0029] The adjustment module is configured to adjust the model parameters of the task processing model according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.

[0030] According to a sixth aspect of the embodiments of this specification, a conference voice separation device is provided, including:

[0031] A third acquisition module is configured to acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects;

[0032] A third extraction module is configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;

[0033] The second processing module is configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing result corresponding to each feature dimension;

[0034] The second separation module is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a conference speech separation result.

[0035] According to a seventh aspect of an embodiment of this specification, there is provided a computing device, including:

[0036] Memory and processor;

[0037] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.

[0038] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.

[0039] According to a ninth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect, the second aspect, or the third aspect above.

[0040] A task processing method provided in an embodiment of the present specification obtains mixed speech data, wherein the mixed speech data includes speech data of multiple objects; performs feature extraction on the mixed speech data to obtain a mixed speech feature sequence; processes the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generates speech separation information corresponding to each object according to the processing results corresponding to each feature dimension; separates the mixed speech feature sequence according to the speech separation information to obtain a task processing result. By processing the mixed speech feature sequence based on the time dimension, the cyclic pattern of the speech signal in the time dimension is fully considered, and the complex time dependency within the speech signal is captured, thereby improving the speech separation performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is an architecture diagram of a task processing system provided by an embodiment of this specification;

[0042] Figure 2 is an architecture diagram of another task processing system provided by an embodiment of this specification;

[0043] Figure 3 is a flowchart of a task processing method provided by an embodiment of this specification;

[0044] Figure 4 It is a schematic diagram of a processing method of an attention processing unit in a task processing method provided by an embodiment of this specification;

[0045] Figure 5 It is a schematic diagram of a processing method of a loop processing unit in a task processing method provided by an embodiment of this specification;

[0046] Figure 6 It is an architecture diagram of a loop processing unit in a task processing method provided by one embodiment of this specification;

[0047] Figure 7 It is an architecture diagram of a convolutional layer in a task processing method provided by an embodiment of this specification;

[0048] Figure 8 It is an architecture diagram of an expanded feedforward layer in a task processing method provided in one embodiment of the present specification;

[0049] Fig. 9 It is an architecture diagram of a two-dimensional dilated convolution layer in a task processing method provided by an embodiment of this specification;

[0050] Fig.10 is a flowchart of a task processing model training method provided by an embodiment of this specification;

[0051] Fig.11 is a flow chart of a conference speech separation method provided by an embodiment of this specification;

[0052] Fig.12 is a processing flow chart of a task processing method provided by an embodiment of this specification;

[0053] Fig.13 is a structural diagram of a task processing device provided by an embodiment of this specification;

[0054] Fig.14 It is a structural diagram of a task processing model training device provided by an embodiment of this specification;

[0055] Fig.15 It is a structural schematic diagram of a conference voice separation device provided by an embodiment of this specification;

[0056] Fig.16 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0057] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0058] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0059] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0060] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0061] First, the terms involved in one or more embodiments of this specification are explained.

[0062] Speech separation technology: Speech separation refers to separating the mixed speech of multiple speakers and obtaining the individual speech of all speakers.

[0063] Self-attention mechanism: Self-attention is a neural network mechanism for processing sequence data. It allows the model to consider every element in the sequence at the same time when processing sequence data, not just the adjacent elements. This mechanism can be used in many natural language processing tasks, such as text classification, machine translation, and question answering systems.

[0064] Deep learning algorithm: Deep learning algorithm is a computational model that imitates the neurons in the human brain and is used to handle complex machine learning tasks. It builds a multi-layer neural network to learn the characteristics and patterns of data, thereby achieving tasks such as data classification, prediction, and generation. Deep learning algorithms have been widely used in many fields, including computer vision, natural language processing, and speech recognition.

[0065] The "cocktail party problem" refers to the problem of how to effectively identify and understand the target speaker's voice from mixed sounds in a noisy environment. This problem is an important issue in the field of computer speech recognition because it requires the model to be able to accurately recognize the target voice in complex background noise. In a scenario where multiple people are communicating at the same time, if the voices of multiple people communicating at the same time are not separated after being collected through a microphone, it will directly affect the speech recognition system or auditory perception and understanding. Therefore, speech separation technology has gradually become a research focus. The purpose of speech separation is to separate the mixed voices of multiple speakers. The separation results can be used as input signals for speech recognition or directly played to the listener, thereby improving the recognition results and auditory perception.

[0066] At present, speech separation can usually be performed in the following ways: the first is to use additional speaker labels during model training, so as to use the trained model for speech separation; the second is to perform speech separation based only on convolutional networks; the third is to truncate long speech sequences into short sequences, and then perform intra-sequence and inter-sequence attention processing. However, the additional speaker labels in the first method increase the model training cost; the second method does not use more advanced attention mechanisms, and cannot process the global information of the sequence; the third method still processes the global information through implicit indirect interactions, which leads to a certain performance degradation and lacks the ability to model locally. At the same time, the third method uses dual channels, which introduces significant processing overhead because there are overlapping blocks and is inefficient in modeling global information through cross processing. In general, the above schemes pay more attention to long-range and coarse-grained dependencies, resulting in speech separation effects that are still not ideal.

[0067] In order to solve the above problems, the embodiments of this specification propose a speech separation solution that focuses on fine-grained cyclic patterns in the speech separation process, which can be applied to the recognition or playback of recorded speech in multi-object speaking scenarios such as conference scenarios and cocktail parties. Specifically, the embodiments of this specification propose a task processing method, which obtains mixed speech data, wherein the mixed speech data includes speech data of multiple objects; performs feature extraction on the mixed speech data to obtain a mixed speech feature sequence; for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension; the mixed speech feature sequence is separated according to the speech separation information to obtain the task processing result. By processing the mixed speech feature sequence based on the time dimension, the cyclic pattern of the speech signal in the time dimension is fully taken into account, and the complex time dependency within the speech signal is captured, thereby improving the speech separation performance.

[0068] In this specification, a task processing method is provided. This specification also involves a task processing model training method, a conference speech separation method, a task processing device, a task processing model training device, a conference speech separation device, a computing device, a computer-readable storage medium and a computer program, which are described in detail one by one in the following embodiments.

[0069] See also Figure 1 , Figure 1 An architecture diagram of a task processing system provided by an embodiment of the present specification is shown, and the task processing system may include a client 100 and a server 200;

[0070] The client 100 is used to send mixed voice data to the server 200, wherein the mixed voice data includes voice data of multiple objects;

[0071] The server 200 is used to extract features from the mixed speech data to obtain a mixed speech feature sequence; for any feature dimension of the mixed speech feature sequence, process the mixed speech feature sequence based on the time dimension, and generate speech separation information corresponding to each object according to the processing results corresponding to each feature dimension; separate the mixed speech feature sequence according to the speech separation information to obtain a task processing result; and send the task processing result to the client 100;

[0072] The client 100 is also used to receive the task processing result sent by the server 200.

[0073] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.

[0074] See also Figure 2 , Figure 2 The architecture diagram of another task processing system provided by an embodiment of the present specification is shown. The task processing system may include multiple clients 100 and a server 200, wherein the client 100 may be a terminal device and the server 200 may be a cloud device. A communication connection may be established between multiple clients 100 through the server 200. In the speech separation task processing scenario, the server 200 is used to provide speech separation services between multiple clients 100. Multiple clients 100 may serve as a sender or a receiver, respectively, and realize communication through the server 200.

[0075] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the speech separation task processing scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 generates a task processing result based on the data stream and pushes the task processing result to other clients that have established communication.

[0076] The client 100 and the server 200 are connected via a network. The network provides a medium for a communication link between the client 100 and the server 200. The network may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, etc. before being released to the server 200.

[0077] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language Version 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 100 can be based on the software development kit (SDK, Software Development Kit) of the corresponding service provided by the server 200, such as based on the real-time communication (RTC, Real Time Communication) SDK development and acquisition. The client 100 can be deployed in an electronic device, and needs to rely on the device to run or some APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0078] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers for background training that provide support for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0079] It is worth noting that the task processing method provided in the embodiments of this specification is generally executed by the server, but in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the task processing method provided in the embodiments of this specification. In other embodiments, the task processing method provided in the embodiments of this specification may also be jointly executed by the client and the server.

[0080] See also Figure 3 , Figure 3 A flowchart of a task processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0081] Step 302: Acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects.

[0082] In one or more embodiments of the present specification, during task processing, mixed voice data corresponding to a target task may be acquired, thereby separating the mixed voice data to obtain separate voice data of each object after separation.

[0083] Specifically, mixed voice data can be data from different scenes, such as mixed voice data from a conference scene, mixed voice data from a cocktail party scene. Mixed voice data refers to voice data that contains two or more voices at the same time, which can be called a mixed voice waveform. The voice data of each object in the mixed voice data may be entangled with each other. Object refers to the source of voice data, which can be a virtual character, a recorder, etc. Voice data can be called a voice signal, including but not limited to spoken data and music data.

[0084] In practical applications, there are many ways to obtain mixed voice data, which can be selected according to actual conditions, and the embodiments of this specification do not limit this. In one possible implementation of this specification, mixed voice data sent by a user through a client can be received. In another possible implementation of this specification, mixed voice data can be read from other data acquisition devices or databases.

[0085] Step 304: extract features from the mixed speech data to obtain a mixed speech feature sequence.

[0086] In one or more embodiments of the present specification, after the mixed speech data is acquired, further, feature extraction may be performed on the mixed speech data to obtain a mixed speech feature sequence.

[0087] Specifically, the mixed speech feature sequence includes a set of feature vectors, each of which represents a small part of the speech signal. The mixed speech feature sequence is used to represent the attribute characteristics of the mixed speech data, such as speech frequency, energy, tone, etc.

[0088] In practical applications, there are many ways to extract features from mixed speech data and obtain mixed speech feature sequences. The specific selection depends on the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, traditional feature extraction methods can be used, such as filters, statistical feature extraction, and wavelet transforms. In another possible implementation of this specification, a deep learning model can be used to automatically learn and generate mixed speech feature sequences from mixed speech data, such as inputting mixed speech data into a coding unit in a task processing model to obtain a mixed speech feature sequence.

[0089] In an optional embodiment of the present specification, the encoding unit can use a one-dimensional (1-D) convolutional layer (Conv1 D) combined with a rectified linear unit (ReLU) to input the mixed speech waveform x∈R 1xT Encoding is performed, and during the encoding process, the mixed speech waveform can be converted into a non-negative speech embedding sequence X∈R N×S , which is the mixed speech feature sequence. Assuming that the convolution kernel size of the encoding unit is K1 and the stride is K1 / 2, the length of the encoded mixed speech feature sequence can be calculated by the following formula (1):

[0090] S=2T / K1-1 (1)

[0091] Where T represents the length of the mixed speech waveform, N represents the embedding dimension, and S represents the length of the mixed speech feature sequence.

[0092] Step 306: for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing result corresponding to each feature dimension.

[0093] In one or more embodiments of the present specification, after obtaining mixed speech data, performing feature extraction on the mixed speech data, and obtaining a mixed speech feature sequence, further, for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension.

[0094] Specifically, the dimensions of speech data include frequency dimension and time dimension, wherein the frequency dimension can be understood as feature dimension. Speech separation information can be understood as mask information, which corresponds to an object one by one and is used to separate the individual speech data of each object from the mixed speech data.

[0095] It should be noted that, for any feature dimension of the mixed speech feature sequence, the process of processing the mixed speech feature sequence based on the time dimension can be understood as fine-grained cyclic processing of the mixed speech feature sequence from both the frequency dimension and the time dimension. After obtaining the processing results corresponding to each feature dimension, the processing results corresponding to each feature dimension can be directly used as the speech separation information corresponding to each object. Furthermore, the data identifier corresponding to each processing result can be obtained from the preset masking identifier library, and the data identifier can be used as the speech separation information.

[0096] In practical applications, there are multiple ways to process a mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, which are selected according to the actual situation, and the embodiments of this specification do not limit this. In a possible implementation of this specification, for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence can be processed based on the time dimension for loop filtering until a preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.

[0097] In another possible implementation of the present specification, the mixed speech feature sequence can be updated by using the joint attention processing mechanism, and then the updated mixed speech feature sequence can be subjected to loop filtering. That is, the mixed speech feature sequence can be processed based on the time dimension for any feature dimension of the mixed speech feature sequence, which can include the following steps:

[0098] Based on the time dimension, the mixed speech feature sequence is processed with joint attention to obtain an updated mixed speech feature sequence;

[0099] For any feature dimension of the updated mixed speech feature sequence, the mixed speech feature sequence is subjected to cyclic filtering processing based on the time dimension until a preset cyclic stopping condition is reached, thereby obtaining processing results corresponding to each feature dimension.

[0100] It should be noted that in order to achieve efficient self-attention on a wide range of mixed speech feature sequences, the embodiments of this specification apply self-attention to the entire mixed speech feature sequence through joint attention processing, wherein the joint attention processing process adopts a joint local-global self-attention strategy.

[0101] In practical applications, when performing joint attention processing on the mixed speech feature sequence based on the time dimension, full computational self-attention can be performed on non-overlapping local segments at the same time, and a linearized and resource-efficient self-attention mechanism can be used on the entire mixed speech feature sequence to obtain an updated mixed speech feature sequence.

[0102] Furthermore, since speech data inherently has cyclic patterns manifested in speech structure, rhythm, and semantic associations, these patterns play an important role in speech separation. Therefore, the complex time dependencies within the speech signal can be modeled. Assuming that different embedding layers retain different cyclic patterns, cyclic filtering learning can be performed on each feature dimension, and all feature dimensions can be learned in parallel until the preset cycle stop condition is reached, and the processing results corresponding to each feature dimension can be obtained, wherein the preset cycle stop condition includes but is not limited to the number of cycles reaching the preset number of cycles, wherein the preset number of cycles, such as 20 times, is selected according to the actual situation, and the embodiments of this specification do not impose any restrictions on this.

[0103] By applying the solution of the embodiment of this specification, based on the time dimension, the mixed speech feature sequence is subjected to joint attention processing to obtain an updated mixed speech feature sequence; for any feature dimension of the updated mixed speech feature sequence, the mixed speech feature sequence is subjected to cyclic filtering processing based on the time dimension until a preset cyclic stop condition is reached to obtain the processing results corresponding to each feature dimension. By leveraging the combined advantages of self-attention and cyclic modeling, it is helpful to capture the extensive dependencies and local cyclic patterns of speech data and improve speech separation performance.

[0104] In an optional embodiment of the present specification, the task processing model can be used to perform joint attention processing and loop filtering processing, that is, the above-mentioned joint attention processing is performed on the mixed speech feature sequence based on the time dimension to obtain an updated mixed speech feature sequence, which can include the following steps:

[0105] Inputting the mixed speech feature sequence into the attention processing unit in the task processing model, and obtaining an updated mixed speech feature sequence based on the time dimension;

[0106] For any feature dimension of the updated mixed speech feature sequence, performing cyclic filtering processing on the mixed speech feature sequence based on the time dimension until a preset cyclic stop condition is reached to obtain processing results corresponding to each feature dimension may include the following steps:

[0107] For any feature dimension of the updated mixed speech feature sequence, the updated mixed speech feature sequence is input into the loop processing unit in the task processing model to obtain the initial processing result corresponding to the feature dimension, and the initial processing result and the mixed speech feature sequence are returned to the attention processing unit in the task processing model until the preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.

[0108] Specifically, the attention processing unit can consider each element in the current layer's mixed speech feature sequence to learn the embedding of each subsequent layer, independent of the loop, where self-attention mainly captures long-range, coarse-grained dependencies. The loop processing unit is used to perform loop learning on each feature dimension, and each feature dimension can be learned in parallel.

[0109] For example, see Figure 4 , Figure 4 A schematic diagram of a processing method of an attention processing unit in a task processing method provided by an embodiment of this specification is shown. Figure 4 As shown, the attention processing unit determines the target local segment ( Figure 4 The embedding of the blocks with black slashes in the figure is done in the time dimension based on the global information in the layers that have completed learning. Figure 5 , Figure 5 A schematic diagram of a processing method of a loop processing unit in a task processing method provided by an embodiment of this specification is shown. Figure 5 As shown, for the target feature dimension ( Figure 5 The loop processing unit can perform loop filtering on the mixed speech feature sequence in the time dimension until a preset loop stop condition is reached to obtain a processing result corresponding to the target feature dimension.

[0110] Applying the scheme of the embodiment of this specification, the mixed speech feature sequence is input into the attention processing unit in the task processing model, and an updated mixed speech feature sequence is obtained based on the time dimension; for any feature dimension of the updated mixed speech feature sequence, the updated mixed speech feature sequence is input into the loop processing unit in the task processing model, and the initial processing result corresponding to the feature dimension is obtained, and the initial processing result and the mixed speech feature sequence are input into the attention processing unit in the task processing model, until the preset loop stop condition is reached, and the processing result corresponding to each feature dimension is obtained. By leveraging the combined advantages of self-attention and loop modeling, it is helpful to capture the extensive dependencies and local loop patterns of speech data and improve speech separation performance.

[0111] In an optional embodiment of the present specification, a recurrent processing unit can be constructed based on a long short-term memory network (LSTM) or a gated recurrent unit (GRU), but since both of the above methods are sequential processing, the amount of calculation is large and the processing speed is slow, resulting in obvious disadvantages in the processing process. Therefore, a recurrent neural network (RNN, Recurrent Neural Network) recurrent processing unit based on a feedforward sequential memory network (FSMN, Feedforward Sequential Memory Networks) can be used. Specifically, the recurrent processing unit includes a bottleneck layer, a gated convolutional layer, and an output layer; the above-mentioned input of the updated mixed speech feature sequence into the recurrent processing unit in the task processing model to obtain the initial processing result corresponding to the feature dimension may include the following steps:

[0112] Inputting the updated mixed speech feature sequence into the bottleneck layer to obtain a first speech feature sequence, wherein the bottleneck layer is used to reduce the feature dimension of the updated mixed speech feature sequence;

[0113] Inputting the first speech feature sequence into a gated convolutional layer to obtain a second speech feature sequence, wherein the gated convolutional layer is used to perform a cyclic filtering process on the first speech feature sequence based on a time dimension;

[0114] The second speech feature sequence is input into the output layer to obtain an initial processing result.

[0115] See also Figure 6 , Figure 6An architecture diagram of a loop processing unit in a task processing method provided by an embodiment of the present specification is shown, wherein the loop processing unit includes a bottleneck layer (Bottleneck layer), a gated convolution layer (GCUlayer, GateControl Unit layer) and an output layer (Output layer); the bottleneck layer includes a 1×1 convolution kernel, a third activation layer (PReLU activation) and a second normalization layer (LayerNorm layer); the gated convolution layer includes a dilated feedforward layer (DiliatedFSMN), two convolution layers (Convolution-U) and a processing layer, in which addition and tensor product operations can be performed; the output layer includes a third normalization layer and a 1×1 convolution kernel. In the loop processing unit, an updated mixed speech feature sequence is input, and the outputs of the bottleneck layer, the gated convolution layer and the output layer are all different expressions of speech features. The bottleneck layer is used to reduce the feature dimension (from N to N'), and the output layer is used to restore the feature dimension (from N' to N).

[0116] It should be noted that, in the bottleneck layer, the updated mixed speech feature sequence is sequentially input into the 1×1 convolution kernel, the third activation layer, and the second normalization layer to obtain the first speech feature sequence. In the output layer, the second speech feature sequence is sequentially input into the third normalization layer and the 1×1 convolution kernel to output the initial processing result.

[0117] By applying the solution of the embodiments of the present specification, the updated mixed speech feature sequence is input into the cyclic processing unit, the feature dimension is reduced in the bottleneck layer while retaining the key features, the first speech feature sequence is subjected to cyclic filtering processing based on the time dimension in the gated convolution layer, and finally the initial processing result is obtained through the output layer, thereby capturing the extensive dependencies and local cyclic patterns of the speech data and ensuring the accuracy of the initial processing result.

[0118] In an optional embodiment of the present specification, the gated convolution layer includes an expanded feedforward layer, a convolution layer, and a processing layer; the above-mentioned inputting the first speech feature sequence into the gated convolution layer to obtain the second speech feature sequence may include the following steps:

[0119] Inputting the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence;

[0120] Input the convolution speech feature sequence into the dilated feedforward layer to obtain the dilated speech feature sequence;

[0121] The first speech feature sequence, the convolution speech feature sequence and the expansion speech feature sequence are input into the processing layer to obtain the second speech feature sequence.

[0122] Specifically, the design of the gated convolution layer is inspired by the effectiveness of the gating mechanism in the gated linear unit (GLU). The gated convolution layer is the main processing part of the recurrent processing unit. The gated convolution layer performs modeling of the recurrent pattern, and the gated convolution layer is obtained based on the gated convolution unit. In the embodiment of this specification, an expansion feedforward layer is combined with the gated convolution layer. The expansion feedforward layer refers to a network structure that adds an "expansion" mechanism to the traditional FSMN. This mechanism can expand the receptive field of the convolution kernel. The expansion feedforward layer is used to enable the FSMN to cover a wider receptive field through expansion operations while reducing the required memory resources. In order to further improve the expansion feedforward layer and allow the embedding dimension to be reduced, convolution units are used instead of linear units in the embodiment of this specification.

[0123] It should be noted that the gated convolutional layer transforms the first speech feature sequence X1∈R N ' ×S Mapping to output O∈R N ' ×S .like Figure 6 As shown in FIG. 1 , the gated convolution layer includes two convolution layers. To facilitate model training, the gated convolution layer includes a skip connection to connect the input of the gated convolution layer to its output. The connection allows the model to learn only the residual part of the input, which is beneficial to gradient back propagation and accelerates model training. First, the convolution speech feature sequence can be obtained by the following formulas (2) and (3); specifically, the first speech feature sequence can be input into the two convolution layers respectively, and the convolution operation (Conv U ) calculates the first speech feature sequence to generate two convolution speech feature sequences (U and V); secondly, the expanded speech feature sequence can be obtained by the following formula (4); specifically, a convolution speech feature sequence (V) can be input into the expanded feedforward layer, and the Dliated FSMN The operation calculates the convolution speech feature sequence (V) to obtain an expanded speech feature sequence (Y) with the same shape as the convolution speech feature sequence; finally, the second speech feature sequence can be obtained by the following formula (5); specifically, the input first speech feature sequence (X1), the convolution speech feature sequence (U) and the expanded speech feature sequence (Y) can be calculated by element-by-element multiplication and element-level addition to obtain a second speech feature sequence (O) with the same shape as the first speech feature sequence, wherein N' is the number of rows of the matrix and S is the number of columns of the matrix.

[0124] U=Conv U (X1)∈R N ' ×S (2)

[0125] V=Conv U (X1)∈RN ' ×S (3)

[0126] Y=Dilated FSMN (V)∈R N ' ×S (4)

[0127]

[0128] By applying the solution of the embodiments of this specification, the first speech feature sequence is input into the convolution layer to obtain a convolution speech feature sequence; the convolution speech feature sequence is input into the dilated feedforward layer to obtain an dilated speech feature sequence; the first speech feature sequence, the convolution speech feature sequence and the dilated speech feature sequence are input into the processing layer to obtain a second speech feature sequence. By using the dilated feedforward layer, a wider receptive field is covered while reducing the required memory resources.

[0129] In an optional embodiment of the present specification, the convolution layer includes a first normalization layer, a first linear layer, a second activation layer, and a deep convolution layer; the above-mentioned inputting the first speech feature sequence into the convolution layer to obtain the convolution speech feature sequence may include the following steps:

[0130] Inputting the first speech feature sequence into the first normalization layer, the first linear layer, and the second activation layer in sequence to obtain an initial convolution speech feature sequence;

[0131] Inputting the initial convolution speech feature sequence into the deep convolution layer to obtain a reference convolution speech feature sequence;

[0132] A convolution speech feature sequence is determined according to an initial convolution speech feature sequence and a reference convolution speech feature sequence.

[0133] See also Figure 7 , Figure 7 The architecture diagram of the convolution layer in a task processing method provided by an embodiment of the present specification is shown. The convolution layer is used to help the gated convolution layer capture the local pattern at the position. The convolution layer includes a first normalization layer, a first linear layer (linear), a second activation layer (SiLU layer), and a deep convolution layer (D-Convolution). In the convolution layer, the first normalization layer is used to normalize each batch of input features, the first linear layer is used to perform linear transformation operations on the features, and the second activation layer uses a nonlinear activation function to process the features. The deep convolution layer is equivalent to performing a 1-dimensional filtering operation on the sequence by sliding a small window, thereby identifying and extracting important local features and outputting a convolution speech feature sequence. It should be noted that after the second activation layer, the model training is promoted by skip connection, and the initial convolution speech feature sequence and the reference convolution speech feature sequence are added to determine the convolution speech feature sequence.

[0134] By applying the solution of the embodiment of this specification, the first speech feature sequence is sequentially input into the first normalization layer, the first linear layer and the second activation layer to obtain an initial convolution speech feature sequence; the initial convolution speech feature sequence is input into the deep convolution layer to obtain a reference convolution speech feature sequence; and the convolution speech feature sequence is determined according to the initial convolution speech feature sequence and the reference convolution speech feature sequence. The speech separation performance is improved.

[0135] In an optional embodiment of the present specification, the dilated feedforward layer includes a feedforward layer and a plurality of two-dimensional dilated convolutional layers connected to each other; the step of inputting the convolution speech feature sequence into the dilated feedforward layer to obtain the dilated speech feature sequence may include the following steps:

[0136] Input the convolution speech feature sequence into the feedforward layer to obtain a feedforward speech feature sequence;

[0137] Inputting the feedforward speech feature sequence into multiple two-dimensional dilated convolutional layers to obtain a reference speech feature sequence;

[0138] An expanded speech feature sequence is determined according to a feedforward speech feature sequence and a reference speech feature sequence.

[0139] It should be noted that multiple interconnected two-dimensional dilated convolutional layers can be called memory layers. The two-dimensional dilated convolutional layer includes a dilation factor, which is used to increase the receptive field and perform context aggregation at different resolutions. The two-dimensional dilated convolutional layer is easy to implement a grouping mechanism. The speech embedding from the feedforward layer is initially two-dimensional, and then reshaped into a three-dimensional sequence, further converting the embedding dimension into input channels, which are then divided into different groups, each of which uses a dedicated filter for convolution. Among them, the dilation factor is the parameter setting of the dilated convolution. The purpose of the dilated convolution is to expand the receptive field of the convolution. The larger the dilation factor, the larger the receptive field obtained. Because different dilation factors obtain different receptive fields, that is, contexts of different resolutions, different dilation factors are used to obtain different resolutions, and the final effect is to perform context aggregation at different resolutions.

[0140] See also Figure 8 , Figure 8 The architecture diagram of the dilated feedforward layer in a task processing method provided by an embodiment of the present specification is shown, wherein the dilated feedforward layer includes a feedforward layer and a memory layer, wherein the memory layer includes a plurality of interconnected two-dimensional dilated convolutional layers, wherein the two-dimensional dilated convolutional layer includes a dilation factor (d=1, d=2, ..., d=2 L-1), the feedforward layer includes a second linear layer, a fourth activation layer and a fifth linear layer. First, the convolution speech feature sequence is sequentially input into the second linear layer, the fourth activation layer and the fifth linear layer to obtain a feedforward speech feature sequence; secondly, the feedforward speech feature sequence is further input into a plurality of two-dimensional dilated convolution layers to obtain a reference speech feature sequence; finally, the feedforward speech feature sequence and the reference speech feature sequence are added to output a dilated speech feature sequence.

[0141] By applying the scheme of the embodiments of this specification, the convolution speech feature sequence is input into the feedforward layer to obtain the feedforward speech feature sequence; the feedforward speech feature sequence is input into multiple two-dimensional dilated convolutional layers to obtain a reference speech feature sequence; based on the feedforward speech feature sequence and the reference speech feature sequence, the dilated speech feature sequence is determined, thereby improving the speech separation performance.

[0142] In an optional embodiment of the present specification, the above-mentioned inputting the feedforward speech feature sequence into multiple two-dimensional dilated convolutional layers to obtain the reference speech feature sequence may include the following steps:

[0143] Obtaining a candidate speech feature sequence corresponding to a target two-dimensional dilated convolutional layer, wherein the target two-dimensional dilated convolutional layer is the last one of multiple two-dimensional dilated convolutional layers, and the candidate speech feature sequence is obtained based on speech feature sequences input by each two-dimensional dilated convolutional layer before the target two-dimensional dilated convolutional layer;

[0144] The candidate speech feature sequence is input into the target two-dimensional dilated convolutional layer to obtain the reference speech feature sequence.

[0145] It should be noted that the use of feed-forward dense connections can be extended in the memory layer to establish interconnections between each 2D dilated convolution layer and all other 2D dilated convolution layers to enhance information flow and promote gradient propagation. Specifically, the input of the L-th 2D dilated convolution layer is the speech embedding of all previous 2D dilated convolution layers, that is, the output of the L-th 2D dilated convolution layer can be determined by the following formula (6):

[0146] X L =H L ([X0,...,X L-1 ]) (6)

[0147] Among them, [X0,...,X L-1 ] refers to the concatenation of the inputs of the two-dimensional dilated convolutional layer preceding the L-th two-dimensional dilated convolutional layer, H L Represents a 2D dilated convolution operation.

[0148] By applying the solution of the embodiment of this specification, a candidate speech feature sequence corresponding to the target two-dimensional dilated convolutional layer is obtained, and the candidate speech feature sequence is input into the target two-dimensional dilated convolutional layer to obtain a reference speech feature sequence, thereby improving the speech separation efficiency.

[0149] In an optional embodiment of the present specification, the target two-dimensional dilated convolution layer includes a constant padding layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer and a connection layer; the above-mentioned inputting the candidate speech feature sequence into the target two-dimensional dilated convolution layer to obtain the reference speech feature sequence may include the following steps:

[0150] Inputting the candidate speech feature sequence into a constant padding layer, a two-dimensional convolutional layer, an instance normalization layer, and a first activation layer in sequence to obtain an initial reference speech feature sequence;

[0151] The initial reference speech feature sequence and the candidate speech feature sequence are input into the connection layer to obtain the reference speech feature sequence.

[0152] It should be noted that the 2D dilated convolutional layer is used to perform convolution operations on features, with the goal of modeling local features. The output of the 2D dilated convolutional layer is concatenated with the input to promote dense connections. Therefore, the dilated feed-forward layer only uses forward connections and convolutional memory layers to capture recurrent patterns.

[0153] See also Fig. 9 , Fig. 9 FIG. 1 shows an architecture diagram of a two-dimensional dilated convolutional layer in a task processing method provided by an embodiment of the present specification, such as Fig. 9 As shown in the figure, the two-dimensional dilated convolution layer includes a constant padding layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer (PReLU activation layer), and a connection layer. The constant padding layer is used to add 0s to the feature sequence so that the sequence length remains unchanged after convolution; the two-dimensional convolution layer is used to perform two-dimensional convolution operations; the instance normalization layer is used to normalize the features on each channel; and the first activation layer is used to perform nonlinear processing on the features. The feature dimension does not change in the two-dimensional dilated convolution layer.

[0154] Step 308: Separate the mixed speech feature sequence according to the speech separation information to obtain a task processing result.

[0155] In one or more embodiments of the present specification, mixed speech data is acquired, features are extracted from the mixed speech data to obtain a mixed speech feature sequence, and for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing results corresponding to each feature dimension. Further, the mixed speech feature sequence can be separated according to the speech separation information to obtain a task processing result, that is, to obtain separate speech data for each object.

[0156] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.

[0157] In an optional embodiment of the present specification, the above separation of the mixed speech feature sequence according to the speech separation information to obtain the task processing result may include the following steps:

[0158] Filtering the target speech feature sequence corresponding to each object from the mixed speech feature sequence according to the speech separation information;

[0159] The target speech feature sequence is decoded to generate target speech data corresponding to each object.

[0160] Specifically, the target voice data refers to voice data of an object alone, that is, the target voice data only includes voice data of one object.

[0161] It should be noted that when the target speech feature sequence corresponding to each object is screened out from the mixed speech feature sequence according to the speech separation information, the tensor product of the speech separation information and the mixed speech feature sequence can be calculated to determine the target speech feature sequence corresponding to each object. There are many ways to decode the target speech feature sequence and generate the target speech data corresponding to each object, which are selected according to the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, a traditional feature decoding method can be used, such as a filter, a spectral inverse transform, and a wavelet transform. In another possible implementation of this specification, a deep learning model can be used to automatically learn the representation of the speech signal from the target speech feature sequence, and then a back propagation algorithm is used to restore the original speech signal, such as inputting the target speech feature sequence corresponding to each object into the decoding unit in the task processing model to obtain the target speech data corresponding to each object, wherein the decoding unit can use the transposed one-dimensional convolution layer with the same convolution kernel size and stride as the encoding unit to reconstruct the original speech signal of length T.

[0162] By applying the solution of the embodiments of this specification, a target speech feature sequence corresponding to each object is screened out from the mixed speech feature sequence according to speech separation information; the target speech feature sequence is decoded to generate target speech data corresponding to each object, thereby improving speech separation performance.

[0163] See also Fig.10 , Fig.10 A flowchart of a task processing model training method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0164] Step 1002: Acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects.

[0165] Step 1004: extract features from the sample mixed speech data to obtain a sample mixed speech feature sequence.

[0166] Step 1006: Input the sample mixed speech feature sequence into the loop processing unit in the task processing model to obtain the predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence.

[0167] Step 1008: Adjust the model parameters of the task processing model according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.

[0168] It should be noted that the implementation of step 1002, step 1004 and step 1006 is the same as the processing method of the above-mentioned task processing method, and will not be described in detail in the embodiment of this specification.

[0169] In practical applications, the speech separation label of the sample object is the speech separation target. The speech separation label can be an independent sample speech label of each sample object, or a sample separation information label corresponding to each sample object, wherein the sample separation information label can be obtained based on the sample speech label encoding.

[0170] By applying the scheme of the embodiments of the present specification, since the predicted speech separation information is obtained by the cyclic processing unit based on the time dimension and the various feature dimensions of the sample mixed speech feature sequence, the cyclic pattern of the speech signal in the time dimension is fully considered, and the complex time dependency relationship within the speech signal is captured, thereby improving the speech separation performance of the task processing model.

[0171] The following combination Fig.11 , taking the application of the task processing method provided in this specification in the intelligent conference scenario as an example, the task processing method is further explained. Fig.11 A flowchart of a conference speech separation method provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0172] Step 1102: Acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects.

[0173] Step 1104: extract features from the mixed speech data to obtain a mixed speech feature sequence.

[0174] Step 1106: for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing result corresponding to each feature dimension.

[0175] Step 1108: Separate the mixed speech feature sequence according to the speech separation information to obtain a conference speech separation result.

[0176] It should be noted that the implementation of steps 1102 to 1108 is the same as the implementation of steps 302 to 308 described above, and the embodiments of this specification do not impose any limitation on this.

[0177] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.

[0178] See also Fig.12 , Fig.12 A processing flow chart of a task processing method provided by an embodiment of the present specification is shown. The goal of the task processing method is to use a task processing model to perform speech separation on given mixed speech data to obtain independent speech data for each object. The task processing model adopts a time domain masking network framework, including an encoding unit, a masking unit and a decoding unit, wherein the masking unit includes an attention processing unit and a loop processing unit, the encoding unit and the decoding unit are responsible for feature extraction and waveform reconstruction of speech signals, and the masking unit is used to map the output of the encoding unit into a set of speech separation information, specifically including:

[0179] Coding unit: inputting the mixed speech data into the coding unit to obtain a mixed speech feature sequence;

[0180] Masking unit: input the mixed speech feature sequence into the attention processing unit, and obtain an updated mixed speech feature sequence based on the time dimension; for any feature dimension of the updated mixed speech feature sequence, input the updated mixed speech feature sequence into the loop processing unit in the task processing model, obtain the initial processing result corresponding to the feature dimension, and return the initial processing result and the mixed speech feature sequence to the attention processing unit in the task processing model until the preset loop stop condition is reached, and the processing result corresponding to each feature dimension is obtained, and the speech separation information corresponding to each object is generated according to the processing result corresponding to each feature dimension; the tensor product after separation is calculated according to the mixed speech feature sequence and the speech separation information to obtain the target speech feature sequence;

[0181] Decoding unit: The target speech feature sequence is input into the decoding unit to generate the target speech data corresponding to each object.

[0182] It should be noted that the dimension of the mixed speech data is Bx1xT, the dimension of the mixed speech feature sequence is BxNxS, the dimension of the updated mixed speech feature sequence is BxNxS, and the dimension of the target speech feature sequence is CxBxNxS, where C is the number of objects, B is the batch size, N is the feature dimension, and S is the number of feature sequence frames.

[0183] By applying the scheme of the embodiments of this specification, the attention processing unit performs joint attention processing on the mixed speech feature sequence, which can fully consider the local feature information of each object and its global feature information in the entire speech, so that the initial speech feature sequence is more accurate, and the loop processing unit performs loop filtering processing on the initial speech feature sequence, which fully considers the loop pattern of the speech signal in the speech structure, rhythm and semantic association, captures the complex time dependency relationship within the speech signal, and improves the speech separation performance. In addition, the loop processing unit uses linear and convolution operations to process the entire sequence through forward connection without block division, thereby ensuring parallel computing, thereby more effectively solving the problems of unsatisfactory separation effect and susceptibility to noise reverberation interference faced by the speech separation task.

[0184] Corresponding to the above-mentioned task processing method embodiment, this specification also provides a task processing device embodiment. Fig.13 FIG. 1 shows a schematic diagram of a task processing device provided by an embodiment of the present specification. Fig.13 As shown, the device comprises:

[0185] A first acquisition module 1302 is configured to acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects;

[0186] The first extraction module 1304 is configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;

[0187] The first processing module 1306 is configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing result corresponding to each feature dimension;

[0188] The first separation module 1308 is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a task processing result.

[0189] Optionally, the first processing module 1306 is further configured to perform joint attention processing on the mixed speech feature sequence based on the time dimension to obtain an updated mixed speech feature sequence; for any feature dimension of the updated mixed speech feature sequence, perform cyclic filtering processing on the mixed speech feature sequence based on the time dimension until a preset loop stop condition is reached to obtain processing results corresponding to each feature dimension.

[0190] Optionally, the first processing module 1306 is further configured to input the mixed speech feature sequence into the attention processing unit in the task processing model, and obtain an updated mixed speech feature sequence based on the time dimension; for any feature dimension of the updated mixed speech feature sequence, input the updated mixed speech feature sequence into the loop processing unit in the task processing model, obtain the initial processing result corresponding to the feature dimension, and return the initial processing result and the mixed speech feature sequence to the attention processing unit in the task processing model until the preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.

[0191] Optionally, the cyclic processing unit includes a bottleneck layer, a gated convolution layer and an output layer; the first processing module 1306 is further configured to input the updated mixed speech feature sequence into the bottleneck layer to obtain a first speech feature sequence, wherein the bottleneck layer is used to reduce the feature dimension of the updated mixed speech feature sequence; input the first speech feature sequence into the gated convolution layer to obtain a second speech feature sequence, wherein the gated convolution layer is used to perform cyclic filtering processing on the first speech feature sequence based on the time dimension; input the second speech feature sequence into the output layer to obtain an initial processing result.

[0192] Optionally, the gated convolution layer includes an expanded feedforward layer, a convolution layer and a processing layer; the first processing module 1306 is further configured to input the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence; input the convolution speech feature sequence into the expanded feedforward layer to obtain an expanded speech feature sequence; input the first speech feature sequence, the convolution speech feature sequence and the expanded speech feature sequence into the processing layer to obtain a second speech feature sequence.

[0193] Optionally, the expanded feedforward layer includes a feedforward layer and multiple two-dimensional expanded convolutional layers connected to each other; the first processing module 1306 is further configured to input the convolutional speech feature sequence into the feedforward layer to obtain a feedforward speech feature sequence; input the feedforward speech feature sequence into multiple two-dimensional expanded convolutional layers to obtain a reference speech feature sequence; and determine the expanded speech feature sequence based on the feedforward speech feature sequence and the reference speech feature sequence.

[0194] Optionally, the first processing module 1306 is further configured to obtain a candidate speech feature sequence corresponding to a target two-dimensional dilated convolution layer, wherein the target two-dimensional dilated convolution layer is the last of multiple two-dimensional dilated convolution layers, and the candidate speech feature sequence is obtained based on speech feature sequences input by each two-dimensional dilated convolution layer before the target two-dimensional dilated convolution layer; the candidate speech feature sequence is input into the target two-dimensional dilated convolution layer to obtain a reference speech feature sequence.

[0195] Optionally, the target two-dimensional dilated convolution layer includes a constant filling layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer and a connection layer; the first processing module 1306 is further configured to input the candidate speech feature sequence into the constant filling layer, the two-dimensional convolution layer, the instance normalization layer and the first activation layer in sequence to obtain an initial reference speech feature sequence; the initial reference speech feature sequence and the candidate speech feature sequence are input into the connection layer to obtain a reference speech feature sequence.

[0196] Optionally, the convolution layer includes a first normalization layer, a first linear layer, a second activation layer, and a deep convolution layer; the first processing module 1306 is further configured to input the first speech feature sequence into the first normalization layer, the first linear layer, and the second activation layer in sequence to obtain an initial convolution speech feature sequence; input the initial convolution speech feature sequence into the deep convolution layer to obtain a reference convolution speech feature sequence; determine the convolution speech feature sequence based on the initial convolution speech feature sequence and the reference convolution speech feature sequence.

[0197] Optionally, the first separation module 1308 is further configured to filter out a target speech feature sequence corresponding to each object from the mixed speech feature sequence according to the speech separation information; decode the target speech feature sequence to generate target speech data corresponding to each object.

[0198] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.

[0199] The above is a schematic scheme of a task processing device of this embodiment. It should be noted that the technical scheme of the task processing device and the technical scheme of the above task processing method belong to the same concept, and the details of the technical scheme of the task processing device that are not described in detail can all be referred to the description of the technical scheme of the above task processing method.

[0200] Corresponding to the above-mentioned task processing model training method embodiment, this specification also provides a task processing model training device embodiment, Fig.14 FIG. 1 is a schematic diagram showing a structure of a task processing model training device provided by an embodiment of the present specification. Fig.14 As shown, the device comprises:

[0201] The second acquisition module 1402 is configured to acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects;

[0202] The second extraction module 1404 is configured to perform feature extraction on the sample mixed speech data to obtain a sample mixed speech feature sequence;

[0203] An input module 1406 is configured to input the sample mixed speech feature sequence into a loop processing unit in the task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence;

[0204] The adjustment module 1408 is configured to adjust the model parameters of the task processing model according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.

[0205] By applying the scheme of the embodiments of the present specification, since the predicted speech separation information is obtained by the cyclic processing unit based on the time dimension and the various feature dimensions of the sample mixed speech feature sequence, the cyclic pattern of the speech signal in the time dimension is fully considered, and the complex time dependency relationship within the speech signal is captured, thereby improving the speech separation performance of the task processing model.

[0206] The above is a schematic scheme of a task processing model training device of this embodiment. It should be noted that the technical scheme of the task processing model training device and the technical scheme of the task processing model training method described above belong to the same concept. For details not described in detail in the technical scheme of the task processing model training device, please refer to the description of the technical scheme of the task processing model training method described above.

[0207] Corresponding to the above-mentioned conference voice separation method embodiment, this specification also provides a conference voice separation device embodiment, Fig.15 FIG. 2 shows a schematic diagram of the structure of a conference voice separation device provided by an embodiment of the present specification. Fig.15 As shown, the device comprises:

[0208] The third acquisition module 1502 is configured to acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects;

[0209] The third extraction module 1504 is configured to perform feature extraction on the mixed speech data to obtain a mixed speech feature sequence;

[0210] The second processing module 1506 is configured to process the mixed speech feature sequence based on the time dimension for any feature dimension of the mixed speech feature sequence, and generate speech separation information corresponding to each object according to the processing result corresponding to each feature dimension;

[0211] The second separation module 1508 is configured to separate the mixed speech feature sequence according to the speech separation information to obtain a conference speech separation result.

[0212] The solution of the embodiment of this specification processes the mixed speech feature sequence based on the time dimension, fully considers the cyclic pattern of the speech signal in the time dimension, captures the complex time dependency within the speech signal, and thus improves the speech separation performance.

[0213] The above is a schematic scheme of a conference voice separation device of this embodiment. It should be noted that the technical scheme of the conference voice separation device and the technical scheme of the conference voice separation method mentioned above belong to the same concept, and the details not described in detail in the technical scheme of the conference voice separation device can be referred to the description of the technical scheme of the conference voice separation method mentioned above.

[0214] Fig.16 The block diagram of a computing device provided by one embodiment of the present specification is shown. The components of the computing device 1600 include but are not limited to a memory 1610 and a processor 1620. The processor 1620 is connected to the memory 1610 via a bus 1630, and the database 1650 is used to store data.

[0215] The computing device 1600 also includes an access device 1640 that enables the computing device 1600 to communicate via one or more networks 1660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0216] In one embodiment of the present specification, the above components of the computing device 1600 and Fig.16 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.16 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0217] The computing device 1600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1600 may also be a mobile or stationary server.

[0218] Among them, the processor 1620 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned task processing method or task processing model training method or conference speech separation method.

[0219] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device belongs to the same concept as the technical schemes of the above-mentioned task processing method, task processing model training method and conference speech separation method. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical schemes of the above-mentioned task processing method, task processing model training method or conference speech separation method.

[0220] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned task processing method or task processing model training method or conference speech separation method.

[0221] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium belongs to the same concept as the technical schemes of the above-mentioned task processing method, task processing model training method and conference speech separation method. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical schemes of the above-mentioned task processing method, task processing model training method or conference speech separation method.

[0222] An embodiment of the present specification also provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned task processing method or task processing model training method or conference speech separation method.

[0223] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of the computer program and the technical schemes of the above-mentioned task processing method, task processing model training method and conference speech separation method belong to the same concept. For details not described in detail in the technical scheme of the computer program, please refer to the description of the technical schemes of the above-mentioned task processing method, task processing model training method or conference speech separation method.

[0224] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0225] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0226] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0227] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0228] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A task processing method, comprising: Acquire mixed voice data, wherein the mixed voice data includes voice data of multiple objects; Extracting features from the mixed speech data to obtain a mixed speech feature sequence; For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing result corresponding to each feature dimension; The mixed speech feature sequence is separated according to the speech separation information to obtain a task processing result.

2. The method according to claim 1, wherein for any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, comprising: Based on the time dimension, performing joint attention processing on the mixed speech feature sequence to obtain an updated mixed speech feature sequence; For any feature dimension of the updated mixed speech feature sequence, a loop filtering process is performed on the mixed speech feature sequence based on the time dimension until a preset loop stop condition is reached, thereby obtaining a processing result corresponding to each feature dimension.

3. The method according to claim 2, wherein the step of performing joint attention processing on the mixed speech feature sequence based on the time dimension to obtain an updated mixed speech feature sequence comprises: Inputting the mixed speech feature sequence into the attention processing unit in the task processing model to obtain an updated mixed speech feature sequence based on the time dimension; The step of performing loop filtering on any feature dimension of the updated mixed speech feature sequence based on the time dimension until a preset loop stop condition is reached to obtain a processing result corresponding to each feature dimension includes: For any feature dimension of the updated mixed speech feature sequence, the updated mixed speech feature sequence is input into the loop processing unit in the task processing model to obtain the initial processing result corresponding to the feature dimension, and the initial processing result and the mixed speech feature sequence are returned to the attention processing unit in the task processing model until the preset loop stop condition is reached to obtain the processing results corresponding to each feature dimension.

4. The method according to claim 3, wherein the loop processing unit comprises a bottleneck layer, a gated convolution layer and an output layer; The step of inputting the updated mixed speech feature sequence into a loop processing unit in the task processing model to obtain an initial processing result corresponding to the feature dimension includes: Inputting the updated mixed speech feature sequence into the bottleneck layer to obtain a first speech feature sequence, wherein the bottleneck layer is used to reduce the feature dimension of the updated mixed speech feature sequence; Inputting the first speech feature sequence into the gated convolution layer to obtain a second speech feature sequence, wherein the gated convolution layer is used to perform a loop filtering process on the first speech feature sequence based on the time dimension; The second speech feature sequence is input into the output layer to obtain an initial processing result.

5. The method according to claim 4, wherein the gated convolutional layer comprises a dilated feed-forward layer, a convolutional layer, and a processing layer; The step of inputting the first speech feature sequence into the gated convolutional layer to obtain a second speech feature sequence comprises: Inputting the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence; Inputting the convolution speech feature sequence into the dilated feedforward layer to obtain a dilated speech feature sequence; The first speech feature sequence, the convolution speech feature sequence and the expansion speech feature sequence are input into the processing layer to obtain a second speech feature sequence.

6. The method according to claim 5, wherein the dilated feed-forward layer comprises a feed-forward layer and a plurality of two-dimensional dilated convolutional layers connected to each other; The step of inputting the convolution speech feature sequence into the dilated feedforward layer to obtain the dilated speech feature sequence comprises: Inputting the convolution speech feature sequence into the feedforward layer to obtain a feedforward speech feature sequence; Inputting the feedforward speech feature sequence into the multiple two-dimensional dilated convolutional layers to obtain a reference speech feature sequence; An expanded speech feature sequence is determined according to the feedforward speech feature sequence and the reference speech feature sequence.

7. The method according to claim 6, wherein the step of inputting the feedforward speech feature sequence into the plurality of two-dimensional dilated convolutional layers to obtain a reference speech feature sequence comprises: Obtaining a candidate speech feature sequence corresponding to a target two-dimensional dilated convolutional layer, wherein the target two-dimensional dilated convolutional layer is the last one of the multiple two-dimensional dilated convolutional layers, and the candidate speech feature sequence is obtained based on speech feature sequences input by each two-dimensional dilated convolutional layer before the target two-dimensional dilated convolutional layer; The candidate speech feature sequence is input into the target two-dimensional dilated convolutional layer to obtain a reference speech feature sequence.

8. The method according to claim 7, wherein the target two-dimensional dilated convolution layer comprises a constant padding layer, a two-dimensional convolution layer, an instance normalization layer, a first activation layer, and a connection layer; The step of inputting the candidate speech feature sequence into the target two-dimensional dilated convolutional layer to obtain a reference speech feature sequence comprises: Inputting the candidate speech feature sequence into the constant padding layer, the two-dimensional convolution layer, the instance normalization layer and the first activation layer in sequence to obtain an initial reference speech feature sequence; The initial reference speech feature sequence and the candidate speech feature sequence are input into the connection layer to obtain a reference speech feature sequence.

9. The method according to claim 5, wherein the convolution layer comprises a first normalization layer, a first linear layer, a second activation layer, and a depth convolution layer; The step of inputting the first speech feature sequence into the convolution layer to obtain a convolution speech feature sequence comprises: Inputting the first speech feature sequence into the first normalization layer, the first linear layer and the second activation layer in sequence to obtain an initial convolution speech feature sequence; Inputting the initial convolution speech feature sequence into the deep convolution layer to obtain a reference convolution speech feature sequence; According to the initial convolution speech feature sequence and the reference convolution speech feature sequence, a convolution speech feature sequence is determined.

10. The method according to claim 1, wherein the mixed speech feature sequence is separated according to the speech separation information to obtain a task processing result, comprising: Filtering the target speech feature sequence corresponding to each object from the mixed speech feature sequence according to the speech separation information; The target speech feature sequence is decoded to generate target speech data corresponding to each object.

11. A task processing model training method, comprising: Acquire a plurality of sample mixed speech data, wherein the sample mixed speech data includes sample speech data of a plurality of sample objects, and the sample mixed speech data carries speech separation labels of the plurality of sample objects; Extracting features from the sample mixed speech data to obtain a sample mixed speech feature sequence; Inputting the sample mixed speech feature sequence into a loop processing unit in a task processing model to obtain predicted speech separation information corresponding to each sample object, wherein the predicted speech separation information is obtained by the loop processing unit based on the time dimension and each feature dimension of the sample mixed speech feature sequence; The model parameters of the task processing model are adjusted according to the speech separation label and the predicted speech separation information to obtain a trained task processing model.

12. A conference voice separation method, comprising: Acquire mixed voice data, wherein the mixed voice data includes conference voice data of multiple objects; Extracting features from the mixed speech data to obtain a mixed speech feature sequence; For any feature dimension of the mixed speech feature sequence, the mixed speech feature sequence is processed based on the time dimension, and speech separation information corresponding to each object is generated according to the processing result corresponding to each feature dimension; The mixed speech feature sequence is separated according to the speech separation information to obtain a conference speech separation result.

13. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 10 or claim 11 or claim 12 are implemented.

14. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1 to 10 or claim 11 or claim 12.