Target Speaker Speech Extraction Method and Device
By introducing a multi-head self-attention and gated recurrent unit layer transformer block into the target speaker speech extraction model, and combining it with speech activity detection and mask extraction models, the problem of poor target speaker extraction under complex mixed speech signals in the prior art is solved, and high-quality speech extraction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2023-11-29
- Publication Date
- 2026-05-26
AI Technical Summary
Existing target speaker extraction methods perform poorly when dealing with mixed speech signals containing multiple speakers, missing target speakers, or low overlap of target speakers.
A target speaker speech extraction method is adopted, which generates the target speaker's speech signal by introducing a first speech extraction model with a transformer block of a multi-head self-attention layer and a gated recurrent unit layer, combined with a target speaker speech activity detection and mask extraction model.
It improves the quality and performance of target speaker speech extraction, and is especially suitable for processing complex mixed speech signals, achieving accurate and efficient target speaker speech extraction.
Smart Images

Figure CN119007728B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and more specifically, to a method, apparatus and device for extracting speech from a target speaker, and a computer-readable storage medium. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0003] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0004] Speech technology is widely used in modern life. Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech being one of the most promising methods. In voiceprint recognition, Target Speaker Extraction (TSE) uses the registered speech information of the target speaker to extract the target speaker's voice from a mixed speech signal containing noise and interference. Existing TSE methods are ineffective for extracting mixed speech signals with multiple speakers, absent target speakers, or low overlap. Therefore, a more effective TSE method for handling such mixed speech signals is needed. Summary of the Invention
[0005] This disclosure presents a method, apparatus and device for extracting speech from a target speaker, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of the present disclosure, a method for extracting speech from a target speaker is provided, comprising: generating a target speaker feature vector based on a reference speech signal from a target speaker; generating a first speech signal of the target speaker using a first speech extraction model and a second speech signal of the target speaker using a second speech extraction model, based at least on a mixed speech signal and the target speaker feature vector, wherein the first speech extraction model includes at least one transformer block having a multi-head self-attention layer and a gated recurrent unit layer; and generating a target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal.
[0007] According to an example of an embodiment of this disclosure, generating a first speech signal of the target speaker using a first speech extraction model includes: generating a first mixed feature vector based on the mixed speech signal; generating a target speaker activity probability based on the target speaker feature vector and the first mixed feature vector using the first speech extraction model, wherein the target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker; and generating the first speech signal based on the target speaker activity probability and the mixed speech signal.
[0008] According to an example of an embodiment of this disclosure, generating the first speech signal based on the target speaker activity probability and the mixed speech signal includes: converting the target speaker activity probability into a binary form of a target speaker activity label based on a predetermined threshold; and filtering the mixed speech signal using the target speaker activity label to generate the first speech signal.
[0009] According to an example of an embodiment of this disclosure, generating a second speech signal of the target speaker using a second speech extraction model includes: generating the second speech signal based on the target speaker feature vector and the first speech signal using the second speech extraction model, wherein generating a target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal includes: determining the second speech signal as the target speech signal.
[0010] According to an example of an embodiment of this disclosure, generating a target speech signal for the target speaker based on at least one of the first speech signal and the second speech signal includes: generating the target speech signal for the target speaker by multiplying the first speech signal and the second speech signal.
[0011] According to an example of an embodiment of this disclosure, generating a second speech signal of the target speaker using a second speech extraction model includes: generating a target speaker mask vector based on the target speaker feature vector and the mixed speech signal using the second speech extraction model, wherein the target speaker mask vector indicates the position of the target speaker speech signal in the mixed speech signal; and generating the second speech signal based on the target speaker mask vector and the mixed speech signal.
[0012] According to an example of an embodiment of this disclosure, generating a first speech signal of the target speaker using a first speech extraction model includes: generating a second mixed feature vector based on the second speech signal; generating a target speaker activity probability based on the target speaker feature vector and the second mixed feature vector, wherein the target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker; generating the first speech signal based on the target speaker activity probability and the mixed speech signal, wherein generating a target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal includes: determining the first speech signal as the target speech signal.
[0013] According to an example of an embodiment of this disclosure, the first speech extraction model further includes at least one convolutional block, and the at least one transformer block further includes a linear layer and at least one residual connection and normalization layer.
[0014] According to an example of an embodiment of this disclosure, the first speech extraction model is a target speaker speech activity detection model for detecting the probability of target speaker activity, and the second speech extraction model is a target speaker mask extraction model for extracting a target speaker mask vector, wherein the target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker, and the target speaker mask vector indicates the position of the target speaker speech signal in the mixed speech signal.
[0015] According to an example of an embodiment of this disclosure, the mixed speech signal includes speech signal segments with multiple speakers and speech signal segments where the target speaker is absent, and the ratio of the speech duration of the target speaker to the total speech duration of the mixed speech signal is variable.
[0016] According to another aspect of the present disclosure, a target speaker speech extraction apparatus is provided, the apparatus comprising: a feature vector extraction unit configured to generate a target speaker feature vector based on a reference speech signal from a target speaker; and a target speech signal generation unit configured to generate a first speech signal of the target speaker using a first speech extraction model and a second speech signal of the target speaker using a second speech extraction model, based at least on a mixed speech signal and the target speaker feature vector, and to generate a target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal, wherein the first speech extraction model includes at least one transformer block, the at least one transformer block having a multi-head self-attention layer and a gated recurrent unit layer.
[0017] According to an example of an embodiment of this disclosure, the first speech extraction model is a target speaker speech activity detection model for detecting the probability of target speaker activity, and the second speech extraction model is a target speaker mask extraction model for extracting a target speaker mask vector, wherein the target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker, and the target speaker mask vector indicates the position of the target speaker speech signal in the mixed speech signal.
[0018] According to another aspect of the present disclosure, a target speaker speech extraction device is provided, comprising: one or more processors; and one or more memories, wherein the memories store computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform the methods described in the above aspects.
[0019] According to another aspect of the present disclosure, a computer-readable storage medium is provided that stores computer-readable instructions thereon, which, when executed by a processor, cause the processor to perform the method as described in any of the foregoing aspects of the present disclosure.
[0020] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer-readable instructions that, when executed by a processor, cause the processor to perform the method as described in any of the foregoing aspects of the present disclosure.
[0021] By utilizing the target speaker speech extraction methods, apparatuses, devices, computer-readable storage media, and computer program products described above, the quality of target speaker speech extraction can be improved compared to conventional methods by introducing at least one transformer block including a multi-head self-attention sublayer and a gated recurrent sublayer into a first speech extraction model such as the TSVAD model. Furthermore, by fusing the first and second speech extraction models in different ways to extract target speaker speech, the characteristics of different speech extraction models can be effectively utilized, further improving the performance of target speaker speech extraction. The target speaker speech extraction method according to embodiments of this disclosure is particularly suitable for processing mixed speech signals with multiple speakers, absent target speakers, and low or variable overlap rates of target speakers, achieving accurate and efficient target speaker speech extraction. Attached Figure Description
[0022] The above and other objects, features, and advantages of the present disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 An exemplary scenario diagram of a target speaker speech extraction system according to an embodiment of the present disclosure is shown.
[0024] Figure 2 A flowchart of a target speaker speech extraction method according to an embodiment of the present disclosure is shown.
[0025] Figure 3 A schematic diagram of the structure of a first speech extraction model according to an embodiment of the present disclosure is shown.
[0026] Figure 4A An example of a first speech extraction model and a first fusion method of a second speech extraction model according to an embodiment of the present disclosure is shown.
[0027] Figure 4B A second fusion method of a first speech extraction model and a second speech extraction model according to an embodiment of the present disclosure is shown.
[0028] Figure 4C A third fusion method of a first speech extraction model and a second speech extraction model according to an embodiment of the present disclosure is shown.
[0029] Figure 5A schematic diagram of the structure of a target speaker speech extraction apparatus according to an embodiment of the present disclosure is shown.
[0030] Figure 6 A schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0031] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0032] As illustrated in the embodiments and claims of this disclosure, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. The terms "first," "second," and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms "connected" or "linked" are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect.
[0033] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0034] Furthermore, flowcharts are used in this disclosure to illustrate the operations performed by the system according to embodiments of this disclosure. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously. Additionally, other operations can be superimposed on these processes, or one or more steps can be removed from these processes.
[0035] In speech technology, voiceprint recognition aims to identify unknown voices by analyzing the features of one or more speech signals. Target Speaker Extraction (TSE) is one such category. Also known as personalized speech enhancement or personalized noise suppression, TSE extracts the voice of the target speaker from a mixed speech signal containing noise and interference, using the registered speech information of the target speaker. Here, the target speaker refers to the speaker of interest or a designated speaker.
[0036] Typically, target speaker extraction models use pre-trained speaker feature extraction models to generate target speaker feature vectors based on reference speech signals from the target speaker, and then further extract the target speaker's speech signal from the mixed speech signal based on these feature vectors. However, most target speaker extraction models are designed for mixed speech signals with a limited number of speakers (e.g., 3 or fewer) where the target speaker is present and the proportion of the target speaker's speech duration to the total speech duration (referred to as the overlap rate) is high. They perform poorly when processing mixed speech signals with a large number of speakers, or where the target speaker is absent (referred to as target speaker absence), or where the target speaker overlap rate is low.
[0037] To address the above issues, this disclosure proposes a target speaker speech extraction model that can effectively extract high-quality target speaker speech signals from mixed speech signals containing multiple speakers, missing target speakers, and low target speaker overlap.
[0038] Figure 1 An exemplary scenario diagram of a target speaker speech extraction system according to an embodiment of this disclosure is shown. Figure 1 As shown, the target speaker voice extraction system 100 may include a user terminal 110, a network 120, a server 130, and a database 140.
[0039] User terminal 110 may be, for example, Figure 1 The computer 110-1 and mobile phone 110-2 are shown in the figure. It is understood that, in fact, the user terminal 110 can be any other type of electronic device capable of performing data processing, which can include, but is not limited to, fixed terminals such as desktop computers, smart TVs, etc., mobile terminals such as smartphones, tablets, portable computers, handheld devices, etc., or any combination thereof, and the embodiments disclosed herein do not impose specific limitations on this.
[0040] According to embodiments of the present disclosure, the user terminal 110 can be used to receive mixed speech signals and generate a target speaker's speech signal using the target speaker speech extraction method provided in the present disclosure. In some embodiments, the target speaker speech extraction method provided in the present disclosure can be executed using the processing unit of the user terminal 110. In some implementations, the user terminal 110 can execute the target speaker speech extraction method provided in the present disclosure using an application program built into the user terminal. In other implementations, the user terminal 110 can execute the target speaker speech extraction method provided in the present disclosure by calling an application program stored externally to the user terminal.
[0041] In other embodiments, the user terminal 110 sends the received mixed speech signal to be processed to the server 130 via the network 120, and the server 130 executes the target speaker speech extraction method. In some implementations, the server 130 can use a built-in application to execute the target speaker speech extraction method. In other implementations, the server 130 can execute the target speaker speech extraction method by calling an application stored externally to the server.
[0042] Network 120 can be a single network or a combination of at least two different networks. For example, network 120 can be one or more of the following: local area network (LAN), wide area network (WAN), public network, private network, etc. Server 130 can be a standalone server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, location services, and big data and artificial intelligence platforms. This disclosure does not impose specific limitations in this regard.
[0043] Database 140 can refer to any device with storage capabilities. Database 140 is primarily used to store various data used, generated, and output by user terminal 110 and server 130 during operation. Database 140 can be local or remote. Database 140 may include various types of memory, such as Random Access Memory (RAM) and Read Only Memory (ROM). The storage devices mentioned above are just a few examples; the storage devices that can be used in this system are not limited to these. Database 140 can be interconnected or communicate with server 130 or a portion thereof via network 120, or directly interconnected or communicate with server 130, or a combination of both.
[0044] The following reference Figure 2 A method for extracting the speech of a target speaker according to embodiments of the present disclosure is described. Figure 2 A flowchart of a target speaker speech extraction method 200 according to an embodiment of the present disclosure is shown. As described above, the target speaker speech extraction method 200 can be executed by a user terminal or a server, and the embodiments of the present disclosure do not impose specific limitations on this.
[0045] In step S210, a target speaker feature vector is generated based on a reference speech signal from the target speaker. The reference speech signal is a signal containing only the target speaker's speech, for example, a 10-second speech signal from the target speaker, used to provide clues for extracting the target speaker's speech from the mixed speech signal. This embodiment does not impose a specific limitation on the length of the reference speech signal; it can be of any length depending on actual needs. The target speaker feature vector, also called the target speaker embedding vector, is a vector that characterizes the audio features of the target speaker. In this embodiment, a speaker encoder can be constructed and trained to extract the feature vector from the speech signal, or an existing speaker encoder model can be used to generate the feature vector, for example, ResNet34 based on a residual neural network, etc. This embodiment does not impose a specific limitation in this regard.
[0046] In step S220, the first speech signal of the target speaker can be extracted using a first speech extraction model, based at least on the target speaker feature vector and the mixed speech signal, and the second speech signal of the target speaker can be extracted using a second speech extraction model. Here, the mixed speech signal refers to a mixed audio including the target speaker signal, interfering speaker signals, background noise signals, etc. In this embodiment of the disclosure, the mixed speech signal includes speech signal segments with multiple speakers and speech signal segments where the target speaker is absent, and the ratio of the target speaker's speech duration to the total speech duration of the mixed speech signal (i.e., the overlap rate) varies, meaning it can simultaneously include speech signal segments with low, medium, and high overlap rates of the target speaker, thereby adapting to various mixed speech signal acquisition scenarios. For example, the mixed speech signal Y can be represented as:
[0047] Y = S + I + N (1)
[0048] Wherein, Y represents the mixed speech signal; S represents the target speaker signal; I represents the interfering speaker signal; and N represents the background noise signal. Generally, the target speaker speech extraction method according to embodiments of this disclosure aims to solve single-channel target speaker speech extraction, that is, the mixed speech signal Y comes from a single audio acquisition channel. However, the target speaker speech extraction method according to embodiments of this disclosure can also be applied to multi-channel target speaker speech extraction with appropriate modifications, provided that there are no contradictions. In the following, embodiments of this disclosure will be described using single-channel target speaker speech extraction as an example.
[0049] In this embodiment of the disclosure, a first speech extraction model and a second speech extraction model based on different principles can be used to obtain the first speech signal and the second speech signal of the target speaker, respectively. Then, in step S230, a target speech signal of the target speaker is generated based on at least one of the first speech signal and the second speech signal. The first speech extraction model and the second speech extraction model can be combined with each other in various different forms to generate the first speech signal, the second speech signal, and the final target speech signal of the target speaker, as will be described in further detail below.
[0050] In this embodiment, the first speech extraction model can be implemented based on a Target Speaker Voice Activity Detection (TSVAD) model, which predicts the probability of the target speaker appearing in each time frame of a mixed speech signal or a processed mixed speech signal with the help of the target speaker's feature vector. This probability is referred to as the target speaker activity probability. The processed mixed speech signal, for example, refers to a mixed speech signal that has already undergone speech extraction processing by other target speakers. Specifically, the first speech extraction model according to this embodiment may include at least one transformer block, which may have neural network layers such as a multi-head self-attention layer and a gated recurrent unit (GRU) layer. The multi-head self-attention layer is configured to enable the first speech extraction model to simultaneously process information at different positions in the audio sequence of the mixed speech signal or the processed mixed speech signal, thereby improving the model's spatial representation ability and training speed. The gated recurrent unit layer is configured to enable the first speech extraction model to better capture the dependencies between time frames with large time step distances in the time frame sequence of the mixed speech signal or the processed mixed speech signal, thereby effectively solving problems such as gradient explosion and gradient decay in long-term memory and backpropagation in neural networks.
[0051] Figure 3 A schematic diagram of the structure of a first speech extraction model according to an embodiment of the present disclosure is shown, wherein TSVAD is used as an example of the first speech extraction model for illustration. Figure 3 As shown, in addition to the transformer block, the first speech extraction model may also include at least one convolutional block, a smoothing block for post-processing, etc.; and in addition to the multi-head self-attention layer and the GRU layer, each transformer block may also include a linear layer and at least one residual connection and normalization layer, wherein the linear layer can be used to introduce linear components, and the residual connection and normalization layer can be used to integrate information, enhance model stability, and accelerate model convergence. Figure 3In the example, the first speech extraction model has two transformer blocks, but the embodiments of this disclosure are not limited to this. The first speech extraction model can also include any number of transformer blocks as needed. Similarly, Figure 3 The number of blocks and layers shown are merely examples and do not constitute any limitation on the embodiments disclosed herein.
[0052] The mixed speech signal Y or the processed mixed speech signal Y' can be divided into frame vectors [y1, y2, ..., y'] by a framing module. t (Where t represents the number of frames). Then, as... Figure 3 As shown, a speaker encoder can be used to extract a mixed feature vector E from a mixed speech signal that has been divided into frame vectors or a processed mixed speech signal. Y =[E y1 E y2 , ..., E yt ]. Mixed feature vector E Y The target speaker feature vector E extracted from the reference speech signal of the target speaker S The inputs are concatenated and fed into the TSVAD model. After processing through at least one convolutional block, at least one transformer block, and a smoothing block, the output is the target speaker activity probability P = [P y1 P y2 , ..., P yt This process can be represented as follows:
[0053] P = TSVAD(E) Y E S (2)
[0054] Where P represents the probability of the target speaker's activity, E Y E represents the mixed feature vector. S Let E represent the target speaker feature vector, and TSVAD(·) represent the target speaker activity detection function, which is based on the mixed feature vector E. Y and the target speaker feature vector E S Generate the probability of target speaker activity.
[0055] After obtaining the target speaker activity probability, the target speaker activity probability can be converted into a binary form of target speaker activity label based on a predetermined threshold, such as a target speaker activity label with a value of 0 or 1. For example, when the target speaker activity label is 1, it can indicate that the current time frame is the target speaker's speech signal; otherwise, when the target speaker activity label is 0, it can indicate that the current time frame is not the target speaker's speech signal. Here, the predetermined threshold can be obtained based on model training, determined based on empirical parameters, or determined based on probability distribution characteristics, and this embodiment does not impose specific limitations on this. Subsequently, the target speaker activity label can be used to filter the mixed speech signal or the processed mixed speech signal, for example, filtering out time frames with a target speaker activity label of 0 and retaining only time frames with a target speaker activity label of 1, thereby obtaining the target speaker's speech signal, i.e., the first speech signal.
[0056] The first speech extraction model according to embodiments of this disclosure improves the performance of target speaker speech extraction by introducing at least one transformer block including a multi-head self-attention layer and a gated recurrent unit layer. The TSVAD model is used as an example of the first speech extraction model for illustration. Table 1 below shows the results of ablation experiments on the TSVAD model, where TSVAD represents a TSVAD model according to embodiments of this disclosure with at least one transformer block, and TSVAD* represents the TSVAD model according to embodiments of this disclosure with the transformer block removed, or the transformer block replaced with another neural network module such as a Long Short Memory network, all other experimental conditions being the same. In Table 1, DER represents the speaker segmentation clustering error rate, and INT represents the energy difference between the mixed speech signal and the extracted target speaker's speech signal. The smaller the DER value and the larger the INT value, the better the target speaker speech extraction performance. As can be seen from Table 1, compared with the TSVAD* model with the transformer block removed or replaced, the TSVAD model according to the embodiments of this disclosure has a smaller DER and a larger INT, indicating that its target speaker speech extraction performance is better. This shows that introducing at least one transformer block plays an important role in improving the target speaker speech extraction performance.
[0057]
[0058] Table 1 Ablation Experiment Results of the TSVAD Model
[0059] In this embodiment of the disclosure, the second speech extraction model can be a target speaker mask extraction model (TSE) for extracting a target speaker mask vector from a mixed speech signal or a processed mixed speech signal. The processed mixed speech signal, for example, refers to a mixed speech signal that has undergone other target speaker speech extraction processing. The target speaker mask vector can indicate the position of the target speaker's speech signal in the mixed speech signal or the processed mixed speech signal. For example, the target speaker mask vector can be a vector consisting of 0s and 1s with the same dimension as the vector of the mixed speech signal, where a value of 1 indicates that the element at the same position in the mixed speech signal vector is the target speaker's speech signal, and a value of 0 indicates that the element at the same position in the mixed speech signal vector is not the target speaker's speech signal. The process of extracting the target speaker mask vector from the mixed speech signal or the processed mixed speech signal can be represented as follows:
[0060] M = TSE(Y, E) S (3)
[0061] Where M represents the target speaker mask vector, Y represents the mixed speech signal or the processed mixed speech signal, and E represents the target speaker mask vector. S TSE(·) represents the target speaker feature vector, and TSE(·) represents the target speaker mask extraction function, which is based on the mixed speech signal and the target speaker feature vector E. S Generate a target speaker mask vector. Based on the target speaker mask vector and the mixed speech signal or processed mixed speech signal, a second speech signal that retains only the speech signal of the target speaker can be generated. For example, the second speech signal can be generated by multiplying the target speaker mask vector with the mixed speech signal or processed mixed speech signal.
[0062] In this embodiment of the disclosure, the first speech extraction model and the second speech extraction model can be fused in different ways to generate the target speech signal of the target speaker. This fusion process can be represented, for example, as follows:
[0063] S'=F(Y,M,P) (4)
[0064] Where S' represents the extracted target speech signal of the target speaker, Y represents the mixed speech signal, M represents the target speaker mask vector, P represents the target speaker activity probability, and F(·) represents the fusion function, which can generate the target speech signal based on the mixed speech signal Y, the target speaker mask vector M, and the target speaker activity probability P.
[0065] The following reference Figures 4A to 4C The different fusion methods of the first speech extraction model and the second speech extraction model according to embodiments of the present disclosure are described. Figure 4AAn example of a first speech extraction model and a first fusion method of a second speech extraction model according to an embodiment of the present disclosure is shown. Figure 4B A second fusion method of a first speech extraction model and a second speech extraction model according to an embodiment of the present disclosure is shown, and Figure 4C A third fusion method of a first speech extraction model and a second speech extraction model according to an embodiment of the present disclosure is shown.
[0066] In one example, the output of the first speech extraction model can be input into the second speech extraction model, and the second speech extraction model can generate the final target speaker's target speech signal. Specifically, as shown... Figure 4A As shown, a first mixed feature vector can be generated based on the mixed speech signal Y. For example, it can be generated using... Figure 4A The speaker encoder shown generates a first mixed feature vector E based on the mixed speech signal Y. Y Subsequently, the first speech extraction model is based on the first mixed feature vector E. Y and the target speaker feature vector E generated in step S210 above based on the reference speech signal from the target speaker. S The target speaker activity probability P is generated. Then, a first speech signal S1 of the target speaker can be generated based on the target speaker activity probability P and the mixed speech signal Y. For example, the target speaker activity probability can be converted into a binary form of a target speaker activity label based on a predetermined threshold, and the mixed speech signal can be filtered using the target speaker activity label to generate the first speech signal S1.
[0067] At this point, the first speech extraction model has performed preliminary target speaker extraction on the mixed speech signal. Therefore, the first speech signal S1 can now be referred to as the processed mixed speech signal. For more accurate target speaker extraction, the first speech signal S1 can be input into the second speech extraction model. The second speech extraction model can be based on the first speech signal S1 and the target speaker feature vector E. S A second speech signal is generated, which serves as the target speech signal of the extracted target speaker.
[0068] In this example, a lower predetermined threshold can be used to generate target speaker activity labels for filtering mixed speech signals, thereby minimizing the risk of erroneously filtering out the speech signals of the target speaker; by combining a second speech extraction model for further target speaker speech extraction, speech signals of the target speaker with higher quality can be effectively extracted.
[0069] In another example, the output of the second speech extraction model can be input into the first speech extraction model, and the first speech extraction model can be used to generate the final target speaker's target speech signal. Specifically, as follows: Figure 4B As shown, the second speech extraction model can be used based on the mixed speech signal Y and the target speaker feature vector E generated in step S210 above based on the reference speech signal from the target speaker. S The target speaker mask vector M is generated, and then a second speech signal S2 is generated based on the target speaker mask vector M and the mixed speech signal Y.
[0070] At this point, preliminary target speaker speech extraction has been performed on the mixed speech signal using the second speech extraction model. Therefore, the second speech signal S2 can now be referred to as the processed mixed speech signal. To perform more accurate target speaker speech extraction, a second mixed feature vector can be generated based on this second speech signal S2, for example, using... Figure 4B The speaker encoder shown generates a second mixed feature vector E based on the second speech signal S2. Y’ Then, the first speech extraction model can be based on this second mixed feature vector E. Y’ and the target speaker feature vector E S The target speaker activity probability P is generated, and then a first speech signal is generated based on the target speaker activity probability P and the mixed speech signal Y, which serves as the target speech signal of the extracted target speaker.
[0071] In this example, a second speech extraction model is first used to extract an initial speech signal with a low overlap rate from the mixed speech signal, and then the second speech extraction model is used to further accurately extract the target speech signal of the target speaker.
[0072] In another example, the first speech extraction model and the second speech extraction model can be connected in parallel. Specifically, as shown below... Figure 4C As shown, a first mixed feature vector can be generated based on the mixed speech signal Y. For example, it can be generated using... Figure 4A The speaker encoder shown generates a first mixed feature vector E based on the mixed speech signal Y. Y Subsequently, the first speech extraction model is based on the first mixed feature vector E. Y and the target speaker feature vector E generated in step S210 above based on the reference speech signal from the target speaker. SThis is used to generate the target speaker activity probability P. Then, a first speech signal S1 of the target speaker can be generated based on the target speaker activity probability P and the mixed speech signal Y. On the other hand, the second speech extraction model can be based on the mixed speech signal Y and the target speaker feature vector E generated in step S210 above based on the reference speech signal from the target speaker. S This process generates a target speaker mask vector M, and then generates a second speech signal S2 based on the target speaker mask vector M and the mixed speech signal Y. Afterwards, the target speaker's target speech signal can be generated based on the first speech signal S1 and the second speech signal S2, for example... Figure 4C As shown, the target speech signal S' of the extracted target speaker can be generated by multiplying the first speech signal S1 and the second speech signal S2.
[0073] In this example, the first speech extraction model and the second speech extraction model are connected in parallel with each other. The interaction between the two models is small, which reduces the reliance on auxiliary information and improves the efficiency of target speaker speech extraction.
[0074] The above combination Figures 4A to 4C Different fusion methods between the first and second speech extraction models are described. It should be noted that... Figures 4A to 4C The fusion methods, specific modules, input parameters, and output parameters shown are merely examples or for illustrative purposes and do not imply that the embodiments of this disclosure are limited thereto. For example, the first speech extraction model and the second speech extraction model can also be fused in other appropriate ways.
[0075] The target speaker speech extraction method according to embodiments of this disclosure improves the quality of target speaker speech extraction compared to conventional methods by introducing at least one transformer block including a multi-head self-attention sublayer and a gated recurrent sublayer into a first speech extraction model such as the TSVAD model. Furthermore, by fusing the first and second speech extraction models in different ways, the characteristics of different speech extraction models can be effectively utilized, further improving the performance of target speaker speech extraction. The target speaker speech extraction method according to embodiments of this disclosure is particularly suitable for processing mixed speech signals with multiple speakers, absent target speakers, and low or variable overlap rates of target speakers, achieving accurate and efficient target speaker speech extraction.
[0076] The following reference Figure 5 A target speaker speech extraction apparatus is described according to embodiments of the present disclosure. Figure 5 A schematic diagram of the structure of a target speaker speech extraction device 500 according to an embodiment of the present disclosure is shown. Figure 5As shown, the target speaker speech extraction device 500 includes a feature vector extraction unit 510 and a target speech signal generation unit 520. Besides these two units, the device 500 may also include other related components, but since these components are not relevant to this disclosure, detailed descriptions of their specific contents are omitted here. Furthermore, details of some functions of the device 500 are referenced... Figure 2 The details of the steps in method 200 are similar, therefore, for the sake of brevity, repeated descriptions of some contents are omitted here. The apparatus 500 according to embodiments of this disclosure can be implemented as a terminal or a server, as referred to above. Figure 1 As described.
[0077] The feature vector extraction unit 510 is configured to generate a target speaker feature vector based on a reference speech signal from the target speaker. The reference speech signal is a signal containing only the speech of the target speaker, for example, a 10-second speech signal from the target speaker, used to provide clues for extracting the target speaker's speech from the mixed speech signal. This embodiment does not impose a specific limitation on the length of the reference speech signal; it can be of any length depending on actual needs. The target speaker feature vector, also called the target speaker embedding vector, is a vector that characterizes the audio features of the target speaker. In this embodiment, the feature vector extraction unit 510 can construct and train a speaker encoder to extract feature vectors from the speech signal, or it can utilize an existing speaker encoder model to generate feature vectors, such as ResNet34 based on a residual neural network, etc. This embodiment does not impose a specific limitation in this regard.
[0078] The target speech signal generation unit 520 is configured to extract a first speech signal of the target speaker using a first speech extraction model and a second speech extraction model, based at least on the target speaker feature vector and the mixed speech signal. Here, the mixed speech signal refers to a mixed audio including the target speaker signal, interfering speaker signals, background noise signals, etc. In this embodiment, the mixed speech signal includes speech signal segments with multiple speakers and speech signal segments where the target speaker is absent. Furthermore, the ratio (i.e., overlap rate) of the target speaker's speech duration to the total speech duration of the mixed speech signal varies, meaning it can simultaneously include speech signal segments with low, medium, and high overlap rates, thereby adapting to various mixed speech signal acquisition scenarios.
[0079] In this embodiment of the disclosure, the target speech signal generation unit 520 can obtain the first speech signal and the second speech signal of the target speaker using a first speech extraction model and a second speech extraction model based on different principles, and then generate the target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal. The first speech extraction model and the second speech extraction model can be combined with each other in various different forms to generate the first speech signal, the second speech signal and the final target speech signal of the target speaker, as will be described in further detail below.
[0080] In this embodiment, the first speech extraction model can be implemented based on a Target Speaker Voice Activity Detection (TSVAD) model for detecting the probability of target speaker activity. This model can predict the probability of the target speaker appearing in each time frame of the mixed speech signal or processed mixed speech signal with the help of the target speaker's feature vector; this probability is called the target speaker activity probability. The processed mixed speech signal refers, for example, a mixed speech signal that has already undergone speech extraction processing by other target speakers. Specifically, the first speech extraction model according to this embodiment may include at least one transformer block, which may have neural network layers such as a multi-head self-attention layer and a gated recurrent unit (GRU) layer. The multi-head self-attention layer is configured to enable the first speech extraction model to simultaneously process information at different positions in the audio sequence of the mixed speech signal or processed mixed speech signal, thereby improving the model's spatial representation ability and training speed. The gated recurrent unit layer is configured to enable the first speech extraction model to better capture the dependencies between time frames with large time step distances in the time frame sequence of the mixed speech signal or processed mixed speech signal, thereby effectively solving problems such as long-term memory and gradient explosion and gradient decay in backpropagation in neural networks.
[0081] like Figure 3 As shown, in addition to the transformer block, the first speech extraction model may also include at least one convolutional block, a smoothing block for post-processing, etc.; and in addition to the multi-head self-attention layer and the GRU layer, each transformer block may also include a linear layer and at least one residual connection and normalization layer, wherein the linear layer can be used to introduce linear components, and the residual connection and normalization layer can be used to integrate information, enhance model stability and accelerate model convergence.
[0082] After obtaining the target speaker activity probability, the target speaker activity probability can be converted into a binary form of target speaker activity label based on a predetermined threshold, such as a target speaker activity label with a value of 0 or 1. For example, when the target speaker activity label is 1, it can indicate that the current time frame is the target speaker's speech signal; otherwise, when the target speaker activity label is 0, it can indicate that the current time frame is not the target speaker's speech signal. Here, the predetermined threshold can be obtained based on model training, determined based on empirical parameters, or determined based on probability distribution characteristics, and this embodiment does not impose specific limitations on this. Subsequently, the target speaker activity label can be used to filter the mixed speech signal or the processed mixed speech signal, for example, removing the time frames with a target speaker activity label of 0 and retaining only the time frames with a target speaker activity label of 1, thereby obtaining the target speaker's speech signal, i.e., the first speech signal.
[0083] The first speech extraction model according to the embodiments of this disclosure improves the performance of target speaker speech extraction by introducing at least one transformer block including a multi-head self-attention layer and a gated recurrent unit layer, as specifically described above with reference to Table 1.
[0084] In this embodiment of the disclosure, the second speech extraction model can be a target speaker mask extraction model (TSE) for extracting the target speaker mask vector from a mixed speech signal or a processed mixed speech signal. The processed mixed speech signal, for example, refers to a mixed speech signal that has undergone other target speaker speech extraction processing. The target speaker mask vector can indicate the position of the target speaker's speech signal in the mixed speech signal or the processed mixed speech signal. For example, the target speaker mask vector can be a vector consisting of 0s and 1s with the same dimension as the vector of the mixed speech signal, where a value of 1 indicates that the element at the same position in the mixed speech signal vector is the target speaker's speech signal, and a value of 0 indicates that the element at the same position in the mixed speech signal vector is not the target speaker's speech signal.
[0085] In this embodiment of the disclosure, the first speech extraction model and the second speech extraction model can be fused in different ways to generate the target speech signal of the target speaker, as referred to above. Figures 4A-4C Specifically described. In one example, the output of the first speech extraction model can be input into the second speech extraction model, and the second speech extraction model can generate the final target speaker's target speech signal, such as... Figure 4A As shown. In another example, the output of the second speech extraction model can be input into the first speech extraction model, and the first speech extraction model can generate the final target speaker's target speech signal, such as... Figure 4BAs shown. In another example, the first speech extraction model and the second speech extraction model can be connected in parallel, and the extracted target speech signal of the target speaker can be generated, for example, by multiplying the first speech signal and the second speech signal, as shown. Figure 4C As shown.
[0086] By employing a target speaker speech extraction apparatus according to embodiments of the present disclosure, the quality of target speaker speech extraction can be improved compared to conventional target speaker speech extraction methods by introducing at least one transformer block including a multi-head self-attention sublayer and a gated recurrent sublayer into a first speech extraction model such as the TSVAD model. Furthermore, by fusing the first and second speech extraction models in different ways to extract target speaker speech, the characteristics of different speech extraction models can be effectively utilized, further improving the performance of target speaker speech extraction. The target speaker speech extraction apparatus according to embodiments of the present disclosure is particularly suitable for processing mixed speech signals with multiple speakers, absent target speakers, and low or variable overlap rates of target speakers, achieving accurate and efficient target speaker speech extraction.
[0087] Furthermore, the device according to embodiments of this disclosure (e.g., a target speaker speech extraction device, etc.) can also be used by means of Figure 6 The architecture of the exemplary computing device shown is used to implement this. Figure 6 A schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure is shown. Figure 6 As shown, computing device 600 may include a bus 610, one or more CPUs 620, read-only memory (ROM) 630, random access memory (RAM) 640, a communication port 650 connected to a network, input / output components 660, a hard disk 670, etc. Storage devices in computing device 600, such as ROM 630 or hard disk 670, may store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU. Computing device 600 may also include a user interface 680. Of course, Figure 6 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 6 One or more components in the computing device shown. The device according to embodiments of this disclosure can be configured to perform a target speaker speech extraction method according to the various embodiments of this disclosure above, or to implement a target speaker speech extraction apparatus according to the various embodiments of this disclosure above.
[0088] The embodiments of this disclosure can also be implemented as a computer-readable storage medium. A computer-readable storage medium according to embodiments of this disclosure stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the target speaker speech extraction method according to embodiments of this disclosure, as described with reference to the above figures, can be performed. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0089] According to embodiments of this disclosure, a computer program product or computer program is also provided, comprising computer-readable instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer-readable instructions from the computer-readable storage medium and execute the computer-readable instructions, causing the computer device to perform the target speaker speech extraction method described in the various embodiments above.
[0090] The program portion of a technology can be considered a "product" or "artifact" existing in the form of executable code and / or related data, and is involved in or implemented through a computer-readable medium. Tangible, permanent storage media can include memory or storage used by any computer, processor, or similar device or related module. For example, various semiconductor memories, tape drives, disk drives, or any similar device capable of providing storage functionality for software.
[0091] All software, or parts thereof, may sometimes communicate via networks, such as the Internet or other communication networks. Such communication can load software from one computer device or processor to another. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., propagated through cables, fiber optic cables, or air. Physical media used for carrier waves, such as cables, wireless connections, or fiber optic cables, can also be considered as media carrying software. In this context, unless limited to tangible "storage" media, the term "readable medium" for a computer or machine refers to the medium involved in the execution of any instructions by the processor.
[0092] This application uses specific terms to describe embodiments of the application. Terms such as "first / second embodiment," "an embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the application. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined.
[0093] Furthermore, those skilled in the art will understand that aspects of this application can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, aspects of this application can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software may be referred to as a “data block,” “module,” “engine,” “unit,” “component,” or “system.” Furthermore, aspects of this application may manifest as a computer product located on one or more computer-readable media, the product including computer-readable program code.
[0094] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in a common dictionary shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not as having an idealized or highly formalized meaning, unless expressly defined herein.
[0095] The foregoing description is illustrative of the invention and should not be construed as limiting it. Although several exemplary embodiments of the invention have been described, those skilled in the art will readily understand that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the invention. Therefore, all such modifications are intended to be included within the scope of the invention as defined in the claims. It should be understood that the foregoing description is illustrative of the invention and should not be construed as limiting it to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The invention is defined by the claims and their equivalents.
Claims
1. A method for extracting the speech of a target speaker, comprising: Generate a target speaker feature vector based on reference speech signals from the target speaker; Based at least on the mixed speech signal and the feature vector of the target speaker, a first speech signal of the target speaker is generated using a first speech extraction model, and a second speech signal of the target speaker is generated using a second speech extraction model. The first speech extraction model includes at least one transformer block, which has a multi-head self-attention layer and a gated recurrent unit layer. The first speech extraction model and the second speech extraction model are integrated to generate a target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal, wherein... The first speech extraction model is a target speaker speech activity detection model used to detect the probability of target speaker activity, and the second speech extraction model is a target speaker mask extraction model used to extract the target speaker mask vector. The target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker. The target speaker mask vector indicates the position of the target speaker's speech signal in the mixed speech signal.
2. The method according to claim 1, wherein, Generating the first speech signal of the target speaker using the first speech extraction model includes: A first hybrid feature vector is generated based on the hybrid speech signal; The first speech extraction model is used to generate the target speaker activity probability based on the target speaker feature vector and the first mixed feature vector, wherein the target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker; The first speech signal is generated based on the target speaker activity probability and the mixed speech signal.
3. The method according to claim 2, wherein, Generating the first speech signal based on the target speaker activity probability and the mixed speech signal includes: Based on a predetermined threshold, the target speaker activity probability is converted into a binary form of the target speaker activity label; The mixed speech signal is filtered using the target speaker activity tag to generate the first speech signal.
4. The method according to claim 2, wherein, Generating the second speech signal of the target speaker using the second speech extraction model includes: The second speech signal is generated using the second speech extraction model based on the target speaker feature vector and the first speech signal. The method of fusing the first speech extraction model and the second speech extraction model to generate the target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal includes: The second speech signal is identified as the target speech signal.
5. The method according to claim 2, wherein, The method of fusing the first speech extraction model and the second speech extraction model to generate a target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal includes: The target speech signal of the target speaker is generated by multiplying the first speech signal and the second speech signal.
6. The method according to claim 1, wherein, Generating the second speech signal of the target speaker using the second speech extraction model includes: A target speaker mask vector is generated based on the target speaker feature vector and the mixed speech signal using a second speech extraction model, wherein the target speaker mask vector indicates the position of the target speaker's speech signal in the mixed speech signal; The second speech signal is generated based on the target speaker mask vector and the mixed speech signal.
7. The method according to claim 6, wherein, Generating the first speech signal of the target speaker using the first speech extraction model includes: A second hybrid feature vector is generated based on the second speech signal; The target speaker activity probability is generated based on the target speaker feature vector and the second mixed feature vector, wherein the target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker; The first speech signal is generated based on the target speaker activity probability and the mixed speech signal. The method of fusing the first speech extraction model and the second speech extraction model to generate the target speech signal of the target speaker based on at least one of the first speech signal and the second speech signal includes: The first speech signal is identified as the target speech signal.
8. The method according to claim 1, wherein, The first speech extraction model further includes at least one convolutional block, and the at least one transformer block further includes a linear layer and at least one residual connection and normalization layer.
9. The method according to claim 1, wherein, The mixed speech signal includes speech signal segments with multiple speakers and speech signal segments where the target speaker is absent, and the ratio of the target speaker's speech duration to the total speech duration of the mixed speech signal is variable.
10. A target speaker speech extraction device, the device comprising: The feature vector extraction unit is configured to generate a target speaker feature vector based on a reference speech signal from the target speaker. A target speech signal generation unit is configured to generate a first speech signal of the target speaker using a first speech extraction model and a second speech signal of the target speaker using a second speech extraction model, based at least on a mixed speech signal and a feature vector of the target speaker. The unit then fuses the first and second speech extraction models to generate a target speech signal of the target speaker based on at least one of the first and second speech signals. The first speech extraction model includes at least one transformer block, which has a multi-head self-attention layer and a gated recurrent unit layer. The first speech extraction model is a target speaker speech activity detection model used to detect the probability of target speaker activity, and the second speech extraction model is a target speaker mask extraction model used to extract the target speaker mask vector. The target speaker activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target speaker. The target speaker mask vector indicates the position of the target speaker's speech signal in the mixed speech signal.
11. A target speaker speech extraction device, comprising: One or more processors; as well as One or more memories, wherein the memories store computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method as described in any one of claims 1-9.
12. A computer-readable storage medium having stored thereon computer-readable instructions, which, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-9.
13. A computer program product comprising computer-readable instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-9.