Speech extraction method and apparatus, and electronic device, computer-readable storage medium and computer program product
By introducing the transformer block of the multi-head self-attention layer and the gated cyclic unit layer into the speech extraction model, and combining it with the second speech extraction model for fusion, the problem of poor extraction effect of target speakers in the mixed speech signals in the prior art is solved, and a higher quality and accurate speech extraction is achieved.
Patent Information
- Application Number
- PCT/CN2024/127398
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art has poor extraction effect when dealing with mixed voice signals with multiple speakers, target speaker absences, and low overlap rate of target speakers.
A speech extraction method is proposed. By introducing a multi-headed self-attention layer and a transformer block of the gated cyclic unit layer in the first speech extraction model, and combining the second speech extraction model, the two models are fused in different ways to generate the speech signal of the target object.
Improves the quality and performance accuracy of the target object's speech extraction, especially for processing complex mixed speech signals.
Smart Images

Figure CN2024127398_05062025_PF_FP_ABST
Abstract
Description
Speech extraction method, device, electronic device, computer-readable storage medium, and computer program product
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on the Chinese patent application with application number 202311626816.0 and application date of November 29, 2023, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field
[0003] The present application relates to the field of artificial intelligence, and relates to, but is not limited to, a speech extraction method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0004] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0005] At present, speech technology has been widely used. The key technologies of speech technology include automatic speech recognition (ASR), text to speech (TTS), and voiceprint recognition. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising human-computer interaction methods in the future. In voiceprint recognition, target speaker extraction (TSE) uses the registered voice information of the target speaker to extract the voice of the target speaker from a mixed speech signal with noise and interference. The target speaker extraction method in the related art has poor extraction effect on mixed speech signals with many speakers, the absence of the target speaker, and a low target speaker overlap rate. Therefore, there is a need for a target speaker extraction method that can effectively process such mixed speech signals.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a speech extraction method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0008] An embodiment of the present application provides a speech extraction method, which is executed by an electronic device and includes: extracting a target object feature vector from a reference speech signal of the target object; using a first speech extraction model to generate a first speech signal of the target object based on a mixed speech signal to be recognized and the target object feature vector; wherein the first speech extraction model includes at least one transformer block, and the transformer block has a multi-head self-attention layer and a gated recurrent unit layer; using a second speech extraction model to generate a second speech signal of the target object based on the mixed speech signal and the target object feature vector; wherein the second speech extraction model is used to extract a target object mask vector of the target object from the mixed speech signal; and determining a target speech signal of the target object based on at least one of the first speech signal and the second speech signal.
[0009] An embodiment of the present application provides a speech extraction device, which includes: a feature vector extraction unit, configured to extract a target object feature vector from a reference speech signal of the target object; a target speech signal generation unit, configured to use a first speech extraction model to generate a first speech signal of the target object based on a mixed speech signal to be recognized and the target object feature vector; wherein the first speech extraction model includes at least one transformer block, and the transformer block has a multi-head self-attention layer and a gated recurrent unit layer; using a second speech extraction model, based on the mixed speech signal and the target object feature vector, generate a second speech signal of the target object; wherein the second speech extraction model is used to extract a target object mask vector of the target object from the mixed speech signal; and based on at least one of the first speech signal and the second speech signal, determine the target speech signal of the target object.
[0010] An embodiment of the present application provides an electronic device, comprising: one or more processors; and one or more memories, wherein the memories store computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the one or more processors execute the above method.
[0011] An embodiment of the present application provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the processor executes the above method.
[0012] An embodiment of the present application provides a computer program product, which includes computer-readable instructions. When the computer-readable instructions are executed by a processor, the processor performs the above method.
[0013] The speech extraction method, apparatus, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present application can improve the quality of target object speech extraction compared with traditional speech extraction methods by introducing at least one transformer block including a multi-head self-attention sublayer and a gated recurrent sublayer into the first speech extraction model; in addition, by simultaneously utilizing the first speech extraction model and the second speech extraction model to perform speech recognition of the target object, that is, fusing the first speech extraction model and the second speech extraction model in different ways to perform target object speech extraction, the respective characteristics of different speech extraction models can be effectively utilized, for example, the characteristics of the first speech extraction model having the transformer block and the characteristics of the second speech extraction model for extracting the target object mask vector of the target object can be effectively utilized, thereby further improving the performance accuracy of target object speech extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and other purposes, features, and advantages of the embodiments of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0015] FIG1 is an exemplary scenario diagram of a speech extraction system provided in an embodiment of the present application.
[0016] FIG2 is a flow chart of the speech extraction method provided in an embodiment of the present application.
[0017] FIG3 is a schematic structural diagram of the first speech extraction model of the example provided in an embodiment of the present application.
[0018] Figure 4A is a schematic diagram of a first fusion method of the first speech extraction model and the second speech extraction model provided in an embodiment of the present application.
[0019] Figure 4B is a schematic diagram of a second fusion method of the first speech extraction model and the second speech extraction model provided in an example embodiment of the present application.
[0020] Figure 4C is a schematic diagram of a third fusion method of the first speech extraction model and the second speech extraction model of the example provided in an embodiment of the present application.
[0021] FIG5 is a schematic diagram of the structure of the speech extraction device provided in an embodiment of the present application.
[0022] FIG6 is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0024] As shown in the embodiments of the present application, unless the context clearly indicates an exception, the words "a", "an", "a kind" and "the" do not specifically refer to the singular, but may also include the plural. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "include" or "comprise" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0025] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0026] In addition, flowcharts are used in the embodiments of the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations do not necessarily need to be performed in exact order. Instead, the various steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more operations can be removed from these processes.
[0027] In voice technology, voiceprint recognition uses feature analysis of one or more voice signals to identify unknown voices. Target speaker extraction (TSE), also known as target object voice extraction, is a subcategory of this technology. TSE, also known as personalized voice enhancement or personalized noise suppression, leverages the target's registered voice information to extract the target's voice from a mixed voice signal containing noise and interfering speech. Here, the target refers to the speaker of interest or a designated speaker.
[0028] Typically, a target object speech extraction model can use a pre-trained speaker feature extraction model to generate a target object feature vector based on a reference speech signal from the target object, and further extract the target object's speech signal from the mixed speech signal based on the target object feature vector. However, most target object speech extraction models are designed for mixed speech with a limited number of speakers (e.g., less than or equal to 3) and the presence of a target object, and the ratio of the target object's speech duration to the total speech duration in the mixed speech (which can be called the overlap rate) is high. For example, the ratio of the target object's speech duration to the total speech duration in the mixed speech is greater than a preset ratio. When processing mixed speech signals with a large number of speakers (e.g., greater than a preset number threshold, which can be 5), or the absence of a target object (which can be called target object absence), or a low target object overlap rate (e.g., the ratio of the target object's speech duration to the total speech duration in the mixed speech is less than a preset ratio), the effect is not good.
[0029] In order to solve the above problems, an embodiment of the present application proposes a speech extraction model (i.e., a target object speech extraction model) that can effectively extract a speech signal of a target object with higher quality from a mixed speech signal in which there are multiple speakers, the target object is absent, and the target object overlap rate is low.
[0030] FIG1 shows an exemplary scenario diagram of a speech extraction system according to an embodiment of the present application. As shown in FIG1 , the speech extraction system 100 may include a user terminal 110 , a network 120 , a server 130 , and a database 140 .
[0031] User terminal 110 may be, for example, computer 110-1 or mobile phone 110-2 shown in FIG1 . It is understood that, in fact, user terminal 110 may be any other type of electronic device capable of performing data processing, including but not limited to fixed terminals such as desktop computers and smart TVs, mobile terminals such as smart phones, tablet computers, portable computers, handheld devices, or any combination thereof, and the present application does not impose specific limitations on this.
[0032] The user terminal 110 of the embodiment of the present application can be used to receive a mixed voice signal and use the voice extraction method provided in the embodiment of the present application to generate a voice signal of a target object. In some embodiments, the voice extraction method provided in the embodiment of the present application can be executed by a processing unit of the user terminal 110. In some implementations, the user terminal 110 can use an application built into the user terminal to execute the voice extraction method provided in the embodiment of the present application. In other implementations, the user terminal 110 can execute the voice extraction method provided in the embodiment of the present application by calling an application stored externally to the user terminal.
[0033] In other embodiments, the user terminal 110 transmits the received mixed speech signal to be processed to the server 130 via the network 120, and the server 130 executes the speech extraction method. In some implementations, the server 130 may execute the speech extraction method using an application built into the server. In other implementations, the server 130 may execute the speech extraction method by calling an application stored externally to the server.
[0034] The network 120 may be a single network or a combination of at least two different networks. For example, the network 120 may include, but is not limited to, a local area network, a wide area network, a public network, a private network, or a combination of one or more of the following. The server 130 may be a standalone server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, positioning services, and big data and artificial intelligence platforms. The embodiments of the present application do not impose specific limitations on this.
[0035] The database 140 can generally refer to a device with a storage function. The database 140 is mainly used to store various data used, generated and output during the work of the user terminal 110 and the server 130. The database 140 can be local or remote. The database 140 can include various memories, such as random access memory (RAM), read-only memory (ROM), etc. The storage devices mentioned above are just some examples, and the storage devices that can be used in the system are not limited to these. The database 140 can be interconnected or communicated with the server 130 or a part thereof via the network 120, or directly interconnected or communicated with the server 130, or a combination of the above two methods.
[0036] The following describes the speech extraction method according to an embodiment of the present application with reference to FIG2 . FIG2 shows a flow chart of speech extraction method 200 according to an embodiment of the present application, which can be a target object speech extraction method. As described above, speech extraction method 200 can be executed by a user terminal or a server, and this embodiment of the present application does not impose any specific limitations on this.
[0037] In step S210 , a target object feature vector is extracted from a reference speech signal of the target object.
[0038] The target object can be any object capable of producing speech, for example, a target speaker, an artificial intelligence device, or a voice device. When the target object is a target speaker, a reference speech signal from the target speaker can be obtained, and the target object feature vector can be extracted from the reference speech signal. When the target object is an artificial intelligence device, the speech produced by the artificial intelligence device can be collected and used as the reference speech signal. When the target object is a voice device, the speech produced by the voice device can be collected and used as the reference speech signal, and the target object feature vector can be extracted from the reference speech signal.
[0039] The reference voice signal is one or more known signals used to assist in extracting the target voice signal of the target object, and may be a signal that only includes the voice of the target object. In an embodiment of the present application, the reference voice signal may include the following types: a pure voice signal, prior information, and an auxiliary microphone signal. Among them, for the pure voice signal, in some application scenarios, there may already be a pure voice recording of the target speaker, which can be used as a reference voice signal to help identify and extract the target voice signal in a noisy environment. For the prior information, in some cases, the reference voice signal may also be based on the prior information of the target speaker, such as specific features of his voice (such as resonance peak frequency, timbre, etc.), which can help the voice extraction model better focus on the target voice signal of the target object. For the auxiliary microphone signal, when using a microphone array for voice extraction, the reference voice signal may come from a specific microphone or a combination of multiple microphones. The signals captured by these microphones can be used to determine the location and direction of the sound source, thereby helping to extract the target voice signal of the target object.
[0040] For example, the reference speech signal can be a 10-second speech signal from the target subject, which is used to provide clues for extracting the target subject's speech from the mixed speech signal. The embodiments of the present application do not impose any specific restrictions on the length of the reference speech signal; it can be a speech signal of any length based on actual needs.
[0041] The embodiments of the present application further provide several methods for extracting reference speech signals. In one extraction method, the reference speech signal refers to a reference speech signal of multiple target speech signals obtained by processing the speech signal captured by the microphone array through algorithms such as sound source localization, beamforming, and wavelet decomposition; these reference speech signals are used as auxiliary information in the Independent Component Analysis (ICA) algorithm to more effectively extract the target speech signal of the target object from background noise and interference sounds. In another extraction method, the reference speech signal can be the distance features and inter-channel features obtained by extracting the time-frequency features of the multi-channel speech signal and mapping it. These features can be regarded as a reference speech signal to help determine and extract the target speech signal of the target object.
[0042] After obtaining the reference speech signal of the target object, a target object feature vector can be generated based on the reference speech signal from the target object. The target object feature vector can be called a target object embedding vector, which is a vector that can characterize the audio features of the target object. In an embodiment of the present application, a speaker encoder can be constructed and trained to extract feature vectors from speech signals, or an existing speaker encoder model can be used to generate feature vectors. The speaker encoder model is a speech signal encoder, which is a type of device or algorithm used to convert analog or digital speech signals into a digital format suitable for storage or transmission. For example, ResNet34 based on a residual neural network can be used as a speaker encoder model, etc., and the embodiment of the present application does not impose specific restrictions on this.
[0043] In step S220, a first speech signal of the target object is generated based on the mixed speech signal to be recognized and the target object feature vector using the first speech extraction model.
[0044] In an embodiment of the present application, a first speech signal of a target object can be extracted using a first speech extraction model based at least on a feature vector of the target object and a mixed speech signal. Here, a mixed speech signal refers to a mixed audio including a target object signal and an interfering speaker signal, a background noise signal, and the like. In an embodiment of the present application, the mixed speech signal includes speech signal segments in which multiple speakers exist and speech signal segments in which the target object is absent, and the ratio of the speech duration of the target object to the total speech duration of the mixed speech signal (i.e., the overlap rate) varies, i.e., in the mixed speech signal, the speech duration of different speech signal segments of the target object and the ratio of the speech duration of the total speech duration of the mixed speech signal are independent of each other, and can simultaneously include speech signal segments with low, medium, and high overlap rates of the target object, thereby adapting to various mixed speech signal acquisition scenarios. For example, in a mixed voice signal with a length of 10 minutes, the voice signal segment from 0 to 1 minute is the first voice signal segment containing the target object, the voice signal segment from 1 to 2 minutes is the second voice signal segment in which the target object is absent, the voice signal segment from 2 to 6 minutes is the third voice signal segment containing the target object, the voice signal segment from 6 to 8 minutes is the fourth voice signal segment in which the target object is absent, and the voice signal segment from 8 to 10 minutes is the fifth voice signal segment containing the target object. It can be seen that the first voice signal segment, the third voice signal segment and the fifth voice signal segment are all voice signal segments containing the target object, but the lengths of these three voice signal segments are different. Therefore, the ratios of these three voice signal segments to the total voice time of the mixed voice signal are different and unrelated.
[0045] For example, the mixed speech signal Y can be expressed as the following formula (1):
[0046] Y=S+I+N (1)
[0047] Wherein, Y represents a mixed speech signal; S represents a target object signal; I represents an interfering speaker signal; and N represents a background noise signal. Generally speaking, the target object speech extraction method according to the embodiment of the present application is dedicated to solving the problem of single-channel target object speech extraction, that is, the mixed speech signal Y comes from a single audio acquisition channel. However, the speech extraction method of the embodiment of the present application can also be applied to multi-channel target object speech extraction with appropriate modifications, as long as there is no contradiction. In the following, the embodiment of the present application is explained by taking single-channel target object speech extraction as an example.
[0048] When generating the first speech signal of the target object, a first speech extraction model can be used. That is, the first speech extraction model can be used to generate the first speech signal of the target object based on the mixed speech signal to be recognized and the target object feature vector. Here, the first speech extraction model includes at least one transformer block, and the transformer block has a multi-head self-attention layer and a gated recurrent unit layer.
[0049] In step S230, a second speech signal of the target object is generated based on the mixed speech signal and the target object feature vector using a second speech extraction model.
[0050] Here, the second speech extraction model is used to extract the target object mask vector of the target object from the mixed speech signal. In an embodiment of the present application, the first speech extraction model and the second speech extraction model based on different principles can be used to obtain the first speech signal and the second speech signal of the target object, respectively. The first speech extraction model and the second speech extraction model can be combined with each other in various forms to generate the first speech signal, the second speech signal, and the final target speech signal of the target object, which will be further described in detail below.
[0051] In an embodiment of the present application, the first speech extraction model can be implemented based on a target speaker voice activity detection (TSVAD) model for detecting the target object activity probability, which can predict the probability of the target object appearing in each time frame in a mixed speech signal or a processed mixed speech signal with the help of a target object feature vector, which is called the target object activity probability.
[0052] Here, the processed mixed speech signal refers to, for example, a mixed speech signal that has been extracted from other target objects, or it may refer to a mixed speech signal obtained by processing the original mixed speech signal (i.e., the originally collected mixed speech signal) using any at least one of the following processing methods: pre-emphasis processing, endpoint detection, denoising processing, speech enhancement processing, frame stacking processing, normalization processing, silence detection and deletion processing, feature conversion processing, speech rate conversion processing, pitch and energy normalization processing, and nonlinear distortion removal processing, etc.
[0053] It should be noted that pre-emphasis is a technique that improves the characteristics of speech signals in the frequency domain. Its primary purpose is to enhance the high-frequency portion of speech, making it easier to analyze and identify in subsequent processing. In other words, since the high-frequency portion of the original mixed speech signal typically decays faster, pre-emphasis can be used to enhance this high-frequency portion, making it easier to extract the mixed speech signal's features in subsequent processing. Endpoint detection determines the beginning and end points of a speech signal to eliminate silence and improve processing efficiency. Denoising uses various algorithms (such as Wiener filtering, wavelet transform, and spectral subtraction) to reduce background noise. Speech enhancement uses algorithms (such as echo cancellation, noise suppression, and nonlinear processing) to improve speech clarity and intelligibility. Frame deinterlacing divides the speech signal into short time frames. For example, window functions such as Hamming or Hanning windows can be used to reduce edge artifacts. Normalization adjusts the amplitude of the speech signal to a uniform scale for easier comparison and further processing. Silence detection and removal identifies and removes silence from the speech signal to reduce unnecessary processing and storage. Feature conversion processing involves converting speech signals from the time domain to the frequency domain or other domains (such as Mel-frequency cepstral coefficients) to extract more abstract features. Speech rate conversion processing involves adjusting the playback speed of speech without changing its pitch. Pitch and energy normalization processing involves normalizing the pitch and energy of speech signals to reduce differences between speakers. Nonlinear distortion removal processing reduces or eliminates nonlinear distortion caused by transmission or processing.
[0054] The first speech extraction model of the embodiment of the present application may include at least one transformer block, which may have a neural network layer such as a multi-head self-attention layer (Multi-head Self-attention), a gated recurrent unit (GRU) layer, etc. Among them, the multi-head self-attention layer is configured to enable the first speech extraction model to simultaneously process information at different positions in the audio sequence of the mixed speech signal or the processed mixed speech signal, thereby improving the spatial expression ability of the model and increasing the training speed of the model. The multi-head self-attention layer processes the input data based on the multi-head self-attention mechanism. The multi-head self-attention mechanism is a key mechanism in the Transformer model, which allows the model to learn information in parallel in different subspaces, thereby improving the expressiveness and flexibility of the model. The multi-head self-attention mechanism is to split the self-attention into multiple "heads", each head has its own weight matrix, so that the self-attention operation is performed in parallel in multiple subspaces. The steps of the working principle of the multi-head self-attention mechanism can be summarized as follows: first, the input is split, and the input sequence can be split into multiple "heads", and each head receives a subset of the sequence; then, the input sequence is linearly transformed, and each head can be linearly transformed to obtain query, key, and value vectors respectively; then, based on the result of the linear transformation, the attention score is calculated, that is, for the query vector of each head, the attention score of the query vector and all key vectors is calculated, and these scores represent the correlation between the query vector and the key vector; then, the attention score of each head is normalized by the softmax function to obtain a probability distribution; finally, the probability distribution obtained by softmax is used to perform weighted summation on the value vector to obtain the output vector of each head, and the output vectors of all heads are merged and linearly transformed again to obtain the final output.
[0055] The gated recurrent unit layer is configured to enable the first speech extraction model to better capture the following information: the dependency between time frames in a time frame sequence of a mixed speech signal or a processed mixed speech signal whose time step distance is greater than a preset distance threshold, thereby effectively solving problems such as gradient explosion and gradient decay in long-term memory and back propagation in neural networks.
[0056] Figure 3 shows a schematic diagram of the structure of the first speech extraction model provided in an embodiment of the present application, in which TSVAD is used as an example of the first speech extraction model for illustration. As shown in Figure 3, in addition to the converter block 301, the first speech extraction model can also include at least one convolution block (such as the convolution block 302 and the convolution block 303 shown in Figure 3), a smoothing block 304 for post-processing, etc.; and in addition to the multi-head self-attention layer 305 and the GRU layer 306, each converter block can also include a linear layer 307 and at least one residual connection and normalization layer (such as the residual connection and normalization layer 308 and the residual connection and normalization layer 309 shown in Figure 3), wherein the linear layer 307 can be used to introduce linear components, and the residual connection and normalization layer can be used to integrate information and enhance model stability and accelerate model convergence. In the example of Figure 3, the first speech extraction model is illustrated as having two converter blocks 301, but the embodiment of the present application is not limited to this. The first speech extraction model can also include any number of converter blocks according to actual needs. Similarly, the numbers of the blocks and layers shown in FIG3 are examples and do not constitute any limitation to the embodiments of the present application.
[0057] The mixed speech signal Y or the processed mixed speech signal Y' can be divided into frame vectors [y1, y2, ..., y t ] (where t represents the number of frames). Then, as shown in FIG3 , a speaker encoder 300 can be used to extract a mixed feature vector E from the mixed speech signal divided into frame vectors or the processed mixed speech signal. Y =[E y1 , E y2 ,...,E yt ]. Mixed eigenvector E Y The target object feature vector E extracted from the reference speech signal of the target object S are connected and input into the TSVAD model. After being processed by at least one convolution block, at least one transformer block and a smoothing block, the target object activity probability P = [P y1 , P y2 ,...,P yt ]. This process can be expressed as the following formula (2):
[0058] P = TSVAD (E Y , E S ) (2)
[0059] Among them, P represents the activity probability of the target object, E Y represents the mixed eigenvector, E S represents the target object feature vector, TSVAD(·) represents the target object activity detection function, which is based on the mixed feature vector EY and the target object feature vector E S Generate target object activity probabilities.
[0060] After obtaining the target object activity probability, the target object activity probability can be converted into a target object activity label in binary form based on a predetermined threshold, such as a target object activity label with a value of 0 or 1. For example, when the target object activity label is 1, it can indicate that the current time frame is a speech signal of the target object; otherwise, when the target object activity label is 0, it can indicate that the current time frame is not a speech signal of the target object. Here, the predetermined threshold can be obtained based on model training, determined based on empirical parameters, or determined based on probability distribution characteristics, and the embodiments of the present application do not impose specific restrictions on this. Afterwards, the target object activity label can be used to filter the mixed speech signal or the processed mixed speech signal, for example, filtering out the time frame with a target object activity label of 0, and retaining only the time frame with a target object activity label of 1, so that the speech signal of the target object, that is, the first speech signal, can be obtained.
[0061] The first speech extraction model of the embodiment of the present application can improve the performance of target object speech extraction by introducing at least one transformer block including a multi-head self-attention layer and a gated recurrent unit layer. Taking the TSVAD model as an example of the first speech extraction model for illustration, the following Table 1 shows the results of the ablation experiment on the TSVAD model, where TSVAD represents the TSVAD model with at least one transformer block of the embodiment of the present application, and TSVAD* represents removing the transformer block from the TSVAD model of the embodiment of the present application, or replacing the transformer block with other neural network modules such as long short-term memory networks, and other experimental conditions are the same. In Table 1, DER represents the target object segmentation clustering error rate, INT represents the energy difference between the mixed speech signal and the extracted target speech signal of the target object, and the smaller the value of DER and the larger the value of INT, the better the target object speech extraction performance. As can be seen from Table 1, compared with the TSVAD* model with the converter block removed or replaced, the TSVAD model of the embodiment of the present application has a smaller DER and a larger INT, indicating that the TSVAD model of the embodiment of the present application has better performance in extracting the target object speech, which shows that the introduction of at least one converter block plays an important role in improving the target object speech extraction performance.
[0062] Table 1 Ablation experiment results of TSVAD model
[0063] In an embodiment of the present application, the second speech extraction model may be a target speaker mask extraction model (TSE) for performing the following processing: extracting a target object mask vector from a mixed speech signal or a processed mixed speech signal, wherein the processed mixed speech signal, for example, refers to a mixed speech signal that has undergone other target object speech extraction processing, and the target object mask vector may indicate the position of the target object speech signal in the mixed speech signal or the processed mixed speech signal. For example, the target object mask vector may be a vector consisting of 0 and 1 having the same dimension as the vector of the mixed speech signal, wherein a certain element value of 1 indicates that the element at the same position of the mixed speech signal vector is the target object speech signal, and a certain element value of 0 indicates that the element at the same position of the mixed speech signal vector is not the target object speech signal. The process of extracting the target object mask vector from a mixed speech signal or a processed mixed speech signal can be expressed as the following formula (3):
[0064] M=TSE(Y,E S ) (3)
[0065] Among them, M represents the target object mask vector, Y represents the mixed speech signal, and E S represents the target object feature vector, TSE(·) represents the target object mask extraction function, which is based on the mixed speech signal and the target object feature vector E S Generate a target object mask vector. Based on the target object mask vector and the mixed speech signal or the processed mixed speech signal, a second speech signal that retains only the speech signal of the target object can be generated. For example, the second speech signal can be generated by multiplying the target object mask vector by the mixed speech signal or the processed mixed speech signal.
[0066] In step S240 , a target speech signal of the target object is determined based on at least one of the first speech signal and the second speech signal.
[0067] In the embodiment of the present application, the first speech extraction model and the second speech extraction model can be fused in different ways to generate a target speech signal of the target object. The fusion process can be expressed as the following formula (4), for example:
[0068] S'=F(Y,M,P) (4)
[0069] Wherein, S' represents the extracted target speech signal of the target object, Y represents the mixed speech signal, M represents the target object mask vector, P represents the target object activity probability, and F(·) represents the fusion function, which can generate the target speech signal based on the mixed speech signal Y, the target object mask vector M and the target object activity probability P.
[0070] The following describes different fusion methods of the first speech extraction model and the second speech extraction model of an embodiment of the present application with reference to Figures 4A to 4C. Figure 4A shows a first fusion method of the first speech extraction model and the second speech extraction model of an embodiment of the present application, Figure 4B shows a second fusion method of the first speech extraction model and the second speech extraction model of an embodiment of the present application, and Figure 4C shows a third fusion method of the first speech extraction model and the second speech extraction model of an embodiment of the present application.
[0071] In one example, the output of the first speech extraction model can be input into the second speech extraction model, and the target speech signal of the final target object can be generated by the second speech extraction model. As shown in FIG4A , a first mixed feature vector can be generated based on the mixed speech signal Y. For example, the speaker encoder 300 shown in FIG4A can be used to generate the first mixed feature vector E based on the mixed speech signal Y. Y , that is, the mixed speech signal Y is input into the speaker encoder 300, and the mixed speech signal Y is encoded by the speaker encoder 300 to obtain the first mixed feature vector E Y Afterwards, the first speech extraction model 401 is based on the first mixed feature vector E Y and the target object feature vector E generated based on the reference speech signal from the target object in the above embodiment S , to generate the target object activity probability P. Furthermore, the first speech signal S1 of the target object can be generated based on the target object activity probability P and the mixed speech signal Y. For example, the target object activity probability can be converted into a binary target object activity label based on a predetermined threshold, and the mixed speech signal can be filtered using the target object activity label to generate the first speech signal S1.
[0072] At this point, the first speech extraction model 401 has performed preliminary target object speech extraction on the mixed speech signal Y. Therefore, the first speech signal S1 can also be called a processed mixed speech signal. In order to perform more accurate target object speech extraction, the first speech signal S1 can be input into the second speech extraction model 402. The second speech extraction model 402 can be based on the first speech signal S1 and the target object feature vector E S A second speech signal is generated and used as the target speech signal S' of the extracted target object.
[0073] In this example, a lower predetermined threshold (for example, the predetermined threshold may belong to a first threshold interval, and the maximum value of the first threshold interval is less than a specific maximum value threshold) can be used to generate a target object activity label for filtering the mixed speech signal, thereby minimizing the risk of erroneously filtering out the speech signal of the target object; by combining the second speech extraction model to perform further target object speech extraction, the speech signal of the target object with higher quality can be effectively extracted.
[0074] In another example, the output of the second speech extraction model 402 can be input to the first speech extraction model 401, and the target speech signal of the final target object can be generated by the first speech extraction model 401. As shown in FIG4B , the target object feature vector E generated by the second speech extraction model 402 based on the mixed speech signal Y and the reference speech signal from the target object in the above embodiment can be used. S , to generate a target object mask vector M, and then generate a second speech signal S2 based on the target object mask vector M and the mixed speech signal Y.
[0075] At this point, the second speech extraction model 402 has performed preliminary target object speech extraction on the mixed speech signal Y. Therefore, the second speech signal S2 can also be called a processed mixed speech signal. In order to perform more accurate target object speech extraction, a second mixed feature vector can be generated based on the second speech signal S2. For example, the speaker encoder 300 shown in FIG4B is used to generate the second mixed feature vector E based on the second speech signal S2. Y’ Afterwards, the first speech extraction model 401 can be based on the second mixed feature vector E Y’ and the target object feature vector E S , generate a target object activity probability P, and then generate a first speech signal based on the target object activity probability P and the mixed speech signal Y, and use the first speech signal as the target speech signal S' of the extracted target object.
[0076] In this example, the second speech extraction model 402 is first used to extract an initial speech signal with a low overlap rate from the mixed speech signal, and then the first speech extraction model 401 is used to achieve further accurate extraction of the target speech signal of the target object.
[0077] In another example, the first speech extraction model 401 and the second speech extraction model 402 can be connected in parallel. As shown in FIG4C , a first mixed feature vector can be generated based on the mixed speech signal Y. For example, the speaker encoder 300 shown in FIG4A can be used to generate the first mixed feature vector E based on the mixed speech signal Y. Y Afterwards, the first speech extraction model 401 is based on the first mixed feature vector EY and the target object feature vector E generated based on the reference speech signal from the target object in the above embodiment S , to generate the target object activity probability P. Then, the first speech signal S1 of the target object can be generated based on the target object activity probability P and the mixed speech signal Y. On the other hand, the second speech extraction model 402 can be based on the mixed speech signal Y and the target object feature vector E generated based on the reference speech signal from the target object in the above embodiment. S , to generate the target object mask vector M, and then generate the second speech signal S2 based on the target object mask vector M and the mixed speech signal Y. Afterwards, the target speech signal of the target object can be generated based on the first speech signal S1 and the second speech signal S2. For example, as shown in FIG4C , the target speech signal S' of the extracted target object can be generated by multiplying the first speech signal S1 and the second speech signal S2.
[0078] In this example, the first speech extraction model 401 and the second speech extraction model 402 are connected in parallel with each other, and the interaction between the two models is small, thereby reducing the dependence on auxiliary information and improving the efficiency of target object speech extraction.
[0079] The above describes different fusion methods of the first speech extraction model and the second speech extraction model in conjunction with Figures 4A to 4C. It should be noted that the fusion methods and specific modules, input parameters, and output parameters shown in Figures 4A to 4C are merely examples or for illustrative purposes, and do not mean that the embodiments of the present application are limited thereto. For example, the first speech extraction model and the second speech extraction model can also be fused in other appropriate ways.
[0080] Utilizing the speech extraction method of the embodiment of the present application, by introducing at least one transformer block including a multi-head self-attention sublayer and a gated recurrent sublayer into a first speech extraction model such as a TSVAD model, the quality of target object speech extraction can be improved compared to traditional target object speech extraction methods. Furthermore, by fusing the first speech extraction model and the second speech extraction model in different ways to perform target object speech extraction, the characteristics of different speech extraction models can be effectively utilized, further improving the performance of target object speech extraction. It should be noted that the target object speech extraction method of the embodiment of the present application is particularly suitable for processing mixed speech signals with multiple speakers, the absence of the target object, and low or variable target object overlap, thereby achieving accurate and efficient target object speech extraction.
[0081] The following describes the speech extraction device for the target object of an embodiment of the present application with reference to Figure 5. Figure 5 shows a schematic structural diagram of the speech extraction device 500 for the target object of an embodiment of the present application. As shown in Figure 5, the speech extraction device 500 includes a feature vector extraction unit 510 and a target speech signal generation unit 520. In addition to these two units, the device 500 may also include other related components, but since these components are not relevant to the content of the embodiment of the present application, a detailed description of their specific contents is omitted here. In addition, since the details of some functions of the speech extraction device 500 are similar to the details of the steps of the speech extraction method 200 described with reference to Figure 2, for the sake of brevity, a repeated description of some contents is omitted here. The speech extraction device 500 of the embodiment of the present application can be implemented as a terminal or a server, as described above with reference to Figure 1.
[0082] The feature vector extraction unit 510 is configured to extract the target object feature vector from the reference speech signal of the target object. In an embodiment of the present application, the feature vector extraction unit 510 can construct and train a speaker encoder to extract feature vectors from the speech signal, or can use an existing speaker encoder model to generate feature vectors. For example, ResNet34 based on a residual neural network can be used, and this embodiment of the present application does not impose any specific restrictions on this. The target speech signal generation unit 520 is configured to use a first speech extraction model to generate a first speech signal of the target object based on the mixed speech signal to be recognized and the target object feature vector; wherein the first speech extraction model includes at least one transformer block, and the transformer block has a multi-head self-attention layer and a gated recurrent unit layer; use a second speech extraction model to generate a second speech signal of the target object based on the mixed speech signal and the target object feature vector; wherein the second speech extraction model is used to extract the target object mask vector of the target object from the mixed speech signal; and determine the target speech signal of the target object based on at least one of the first speech signal and the second speech signal.
[0083] In some embodiments, the target speech signal generating unit 520 is further configured to extract a first mixed feature vector from the mixed speech signal; generate a target object activity probability based on the target object feature vector and the first mixed feature vector using the first speech extraction model, wherein the target object activity probability represents the probability that each time frame in the mixed speech signal is a speech signal of the target object; and generate the first speech signal based on the target object activity probability and the mixed speech signal.
[0084] In some embodiments, the target speech signal generating unit 520 is further configured to convert the target object activity probability into a binary target object activity label based on a predetermined threshold; and use the target object activity label to filter the time frames in the mixed speech signal to obtain the first speech signal.
[0085] In some embodiments, the target speech signal generation unit 520 is further configured to obtain the first speech signal generated based on the mixed speech signal and the target object feature vector; and generate the second speech signal based on the target object feature vector and the first speech signal using the second speech extraction model.
[0086] In some embodiments, the target speech signal generating unit 520 is further configured to determine the second speech signal as the target speech signal of the target object.
[0087] In some embodiments, the target speech signal generating unit 520 is further configured to multiply the first speech signal and the second speech signal to obtain the target speech signal of the target object.
[0088] In some embodiments, the target speech signal generating unit 520 is further configured to generate a target object mask vector based on the target object feature vector and the mixed speech signal using the second speech extraction model, wherein the target object mask vector indicates the position of the target object speech signal in the mixed speech signal; and generate the second speech signal based on the target object mask vector and the mixed speech signal.
[0089] In some embodiments, the target speech signal generating unit 520 is further configured to generate a second mixed feature vector based on the second speech signal; generate a target object activity probability based on the target object feature vector and the second mixed feature vector, wherein the target object activity probability represents the probability that each time frame in the mixed speech signal is a speech signal of the target object; and generate the first speech signal based on the target object activity probability and the mixed speech signal.
[0090] In some embodiments, the target speech signal generating unit 520 is further configured to determine the first speech signal as the target speech signal of the target object.
[0091] In some embodiments, the first speech extraction model further comprises at least one convolution block; and the at least one transformer block further comprises a linear layer, at least one residual connection, and a normalization layer.
[0092] In some embodiments, the first speech extraction model is a target object speech activity detection model for detecting the target object activity probability, and the second speech extraction model is a target object mask extraction model for extracting a target object mask vector; the target object activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target object; the target object mask vector indicates the position of the target object speech signal in the mixed speech signal.
[0093] In some embodiments, the mixed speech signal includes speech signal segments in which multiple objects exist and speech signal segments in which the target object is absent; in the mixed speech signal, the ratios of the speech durations of different speech signal segments of the target object to the total speech duration of the mixed speech signal are independent of each other.
[0094] It should be noted that in an embodiment of the present application, the first speech extraction model can be implemented based on a target object speech activity detection (TSVAD) model for detecting the probability of target object activity. It can predict the probability of the target object appearing in each time frame of a mixed speech signal or a processed mixed speech signal with the help of a feature vector of the target object, which is called the target object activity probability. Wherein, the processed mixed speech signal refers to, for example, a mixed speech signal that has undergone other target object speech extraction processes. The first speech extraction model of an embodiment of the present application may include at least one transformer block, which may have a neural network layer such as a multi-head self-attention layer and a gated recurrent unit layer. Wherein, the multi-head self-attention layer is configured to enable the first speech extraction model to simultaneously process information at different positions in the audio sequence of the mixed speech signal or the processed mixed speech signal, thereby improving the spatial expression ability of the model and increasing the training speed of the model; the gated recurrent unit layer is configured to enable the first speech extraction model to better capture the dependencies between time frames with large time step distances in the time frame sequence of the mixed speech signal or the processed mixed speech signal, thereby effectively solving problems such as gradient explosion and gradient decay in long-term memory and backpropagation in neural networks. After obtaining the target object activity probability, the target object activity probability can be converted into a target object activity label in binary form based on a predetermined threshold, such as a target object activity label with a value of 0 or 1. For example, when the target object activity label is 1, it can indicate that the current time frame is a speech signal of the target object; otherwise, when the target object activity label is 0, it can indicate that the current time frame is not a speech signal of the target object. Here, the predetermined threshold can be obtained based on model training, determined based on empirical parameters, or determined based on probability distribution characteristics, and the embodiments of the present application do not impose specific restrictions on this. Afterwards, the target object activity label can be used to filter the mixed speech signal or the processed mixed speech signal, for example, removing the time frame with a target object activity label of 0 and retaining only the time frame with a target object activity label of 1, so that the speech signal of the target object, that is, the first speech signal, can be obtained. The first speech extraction model of the embodiment of the present application can improve the performance of target object speech extraction by introducing at least one transformer block including a multi-head self-attention layer and a gated recurrent unit layer.
[0095] In an embodiment of the present application, the second speech extraction model may be a target object mask extraction model for extracting a target object mask vector from a mixed speech signal or a processed mixed speech signal, wherein the processed mixed speech signal, for example, refers to a mixed speech signal that has undergone other target object speech extraction processing, and the target object mask vector may indicate the position of the target object speech signal in the mixed speech signal or the processed mixed speech signal. For example, the target object mask vector may be a vector composed of 0 and 1 having the same dimension as the vector of the mixed speech signal, wherein a certain element value of 1 indicates that the element at the same position of the mixed speech signal vector is the target object speech signal, and a certain element value of 0 indicates that the element at the same position of the mixed speech signal vector is not the target object speech signal.
[0096] By utilizing the target object speech extraction device of the embodiment of the present application, the quality of target object speech extraction can be improved compared with the traditional target object speech extraction method by introducing at least one transformer block including a multi-head self-attention sublayer and a gated recurrent sublayer in a first speech extraction model such as a TSVAD model; in addition, by fusing the first speech extraction model and the second speech extraction model in different ways to perform target object speech extraction, the characteristics of different speech extraction models can be effectively utilized to further improve the performance of target object speech extraction. The target object speech extraction device according to the embodiment of the present application is particularly suitable for processing mixed speech signals in which there are multiple speakers, the target object is absent, and the target object overlap rate is low or has a variable overlap rate, thereby achieving accurate and efficient target object speech extraction.
[0097] In addition, an embodiment of the present application also provides an electronic device (for example, the electronic device can be a target object voice extraction device, etc.) which can also be implemented with the help of the architecture of the exemplary electronic device shown in Figure 6. Figure 6 shows a schematic diagram of the architecture of the exemplary electronic device of the embodiment of the present application. As shown in Figure 6, the electronic device 600 may include a bus 610, one or more CPUs 620, a read-only memory (ROM) 630, a random access memory (RAM) 640, a communication port 650 connected to a network, an input / output component 660, a hard disk 670, etc. The storage device in the electronic device 600, such as the ROM 630 or the hard disk 670, can store various data or files used for computer processing and communication, as well as program instructions executed by the CPU. The electronic device 600 may also include a user interface 680. Of course, the architecture shown in Figure 6 is only exemplary. When implementing different electronic devices, one or more components in the electronic device shown in Figure 6 can be omitted according to actual needs. The electronic device of the embodiment of the present application can be configured to perform the voice extraction method according to the above-mentioned various embodiments of the present application, or to implement the voice extraction device according to the above-mentioned various embodiments of the present application.
[0098] The embodiments of the present application may also be implemented as a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the speech extraction method according to the embodiments of the present application described with reference to the above figures may be executed. The computer-readable storage medium includes, but is not limited to, at least one of a volatile memory and a non-volatile memory. For example, the volatile memory may include at least one of a random access memory and a cache memory. For example, the non-volatile memory may include a read-only memory, a hard disk, a flash memory, etc.
[0099] The present application also provides a computer program product or computer program, which includes computer-readable instructions stored in a computer-readable storage medium. A processor of an electronic device can read the computer-readable instructions from the computer-readable storage medium and execute the computer-readable instructions, causing the electronic device to perform the speech extraction method described in each of the above embodiments.
[0100] It should be noted that the program portion of the technology can be considered a "product" or "manufactured good" in the form of executable code and related data, implemented or implemented through computer-readable storage media. Tangible, permanent storage media can include any memory or storage used by a computer, processor, or similar device or related module. For example, various semiconductor memories, tape drives, disk drives, or any similar device capable of providing storage for software.
[0101] All or part of the software may sometimes be communicated over a network, such as the Internet or other communications network. Such communications can load the software from one computer device or processor to another. Therefore, another medium capable of transmitting software elements may also be used as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., which are transmitted through cables, optical cables or air. Physical media used to carry data, such as cables, wireless connections or optical cables and the like, can also be considered as the medium that carries the software. As used herein, unless limited to tangible "storage" media, other terms referring to computer or machine "readable media" refer to media that participate in the process of executing any instructions by the processor.
[0102] This application uses specific terms to describe the embodiments of this application. For example, "first / second embodiment," "one embodiment," and "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this application. Therefore, it should be emphasized and noted that "one embodiment," "an embodiment," or "an alternative embodiment" mentioned twice or multiple times in different locations in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application may be appropriately combined.
[0103] In addition, it will be appreciated by those skilled in the art that the various aspects of the embodiments of the present application can be illustrated and described by several patentable species or situations, including any new and useful process, machine, product or material combination, or any new and useful improvement thereof. Accordingly, the various aspects of the embodiments of the present application can be performed entirely by hardware, can be performed entirely by software (including firmware, resident software, microcode, etc.), or can be performed by a combination of hardware and software. The above hardware or software can all be referred to as "data block", "module", "engine", "unit", "component" or "system". In addition, the various aspects of the embodiments of the present application may be expressed as a computer product in one or more computer-readable media, which includes computer-readable program coding.
[0104] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the embodiments of the present application belong. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or highly formalized sense, unless explicitly defined as such herein.
[0105] The above is an explanation of the embodiment of the present application and should not be considered as a limitation thereto. Although several exemplary embodiments of the embodiment of the present application have been described, it will be readily understood by those skilled in the art that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the embodiment of the present application. Therefore, all such modifications are intended to be included within the scope of the present application as defined by the claims. It should be understood that the above is an explanation of the embodiment of the present application and should not be considered as being limited to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims.
Claims
1. A speech extraction method, the method being performed by an electronic device, the method comprising: Extracting a target object feature vector from a reference speech signal of the target object; Using a first speech extraction model, based on the mixed speech signal to be recognized and the target object feature vector, a first speech signal of the target object is generated; wherein the first speech extraction model includes at least one transformer block, and the transformer block has a multi-head self-attention layer and a gated recurrent unit layer; Using a second speech extraction model, based on the mixed speech signal and the target object feature vector, a second speech signal of the target object is generated; wherein the second speech extraction model is used to extract a target object mask vector of the target object from the mixed speech signal; Based on at least one of the first speech signal and the second speech signal, a target speech signal of the target object is determined.
2. The method according to claim 1, wherein: The method of using the first speech extraction model to generate a first speech signal of the target object based on the mixed speech signal to be recognized and the target object feature vector includes: Extracting a first mixed feature vector from the mixed speech signal; Using the first speech extraction model, generating a target object activity probability based on the target object feature vector and the first mixed feature vector, wherein the target object activity probability represents the probability that each time frame in the mixed speech signal is a speech signal of the target object; The first speech signal is generated based on the target object activity probability and the mixed speech signal.
3. The method according to claim 1 or 2, wherein: The generating the first speech signal based on the target object activity probability and the mixed speech signal comprises: Converting the target object activity probability into a target object activity label in a binary form based on a predetermined threshold; The time frames in the mixed speech signal are filtered using the target object activity label to obtain the first speech signal.
4. The method according to any one of claims 1 to 3, wherein: The step of using the second speech extraction model to generate a second speech signal of the target object based on the mixed speech signal and the target object feature vector comprises: Acquire the first speech signal generated based on the mixed speech signal and the target object feature vector; The second speech signal is generated based on the target object feature vector and the first speech signal using the second speech extraction model.
5. The method according to any one of claims 1 to 4, wherein: The step of determining the target speech signal of the target object based on at least one of the first speech signal and the second speech signal includes: The second speech signal is determined as a target speech signal of the target object.
6. The method according to any one of claims 1 to 5, wherein: The step of determining the target speech signal of the target object based on at least one of the first speech signal and the second speech signal includes: The first speech signal and the second speech signal are multiplied to obtain a target speech signal of the target object.
7. The method according to any one of claims 1 to 6, wherein: The step of using the second speech extraction model to generate a second speech signal of the target object based on the mixed speech signal and the target object feature vector comprises: Generate a target object mask vector based on the target object feature vector and the mixed speech signal using the second speech extraction model, wherein the target object mask vector indicates a position of the target object speech signal in the mixed speech signal; The second speech signal is generated based on the target object mask vector and the mixed speech signal.
8. The method according to any one of claims 1 to 7, wherein: The method of using the first speech extraction model to generate a first speech signal of the target object based on the mixed speech signal to be recognized and the target object feature vector includes: generating a second mixed feature vector based on the second speech signal; Generate a target object activity probability based on the target object feature vector and the second mixed feature vector, wherein the target object activity probability represents the probability that each time frame in the mixed speech signal is a speech signal of the target object; The first speech signal is generated based on the target object activity probability and the mixed speech signal.
9. The method according to any one of claims 1 to 8, wherein: The step of determining the target speech signal of the target object based on at least one of the first speech signal and the second speech signal includes: The first speech signal is determined as a target speech signal of the target object.
10. The method according to any one of claims 1 to 9, wherein: The first speech extraction model also includes at least one convolution block; the at least one transformer block also includes a linear layer, at least one residual connection and a normalization layer.
11. The method according to any one of claims 1 to 10, wherein: The first speech extraction model is a target object speech activity detection model for detecting a target object activity probability, and the second speech extraction model is a target object mask extraction model for extracting a target object mask vector; The target object activity probability represents the probability that each time frame in the mixed speech signal is the speech signal of the target object; The target object mask vector indicates a position of the target object speech signal in the mixed speech signal.
12. The method according to any one of claims 1 to 11, wherein: The mixed speech signal includes speech signal segments where multiple objects exist and speech signal segments where the target object is absent; In the mixed speech signal, the ratios of speech durations of different speech signal segments of the target object to the total speech duration of the mixed speech signal are independent of each other.
13. A speech extraction device, comprising: A feature vector extraction unit configured to extract a target object feature vector from a reference speech signal of the target object; A target speech signal generating unit is configured to generate a first speech signal of the target object based on a mixed speech signal to be recognized and a feature vector of the target object by using a first speech extraction model; wherein the first speech extraction model includes at least one transformer block having a multi-head self-attention layer and a gated recurrent unit layer; generate a second speech signal of the target object based on the mixed speech signal and the feature vector of the target object by using a second speech extraction model; wherein the second speech extraction model is used to extract a target object mask vector of the target object from the mixed speech signal; and determine a target speech signal of the target object based on at least one of the first speech signal and the second speech signal.
14. An electronic device comprising: one or more processors; as well as One or more memories, wherein the memories store computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 12. 15 . A computer-readable storage medium having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by a processor, the processor is caused to perform the method according to claim 1 .
16. A computer program product comprising computer readable instructions, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Speech separation method and module based on multi-layer attention mechanism
CN110675891A
Speaker recognition method and device, computer equipment and storage medium
CN113436633A
Real-time voice noise reduction method and device based on target person and electronic equipment
CN114898762A
Voice processing method and device, equipment and storage medium
CN115954013A
Speech processing method and related equipment thereof
CN116259311A