End-to-End Speaker Diarization by Iterative Speaker Embedding
The DIVE system addresses speaker ambiguity and real-time processing challenges by using a neural diarization approach with a time encoder, iterative speaker selector, and voice activity detector, enhancing accuracy and efficiency in multi-speaker scenarios.
Patent Information
- Application Number
- JP2023570013
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-11
- Filing Date
- 2021-06-22
- Publication Date
- 2025-07-16
- Estimated Expiration
- 2041-06-22
AI Technical Summary
Existing speaker diarization systems struggle with speaker ambiguity in the presence of overlaps and require labeled training data for accuracy, and they are often executed offline, making real-time processing difficult.
An end-to-end neural diarization system (DIVE) that includes a time encoder, iterative speaker selector, and voice activity detector to predict voice activity indicators for each speaker at each time step, using a multi-class linear classifier and fully connected neural networks for speaker embedding and voice activity detection.
The DIVE system provides accurate and real-time speaker diarization results by iteratively selecting speaker embeddings and predicting voice activity, improving accuracy and efficiency in multi-speaker environments.
Smart Images

Figure 0007709552000019 
Figure 0007709552000020 
Figure 0007709552000021
Abstract
Description
Technical Field
[0001] The present disclosure relates to end-to-end speaker diarization by iterative speaker embedding.
Background Art
[0002] Speaker diarization is the process of partitioning an input audio stream into homogeneous segments according to speaker identification. In an environment with multiple speakers, speaker diarization answers the question of "who is speaking when", and has various applications including, for example, multimedia information retrieval, speaker turn analysis, speech processing, and automatic transcription of conversational speech. For example, speaker diarization involves the task of annotating the turns of speakers in a conversation by identifying that the first segment of the input audio stream is due to a first human speaker (without specifically identifying who the first human speaker is), the second segment of the input audio stream is due to a different second human speaker (without specifically identifying who the second human speaker is), the third segment of the input audio stream is due to the first human speaker, and so on.
Summary of the Invention
Means for Solving the Problems
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations. The operations include receiving an input audio signal corresponding to utterances spoken by a plurality of speakers. The operations also include encoding the input audio signal into a sequence of T temporal embeddings. Each temporal embedding is associated with a corresponding time step and represents the audio content extracted from the input audio signal at the corresponding time step. During each of a plurality of iterations, each corresponding to a respective one of the plurality of speakers, the operations include selecting a respective speaker embedding for each speaker. For each temporal embedding in the sequence of T temporal embeddings, the operations include selecting each speaker embedding by determining the probability that the corresponding temporal embedding includes the presence of audio activity by a new speaker for which a speaker embedding was not previously selected during a previous iteration. The operations also include selecting each speaker embedding by selecting the respective speaker embedding for each speaker as the temporal embedding in the sequence of T temporal embeddings that is associated with the highest probability regarding the presence of audio activity by a new speaker. The operations also include predicting, at each time step, a respective audio activity indicator for each speaker of the plurality of speakers based on each speaker embedding selected during the plurality of iterations and the temporal embedding associated with the corresponding time step. Each audio activity indicator indicates whether the voice of the respective speaker is active or inactive at the corresponding time step.
[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, at least a portion of the utterances in the received input audio signal overlap. In some examples, when the input audio signal is received, the number of multiple speakers is unknown. The operation may further include projecting a sequence of T time embeddings encoded from the input audio signal into a downsampled embedding space while encoding the input audio signal.
[0005] In some implementations, for each of the multiple repetitions of each time embedding in the sequence of time embeddings, determining the probability that the corresponding time embedding includes the presence of voice activity by a new speaker includes determining the probability distribution of possible event types for the corresponding time embedding. Possible event types include the presence of voice activity by a new speaker, the presence of voice activity of a previous speaker where each respective speaker embedding was previously selected during a previous repetition, the presence of overlapping voices, and the presence of silence. In some implementations, determining the probability distribution of possible event types for the corresponding time embedding may include receiving, as input to a multi-class linear classifier having a fully connected network, the corresponding time embedding and previously selected speaker embeddings including the average of each respective speaker embedding previously selected during a previous repetition, and using a multi-class linear classifier having a fully connected network to map the corresponding time embedding to each of the possible event types. The multi-class linear classifier may be trained on a corpus of training audio signals each encoded with a training sequence of time embeddings. Here, each training time embedding includes a respective speaker label.
[0006] In some examples, during each iteration following the first iteration, determining the probability that the corresponding temporal embedding includes the presence of voice activity by one new speaker is based on each of the other previously selected speaker embeddings during each iteration preceding the corresponding iteration. In some implementations, the operation further includes, during each of a plurality of iterations, determining whether the probability of the corresponding temporal embedding in the sequence of temporal embeddings associated with the highest probability of the presence of voice activity by one new speaker satisfies a confidence threshold. Here, selecting each speaker embedding is conditioned on the probability of the corresponding temporal embedding in the sequence of temporal embeddings associated with the highest probability of the presence of voice activity by one new speaker that satisfies the confidence threshold. In these implementations, the operation may further include, during each of a plurality of iterations, bypassing the selection of each speaker embedding during the corresponding iteration when the probability of the corresponding temporal embedding in the sequence of temporal embeddings associated with the highest probability of the presence of voice activity by one new speaker does not satisfy the confidence threshold. Optionally, after bypassing the selection of each speaker embedding during the corresponding iteration, the operation may further include determining the number N of a plurality of speakers based on the number of speaker embeddings previously selected during the iteration preceding the corresponding iteration.
[0007] Predicting the respective voice activity indicators for each speaker of a plurality of speakers at each time step can be based on a temporal embedding associated with the corresponding time step, each speaker embedding selected for each speaker, and the average of all speaker embeddings selected over a plurality of iterations. In some examples, predicting the respective voice activity indicators for each speaker of a plurality of speakers at each time step includes using a voice activity detector having first and second fully-connected neural networks in parallel. In these examples, the first fully-connected neural network of the voice activity detector is configured to project a temporal embedding associated with the corresponding time step, and the second fully-connected neural network of the voice activity detector is configured to project a concatenation of each speaker embedding selected for each speaker and the average of all speaker embeddings selected over a plurality of iterations.
[0008] The training process can train the voice activity indicators on a corpus of training audio signals each encoded with a sequence of temporal embeddings, where each temporal embedding includes a corresponding speaker label. Optionally, the training process can include a color-aware training process that removes losses associated with any of the training temporal embeddings that fall within a radius around speaker turn boundaries.
[0009] Another aspect of the present disclosure provides a system that includes data processing hardware and memory hardware that stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving an input audio signal corresponding to utterances spoken by a plurality of speakers. The operations also include encoding the input audio signal into a sequence of T temporal embeddings. Each temporal embedding is associated with a corresponding time step and represents the audio content extracted from the input audio signal at the corresponding time step. During each of a plurality of iterations, each corresponding to a respective speaker of the plurality of speakers, the operations include selecting a respective speaker embedding for each speaker. For each temporal embedding in the sequence of T temporal embeddings, the operations include selecting a respective speaker embedding by determining the probability that the corresponding temporal embedding includes the presence of voice activity by a new speaker for which a speaker embedding has not been previously selected during a previous iteration. The operations also include selecting a respective speaker embedding by selecting the respective speaker embedding for each speaker as the temporal embedding in the sequence of T temporal embeddings that is associated with the highest probability regarding the presence of voice activity by a new speaker. The operations also include predicting, at each time step, a respective voice activity indicator for each speaker of the plurality of speakers based on each respective speaker embedding selected during the plurality of iterations and the temporal embedding associated with the corresponding time step. Each voice activity indicator indicates whether the voice of the corresponding speaker is active or inactive at the corresponding time step.
[0010] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, at least a portion of the utterances in the received input audio signal are overlapping. In some examples, when the input audio signal is received, the number of speakers is unknown. The operation may further include projecting a sequence of T time embeddings encoded from the input audio signal into a downsampled embedding space while encoding the input audio signal.
[0011] In some implementations, for each of a plurality of repetitions for each time embedding in the sequence of time embeddings, determining the probability that the corresponding time embedding includes the presence of voice activity by a new speaker includes determining a probability distribution of possible event types for the corresponding time embedding. Possible event types include the presence of voice activity by a new speaker, the presence of voice activity of a previous speaker where each other speaker embedding was previously selected during a previous repetition, the presence of overlapping voices, and the presence of silence. In some implementations, determining the probability distribution of possible event types for the corresponding time embedding includes receiving, as input to a multi-class linear classifier having a fully connected network, the corresponding time embedding and previously selected speaker embeddings including the average of each previously selected speaker embedding during a previous repetition, and using a multi-class linear classifier having a fully connected network to map the corresponding time embedding to each of the possible event types. The multi-class linear classifier may be trained on a corpus of training audio signals each encoded with a training sequence of time embeddings. Here, each training time embedding includes a respective speaker label.
[0012] In some examples, during each iteration following the first iteration, determining the probability that the corresponding temporal embedding includes the presence of voice activity by a new speaker is based on each of the other previously selected speaker embeddings during each iteration preceding the corresponding iteration. In some implementations, the operation further includes, during each of a plurality of iterations, determining whether the probability of the corresponding temporal embedding in the sequence of temporal embeddings associated with the highest probability of the presence of voice activity by a new speaker satisfies a confidence threshold. Here, selecting each speaker embedding is conditioned on the probability of the corresponding temporal embedding in the sequence of temporal embeddings associated with the highest probability of the presence of voice activity by a new speaker that satisfies the confidence threshold. In these implementations, the operation may further include, during each of a plurality of iterations, bypassing the selection of each speaker embedding during the corresponding iteration when the probability of the corresponding temporal embedding in the sequence of temporal embeddings associated with the highest probability of the presence of voice activity by a new speaker does not satisfy the confidence threshold. Optionally, after bypassing the selection of each speaker embedding during the corresponding iteration, the operation may further include determining the number N of a plurality of speakers based on the number of speaker embeddings previously selected during the iteration preceding the corresponding iteration.
[0013] Predicting each speaker's respective voice activity indicator for a plurality of speakers at each time step can be based on a temporal embedding associated with the corresponding time step, each speaker embedding selected for each speaker, and an average of all speaker embeddings selected during a plurality of iterations. In some examples, predicting each speaker's respective voice activity indicator for a plurality of speakers at each time step includes using a voice activity detector having first and second parallel fully-connected neural networks. In these examples, the first fully-connected neural network of the voice activity detector is configured to project a temporal embedding associated with the corresponding time step, and the second fully-connected neural network of the voice activity detector is configured to project a concatenation of each speaker embedding selected for each speaker and the average of all speaker embeddings selected during a plurality of iterations.
[0014] The training process can train voice activity indicators on a corpus of training audio signals each encoded with a sequence of temporal embeddings. Here, each temporal embedding includes a corresponding speaker label. Optionally, the training process can include a color-aware training process that removes losses associated with any of the training temporal embeddings that fall within a radius around a speaker turn boundary.
[0015] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the description and drawings, as well as from the claims.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0017] Like reference numerals in the various drawings indicate like elements.
[0018] An automatic speech recognition (ASR) system generally relies on speech processing algorithms that assume there is only one speaker in a given input audio signal. Input audio signals containing multiple speakers can disrupt these speech processing algorithms, thereby potentially making the speech recognition results output by the ASR system inaccurate. Therefore, speaker diarization is not the process of specifically determining who is speaking (speaker recognition / identification), but rather the process of segmenting the speech of the same speaker in a larger conversation in order to determine when someone is speaking. Put another way, speaker diarization involves a series of speaker recognition tasks based on short utterances, judging whether two segments of a given conversation were spoken by the same individual or by different individuals, and this is repeated for all segments of the conversation.
[0019] Existing speaker diarization systems generally include a plurality of relatively independent components such as, but not limited to, a voice segmentation module, an embedding extraction module, and a clustering module. The voice segmentation module is generally configured to remove non-speech parts from the input utterance and divide the input utterance into small fixed-length segments, and the embedding extraction module is configured to extract corresponding speaker identification embeddings from each fixed-length segment. The speaker identification embeddings can include i-vectors or d-vectors. The clustering module employed by existing speaker diarization systems plays a role in determining the number of speakers present in the input utterance and assigning speaker identifications (such as labels) to each fixed-length segment. These clustering modules can use common clustering algorithms including Gaussian mixture models, mean shift clustering, agglomerative hierarchical clustering, k-means clustering, link clustering, and spectral clustering. The speaker diarization system can also use an additional re-segmentation module to further refine the diarization results output from the clustering module by imposing additional constraints.
[0020] These existing speaker diarization systems are limited by the fact that the extracted speaker identification embeddings are not optimized for diarization and thus may not necessarily extract relevant features to resolve speaker ambiguity in the presence of overlaps. Furthermore, the clustering module operates in an unsupervised manner, and thus all speakers are assumed to be unknown, and the clustering algorithm needs to generate new "clusters" for new / unknown speakers for each new input utterance. The drawbacks of these unsupervised frameworks are that they cannot be improved by learning from large sets of labeled training data, including fine-grained annotations of speaker turns (i.e., speaker changes), timestamped speaker labels, and ground truths. Since this labeled training data is readily available in many domain-specific applications and diarization training data sets, speaker diarization systems may benefit from the labeled training data by being more robust and accurate in generating diarization results. Additionally, most existing state-of-the-art clustering algorithms are executed offline, which makes it difficult to generate diarization results by clustering in real-time scenarios. Speaker diarization systems also need to be executed on long audio sequences (i.e., several minutes etc.), but training the speaker diarization system over long audio sequences with large batch sizes can be difficult due to memory constraints.
[0021] To overcome the limitations of the typical dialyzation systems described above, the implementations of this specification target a Dialyzation by Iterative Voice Embedding (DIVE) system. The DIVE system includes an end-to-end neural dialyzation system that jointly trains a time encoder that projects an input audio signal into a downsampled embedding space containing a sequence of temporal embeddings each representing the current audio content at the corresponding time step, a speaker selector that performs an iterative speaker selection process to select long-term speaker vectors for all speakers in the input audio stream, and a Voice Activity Detector (VAD) that detects the voice activity of each speaker at each of a plurality of time steps.
[0022] Referring to FIG. 1, system 100 includes a user device 110 that captures an audio utterance 120 from a group of speakers (e.g., users) 10, 10a - n and communicates with a remote system 140 via a network 130. The remote system 140 may be a distributed system (e.g., a cloud computing environment) having expandable / elastic resources 142. The resources 142 include computing resources 144 (e.g., data processing hardware) and / or storage resources 146 (e.g., memory hardware). In some implementations, the user device 110 and / or the remote system 140 execute a DIVE system 200 (also referred to as an end - to - end neural diarization system 200) configured to receive input audio signals (i.e., audio data) 122 corresponding to the captured utterance 120 from a plurality of speakers 10. The DIVE system 200 encodes the input audio signal 122 into a sequence of T temporal embeddings 220, 220a - t and iteratively selects respective speaker embeddings 240, 240a - n for each speaker 10. Using the sequence of T temporal embeddings 220 and each selected speaker embedding 240, the DIVE system 200 predicts respective voice activity indicators 262 for each speaker 10 between each of a plurality of time steps. Here, the voice activity indicator 262 indicates whether the voice of each speaker is active or non - active at each time step. The respective voice activity indicators 262 predicted for each speaker 10 between each of the plurality of time steps can provide a diarization result 280 indicating when the voice of each speaker is active (or non - active) in the input audio signal 122. Each time step may correspond to each of the temporal embeddings. In some examples, each time step includes a duration of 1 millisecond.Therefore, the diarization result 280 can provide timestamped speaker labels based on the per-speaker voice activity indicators 262 predicted at each time step, which labels not only identify who is speaking at a given time but also identify when speaker changes (e.g., speaker turns) occur between adjacent time steps.
[0023] In some examples, the remote system 140 further executes an automatic speech recognition (ASR) module 150 configured to receive the audio data 122 and transcribe it into the corresponding ASR results 152. Similarly, the user device 110 can execute the ASR module 150 on the device instead of the remote system 140, which is useful when a network connection is unavailable or when quick (albeit less accurate) transcription is desired. Additionally or alternatively, both the user device 110 and the remote system 140 can execute the corresponding ASR module 150 such that the audio data 122 can be transcribed on the device, via the remote system 140, or some combination thereof. In some implementations, both the ASR module 150 and the DIVE system 200 are executed entirely on the user device 110 and do not require any network connection to the remote system 140. The ASR results 152 may also be referred to as "transcriptions" or simply "text". The ASR module 150 can communicate with the DIVE system 200 to utilize the diarization results 280 associated with the audio data 122 in order to improve speech recognition in the audio data 122. For example, the ASR module 150 can apply different speech recognition models (e.g., language models, prosody models) for different speakers identified from the diarization results 280. Additionally or alternatively, the ASR module 150 and / or the DIVE system 200 (or some other component) can use the per-speaker, per-time-step voice activity indicators 262 to index the transcription 152 of the audio data 122. For example, the transcription of a conversation among multiple colleagues (e.g., speaker 10) during a work meeting can be indexed per speaker to identify what each speaker said and to associate portions of the transcription with each speaker.
[0024] User device 110 includes data processing hardware 112 and memory hardware 114. User device 110 may include an audio capture device (e.g., a microphone) for capturing voice utterance 120 from speaker 10 and converting it into audio data 122 (e.g., an electrical signal). In some implementations, data processing hardware 112 is configured to locally execute a part of DIVE system 200 while the remaining part of the diarization system 200 is executed on remote system 140. Alternatively, data processing hardware 112 may execute DIVE system 200 instead of executing the DIVE system 200 on remote system 140. User device 110 can be any computing device that can communicate with remote system 140 via network 130. User device 110 includes, but is not limited to, desktop computing devices, as well as mobile computing devices such as laptops, tablets, smartphones, smart speakers / displays, smart home appliances, Internet of Things (IoT) devices, and wearable computing devices (e.g., headsets and / or wristwatches). User device 110 can optionally execute ASR module 150 to transcribe audio data 122 into corresponding text 152. For example, when network communication is down or unavailable, user device 110 can locally execute the diarization system 200 and / or ASR module 150 to generate a diarization result of audio data 122 and / or generate a transcription 152 of audio data 122.
[0025] In the illustrated example, the speaker 10 and the user device 110 can be arranged within an environment (e.g., a room) configured such that the user device 110 captures the voice utterance 120 spoken by the speaker 10 and converts it into an audio signal 122 (also referred to as audio data 122). For example, the speaker 10 is responding to a colleague during a meeting, and the user device 110 can record the voice utterance 120 and convert it into an audio signal 122. Next, the user device 110 can provide the audio signal 122 to the DIVE system 200 to predict a voice activity indicator 262 for each of the speakers 10 during each of a plurality of time steps. Thus, the DIVE system 200 is responsible for processing the audio signal 122 to determine when someone is speaking without specifically determining who is speaking through speaker recognition / identification.
[0026] In some examples, at least a portion of the utterance 120 conveyed by the audio signal 122 overlaps such that the voices of two or more speakers 10 are active at a given instant. It should be noted that when the input audio signal 122 is provided as an input to the DIVE system 200, the number N of the plurality of speakers 10 may be unknown, and the DIVE system 200 can predict the number N of the plurality of speakers 10. In some implementations, the user device 110 is located away from the speaker 10. For example, the user device 110 can include a remote device (e.g., a network server) that captures the voice utterance 120 from a speaker who is a participant in a phone or video conference. In this scenario, each speaker 10 will speak towards their own device (e.g., a phone, radio, computer, smartwatch, etc.) that captures the voice utterance 120 and provides it to the remote user device 110 to convert the voice utterance 120 into audio data 122. Of course, in this scenario, the utterance 120 is processed at each user device and converted into a corresponding audio signal 122 that is transmitted to the remote user device 110, and the remote user device 110 can further process the audio signal 122 provided as an input to the DIVE system 200.
[0027] In the illustrated example, the DIVE system 200 includes a time encoder 210, an iterative speaker selector 230, and a voice activity detector (VAD) 260. The time encoder 210 is configured to receive an audio signal 122 and encode the input audio signal 122 into a sequence of temporal embeddings h220, 220a~t. Each temporal embedding h220 is associated with a corresponding time step t and may represent the audio content extracted from the input audio signal 122 during the corresponding time step t. The time encoder 210 transmits the sequence of temporal embeddings 220 to the iterative speaker selector 230 and the VAD 260.
[0028] Between each of a plurality of iterations i, each corresponding to a respective one of a plurality of speakers 10, the iterative speaker selector 230 is configured to select respective speaker embeddings 240, 240a-n for each respective speaker 10. For simplicity, the audio signal 122 of FIG. 1 includes utterances 120 spoken by only two different speakers 10, but the iterative speaker selector 230 can select speaker embeddings 240 for any number N of different speakers 10 present in the input audio signal 122. Thus, in an exemplary two-speaker scenario, the iterative speaker selector 230 selects a first speaker embedding s1240 for a first speaker 10a during a first first iteration (i = 1), and during a subsequent second iteration (i = 2), the iterative speaker selector 230 selects a second speaker embedding s2240 for a second speaker 10b. During each iteration i, for each temporal embedding 220 in a sequence of T temporal embeddings 220, the iterative speaker selector 230 determines the probability that the corresponding temporal embedding 220 includes the presence of voice activity by a new speaker 10 for whom a speaker embedding 240 was not previously selected during a previous iteration, and then selects the respective speaker embedding 240. Then, during the corresponding iteration i, the iterative speaker selector 230 selects the respective speaker embedding 240 for each respective speaker 10 as the temporal embedding 220 in the sequence of T temporal embeddings 220 associated with the highest probability regarding the presence of voice activity by a new speaker 10. That is, the iterative speaker selector 230 selects the respective speaker embedding 240 having the highest probability associated with the audio content of each of the T temporal embeddings 220.
[0029] VAD260 receives the temporal embedding 220 and the speaker embedding 240 (e.g., s1 and s2 in the two-speaker scenario of FIG. 1), and at each time step, predicts a respective voice activity indicator 262 for each of the plurality of N speakers. In particular, VAD260 predicts the voice activity indicator 262 based on the temporal embedding 220 representing the audio content at each time step t, the speaker embedding 240 representing the identification of the target speaker, and another speaker embedding 240 representing all of the speakers 10. Here, each voice activity indicator 262 indicates whether the voice of each speaker 10 is active or inactive at the corresponding time step. It should be noted that VAD260 predicts the voice activity indicator 262 without specifically identifying each speaker 10 from the plurality of speakers 10. The DIVE system 200 can use the voice activity indicator 262 at each time step to provide a diarization result 280. As shown in FIG. 1, the diarization result 280 includes the voice activity indicator y i,t for speaker i at time step t. Thus, the voice activity indicator y i,t 262 of the diarization result 280 provides a VAD result for each speaker and each time step, having a value of "0" when the speaker 10 is inactive and a value of "1" when the speaker 10 is active during time step t. As shown at time step (t = 4), multiple speakers 10 may become active simultaneously.
[0030] Figure 2 shows the time encoder 210, the iterative speaker selector 230, and the VAD 260 of the DIVE system 200. The time encoder 210 encodes the input audio signal 122 into a sequence of temporal embeddings 220, 220a~t, each associated with a corresponding time step t. The time encoder 210 may project the sequence of encoded temporal embeddings 220 from the input audio signal 122 into a downsampled embedding space. The time encoder 210 can perform downsampling by cascading dilated 1D-convolutional residual blocks with parametric rectified linear unit (PReLU) activation and layer normalization, and introducing a 1D average pooling layer between the residual blocks. Thus, the input audio signal 122 may correspond to the input waveform x such that the time encoder 210 generates T temporal embeddings 220 (e.g., latent vectors), each having dimension D. Thus, each temporal embedding 220 may be represented as follows.
[0031] [Number]
[0032] For each speaker 10 detected in the input audio signal 122, the iterative speaker selector 230 outputs respective speaker embeddings 240, 240a - n. Between each of the multiple iterations i, the iterative speaker selector 230 receives the temporal embedding 220 and selects each speaker embedding 240 that was not selected in the previous iteration i (i.e., the new speaker embedding 240). In some examples, the speaker selector 230 receives, during each iteration i, the previously selected speaker embedding 240 as an input along with a sequence of T temporal embeddings 220, and outputs, for each corresponding temporal embedding 220, a confidence c that the presence of voice activity by one new speaker is included. The previously selected speaker embedding 240 may include the average of each previously selected speaker embedding 240 during the previous iteration. It should be noted that, since there is no previously selected speaker embedding 240, during the first iteration 1, the previously selected speaker embedding 240 is zero. Advantageously, the iterative process performed by the iterative speaker selector 230 does not require a specific speaker order for training to select the speaker embedding 240, and thus does not require permutation invariant training (PIT) to avoid a penalty for selecting the speaker order. PIT has a problem that when applied to a long audio sequence, the assignment becomes inconsistent, and thus it is not preferable to use for long - term speaker representation / embedding learning.
[0033] For simplicity, FIG. 2 shows an audio signal 122 that includes utterances 120 spoken by only two different speakers 10, but this is a non-limiting example, and the audio signal 122 can include utterances spoken by any number of different speakers 10. In the illustrated example, the iterative speaker selector 230 includes first speaker selectors 230, 230a that receive a sequence of T temporal embeddings 220 in a first iteration (i = 1) and select first speaker embeddings s1240, 240a. The first speaker embedding s1240a includes a first confidence level c1 that indicates the likelihood that the temporal embedding 220 includes the first speaker embedding s1240a. Here, since there is no previously selected speaker embedding 240, the first speaker selector 230a can select any speaker embedding 240. Continuing with this example, in a subsequent iteration (i = 2), the iterative speaker selector 230 receives a sequence of T temporal embeddings 220 and the previously selected first speaker embedding s1240a and includes second speaker selectors 230, 230b that select second speaker embeddings s2240, 240b. The second speaker embedding s2240b includes a second confidence level c2 that indicates the likelihood that the temporal embedding 220 includes the second speaker embedding s2240b. Here, the second speaker selector 230b can select any speaker embedding 240 other than the previously selected speaker embedding (e.g., the first speaker embedding s1240a).
[0034] The iterative speaker selector 230 can include any number of speaker selectors 230 for selecting the speaker embedding 240. In some examples, the iterative speaker selector 230 determines whether the confidence level c associated with the speaker embedding 240 for the corresponding temporal embedding 220 associated with the highest probability of the presence of voice activity by one new speaker meets a confidence level threshold. The iterative speaker selector 230 can continue to iteratively select the speaker embedding 240 until the confidence level c no longer meets the confidence level threshold.
[0035] In some implementations, the iterative speaker selector 230 includes a multi-class linear classifier having a fully connected network configured to determine, for each iteration i, the probability distribution of the possible event types e of each corresponding temporal embedding 220. The possible event types e t may include four possible types: the presence of voice activity by a single new speaker 10, the presence of voice activity by a single previous speaker 10 where each respective speaker embedding 240 was previously selected during a previous iteration, the presence of overlapping voices, and the presence of silence. Thus, the multi-class linear classifier having a fully connected network is a 4×D matrix g t representing a 4-class linear classifier that maps each temporal embedding h t 220 to one of four possible event types e μ (μ i ). Here, each temporal embedding 220 may be mapped to the event type having the highest probability in the probability distribution of possible event types during each iteration i. The probability distribution may be expressed as follows. P(e t |h t ,u i ) = softmax(g u (μ i )g h (h t )) (1) In Equation 1, e t represents the event type, h t represents each respective temporal embedding at time t, u i represents the average embedding of each previously selected speaker in iteration i, and g h represents the fully connected neural network. During inference, the confidence c for each speaker embedding 240 may be expressed as follows.
[0036]
Equation
[0037] In the equation,
[0038] [Number]
[0039] corresponds to the temporal embedding 220 associated with the highest probability regarding the presence of voice activity by a single new speaker 10. Therefore, the speaker embedding 240 selected during each iteration corresponds to the temporal embedding that reaches the maximum confidence (i.e., the highest probability) according to Equation 1. Selecting the speaker embedding 240
[0040] [Number]
[0041] may be conditional on satisfying a confidence threshold. If the confidence threshold is not satisfied, the iterative speaker selector 230 may bypass the selection during the corresponding iteration and not perform subsequent iterations. In this scenario, the DIVE system 200 can determine the number N of multiple speakers 10 based on the number of previously selected speaker embeddings 240 during the iteration preceding the iteration in which the selection of the speaker embedding 240 is bypassed. During training
[0042] [Number]
[0043] is not output by the iterative speaker selector 230. Instead, the temporal embedding h t 220 is uniformly sampled from the time when a new speaker is marked as active in the labeled training data. The iterative speaker selector 230 is trained in a supervised manner by the training process, and the parameters of the training process are learned to minimize the negative log-likelihood of a four-way linear classifier as follows.
[0044] [Number]
[0045] After speaker embedding 240 is selected (e.g., s1 and s2 in the two - speaker scenario of FIG. 2), for each time step, the VAD 260 predicts, for each speaker of the plurality of N speakers, each respective voice activity indicator 262 based on each respective speaker embedding 240, the average of all previously selected speaker embeddings 240, and the temporal embedding 220 associated with the corresponding time step. y i ∈{0,1} T (4) where i = 1, 2,...N. Each voice activity indicator (y i,t ) 262 indicates whether the voice of each respective speaker (indexed by iteration i) is active (y i,t = 1) or inactive (y i,t = 0) at the corresponding time step (indexed by time step t). Each voice activity indicator 262 may correspond to a binary speaker - by - speaker voice activity mask that provides a value of "0" when each respective speaker is inactive during time step t and a value of "1" when each respective speaker is active. The predicted voice activity indicator (y i,t ) 262 for each speaker i at each time step t is based on the temporal embedding h t associated with the corresponding time step, each respective speaker embedding s i selected for each respective speaker, and the average of all speaker embeddings selected during the plurality of iterations
[0046]
Number
[0047] and may be based on.
[0048] In some implementations, VAD 260 includes two parallel fully-connected neural networks f with PReLU activation by layer normalization, excluding the last linear projection layer that includes linear projection. h and f s In these implementations, to predict the voice activity indicator y i,t of speaker i at time step t, f h and f s project the corresponding temporal embedding
[0049]
Number
[0050] and speaker embedding
[0051]
Number
[0052] as follows.
[0053]
Number
[0054] In Equation 5,
[0055]
Number
[0056] are the respective speaker embeddings s i and
[0057]
Number
[0058] represents the concatenation along the channel axis of the average value of all speaker embeddings 240. It should be noted that the average speaker embedding calls the VAD 260 to utilize the contrast between each speaker embedding 240 associated with the target speaker i and all other speakers present within the sequence of temporal embeddings 220. In the illustrated example, the VAD 260 predicts the voice activity indicators 262 of the first and second speakers 10 at time step (t = 2). The voice activity indicator 262 can provide a diarization result 280 indicating that the first speaker 10 is active at time step (t = 2) (e.g., y 1,2 = 1), and the second speaker 10 is inactive at time step (t = 2) (e.g., y 2,2 = 0).
[0059] Referring to FIG. 3, schematic diagram 300 shows an exemplary training process 301 and inference 304 of DIVE system 200. In some implementations, the training process 301 jointly trains the time encoder 210, the iterative speaker selector 230 including a multi-class linear classifier with a fully connected network, and the VAD 260 on the fully labeled training data 302 including a corpus of training audio signals x* each containing an utterance 120 spoken by a plurality of different speakers 10. The training audio signal x* may include a long audio sequence representing several minutes of speech. In some examples, the training process 301 samples W fixed-length windows for each training audio signal x*, encodes the training audio signal x* using the time encoder 210, and concatenates the W fixed-length windows along the time axis. By concatenating the W fixed-length windows, the fully labeled training data 302 increases the speaker diversity and speaker turns of each training audio signal x* while keeping the memory usage low. That is, the training audio signal x* may represent the same speaker on distant windows during a long audio sequence. Some of the training audio signals x* may include portions where utterances 120 spoken by two or more different speakers 10 overlap. Each training audio signal x* is encoded by the time encoder 210 into a sequence of training temporal embeddings 220 each assigned a respective speaker label 350 indicating an active speaker or silence.
[0060] The speaker labels 350 can be represented as a sequence of training speaker labels
[0061]
Number
[0062] and can be expressed as, where the entry in the sequence
[0063]
Number
[0064] represents the speaker label 350 assigned to the training time embedding 220 at time step t. In the illustrated example, the training process 301 provides, during each of a plurality of i iterations, a sequence of training time embeddings 220T encoded by the time encoder 210, and the assigned speaker labels 350 for training the iterative speaker selector 230, and then provides the VAD 260 based on the speaker embeddings 240 selected by the iterative speaker selector 230 during multiple iterations.
[0065] When the time encoder 210, the iterative speaker selector 230, and the VAD 260 are jointly trained, the VAD 260 is also trained on a corpus of training audio signals x*, where each training audio signal x* has a corresponding voice activity indicator (i.e., speaker label) indicating which voice is present / active in the corresponding training time embedding.
[0066]
Number
[0067] is encoded into a sequence of training time embeddings each containing one. The training process can train the VAD 260 for the following VAD loss.
[0068]
Number
[0069] Here, the training process backpropagates the speaker-by-speaker, time-step-by-time-step VAD loss of Equation 6 as an independent binary classification task. The DIVE system 200 can be evaluated from the perspective of the diarization error rate (DER) of the diarization result 280. In some examples, the training process applies a color that provides a tolerance near the speaker boundaries so that the training VAD loss of Equation 6 does not penalize the VAD 260 for small annotation errors in the training data. In some examples, a typical value of the tolerance representing the color is about 250 ms on both sides of the speaker turn boundary (500 ms) specified in the labeled training data. Thus, the training process can calculate the masked VAD loss by removing from the total loss the VAD loss associated with frames / time steps that fall within the color as follows.
[0070] [Number]
[0071] where B r includes the set of audio frames / time steps within a radius r around the speaker turn boundary. The training process can backpropagate the masked VAD loss calculated by Equation 7. During training, the total loss of the DIVE system 200 is calculated as follows to jointly train the time encoder 210, the iterative speaker selector 230, and the VAD 260.
[0072] [Number]
[0073] The total loss can be calculated similarly without applying the color loss by substituting the VAD loss of Equation 7.
[0074] The iterative speaker selector 230 is trained based on the speaker selector loss represented by Equation 3, and the VAD 260 can be trained based on the VAD loss represented by Equation 6 or Equation 7 when the training color is applied. That is, the training process 301 may include a color-aware training process that removes the loss associated with any of the training temporal embeddings 220 that fall within the radius around the speaker turn boundary. The color-aware training process does not penalize or train the DIVE system 200 for small annotation errors. For example, the radius around the speaker turn boundary may include 250 ms (a total of 500 ms) on both sides of the speaker turn boundary. Thus, the DIVE system 200 can be trained with the total loss calculated by Equation 8.
[0075] Separate components 210, 230, 260 of the DIVE system 200 may each include a neural network such that the training process 201 generates the weights of the connections between the hidden nodes, the hidden nodes corresponding to the fully labeled training data 302, and the input nodes, the weights of the connections between the hidden nodes and the output nodes, and the weights of the connections between the layers of the hidden nodes themselves so as to minimize the losses of Equations 3, 6, 7, and 8. Thereafter, during inference 304, the fully trained DIVE system 200 can be employed on input data (e.g., raw audio signal 122) to generate unknown output data corresponding to the diarization result 280 (e.g., voice activity indicator 262).
[0076] Figure 4 shows a plot 400 of the raw diarization error rate (DER) (%) evaluations for the standard training process and the color-aware training process used to train the DIVE system 200. In plot 400, the standard training process outperforms the color-aware training process when using the raw DER (%) evaluations. Figure 5 shows a plot 500 of the color-aware DER evaluations with 250 ms of color applied on both sides of the speaker turn boundary for the standard training process and the color-aware training process. Here, 250 ms is applied on both sides of the speaker turn boundary according to Equation 7. Of note is that the color-aware training process outperforms the standard training process when evaluated using the color-aware DER evaluations. Thus, Figure 5 shows that it is beneficial to integrate color-aware training for training the DIVE system 200 when the evaluation technique includes the color-aware DER evaluation.
[0077] Figure 6 is a flowchart of an exemplary operational arrangement of a method 600 for performing speaker diarization on a received utterance 120. Data processing hardware 112, 144 may perform operations for method 600 by executing instructions stored on memory hardware 114, 146. In operation 602, method 600 includes receiving an input audio signal 122 corresponding to an utterance 120 spoken by a plurality of speakers 10, 10a-n. In operation 604, method 600 includes encoding the input audio signal 122 into a sequence of T temporal embeddings 220, 220a-t. Here, each temporal embedding 220 is associated with a corresponding time step t and represents the audio content extracted from the input audio signal 122 at the corresponding time step t.
[0078] Between each of a plurality of iterations i, each corresponding to a respective one of a plurality of speakers 10, the method 600 includes, at operation 606, selecting respective speaker embeddings 240, 240a - n for each respective speaker 10. For each temporal embedding 220 within a sequence of T temporal embeddings 220, the method 600 includes, at operation 608, determining a probability (e.g., confidence c) that the corresponding temporal embedding 220 includes the presence of voice activity by one new speaker 10 for which a speaker embedding 240 was not previously selected during a previous iteration i. At operation 610, the method 600 includes selecting, as the temporal embedding (220) in the sequence of T temporal embeddings 220 associated with the highest probability for the presence of voice activity by a single new speaker 10, the respective speaker embedding 240 for each respective speaker 10. Operation 612 includes the method 600 predicting, for each respective speaker 10 of the plurality of speakers 10, a respective voice activity indicator 262 based on each respective speaker embedding 240 selected between the plurality of iterations i and the temporal embedding 220 associated with the corresponding time step t. Here, each voice activity indicator 262 indicates whether the voice of each respective speaker 10 is active or inactive at the corresponding time step t.
[0079] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application", an "app", or a "program". Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, document processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0080] A non-transitory memory can be a physical device used to temporarily or permanently store a program (e.g., a sequence of instructions) or data (e.g., program state information) for use by a computing device. The non-transitory memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0081] FIG. 7 is a schematic diagram of an exemplary computing device 700 that can be used to implement the systems and methods described herein. Computing device 700 is intended to represent various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only, and are not intended to limit the implementations of the invention described and / or claimed herein.
[0082] The computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and the storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 may be interconnected using various buses and may be mounted on a common motherboard or in other manners as appropriate. The processor 710, i.e., the data processing hardware 112, 144 of FIG. 1, can process instructions for execution within the computing device 700, including instructions stored in the memory 720, i.e., the memory hardware 114, 146 of FIG. 1, or in the storage device 730, i.e., the memory hardware 144, 146 of FIG. 1, for displaying graphical information about a graphical user interface (GUI) on an external input / output device such as a display 780 coupled to the high-speed interface 740. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as necessary. Also, multiple computing devices 700 may be connected, and each device may provide a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0083] Memory 720 stores information non-temporarily within computing device 700. Memory 720 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-temporary memory 720 may be a physical device used to store a program (e.g., a sequence of instructions) or data (e.g., program state information) temporarily or persistently for use by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0084] Storage device 730 is capable of providing mass storage for computing device 700. In some implementations, storage device 730 is a computer-readable medium. In various different implementations, storage device 730 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including devices in a storage area network or other configuration. In additional implementations, a computer program product is tangibly implemented in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer-readable medium or a machine-readable medium such as memory 720, storage device 730, or memory on processor 710.
[0085] The high-speed controller 740 manages bandwidth-intensive operations for the computing device 700, while the low-speed controller 760 manages low-bandwidth-intensive operations. Such an assignment of duties is merely exemplary. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., through a graphics processor or accelerator), and the high-speed expansion port 750 that may receive various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and the low-speed expansion port 790. The low-speed expansion port 790 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), but may be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, etc., or a network device such as a switch or router, e.g., through a network adapter.
[0086] As shown in the figure, the computing device 700 can be implemented in several different forms. For example, the computing device 700 can be implemented as a standard server 700a, or multiple times as a group of such servers 700a, or as a laptop computer 700b, or as part of a rack server system 700c.
[0087] The various implementations of the systems and techniques described in this specification can be realized in digital electronics and / or optical circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which programmable processor is coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device, and can be either special purpose or general purpose.
[0088] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. The terms "machine-readable medium" and "computer-readable medium" as used herein refer to any computer program product, non-transitory computer-readable medium, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0089] The processes and logical flows described in this specification can be executed by one or more programmable processors, also called data processing hardware, that execute one or more computer programs to operate on input data and generate output. The processes and logical flows can also be executed by dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or receive data from, transfer data to, or both, mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0090] To enable interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, and optionally a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user may be received in any form, including acoustic, voice, or tactile input. Further, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user, such as by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0091] Some implementations have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.
Explanation of Signs
[0092] 100 System 110 User Device 112 Data Processing Hardware 114 Memory Hardware 120 Utterance 122 Input Audio Signal 130 Fully Connected Neural Network 140 Remote System 142 Expandable / Elastic Resource 144 Computing Resource 146 Storage Resource 150 Automatic Speech Recognition (ASR) Module 152 ASR Result 152 Transcription 200 DIVE System 210 Time Encoder 220 Temporal Embedding 230 Iterative Speaker Selector 240 Speaker Embedding 260 Voice Activity Detector 262 Voice Activity Indicator 280 Dialyzation Result 301 Training Process 302 Training Data 304 Inference 350 Speaker Label 600 Computer-Implemented Method 700 Computing Device 710 Data Processing Hardware 720 Memory Hardware 730 Storage Device 740 High-Speed Interface 750 High-Speed Expansion Port 760 Low-Speed Controller 780 Display 790 Low-Speed Expansion Port
Claims
1. A computer-implemented method (600) which, when executed on data processing hardware (710), causes the data processing hardware (710) to receive an input audio signal (122) corresponding to an utterance (120) spoken by a plurality of speakers (10); encode the input audio signal (122) into a sequence of T temporal embeddings (220), each temporal embedding (220) being associated with a corresponding time step and representing audio content extracted from the input audio signal (122) at the corresponding time step; for each of a plurality of repetitions, each corresponding to a respective one of the plurality of speakers (10), for each temporal embedding (220) in the sequence of T temporal embeddings (220), determine the probability that the corresponding temporal embedding (220) includes the presence of voice activity by a new speaker for whom a speaker embedding (240) has not been previously selected during a previous repetition; select, for each respective speaker, the speaker embedding (240) as the temporal embedding (220) in the sequence of T temporal embeddings (220) that is associated with the highest probability regarding the presence of voice activity by the new speaker; thereby selecting, for each respective speaker, the respective speaker embedding (240); at each time step, predict, for each respective speaker of the plurality of speakers (10), a respective voice activity indicator (262) based on the respective speaker embedding (240) selected during the plurality of repetitions and the temporal embedding (220) associated with the corresponding time step, each respective voice activity indicator (262) indicating whether the voice of the respective speaker is active or inactive at the corresponding time step; A computer-implemented method (600) that performs operations including the above.
2. The computer-implemented method (600) according to claim 1, wherein at least a part of the utterance (120) in the received input audio signal (122) overlaps.
3. The computer-implemented method (600) according to claim 1 or 2, wherein the number of the plurality of speakers (10) is unknown when the input audio signal (122) is received.
4. The computer-implemented method (600) according to any one of claims 1 to 3, wherein the operation further includes projecting the sequence of T time embeddings (220) encoded from the input audio signal (122) into a downsampled embedding space while encoding the input audio signal (122).
5. For each of the plurality of repetitions for each time embedding (220) in the sequence of time embeddings (220), determining the probability that the corresponding time embedding (220) includes the presence of voice activity by the one new speaker, including determining a probability distribution of possible event types for the corresponding time embedding (220), wherein the possible event types are the presence of voice activity by the one new speaker, the presence of voice activity of one previous speaker, where a respective speaker embedding (240) was previously selected during a previous iteration, the presence of overlapping voices, and the presence of silence The computer-implemented method (600) according to any one of claims 1 to 4.
6. Determining the probability distribution of possible event types for the corresponding time embedding (220) includes receiving, as an input to a multi-class linear classifier having a fully connected network, the corresponding time embedding (220) and the previously selected speaker embeddings (240) including the average of the respective speaker embeddings (240) previously selected during a previous iteration, using the multi-class linear classifier having a fully connected network to map the corresponding time embedding (220) to each of the possible event types The computer-implemented method (600) according to claim 5.
7. The computer-implemented method (600) according to claim 6, wherein the multi-class linear classifier is trained on a corpus of training audio signals (122), each training audio signal (122) being encoded into a sequence of training temporal embeddings (220), each training temporal embedding (220) including a respective speaker label (350).
8. The computer-implemented method (600) according to any one of claims 1 to 7, wherein during each iteration subsequent to the first iteration, determining the probability that the corresponding temporal embedding (220) includes the presence of voice activity by the one new speaker is based on each of the other respective speaker embeddings (240) previously selected during each iteration preceding the corresponding iteration.
9. During each of the plurality of iterations, the operation further includes determining whether the probability of the corresponding temporal embedding (220) in the sequence of temporal embeddings (220) associated with the highest probability of the presence of voice activity by the one new speaker meets a confidence threshold, wherein selecting each of the speaker embeddings (240) is conditioned on the probability of the corresponding temporal embedding (220) in the sequence of temporal embeddings (220) associated with the highest probability of the presence of voice activity by the one new speaker that meets the confidence threshold. The computer-implemented method (600) according to any one of claims 1 to 8.
10. The computer-implemented method (600) according to claim 9, further including bypassing the selection of each of the speaker embeddings (240) during the corresponding iteration during each of the plurality of iterations when the probability of the corresponding temporal embedding (220) in the sequence of temporal embeddings (220) associated with the highest probability of the presence of voice activity by the one new speaker does not meet the confidence threshold.
11. The computer-implemented method (600) of claim 10, wherein the operation further comprises determining the number N of the plurality of speakers (10) based on the number of speaker embeddings (240) previously selected during the iteration prior to the corresponding iteration, after bypassing the selection of the respective speaker embeddings (240) during the corresponding iteration.
12. Predicting the respective voice activity indicators (262) for each of the plurality of speakers (10) at each time step is based on the temporal embedding (220) associated with the corresponding time step, the respective speaker embedding (240) selected for each speaker, and the average of all the speaker embeddings (240) selected during the plurality of iterations, according to the computer-implemented method (600) of any one of claims 1 to 11.
13. Predicting the respective voice activity indicators (262) for each of the plurality of speakers (10) at each time step includes using a voice activity detector (260) having parallel first and second fully connected neural networks, wherein the first fully connected neural network (130) of the voice activity detector (260) is configured to project the temporal embedding (220) associated with the corresponding time step, and the second fully connected neural network (130) of the voice activity detector (260) is configured to project the concatenation of the respective speaker embedding (240) selected for each speaker and the average of all the speaker embeddings (240) selected during the plurality of iterations, according to the computer-implemented method (600) of any one of claims 1 to 12. The computer-implemented method (600) of any one of claims 1 to 12.
14. In a training process (301), the voice activity indicator (262) is trained on a corpus of training audio signals (122), each training audio signal (122) being encoded with a sequence of training temporal embeddings (220), each training temporal embedding (220) including a corresponding speaker label (350), according to the computer-implemented method (600) of any one of claims 1 to 13.
15. The computer-implemented method (600) according to claim 14, wherein the training process (301) includes a color-aware training process (301) that removes losses associated with any of the training temporal embeddings (220) that fall within a radius around a speaker turn boundary.
16. A system (100) comprising: data processing hardware (710); and memory hardware (720) in communication with the data processing hardware (710), wherein when executed by the data processing hardware (710), the data processing hardware (710) stores instructions that cause the data processing hardware (710) to perform operations, and the operations are receiving an input audio signal (122) corresponding to an utterance (120) spoken by a plurality of speakers (10); encoding the input audio signal (122) into a sequence of T temporal embeddings (220), each temporal embedding (220) being associated with a corresponding time step and representing audio content extracted from the input audio signal (122) at the corresponding time step; during each of a plurality of iterations, each corresponding to a respective one of the plurality of speakers (10), for each temporal embedding (220) in the sequence of T temporal embeddings (220), determining a probability that the corresponding temporal embedding (220) includes the presence of voice activity by a single new speaker for which a speaker embedding (240) was not previously selected during a previous iteration; selecting, for each speaker embedding (240) for each respective speaker, the temporal embedding (220) in the sequence of T temporal embeddings (220) that is associated with the highest probability regarding the presence of voice activity by the single new speaker; thereby selecting each speaker embedding (240) for each respective speaker. At each time step, for each speaker of the plurality of speakers (10), predicting respective voice activity indicators (262) based on the respective speaker embeddings (240) selected during the plurality of iterations and the temporal embeddings (220) associated with the corresponding time step, wherein the respective voice activity indicators (262) indicate whether the voice of the respective speaker is active or inactive at the corresponding time step, and predicting A system (100) comprising. **Claim 17** The system (100) according to claim 16, wherein at least a part of the utterance (120) in the received input audio signal (122) overlaps. **Claim 18** The system (100) according to claim 16 or 17, wherein the number of the plurality of speakers (10) is unknown when the input audio signal (122) is received. **Claim 19** The operation further includes projecting the sequence of T temporal embeddings (220) encoded from the input audio signal (122) into a downsampled embedding space while encoding the input audio signal (122), the system (100) according to any one of claims 16 to 18. **Claim 20** During each of the plurality of iterations for each temporal embedding (220) in the sequence of temporal embeddings (220), determining the probability that the corresponding temporal embedding (220) includes the presence of voice activity by the one new speaker, including determining a probability distribution of possible event types for the corresponding temporal embedding (220), wherein the possible event types are The presence of voice activity by the one new speaker, The presence of voice activity of one previous speaker, where another respective speaker embedding (240) was previously selected during a previous iteration, The presence of overlapping voices, and The presence of silence The system (100) according to any one of claims 16 to 19. **Claim 21** Determining the probability distribution of possible event types for the corresponding temporal embedding (220) is As an input to a multi-class linear classifier having a fully connected network (130), receiving the corresponding temporal embedding (220) and the previously selected speaker embeddings (240) including the average of each previously selected speaker embedding (240) during a previous iteration Using the multi-class linear classifier having the fully connected network (130) to map the corresponding temporal embedding (220) to each of the possible event types The system (100) according to claim 20, comprising **Claim 22** The system (100) according to claim 21, wherein the multi-class linear classifier is trained on a corpus of training audio signals (122), each training audio signal (122) is encoded into a sequence of training temporal embeddings (220), and each training temporal embedding (220) includes a respective speaker label (350). **Claim 23** During each iteration following the first iteration, determining the probability that the corresponding temporal embedding (220) includes the presence of voice activity by the one new speaker, based on each of the other previously selected speaker embeddings (240) during each iteration preceding the corresponding iteration, the system (100) according to any one of claims 16 to 22. **Claim 24** The operation is, during each of the plurality of iterations, further comprising determining whether the probability of the corresponding temporal embedding (220) in the sequence of temporal embeddings (220) associated with the highest probability regarding the presence of voice activity by the one new speaker meets a confidence threshold, wherein selecting each of the speaker embeddings (240) is conditioned on the probability of the corresponding temporal embedding (220) in the sequence of temporal embeddings (220) associated with the highest probability regarding the presence of voice activity by the one new speaker that meets the confidence threshold, The system (100) according to any one of claims 16 to 23. **Claim 25** When the probability of the corresponding temporal embedding (220) in the sequence of temporal embeddings (220) associated with the highest probability regarding the presence of voice activity by the new speaker of the one person does not satisfy the confidence threshold during each of the plurality of iterations, the system (100) according to claim 24 further includes bypassing the selection of the respective speaker embedding (240) during the corresponding iteration.
26. The system (100) according to claim 25 further includes determining the number N of the plurality of speakers (10) based on the number of previously selected speaker embeddings (240) during the iteration prior to the corresponding iteration after bypassing the selection of the respective speaker embedding (240) during the corresponding iteration.
27. Predicting the respective voice activity indicators (262) for each of the plurality of speakers (10) at each time step is based on the temporal embedding (220) associated with the corresponding time step, the respective speaker embedding (240) selected for each of the respective speakers, and the average of all the speaker embeddings (240) selected during the plurality of iterations. The system (100) according to any one of claims 16 to 26.
28. Predicting the respective voice activity indicators (262) for each of the plurality of speakers (10) at each time step includes using a voice activity detector (260) having first and second fully connected neural networks in parallel. The first fully connected neural network (130) of the voice activity detector (260) is configured to project the temporal embedding (220) associated with the corresponding time step. The second fully connected neural network (130) of the voice activity detector (260) is configured to project the concatenation of the respective speaker embedding (240) selected for each of the respective speakers and the average of all the speaker embeddings (240) selected during the plurality of iterations. The system (100) according to any one of claims 16 to 27.
29. In the training process (301), the voice activity indicator (262) is trained on a corpus of training audio signals (122), each training audio signal (122) being encoded in a sequence of training temporal embeddings (220), each training temporal embedding (220) including a corresponding speaker label (350), the system (100) according to any one of claims 16 to 28.
30. The system (100) according to claim 29, comprising a color-aware training process (301) in which the training process (301) removes losses associated with any of the training temporal embeddings (220) that fall within a radius around speaker turn boundaries.
Citation Information
Patent Citations
Multi-speaker diarization of speech input using neural networks
JP2022541380A