Accelerating speaker classification using multi-stage clustering
Through the multi-stage clustering method, the embedding of the speaker fragments is first preclustered, and then the centroid values are spectral clustered, which solves the problem of high computational complexity in the prior art and improves the performance of speaker classification.
Patent Information
- Application Number
- CN202280101406.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-05
- Publication Date
- 2025-06-06
AI Technical Summary
When the existing speaker classification system processes audio signals of multiple speakers, the calculation complexity is high, making it difficult to effectively accelerate the speaker classification performance.
By using the multi-stage clustering method, preclustering is first performed on speaker discriminative embeddings extracted from the speaker segment, clustering as a target number of precluster clusters, and then spectral clustering is performed on the centroid values of the precluster clusters, reducing the computational cost.
It effectively reduces the computational complexity of spectral clustering and improves the performance of speaker classification system, especially when processing long-term audio.
Smart Images

Figure CN120112993A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to accelerating speaker classification using multi-stage clustering. Background Art
[0002] Speaker classification is the process of segmenting an input audio stream into homogeneous segments based on speaker identity. In an environment with multiple speakers, speaker classification answers the question "who is speaking when" and has a variety of applications, including multimedia information retrieval, speaker turn analysis, audio processing, and automatic transcription of conversational speech, among others. For example, speaker classification involves the task of annotating speaker turns in a conversation by identifying that a first segment of an input audio stream is attributable to a first human speaker (without specifically identifying who the first human speaker is), a second segment of the input audio stream is attributable to a different second human speaker (without specifically identifying who the second human speaker is), a third segment of the input audio stream is attributable to the first human speaker, and so on. Summary of the invention
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for accelerating speaker classification. The operations include: receiving an input audio signal corresponding to an utterance spoken by one or more speakers; and processing the input audio signal using a speech recognition model to jointly generate the following as output from the speech recognition model: a transcription of the utterance; and one or more speaker turn markers, each speaker turn marker indicating the position of a corresponding speaker turn detected between a pair of corresponding adjacent terms in the transcription. The input audio signal includes N audio frames of fixed length. The operations also include: segmenting the input audio signal into a plurality of N speaker segments based on the one or more speaker turn markers generated as output from the speech recognition model; and for each speaker segment in the plurality of N speaker segments, extracting a corresponding speaker discriminative embedding from the speaker segment. Based on determining that the number of N speaker segments is greater than a threshold number M, the operations also include: performing pre-clustering on speaker discriminative embeddings extracted from the N speaker segments to cluster the N speaker segments into a target number of pre-clustering clusters; for each corresponding pre-clustering cluster in the target number of pre-clustering clusters, determining corresponding centroid values based on speaker discriminative embeddings extracted from speaker segments clustered into the corresponding pre-clustering cluster; performing spectral clustering on the centroid values determined for the target number of pre-clustering clusters to cluster the centroid values into k classes; and for each corresponding class in the k classes, assigning a corresponding speaker label to each centroid value clustered into the corresponding class, which speaker label is different from the corresponding speaker label assigned to the centroid value in each other class clustered into the k classes.
[0004] Implementations of this aspect include one or more of the following optional features. In some implementations, the operations further include annotating a transcription of the utterance based on the speaker label assigned to each centroid value. In additional implementations, the operations further include setting a target number of pre-clustering clusters equal to a threshold number M. The target number of pre-clustering clusters may be less than the number of N speaker segments.
[0005] In some examples, the operations also include: for each of the one or more speaker turn markers generated as output from the speech recognition model, predicting a corresponding confidence value for a corresponding speaker turn detected in the transcription; and determining a threshold number of the one or more speaker markers that each have a corresponding confidence value that satisfies a confidence value threshold. Here, segmenting the input audio signal into a plurality of N speaker segments is based on determining a threshold number of the one or more speaker markers that each have a corresponding confidence value that satisfies the confidence value threshold. In these examples, the operations may also include determining a pairwise constraint based on the confidence values predicted for the speaker turn markers, wherein the spectral clustering performed on the centroid values determined for the target number of pre-clustered clusters is constrained by the pairwise constraint.
[0006] In some implementations, each speaker turn marker in the speaker turn marker sequence has a corresponding timestamp, and segmenting the input audio signal into a plurality of N speaker segments based on the speaker turn marker sequence includes segmenting the input audio signal into initial speaker segments, each initial speaker segment being delimited by corresponding timestamps of a pair of corresponding adjacent speaker turn markers in the speaker turn marker sequence. In these implementations, the operations may also include, for each initial speaker segment having a corresponding duration exceeding a segment duration threshold, further segmenting the initial speaker segment into two or more shortened speaker segments having corresponding durations less than or equal to the segment duration threshold, wherein the plurality of N speaker segments segmented from the input audio signal include: initial speaker segments having corresponding durations less than or equal to the segment duration threshold; and shortened speaker segments further segmented from any of the initial speaker segments having corresponding durations exceeding the segment duration threshold.
[0007] Extracting corresponding speaker discriminative embeddings from speaker segments may include: receiving the speaker segments as input to a speaker encoder model; and generating corresponding speaker discriminative embeddings as output from the speaker encoder model. The speaker encoder model may include a long short-term memory (LSTM-based) speaker encoder model configured to extract corresponding speaker discriminative embeddings from each speaker segment.
[0008] In some implementations, the speech recognition model includes a streaming transducer-based speech recognition model, which includes an audio encoder, a label encoder, and a joint network. The audio encoder is configured to: receive a sequence of acoustic frames as input; and at each of a plurality of time steps, generate a high-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The label encoder is configured to: receive a sequence of non-empty symbols output by the last softmax layer as input; and generate a dense representation at each of the plurality of time steps. The joint network is configured to: receive a high-order feature representation generated by the audio encoder at each of the plurality of time steps and a dense representation generated by the label encoder at each of the plurality of time steps as input; and at each of the plurality of time steps, generate a probability distribution of possible speech recognition hypotheses at the corresponding time step. In these implementations, the audio encoder may include a neural network with a plurality of multi-head attention layers and / or the label encoder may include a bigram embedding lookup decoder model.
[0009] The speech recognition model may be trained on training samples, each training sample comprising a pairing of training utterances spoken by two or more different speakers and corresponding ground-truth transcriptions of the training utterances, each ground-truth transcription being injected with ground-truth speaker turn markers indicating locations in the ground-truth transcription where speaker turns occur. The corresponding ground-truth transcription of each training sample may not be annotated with any timestamp information.
[0010] Another aspect of the present disclosure includes a system comprising: data processing hardware; and memory hardware that communicates with the data processing hardware and stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations. The operations include: receiving an input audio signal corresponding to an utterance spoken by one or more speakers; and processing the input audio signal using a speech recognition model to jointly generate the following as output from the speech recognition model: a transcription of the utterance; and one or more speaker turn markers, each speaker turn marker indicating the position of a corresponding speaker turn detected in the transcription between a pair of corresponding adjacent terms. The input audio signal includes N fixed-length audio frames. The operations also include: segmenting the input audio signal into a plurality of N speaker segments based on the one or more speaker turn markers generated as output from the speech recognition model; and for each speaker segment in the plurality of N speaker segments, extracting a corresponding speaker discriminative embedding from the speaker segment. Based on determining that the number of N speaker segments is greater than a threshold number M, the operations also include: performing pre-clustering on speaker discriminative embeddings extracted from the N speaker segments to cluster the N speaker segments into a target number of pre-clustering clusters; for each corresponding pre-clustering cluster in the target number of pre-clustering clusters, determining corresponding centroid values based on speaker discriminative embeddings extracted from speaker segments clustered into the corresponding pre-clustering cluster; performing spectral clustering on the centroid values determined for the target number of pre-clustering clusters to cluster the centroid values into k classes; and for each corresponding class in the k classes, assigning a corresponding speaker label to each centroid value clustered into the corresponding class, which speaker label is different from the corresponding speaker label assigned to the centroid value in each other class clustered into the k classes.
[0011] This aspect may include one or more of the following optional features. In some implementations, the operations further include annotating a transcription of the utterance based on the speaker label assigned to each centroid value. In additional implementations, the operations further include setting a target number of pre-clustering clusters equal to a threshold number M. The target number of pre-clustering clusters may be less than the number of N speaker segments.
[0012] In some examples, the operations also include: for each of the one or more speaker turn markers generated as output from the speech recognition model, predicting a corresponding confidence value for a corresponding speaker turn detected in the transcription; and determining a threshold number of the one or more speaker markers that each have a corresponding confidence value that satisfies a confidence value threshold. Here, segmenting the input audio signal into a plurality of N speaker segments is based on determining a threshold number of the one or more speaker markers that each have a corresponding confidence value that satisfies the confidence value threshold. In these examples, the operations may also include determining a pairwise constraint based on the confidence values predicted for the speaker turn markers, wherein the spectral clustering performed on the centroid values determined for the target number of pre-clustered clusters is constrained by the pairwise constraint.
[0013] In some implementations, each speaker turn marker in the speaker turn marker sequence has a corresponding timestamp, and segmenting the input audio signal into a plurality of N speaker segments based on the speaker turn marker sequence includes segmenting the input audio signal into initial speaker segments, each initial speaker segment being delimited by corresponding timestamps of a pair of corresponding adjacent speaker turn markers in the speaker turn marker sequence. In these implementations, the operations may also include, for each initial speaker segment having a corresponding duration exceeding a segment duration threshold, further segmenting the initial speaker segment into two or more shortened speaker segments having corresponding durations less than or equal to the segment duration threshold, wherein the plurality of N speaker segments segmented from the input audio signal include: initial speaker segments having corresponding durations less than or equal to the segment duration threshold; and shortened speaker segments further segmented from any of the initial speaker segments having corresponding durations exceeding the segment duration threshold.
[0014] Extracting corresponding speaker discriminative embeddings from speaker segments may include: receiving the speaker segments as input to a speaker encoder model; and generating corresponding speaker discriminative embeddings as output from the speaker encoder model. The speaker encoder model may include a long short-term memory (LSTM-based) speaker encoder model configured to extract corresponding speaker discriminative embeddings from each speaker segment.
[0015] In some implementations, the speech recognition model includes a streaming transducer-based speech recognition model, which includes an audio encoder, a label encoder, and a joint network. The audio encoder is configured to: receive a sequence of acoustic frames as input; and at each of a plurality of time steps, generate a high-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The label encoder is configured to: receive a sequence of non-empty symbols output by the last softmax layer as input; and generate a dense representation at each of the plurality of time steps. The joint network is configured to: receive a high-order feature representation generated by the audio encoder at each of the plurality of time steps and a dense representation generated by the label encoder at each of the plurality of time steps as input; and at each of the plurality of time steps, generate a probability distribution of possible speech recognition hypotheses at the corresponding time step. In these implementations, the audio encoder may include a neural network with a plurality of multi-head attention layers and / or the label encoder may include a bigram embedding lookup decoder model.
[0016] The speech recognition model may be trained on training samples, each training sample comprising a pairing of training utterances spoken by two or more different speakers and corresponding ground-truth transcriptions of the training utterances, each ground-truth transcription being injected with ground-truth speaker turn markers indicating locations in the ground-truth transcription where speaker turns occur. The corresponding ground-truth transcription of each training sample may not be annotated with any timestamp information.
[0017] The details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of an example speaker diarization system for performing speaker diarization.
[0019] Figure 2 is a schematic diagram of an example transcription output from a speech recognition model, the transcription including speaker turn markers indicating the locations of predicted speaker turns in the transcription.
[0020] Figure 3 is a diagram of an example automatic speech recognition model with a transducer-based architecture.
[0021] Figure 4 yes Figure 1 Schematic diagram of an example cluster selector for a speaker diarization system.
[0022] Figure 5is a flow diagram of an example arrangement of operations of a computer-implemented method of performing speaker classification on an input audio signal containing utterances of speech spoken by multiple different speakers.
[0023] Figure 6 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0024] Like reference numerals in the various drawings indicate like elements. DETAILED DESCRIPTION
[0025] Automatic speech recognition (ASR) systems typically rely on speech processing algorithms that assume that there is only one speaker in a given input audio signal. Input audio signals that include the presence of multiple speakers may potentially disrupt these speech processing algorithms, thereby causing the ASR system to output incorrect speech recognition results. These ASR systems include speaker classification systems to answer the question of "who is speaking when". Therefore, speaker classification is the process of segmenting speech from multiple speakers participating in a larger conversation, not to specifically determine who is speaking (speaker identification / recognition), but to determine when someone is speaking. In other words, speaker classification includes a series of speaker identification tasks with short utterances, and determines whether two segments of a given conversation are spoken by the same person or different people, and repeats for all segments of the conversation. Therefore, speaker classification detects speaker turns from a conversation involving multiple speakers. As used herein, the term 'speaker turn' refers to the transition from one person speaking to another person speaking in a larger conversation.
[0026] Existing speaker classification systems typically include multiple relatively independent components, such as but not limited to a speech segmentation module, an embedding extraction module, and a clustering module. The speech segmentation module is typically configured to remove non-speech parts from the input utterance and divide the entire input utterance into fixed-length segments and / or word-length segments. Although it is easy to divide the input utterance into fixed-length segments, it is often difficult to find a good segment length. That is, long fixed-length segments may include multiple speaker turns, while short segments include insufficient speaker information. In addition, the ASR model generates word-length segments that are typically spoken by a single speaker, however, a single word also includes insufficient speaker information. The embedding extraction module is configured to extract a corresponding speaker-discriminative embedding from each segment. The speaker-discriminative embedding may include an i-vector or a d-vector.
[0027] The task of the clustering module used by existing speaker classification systems is to determine the number of speakers present in the input utterance and assign speaker identities (e.g., labels) to each segment. These clustering modules can use popular clustering algorithms, including Gaussian mixture models, naive clustering, link clustering, agglomerative hierarchical clustering (AHC), and spectral clustering. Speaker classification systems can also use additional re-segmentation modules to further refine the classification results output from the clustering module by enforcing additional constraints. The clustering module can execute an online clustering algorithm that usually has low quality or an offline clustering algorithm that can only return classification results at the end of the entire input sequence. In some examples, in order to achieve high quality while minimizing latency, the clustering algorithm is run offline in an online manner. For example, in response to receiving each speaker discriminative embedding, the clustering algorithm is run offline for the entire sequence of all existing embeddings. However, if the sequence of speaker discriminative embeddings is long, the computational cost of implementing these examples may be very high.
[0028] For unsupervised speaker classification systems where the number of distinct speakers in the input audio is unknown, state-of-the-art spectral clustering algorithms are computationally expensive to implement in production applications when the sequence of speaker discriminative embeddings extracted from the corresponding audio segments is large. For example, assuming there are N speaker segments, each with a corresponding speaker discriminative embedding extracted from it, the spectral clustering algorithm needs to compute the Laplacian matrix of the N*N affinity matrix and perform eigendecomposition on the computed Laplacian matrix. Since both the Laplacian matrix and the eigendecomposition have ~O(N 2.7 ), so the computational complexity of performing spectral clustering to classify long-duration audio is usually unacceptable because the number of N speaker segments can be large. Long-duration audio can include, but is not limited to, conference audio recordings, podcasts, and videos.
[0029] In contrast, when the sequence of speaker-discriminative embeddings extracted from the corresponding audio clips is small, spectral clustering will produce results of reduced quality relative to other types of clustering algorithms, because there is not enough information to perform the graph cut technique required by spectral clustering. In addition, since spectral clustering uses a feature gap criterion to ultimately determine the number of clusters, this criterion is only valid if it is known in advance that at least two speakers are captured in the input audio data. That is, if it is not known whether there are at least two speakers, spectral clustering will often predict the wrong number of speakers.
[0030] Implementations herein are directed to accelerating speaker classification performance on a speaker classification system that includes a speech recognition model that performs speech recognition and speaker turn detection (i.e., when an active speaker changes) on received utterances spoken by multiple speakers. The speaker classification system segments the utterance into speaker segments based on detected speaker turns and extracts speaker discriminative embeddings therefrom. Advantageously, each speaker segment segmented from the utterance based on speaker turn detection includes continuous speech from the speaker that carries sufficient information to extract a robust speaker discriminative embedding.
[0031] For long duration audio characterized by a large number N of speaker segments exceeding a threshold, implementations herein are specifically directed to reducing the computational cost of performing spectral clustering on speaker discriminative embeddings extracted from the number N speaker segments by first performing pre-clustering on the speaker discriminative embeddings extracted to cluster the speaker segments into a target number M of pre-clustered clusters. Thereafter, a speaker classification system determines corresponding centroid values for each of the target number M pre-clustered clusters based on the speaker discriminative embeddings extracted from the speaker segments clustered into the corresponding pre-clustered clusters, and then performs spectral clustering on the centroid values determined for the target number M pre-clustered clusters. Here, the target number M pre-clustered clusters is less than the number N speaker segments, thereby limiting the computational cost to the value of M specified for the target number of pre-clustered clusters, regardless of the actual number of N speaker segments, each speaker segment having a corresponding speaker embedding extracted therefrom. Advantageously, the value of M can be manually specified based on the availability of computing resources for an application running the speaker classification system. For example, a larger value of M may be specified for server-side applications, while a smaller value of M may be specified for on-device applications with a smaller computational budget. Therefore, since the number of speaker turns (i.e., the number of speaker changes) is typically much smaller than the number of fixed-length segments, speaker discriminative embeddings are extracted only from speaker segments delimited by speaker turns.
[0032] After performing pre-clustering, the speaker classification system performs spectral clustering on the centroid values determined for the target number of M pre-clustering clusters to cluster the centroid values into k classes, and for each respective class in the k classes, the speaker classification system assigns a respective speaker label to each centroid value clustered into the respective class that is different from the respective speaker label assigned to the centroid value in each other class clustered into the k classes. Based on the pre-clustering information indicating which pre-clustering clusters of the target number of M pre-clustering clusters contain which speaker segments among the number N speaker segments, the speaker classification system can map the speaker labels assigned to the M centroid values back to the number N speaker segments and annotate the transcription of the utterance based on the speaker labels now assigned to each speaker segment. For example, a transcription of a conversation between multiple speakers can be indexed by speaker to associate portions of the transcription with respective speakers, thereby identifying which portions each speaker said in the transcription.
[0033] It is worth noting that when the number N of speaker segments does not exceed a threshold, the speaker classification system can bypass pre-clustering and perform spectral clustering on speaker discriminative embeddings extracted from the number N speaker segments to cluster the plurality of speaker segments into k classes. Here, for each respective class in the k classes, the speaker classification system assigns a respective speaker label to each speaker segment clustered into the respective class, the speaker label being different from the respective speaker labels assigned to speaker segments in each other class clustered into the k classes.
[0034] As an additional measure to help reduce the computational cost of performing spectral clustering, the number of speaker turns (i.e., the number of times a speaker changes) is typically much smaller than the number of fixed-length audio segments that are additionally input to the speaker classification system. In other words, since speaker discriminative embeddings are extracted only from speaker segments delimited by speaker turns, the computational cost of performing the pre-clustering algorithm and / or the spectral clustering algorithm is reduced. Advantageously, since the per-turn speaker discriminative embeddings are extracted sparsely from the speaker segments (i.e., only after a speaker turn), the sequence of all existing speaker discriminative embeddings is relatively short even for relatively long conversations (i.e., several hours).
[0035] In addition, training time is significantly reduced because human annotators are not required to assign accurate timestamps to speaker turns and manually identify different speakers across these turns. Annotating timestamps and identifying speakers across turns is a time-consuming process, and it may take a single annotator about two hours to annotate 10 minutes of audio at one time. Instead, the speech recognition model is trained to detect speaker turns from the semantic information conveyed in the speech recognition results, so that each detected speaker turn is associated with a corresponding timestamp known to the speech recognition model. Therefore, these timestamps are not annotated by humans and can be used to segment the training audio data into corresponding speaker segments.
[0036] refer to Figure 1 , the system 100 includes a user device 110 that captures speech utterances 120 from a group of speakers (e.g., users) 10 (i.e., 10a-n) and communicates with a cloud computing environment 140 via a network 130. The cloud computing environment 140 can be a distributed system with scalable / elastic resources 142. The resources 142 include computing resources 142 (e.g., data processing hardware) and / or storage resources 146 (e.g., memory hardware). In some implementations, the user device 110 and / or the cloud computing environment 140 executes a classification system 150 that is configured to receive an input audio signal (i.e., audio data) 122 corresponding to the captured utterances 120 from a plurality of speakers 10. The classification system 150 processes the input audio signal 122 and generates a transcription 200 of the captured utterance 120 and one or more speaker turn markers 224 (i.e., 224a-n). The speaker turn markers 224 indicate a speaker turn (e.g., a speaker change) detected between a pair of corresponding adjacent terms in the transcription 200. Using one or more speaker turn markers 224, the classification system 150 segments the input audio signal 122 into a plurality of N speaker segments 225 (i.e., 225a-N), each speaker segment being associated with a corresponding speaker discriminative embedding 240 extracted therefrom. Thereafter, the classification system 150 generates a classification result 280 based on the speaker discriminative embeddings 240 and the pairwise constraints 226. The classification result 280 includes a corresponding speaker label 250 assigned to each speaker segment 225.
[0037] The user device 110 includes data processing hardware 112 and memory hardware 114. The user device 110 may include an audio capture device (e.g., a microphone) for capturing speech 120 from the speaker 10 and converting the speech into audio data 122 (e.g., an electrical signal). In some implementations, the data processing hardware 112 is configured to execute a portion of the classification system 150 locally, while the rest of the classification system 150 is executed on the cloud computing environment 140. Alternatively, the data processing hardware 112 may execute the classification system 150 instead of executing the classification system 150 on the cloud computing environment 140. The user device 110 may be any computing device capable of communicating with the cloud computing environment 140 via the network 130. The user device 110 includes, but is not limited to, desktop computing devices and mobile computing devices, such as laptop computers, tablet computers, smart phones, smart speakers / displays, smart appliances, Internet of Things (IoT) devices, and wearable computing devices (e.g., headphones and / or watches).
[0038] Although the present disclosure generally depicts audio data 122 representing speech captured in real time between one or more speakers, audio data 122 may be derived from recorded media content and streaming media content. For example, audio data 122 may be derived from any audio and / or audiovisual source, such as broadcast content (e.g., television programs), podcasts, web-based audiovisual content, pre-recorded audio and / or audiovisual content (such as a recording of a telephone conference between two or more participants), and streaming audio and / or audiovisual content captured in real time during, for example, a telephone conference between two or more participants.
[0039] In the example shown, the speaker 10 and the user device 110 may be located within an environment (e.g., a room), wherein the user device 110 is configured to capture speech utterances 120 spoken by the speaker 10 and convert the speech utterances into input audio signals 122 (also referred to as audio data 122). For example, the speaker may correspond to a colleague having a conversation during a meeting, and the user device 110 may record the speech utterances and convert the speech utterances into the input audio signals 122. The user device 110 may then provide the input audio signals 122 to the classification system 150 for predicting which speaker 10 is speaking for each speech segment. Thus, the task of the classification system 150 is to process the input audio signals 122 to determine when someone is speaking without specifically determining who is speaking via speaker recognition / identification.
[0040] In some examples, at least a portion of the utterances 120 conveyed in the input audio signal 122 overlap, such that at a given instant in time, the speech of two or more of the speakers 10 is active. Notably, the number of the plurality of speakers 10 may be unknown when the input audio signal 122 is provided as input to the classification system 150, and the classification system may predict the number of the plurality of speakers 10. In some implementations, the user device 110 is located remotely from the speakers 10. For example, the user device may include a remote device (e.g., a network server) that captures voice utterances 120 from speakers who are participants in a phone call or video conference. In such a scenario, each speaker 10 (or a group of multiple speakers 10) will speak to their own device (e.g., a phone, radio, computer, smart watch, etc.), which captures the voice utterances 120 and provides the voice utterances to the remote user device 110 for conversion of the voice utterances 120 into audio data 122. Of course, in such a scenario, the speech 120 may be processed at each of the user devices and converted into corresponding input audio signals 122, which are transmitted to the remote user device 110, which may further process the input audio signals 122 provided as input to the classification system 150.
[0041] In the example shown, the classification system 150 includes an ASR model 300, a segmentation module 210, a speaker encoder 230, a cluster selector 400, and a clustering module 260. Figure 4 Described in more detail, the clustering module 260 may execute a backup clustering algorithm 260a, a spectral clusterer algorithm 260b, and a pre-clusterer algorithm 260c. The ASR model 300 is configured to receive an input audio signal 122 and process the input audio signal 122 to jointly generate a transcription 200 of the utterance 120 and a sequence of speaker turn markers 224 (i.e., 224a–n). The ASR model 300 may include a streaming ASR model 300 that jointly generates the transcription 200 and the speaker turn markers 224 in a streaming manner upon receiving the input audio signal 122. The transcription 200 includes a sequence of speaker turn markers 224 that indicates the locations of corresponding speaker turns detected between a pair of corresponding adjacent terms in the transcription 200. For example, the utterance 120 may include “Hello how are you I’m fine”, and the ASR model 300 generates the transcription 200 “Hello how are you <st>I'm fine." In this example, <st>represents a speaker turn marker 224 indicating a speaker turn between adjacent terms 'you' and 'I'. Each speaker turn marker 224 in the sequence of speaker turn markers 224 may also include a corresponding timestamp 223.
[0042] Figure 2 An example transcription 200 of an utterance 120 represented by an input audio signal 122 and a Figure 1 1. The ASR model 300 of FIG. 10 shows a sequence of speaker turn markers 224 output by an ASR model 300 of FIG. 10. The transcription 200 includes one or more terms 222 corresponding to words spoken by one or more speakers. The sequence of speaker turn markers 224 indicates the positions of corresponding speaker turns detected between a pair of corresponding adjacent terms 222 in the transcription 200. In the example shown, the input audio signal 122 may include an utterance in which the first and second terms 222 are spoken by a first speaker 10, the third and fourth terms 222 are spoken by a second speaker 10, and the fifth and sixth terms 222 are spoken by a third speaker. Here, the ASR model 300 generates a first speaker marker 224 between the second term 222 and the third term 222 to indicate a speaker turn from the first speaker to the second speaker, and generates a second speaker marker 224 between the fourth term 222 and the fifth term 222 to indicate a speaker turn from the second speaker to the third speaker. Additionally, in some examples, the ASR model 300 generates a start of speech (SOS) marker 227 indicating the start of an utterance and an end of speech (EOS) marker 229 indicating the end of an utterance.
[0043] In some implementations, the ASR model 300 processes acoustic information and / or semantic information to detect speaker turns in the input audio signal 122. That is, using natural language understanding (NLU), the ASR model 300 can determine, for the utterance “how are you I’m fine,” that “how are you” and “I’m fine” are likely spoken by different users, independent of any acoustic processing of the input audio signal 122. This semantic interpretation of the transcription 200 can be used independently or in conjunction with acoustic processing of the input audio signal 122.
[0044] Optionally, the ASR model 300 can utilize the classification results 280 to improve speech recognition of the audio data 122. For example, the ASR model 300 can apply different speech recognition models (e.g., language models, prosody models) for different speakers identified from the classification results 280. Additionally or alternatively, the ASR model 300 and / or the classification system 150 (or some other component) can index the transcription 200 of the audio data 122 using the speaker tags 250 for each speaker segment 225. For example, a transcription of a conversation between multiple colleagues (e.g., speakers 10) during a business meeting can be indexed by speaker to associate portions of the transcription 200 with the corresponding speakers 10, thereby identifying what each speaker said.
[0045] The ASR model 300 may include any transducer-based architecture, including but not limited to transformer-transducer (TT), recurrent neural network transducer (RNN-T) and / or conformer-transducer (CT). The ASR model 300 is trained on training samples, each of which includes a pairing of training utterances spoken by two or more different speakers 10 and the corresponding true value transcriptions of the training utterances. Each true value transcription is injected with true value speaker turn markers, which indicate the locations in the true value transcription where speaker turns occur. Here, the corresponding true value transcription of each training sample is not annotated with any timestamp information.
[0046] refer to Figure 3 , the ASR model 300 can provide end-to-end (E2E) speech recognition by integrating acoustic, pronunciation, and language models into a single neural network, and does not require a dictionary or a separate text normalization component. Various structures and optimization mechanisms can provide increased accuracy and reduced model training time. The ASR model 300 may include a streaming Transformer-Transducer (TT) model architecture that complies with delay constraints associated with interactive applications. The ASR model 300 may similarly include an RNN-T model architecture or a Conformer-Transducer (CT) model architecture. In addition to the TT and CT model architectures, the ASR model 300 may also include other types of Transducer model architectures that have an audio encoder 310 including multiple multi-head attention layers. The ASR model 300 provides a small computational footprint and uses less memory requirements than conventional ASR architectures, making the TT model architecture suitable for performing speech recognition entirely on the user device 110 (e.g., without the need to communicate with the cloud computing environment 140). The ASR model 300 includes an audio encoder 310, a label encoder 320, and a joint network 330. The audio encoder 310, which is generally similar to the acoustic model (AM) in a conventional ASR system, includes a neural network having a plurality of transformer layers. For example, the audio encoder 310 reads a d-dimensional feature vector (e.g., a speaker segment 225 ( Figure 1 )) of the sequence x=(x 1 ,x 2 ,...,x T ),in And generate high-level feature representation 312 at each time step. Here, each speaker segment 225 ( Figure 1 ) includes corresponding speaker segment 225 ( Figure 1 ) corresponds to an acoustic frame sequence (e.g., audio data 122). This high-level feature representation is represented as ah 1 ,……,ah T .
[0047] Similarly, the label encoder 320 may also include a neural network or a lookup table embedding model of a transformer layer, which, like a language model (LM), converts the sequence y of non-empty symbols outputted so far by the last Softmax layer 340 into 0 ,……,y ui-1 (For example, Figure 2 As shown, one or more terms 222 including speaker turn markers 224 are processed into a dense representation 322 (represented by Ih u denoted). In implementations where the label encoder 320 includes a neural network of transformer layers, each transformer layer may include a normalization layer, a masked multi-head attention layer with relative position encoding, a residual connection, a feedforward layer, and a dropout layer. In these implementations, the label encoder 320 may include two transformer layers. In implementations where the label encoder 320 includes a lookup table embedding model with a two-tuple label context, the embedding model is configured to learn a d-dimensional weight vector for each possible two-tuple label context, where d is the dimension of the output of the audio encoder 310 and the label encoder 320. In some examples, the total number of parameters in the embedding model is N 2 xd, where N is the vocabulary size of the tag. Here, the learned weight vector is then used as the embedding of the bigram tag context in the ASR model 300 to produce the fast tag encoder 320 runtime.
[0048] Finally, using the TT model architecture, the joint network 330 uses dense layers J u,t The representations produced by the audio encoder 310 and the label encoder 320 are combined. Then, the joint network 330 predicts P(z u,t |x,t,y 1 ,...,y u-1 ), which is the distribution of the next output symbol. In other words, for one or more terms 222 ( Figure 2 ), the joint network 330 generates a probability distribution of possible speech recognition hypotheses 342 at each output step (e.g., time step). Here, "possible speech recognition hypotheses" correspond to a set of output labels (also called "speech units"), each of which represents a grapheme (e.g., symbol / character), a lexical item 222 ( Figure 2 ) or word fragments. For example, when the natural language is English, the set of output labels may include twenty-seven (27) symbols, for example, one label for each of the 26 letters in the English alphabet, and one label for a space. Thus, the joint network 330 may output a set of values indicating the likelihood of each of a set of predetermined output labels occurring. The set of values may be a vector (e.g., a one-hot vector) and may indicate a probability distribution of the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and possibly punctuation and other symbols), but the set of output labels is not limited to this. For example, in addition to or in lieu of graphemes, the set of output labels may also include word fragments and / or entire words. The output distribution of the joint network 330 may include a posterior probability value for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output z of the joint network 330 may be 100. u,t 100 different probability values may be included, one for each output label. The probability distribution may then be used (e.g., by a Softmax layer 340) to select candidate orthographic elements (e.g., graphemes, word fragments, and / or words) and assign scores to them during a beam search process for use in determining a transcription.
[0049] The Softmax layer 340 may employ any technique to select the output label / symbol with the highest probability in the distribution as the next output symbol predicted at the corresponding output step by the ASR model 300. In this manner, the ASR model 300 does not make conditional independence assumptions, but rather the prediction of each symbol is conditioned not only on the acoustics, but also on the sequence of labels output so far.
[0050] Return to reference Figure 1 In the speaker classification system 150 of the present invention, the segmentation module 210 is configured to receive audio data 122 corresponding to a speech utterance 120 (also referred to as an 'utterance of speech') and segment the audio data 122 into a plurality of N speaker segments 225 (i.e., 225a-N). The segmentation module 210 receives the audio data 122 and a transcription 200, the transcription comprising a sequence of speaker turn markers 224 with corresponding timestamps 223 to segment the audio data 122 into a plurality of N speaker segments 225. Here, each speaker segment 225 corresponds to audio data between two adjacent speaker turn markers 224. Optionally, the segmentation module 210 may also remove non-speech portions from the audio data 122 (e.g., by applying a voice activity detector). In some examples, the segmentation module 210 further segments the speaker segments 225 that exceed a segment duration threshold, which is described in more detail below.
[0051] The segmentation module 210 segments the input audio signal 122 into a plurality of N speaker segments 225 by segmenting the input audio signal 122 into initial speaker segments 225, each of which is delimited by a corresponding time stamp 223 of a pair of corresponding adjacent speaker turn markers 224. For example, the input audio signal 122 may include fifteen seconds of audio, wherein a sequence of speaker turn markers 224 has time stamps 223 at three seconds, six seconds, and fourteen seconds. In this case, the segmentation module 210 segments the input audio signal into three initial speaker segments 225, which are delimited by speaker turn markers 224 with time stamps 223 at three seconds, six seconds, and fourteen seconds.
[0052] In some implementations, one or more of the initial speaker segments 225 have a corresponding duration that exceeds a segment duration threshold. In these implementations, the segmentation module 210 further segments the initial speaker segment 225 into two or more shortened duration speaker segments 225 having corresponding durations less than or equal to the segment duration threshold. Continuing with the above example, the segmentation module may determine that the initial speaker segment 225 delimited by the speaker turn markers 224 with timestamps at six seconds and fourteen seconds (e.g., having a duration of eight seconds) exceeds the segment duration threshold of six seconds. In this scenario, the segmentation module 210 may further segment the initial speaker segment 225 into two or more shortened duration speaker segments 225 having corresponding durations less than or equal to the segment duration threshold. Here, the segmentation module 210 may segment the eight-second initial speaker segment 225 into a first shortened duration speaker segment 225 having a duration of six seconds and a second shortened duration speaker segment 225 having a duration of two seconds. Therefore, the multiple speaker segments 225 segmented from the input audio signal 122 may include both initial speaker segments 225 having corresponding durations less than or equal to the segment duration threshold and speaker segments 225 of shortened duration further segmented from any of the initial speaker segments 225 having corresponding durations exceeding the segment duration threshold.
[0053] The speaker encoder 230 is configured to receive a plurality of speaker segments 225 and, for each speaker segment 225 in the plurality of speaker segments 225, extract a corresponding speaker discriminative embedding 240 from the speaker segment 225 as an output. Thereafter, the speaker encoder embeds the observation sequence X=(x 1 ,x 2 ,...,x T ) is provided to the clustering module 260, where the entry x in the sequence T 1 represents a real-valued speaker discriminative embedding 240 associated with a corresponding speaker segment 225 in the audio data 122 of the original utterance 120. The speaker discriminative embedding 240 may include a speaker vector, such as a d-vector or an i-vector.
[0054] In some examples, the speaker encoder 230 includes a text-independent speaker encoder model trained with a generalized end-to-end extended set softmax loss. The speaker encoder can include a long short-term memory (LSTM-based) speaker encoder model that is configured to extract a corresponding speaker discriminative embedding 240 from each speaker segment 225. In particular, the speaker encoder 230 includes (3) long short-term memory (LSTM) layers with 768 nodes and a projection size of 256. Here, the output of the last LSTM is transformed into a final 256-dimensional d-vector. In some configurations, the final dimensionality of the speaker discriminative embedding 240 output from the speaker encoder 230 is reduced to 64 dimensions to speed up the affinity matrix calculation performed by the clustering module 260.
[0055] In some implementations, each speaker turn marker 224 in the sequence of speaker turn markers 224 resets the LSTM state of the speaker encoder 230 so that the speaker discriminative embedding 240 does not include information from other speaker segments 225. For example, the speaker encoder 230 may extract a speaker discriminative embedding 240 corresponding to only a portion of the speaker segment 225. Thus, the speaker discriminative embedding 240 includes sufficient information from the speaker segment 225, but not so close to the speaker turn boundary that the speaker discriminative embedding 240 may include inaccurate information or contain overlapping speech from another speaker 10.
[0056] Based on the total number of speaker turn markers 224 and the corresponding confidence value 331 of each speaker turn marker 224 ( Figure 4 ), and the total number N of the plurality of N speaker segments 225, the cluster selector 400 determines clustering instructions 402 that instruct the clustering module 260 how to cluster the observation sequence of the embedding X. For example, when at least a threshold number of speaker turn markers among the one or more speaker turn markers do not have corresponding confidence values that satisfy a confidence value threshold, the cluster selector 400 may simply output a single speaker indication 412 indicating that speaker classification does not need to be performed, and therefore, provide clustering instructions 402 that instruct the clustering module 260 not to perform any clustering algorithm on the observation sequence of the embedding X. Otherwise, and with reference to Figure 4 Described in more detail, when the cluster selector 402 determines that a threshold number of one or more speaker tags each having a corresponding confidence value that satisfies a confidence value threshold is met, the cluster selector 402 provides clustering instructions 402 based on the number of N speaker segments 225, which instruct the clustering module 260 to execute at least one of the backup clustering algorithm 260a, the spectral clustering algorithm 260b, 260d, or the pre-clustering algorithm 260c for generating a classification result 280.
[0057] Figure 4 A schematic diagram of a cluster selector 400 is shown, depicting a process flow for determining a cluster instruction 402 based on a pairwise constraint 226 and a total number N of the plurality of N speaker segments 225. Initially, a speaker turn counter 410 receives a pairwise constraint 226 indicating a total number of speaker turn tokens 224 and a corresponding confidence value 331 ( Figure 4 ). The speaker turn counter 410 determines whether there are at least a threshold number of speaker turn markers 224 having corresponding confidence values 331 that satisfy a confidence threshold. In some examples, the threshold number of speaker turn markers is equal to one (1). In other examples, the threshold number of speaker turn markers may include any integer greater than one. The threshold number may be manually selected and changed. Similarly, the confidence value threshold may be adjusted to change the sensitivity. When the speaker turn counter 410 determines that there are not at least a threshold number of speaker turn markers 224 having corresponding confidence values 331 that satisfy the confidence threshold, the speaker turn counter 410 outputs a single speaker indication 412 ("Single Speaker") indicating that only a single speaker is predicted to be present in the input audio signal. In this case, the clustering module 260 is instructed not to execute any clustering algorithm for classifying the input audio signal 120. On the other hand, if the speaker turn counter 410 determines that there are at least a threshold number of speaker turn markers 224 having corresponding confidence values 331 that satisfy the confidence threshold (i.e., indicating that there are at least two different speakers), the speaker turn counter 410 outputs a multi-speaker indication 413 ("T turns detected"), which causes the cluster selector 400 to evaluate the total number N of multiple N speaker segments 225 for determining the clustering instructions 402.
[0058] At decision step 420, cluster selector 400 determines whether the number of plurality of N speaker segments 225 is greater than or equal to a minimum threshold number L. When cluster selector 400 determines that the number of plurality of N speaker segments is less than the minimum threshold number L (i.e., decision step 420 is “no”), cluster selector 400 determines clustering instructions 402 that instruct clustering module 260 to execute backup clusterer algorithm 260a for clustering the observed sequence of embedding X (i.e., N speaker embeddings 240) to obtain classification results 280 ( Figure 1 )。Rather than performing spectral clustering, the fallback clustering algorithm 260 performs another type of clustering algorithm that is more suitable for clustering the embedded observation sequences when the number of N speaker segments 225 is small (e.g., N < L). The fallback clustering algorithm may require specifying a threshold parameter for the similarity scores used for clustering. Notably, appropriately choosing the threshold parameter for the similarity scores can result in the fallback clustering algorithm 260a performing significantly better than spectral clustering when the number of N speaker segments 225 is small. The fallback clustering algorithm 260a may include a naive clustering algorithm. The fallback clustering algorithm 260a may include a linkage clustering algorithm. The fallback clustering algorithm 260a may include an agglomerative hierarchical clustering (AHC) algorithm.
[0059] Conversely, when the clustering selector 400 determines that the number of the plurality of N speaker segments 225 is greater than or equal to the minimum threshold number L (i.e., the decision step 420 is "yes"), the clustering selector 400 proceeds to the decision step 430 to determine whether the number of the plurality of N speaker segments 225 is less than or equal to the maximum threshold number M. When the clustering selector 400 determines that the number of the plurality of N speaker segments 225 is less than or equal to the maximum threshold number M (i.e., the decision step 430 is "yes"), the clustering selector 400 determines clustering instructions 402 that direct the clustering module 260 to perform the spectral clustering algorithm 260b, and thus perform spectral clustering on the speaker-discriminative embeddings 240 extracted from the plurality of N speaker segments to cluster the plurality of N speaker segments into k classes 262. The k classes 262 represent the predicted number of active speakers included in the received utterance 120. Thereafter, for each respective class 262 among the k classes 262, the clustering module 260 assigns a respective speaker label 250 to each speaker segment 225 clustered into the respective class 262, the speaker label being different from the respective speaker label 250 assigned to the speaker segments 225 clustered into each other class 262 among the k classes 262. Here, the spectral clustering algorithm 260b receives the speaker-discriminative embeddings 240 of each speaker segment 225 and the pairwise constraints 226, and is configured to predict the speaker label 250 of each speaker-discriminative embedding 240. In short, the clustering module 260 performs the spectral clustering algorithm 260b to predict which speaker 10 each speaker segment 225 is from.
[0060] The clustering module 260 receives the speaker discriminative embedding 240 and the pairwise constraints 226 of each speaker segment 225, and is configured to predict a speaker label 250 for each speaker discriminative embedding 240. In short, the clustering module 260 predicts which speaker 10 said each speaker segment 225. More specifically, the clustering module 260 performs spectral clustering on the speaker discriminative embedding 240 extracted from the plurality of speaker segments 225 to cluster the plurality of speaker segments 225 into k clusters 262. The k clusters 262 represent the predicted number of active speakers included in the received utterance 120. Thereafter, for each respective class 262 in the k classes 262 , the clustering module 260 assigns a respective speaker label 250 to each speaker segment 225 clustered into the respective class 262 , which is different from the respective speaker label 250 assigned to the speaker segments 225 in each other class 262 clustered into the k classes 262 .
[0061] refer to Figure 1 and Figure 4 In some implementations, the classification system 150 annotates the transcription 200 of the utterance 120 based on the speaker tags 250 (i.e., the classification results 280) assigned to each speaker segment 225. For example, the transcription 200 of a conversation between multiple speakers 10 can be indexed by speaker to associate portions of the transcription 200 with the corresponding speakers 10, thereby identifying what each speaker 10 said in the transcription 200. The annotated transcription 200 can be stored in the memory hardware 114, 146 of the user device 110 or the cloud computing environment 140 for later access by one of the speakers 10.
[0062] The pairwise constraints 226 generated by the ASR model 300 can further constrain the spectral clustering performed on the speaker discriminative embeddings 240. In addition to the confidence value 331 of each speaker turn marker 224, the pairwise constraints 226 can also indicate contextual information about the adjacent speaker segments 225. For example, the adjacent speaker segments 225 can include any combination of the following: both speaker segments 225 have a duration less than a segment duration threshold; one speaker segment 225 has a duration less than the segment duration threshold and one speaker segment 225 has a shortened duration (i.e., an initial speaker segment that exceeds the segment duration threshold); or both speaker segments 225 have a shortened duration. The confidence values 331 and the contextual information (collectively referred to as the constraints 226) of the corresponding speaker turns detected in the transcription 200 are used to further constrain the spectral clustering performed by the spectral clustering module 260 that executes the spectral clustering algorithm 260b.
[0063] In some implementations, the clustering module 260 executes the spectral clustering algorithm 260a to perform spectral clustering on the speaker discriminative embedding 240, the spectral clustering being constrained by the pairwise constraints 226 received from the ASR model 300. For example, when two adjacent speaker segments 225 both have a duration less than a segment duration threshold, the spectral clustering is constrained to force the speaker labels 250 of the adjacent speaker segments 225 separated by a speaker turn marker with high confidence to be different. In other cases, when two adjacent speaker segments 225 both have a shortened duration, the spectral clustering is constrained to force the speaker labels of the adjacent speaker segments 225 to be the same. That is, since the adjacent shortened duration speaker segments 225 are divided based on exceeding the segment duration threshold instead of the speaker turn marker 224, it is very likely that the adjacent shortened duration speaker segments 225 are spoken by the same speaker 10. In some examples, when one speaker segment 225 having a duration less than a segment duration threshold is adjacent to another speaker segment 225 having a shortened duration, spectral clustering is constrained based on the confidence of the speaker turn marker 224. Here, when the speaker turn marker 224 has a high confidence value 331, the clustering module 260 is constrained to promote different speaker labels 250. Alternatively, when the speaker turn marker 224 has a low confidence value, the clustering module 260 can be constrained to promote the same speaker label 250.
[0064] The spectral clustering algorithm 260b receives the speaker discriminative embedding 240 and the pairwise constraints 226 for each speaker segment 225 and is configured to predict a speaker label 250 for each speaker discriminative embedding 240. 1 、x 2 ,……,x T ), the spectral clustering algorithm 260b calculates the pairwise similarity a ij To construct a similarity graph, where A represents the affinity matrix of the similarity graph In addition, the two samples x i and x j The affinity can be determined by Representation. The spectral clustering algorithm 260b identifies partitions such that edges connecting different clusters have low weights, while edges within a cluster have high weights. In general, the similarity graph is connected or includes only a few connected components and very few isolated vertices. Spectral clustering is sensitive to the quality and noise of the similarity graph, so the spectral clustering algorithm 260b performs several refinement operations on the affinity matrix to model the local neighborhood relationships between data samples. One refinement operation includes row-wise thresholding with p percentiles, which sets the diagonal values of the affinity matrix to 0, sets affinity values greater than the p percentile value to 1, multiplies affinity values less than the p percentile of the row by 0.01, and resets the diagonal values of the affinity matrix to 1. Another refinement operation includes applying an average summary operation to use the following equation The affinity matrix becomes positive semidefinite. The classification error rate (DER) is significantly affected by the hyperparameter p of the p percentile. Therefore, the ratio value r(p) is a good proxy for DER, making the maximum feature gap large while not generating excessive connections in the similarity graph.
[0065] Given an affinity matrix A, the unnormalized Laplacian matrix L is defined by L = D – A, while the normalized Laplacian matrix Depend on Definition. Here, D is defined as To perform spectral clustering, the spectral clustering algorithm 260b applies eigendecomposition to estimate the number of k clusters 262 using the maximum eigengap method. The spectral clustering algorithm 260b selects the first cluster k 262 of the eigenvectors and applies row-wise renormalization of the spectral embeddings and applies a k-means algorithm on the spectral embeddings to predict the speaker labels 250.
[0066] In some examples, the spectral clustering algorithm 260b receives the confidence values 331 indicating the speaker turn markers 224 and the pairwise constraints 226 of the contextual information to constrain the spectral clustering. The pairwise constraints 226 are configured to promote different speaker labels 250 for adjacent speaker segments 225 with high confidence speaker turn markers 224 and promote the same speaker labels 250 for adjacent speaker segments 225 with low confidence speaker turn markers 224. Using the pairwise constraints 226Q, the constrained spectral clustering identifies one or more partitions that maximize constraint satisfaction and minimize the cost of the similarity graph G. The pairwise constraints 226 can be represented by The spectral clustering algorithm 260b processes the constraint matrix Q in the following way:
[0067]
[0068] Here, if there is a speaker change between speaker segments 225i and i+1, and the speaker change marker c( <st>) is greater than the threshold σ, the spectral clustering algorithm 260b defines the adjacent speaker segments 225 as "cannot be linked" (CL). The CL definition indicates that the speaker labels 250 between the adjacent speaker segments 225 are likely to be different. If there are no speaker turn marks 224 between the adjacent speaker segments 225, the clustering module defines the adjacent speaker segments as "must be linked" (ML). The ML definition indicates that the speaker labels 250 between the adjacent speaker segments 225 are likely to be the same.
[0069] The adjacent speaker segments 225 defined by ML are considered as positive classes, while the adjacent speaker segments 225 defined by CL are considered as negative classes. The class labels (i.e., positive and negative) are given in the affinity matrix In each iteration t, the initial constraint matrix is added to adjust Q(t). In addition, the parameter α is used to control the relative amount of constraint information from the neighboring speaker segments 225 and the initial constraints 226. The spectral clustering algorithm 260b first performs vertical propagation until convergence and then performs horizontal propagation by the following algorithm:
[0070] Algorithm 1: Exhaustive Efficient Constraint Propagation (E2CP) method
[0071] Requirements: initial constraint matrix Z=Q(0), matrix, parameters.
[0072]
[0073] Output As the final converged pairwise constraint matrix
[0074] Q * It has a closed-form solution, as follows:
[0075]
[0076] Using the propagated constraint matrix Q * , the spectral clustering algorithm 260b obtains the adjusted affinity matrix by
[0077]
[0078] For the constraint Q ij > 0, the affinity matrix adds samples x i With x j Alternatively, for Q ij <0, affinity matrix reduced by x i With x j After this operation, the spectral clustering algorithm 260b performs spectral clustering based on the normalized Laplacian matrix to predict the speaker labels 250 of the speaker segments 225. The spectral clustering algorithm 260b generates classification results 280, which may include an indication that the first speaker uttered the first speaker segment 225a (i.e., Figure 2 The first speaker tag 250a of the first and second terms 222 of the transcription 200 indicates that the second speaker said the second speaker segment 225b (ie, Figure 2 The third and fourth terms 222 of the transcription of the second speaker label 250b, and the third speaker segment 225c indicating that the third speaker said the third speaker segment 225c (ie, Figure 2 The third speaker tag 250c of the fifth and sixth terms 222 of the transcription 200).
[0079] Continue to refer Figure 4 , when the cluster selector 400 determines that the number of the plurality of N speaker segments 225 is greater than the maximum threshold number M (i.e., the decision step 430 is “no”), the cluster selector 400 determines clustering instructions 402 that instruct the clustering module 260 to execute the pre-clustering algorithm 260 c and, therefore, cause the clustering module 260 to perform pre-clustering on the speaker discriminative embeddings 240 extracted from the N speaker segments 225 to cluster the N speaker segments 225 into a target number of pre-clustering clusters 264 (i.e., 264 a–M). Here, the target number of pre-clustering clusters is less than the number of the N speaker segments. In some examples, the target number of pre-clustering clusters is set equal to the maximum threshold number M. The value of the maximum threshold number M may be manually set based on the availability of computing resources. For example, the maximum threshold number M may be set lower when the classification system 150 is executed on the user device 110 than when the classification system 150 is executed on the distributed system 140.
[0080] After the pre-clustering algorithm 260c clusters the N speaker segments 225 into the target number of pre-clustering clusters 264, the clustering module determines a corresponding centroid value 265 for each corresponding pre-clustering cluster 264 in the target number of M pre-clustering clusters 264a-M based on the speaker discriminative embeddings 240 extracted from the speaker segments 225 clustered into the corresponding pre-clustering clusters 264. Thereafter, the clustering module 260 executes the spectral clustering algorithm 260d to perform spectral clustering on the M centroid values 265 determined for the target number of M pre-clustering clusters 264 to cluster the centroid values 265 into k clusters 266.
[0081] The k classes 266 represent the predicted number of active speakers included in the received utterance 120. Thereafter, for each respective class 266 in the k classes 262, the clustering module 260 assigns a respective speaker label 250 to each centroid value 265 clustered into the respective class 266, which is different from the respective speaker label 250 assigned to the centroid value 265 in each other class clustered into the k classes 266. Notably, by pre-clustering the N speaker discriminative embeddings 240 into a target number of M pre-clustered clusters 264, the computational cost for executing the spectral clustering algorithm 260d is limited to the value of M specified for the target number of pre-clustered clusters, regardless of the total number of N speaker segments 225, each having a corresponding speaker discriminative embedding 240 extracted therefrom.
[0082] Based on the pre-clustering information indicating which of the target number of M pre-clustering clusters 264 contain which speaker segments 225 of the number N speaker segments 225, the mapper 270 can map the speaker labels 225 assigned to the M centroid values 265 back to the number N speaker segments 225 and annotate the transcription 200 of the utterance 120 based on the speaker labels 250 now assigned to each speaker segment 225. For example, a transcription of a conversation between multiple speakers can be indexed by speaker to associate portions of the transcription with corresponding speakers to identify which portions of the transcription were spoken by each speaker.
[0083] Although the use of pre-clustering limits the computational cost for executing the spectral clustering algorithm 260d to the value of M specified for the target number of pre-clustering clusters, for very long audio files, such as audio files of more than several hours in duration, the computational cost for executing the pre-clustering algorithm 260c itself may be unacceptable. In these scenarios, an upper limit U for the pre-clustering algorithm 260c can also be set so that the first time N speaker discriminative embeddings 240 are observed, where N ≥ U, the pre-clustering algorithm 260c can be run to obtain and cache (e.g., in memory hardware 114, 146) the target number of M pre-clustering clusters 264, which have associated M centroid values 265 that map back to the number N speaker segments 225. Thereafter, once each new (N+1) speaker discriminative embedding 240 is observed, the first N speaker discriminative embeddings 240 are replaced with the cached M centroid values 264, and the pre-clustering algorithm 260c is run on the M+1 embeddings including the M centroid values 264 and the new speaker discriminative embeddings 240. After having N' embeddings (where M+(N'-N) ≥ U), the clustering algorithm 260c can be run on M+(N'-N) to obtain and cache an updated target number of M pre-clustering clusters 264 having associated M centroid values 265 that map back to the N number of speaker segments 225. Therefore, by setting an upper limit U for the pre-clustering algorithm 260c (i.e., based on available computing resources), the pre-clustering algorithm is never run on more than the upper limit number U of embeddings.
[0084] Figure 5 is a flow chart of an exemplary arrangement of operations of a computer-implemented method 500 for performing speaker classification on a received speech utterance 120. Figure 1 The data processing hardware 610 of any of the data processing hardware 112, 144 may be implemented by executing the data stored in Figure 1 The method 500 is executed by instructions on the memory hardware 114, 146 of the memory hardware 620. At operation 502, the method 500 includes receiving an input audio signal 122 corresponding to an utterance 120 spoken by one or more speakers 10 (i.e., 10a-n). At operation 504, the method 500 includes processing the input audio signal 122 using a speech recognition model (e.g., an ASR model) 300 to jointly generate a transcription 200 of the utterance 120 and one or more speaker turn markers 224 (i.e., 224a-n) as output from the speech recognition model 300. Each speaker turn marker 224 indicates the position of a corresponding speaker turn detected between a pair of corresponding adjacent terms 222 in the transcription 200.
[0085] At operation 506, the method 500 includes segmenting the input audio signal 122 into a plurality of N speaker segments 225 based on one or more of the speaker tags 224. At operation 508, the method 500 includes, for each speaker segment 225 in the plurality of N speaker segments 225, extracting a corresponding speaker discriminative embedding 240 from the speaker segment 225.
[0086] Based on determining that the number of N speaker segments 225 is greater than a threshold number M (e.g., Figure 4 4 ), method 500 performs operations 510 to 514. At operation 510, method 500 includes performing pre-clustering on speaker discriminative embeddings 240 extracted from N speaker segments 225 to cluster the N speaker segments into a target number of pre-clustering clusters 264 (i.e., 264a-M). Here, the target number of pre-clustering clusters 264 is less than the number of N speaker segments. The target number of pre-clustering clusters 264 can be equal to or less than the maximum threshold number M. At operation 510, method 500 also includes determining a corresponding centroid value 265 for each corresponding pre-clustering cluster 264.
[0087] At operation 512, the method 500 includes performing spectral clustering (e.g., by executing the spectral clustering algorithm 260d) on the centroid values 265 determined for the target number of pre-clustered clusters 264 to cluster the centroid values 265 into k classes 266. At operation 514, for each respective class in the k classes 166, the method 500 includes assigning a respective speaker label 250 to each centroid value 265 clustered into the respective class that is different from the respective speaker label 250 assigned to the centroid value 265 in each other class clustered into the k classes 166. In some examples, the method also includes mapping the speaker labels 225 assigned to the M centroid values 265 back to the number N speaker segments 225 and annotating the transcription 200 of the utterance 120 based on the speaker labels 250 now assigned to each speaker segment 225.
[0088] A software application (i.e., software resource) may refer to computer software that enables a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0089] Non-transitory memory can be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by a computing device. Non-transitory memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware, such as bootloaders). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0090] Figure 6 600 that can be used to implement the systems and methods described in this document. Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit implementations of the inventions described and / or claimed in this document.
[0091] The computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connected to the memory 620 and a high-speed expansion port 650, and a low-speed interface / controller 660 connected to a low-speed bus 670 and the storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 is interconnected using various buses and can be installed on a common motherboard or installed in other ways as appropriate. The processor 610 can process instructions for execution within the computing device 600, including instructions stored in the memory 620 or on the storage device 630, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 680 coupled to the high-speed interface 640. In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memories can be used as appropriate. In addition, multiple computing devices 600 can be connected, each of which provides a portion of the necessary operations (for example, as a server group, a blade server group, or a multi-processor system).
[0092] Memory 620 stores information non-temporarily within computing device 600. Memory 620 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-temporary memory 620 may be a physical device for temporarily or permanently storing programs (e.g., instruction sequences) or data (e.g., program state information) for use by computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware, such as bootloaders). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0093] The storage device 630 is capable of providing mass storage for the computing device 600. In some implementations, the storage device 630 is a computer-readable medium. In various implementations, the storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices (including devices in a storage area network or other configurations). In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as the memory 620, the storage device 630, or a memory on the processor 610.
[0094] The high-speed controller 640 manages bandwidth-intensive operations of the computing device 600, while the low-speed controller 660 manages less bandwidth-intensive operations. Such a division of responsibilities is exemplary only. In some implementations, the high-speed controller 640 is coupled to a memory 620, a display 680 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 650 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to a storage device 630 and a low-speed expansion port 690. The low-speed expansion port 690, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device, such as a switch or a router, for example, through a network adapter.
[0095] The computing device 600 can be implemented in many different forms, as shown. For example, it can be implemented as a standard server 600a or multiple times as a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.
[0096] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuit systems, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be special purpose or general purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0097] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0098] The processes and logic flows described in this specification can be performed by one or more programmable processors (also referred to as data processing hardware), which execute one or more computer programs to perform functions by operating on input data and generating outputs. The processes and logic flows can also be performed by a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). For example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as a magnetic disk, a magneto-optical disk, or an optical disk, or be operably coupled to receive data from one or more mass storage devices or to transfer data to one or more mass storage devices, or both. However, a computer does not have to have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0099] To provide interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen) for displaying information to the user, and optionally a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, speech, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.
[0100] A variety of implementations have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Therefore, other implementations are within the scope of the following claims.< / st> < / st> < / st>
Claims
1. A computer-implemented method (500), It is characterized in that The computer-implemented method, when executed on data processing hardware (610), causes the data processing hardware (610) to perform operations comprising: Receiving an input audio signal (122) corresponding to an utterance (120) spoken by one or more speakers, the input audio signal (122) comprising N fixed-length audio frames; The input audio signal (122) is processed using a speech recognition model to jointly generate the following as output from the speech recognition model: a transcription (200) of the utterance (120); and one or more speaker turn markers (224), each speaker turn marker indicating a position of a respective speaker turn detected in the transcription (200) between a pair of respective adjacent terms (222); segmenting the input audio signal (122) into a plurality of N speaker segments (225) based on the one or more speaker turn markers (224) generated as output from the speech recognition model; For each speaker segment (225) of the plurality of N speaker segments (225), extracting a corresponding speaker discriminative embedding from the speaker segment (225); and Based on determining that the number of the N speaker segments (225) is greater than a threshold number M: performing pre-clustering on the speaker discriminative embeddings extracted from the N speaker segments (225) to cluster the N speaker segments (225) into a target number of pre-clustered clusters; For each corresponding pre-clustered cluster of the target number of pre-clustered clusters, determining a corresponding centroid value (265) based on the speaker discriminative embedding extracted from the speaker segments (225) clustered into the corresponding pre-clustered cluster; performing spectral clustering on the centroid values (265) determined for the target number of pre-clustered clusters to cluster the centroid values (265) into k clusters; and For each respective class (262) of the k classes, each centroid value (265) clustered into the respective class (262) is assigned a respective speaker label (250), the speaker label being different from the respective speaker label (250) assigned to the centroid value (265) clustered into each other class (262) of the k classes.
2. The computer-implemented method (500) of claim 1, It is characterized in that Wherein the operations further include annotating the transcription (200) of the utterance (120) based on the speaker label (250) assigned to each centroid value (265).
3. The computer-implemented method (500) of claim 1 or 2, It is characterized in that The operation further includes setting the target number of pre-clustering clusters to be equal to the threshold number M.
4. The computer-implemented method (500) of any one of claims 1 to 3, It is characterized in that Wherein the target number of pre-clustering clusters is less than the number of N speaker segments (225).
5. The computer-implemented method (500) of any one of claims 1 to 4, It is characterized in that The operations also include: For each of the one or more speaker turn tokens (224) generated as output from the speech recognition model, predicting a corresponding confidence value (331) for the corresponding speaker turn detected in the transcription (200); and determining a threshold number of said one or more speaker tags (224) each having said corresponding confidence value (331) that satisfies a confidence value (331) threshold, Wherein segmenting the input audio signal (122) into the plurality of N speaker segments (225) is based on determining a threshold number of the one or more speaker tags (224) each having a corresponding confidence value (331) that satisfies the confidence value (331) threshold.
6. The computer-implemented method (500) of claim 5, It is characterized in that The operations also include: determining a pairwise constraint (226) based on the confidence value (331) predicted for the speaker turn marker (224), wherein the spectral clustering performed on the centroid values (265) determined for the target number of pre-clustered clusters is subject to the pairwise constraints (226).
7. The computer-implemented method (500) of any one of claims 1 to 6, It is characterized in that in: Each speaker turn marker in the sequence of speaker turn markers (224) has a corresponding timestamp; and Segmenting the input audio signal (122) into the plurality of N speaker segments (225) based on the sequence of speaker turn markers (224) includes segmenting the input audio signal (122) into initial speaker segments (225), each initial speaker segment being delimited by the corresponding timestamps (223) of a pair of corresponding adjacent speaker turn markers (224) in the sequence of speaker turn markers (224).
8. The computer-implemented method (500) of claim 7, It is characterized in that The operations also include: for each initial speaker segment (225) having a corresponding duration exceeding a segment duration threshold, further segmenting the initial speaker segment (225) into two or more shortened duration speaker segments (225) having corresponding durations less than or equal to the segment duration threshold, The plurality of N speaker segments (225) obtained by segmenting the input audio signal (122) include: the initial speaker segments (225) having respective durations less than or equal to the segment duration threshold; and The shortened duration speaker segments (225) are further segmented from any of the initial speaker segments (225) having a corresponding duration exceeding the segment duration threshold.
9. The computer-implemented method (500) of any one of claims 1 to 8, It is characterized in that Wherein extracting the corresponding speaker discriminative embedding from the speaker segment (225) comprises: receiving the speaker segments (225) as input to a speaker encoder model (230); and The corresponding speaker discriminative embedding is generated as an output from the speaker encoder model (230).
10. The computer-implemented method (500) of claim 9, It is characterized in that The speaker encoder model (230) comprises a long short-term memory (LSTM) based speaker encoder model (230), wherein the LSTM based speaker encoder model is configured to extract the corresponding speaker discriminative embedding from each speaker segment (225).
11. The computer-implemented method (500) of any one of claims 1 to 10, It is characterized in that The speech recognition model includes a streaming transducer-based speech recognition model, and the streaming transducer-based speech recognition model includes: An audio encoder (310), the audio encoder being configured to: receiving as input a sequence of acoustic frames; and generating, at each of a plurality of time steps, a high-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames; A tag encoder (320), wherein the tag encoder is configured to: receiving as input a sequence of non-null symbols output by the last softmax layer (340); and generating a dense representation at each of the plurality of time steps (322); and a joint network (330), the joint network being configured to: receiving as input the high-level feature representation generated by the audio encoder (310) at each of the plurality of time steps and the dense representation (322) generated by the label encoder (320) at each of the plurality of time steps; and At each of the plurality of time steps, a probability distribution of possible speech recognition hypotheses (342) at the corresponding time step is generated.
12. The computer-implemented method (500) of claim 11, It is characterized in that The audio encoder (310) comprises a neural network having multiple multi-head attention layers.
13. The computer-implemented method (500) of claim 11 or 12, It is characterized in that Wherein the label encoder (320) comprises a bigram embedding lookup decoder model.
14. The computer-implemented method (500) of any one of claims 1 to 13, It is characterized in that The speech recognition model is trained on training samples, each training sample comprising a pairing of training utterances (120) spoken by two or more different speakers and corresponding true value transcriptions (200) of the training utterances (120), each true value transcription (200) being injected with a true value speaker turn marker (224), the true value speaker turn marker indicating the position in the true value transcription (200) where a speaker turn occurs.
15. The computer-implemented method (500) of claim 14, It is characterized in that The corresponding true value transcription (200) of each training sample is not labeled with any timestamp information.
16. A system (100), It is characterized in that include: Data processing hardware (610); Memory hardware (620) in communication with the data processing hardware (610) and storing instructions that, when executed by the data processing hardware (610), cause the data processing hardware (610) to perform operations, the operations comprising: Receiving an input audio signal (122) corresponding to an utterance (120) spoken by one or more speakers, the input audio signal (122) comprising N fixed-length audio frames; The input audio signal (122) is processed using a speech recognition model to jointly generate the following as output from the speech recognition model: a transcription (200) of the utterance (120); and one or more speaker turn markers (224), each speaker turn marker indicating a position of a respective speaker turn detected in the transcription (200) between a pair of respective adjacent terms (222); segmenting the input audio signal (122) into a plurality of N speaker segments (225) based on the one or more speaker turn markers (224) generated as output from the speech recognition model; For each speaker segment (225) of the plurality of N speaker segments (225), extracting a corresponding speaker discriminative embedding from the speaker segment (225); and Based on determining that the number of the N speaker segments (225) is greater than a threshold number M: performing pre-clustering on the speaker discriminative embeddings extracted from the N speaker segments (225) to cluster the N speaker segments (225) into a target number of pre-clustered clusters; For each corresponding pre-clustered cluster of the target number of pre-clustered clusters, determining a corresponding centroid value (265) based on the speaker discriminative embedding extracted from the speaker segments (225) clustered into the corresponding pre-clustered cluster; performing spectral clustering on the centroid values (265) determined for the target number of pre-clustered clusters to cluster the centroid values (265) into k clusters; and For each respective class (262) of the k classes, each centroid value (265) clustered into the respective class (262) is assigned a respective speaker label (250), the speaker label being different from the respective speaker label (250) assigned to the centroid value (265) clustered into each other class (262) of the k classes.
17. The system (100) of claim 16, It is characterized in that Wherein the operations further include annotating the transcription (200) of the utterance (120) based on the speaker label (250) assigned to each centroid value (265).
18. The system (100) according to claim 16 or 17, It is characterized in that The operation further includes setting the target number of pre-clustering clusters to be equal to the threshold number M.
19. The system (100) according to any one of claims 16 to 18, It is characterized in that Wherein the target number of pre-clustering clusters is less than the number of N speaker segments (225).
20. The system (100) according to any one of claims 16 to 19, It is characterized in that The operations also include: For each of the one or more speaker turn tokens (224) generated as output from the speech recognition model, predicting a corresponding confidence value (331) for the corresponding speaker turn detected in the transcription (200); and determining a threshold number of said one or more speaker tags (224) each having said corresponding confidence value (331) that satisfies a confidence value (331) threshold, Wherein segmenting the input audio signal (122) into the plurality of N speaker segments (225) is based on determining a threshold number of the one or more speaker tags (224) each having a corresponding confidence value (331) that satisfies the confidence value (331) threshold.
21. The system (100) of claim 20, It is characterized in that The operations also include: determining a pairwise constraint (226) based on the confidence value (331) predicted for the speaker turn marker (224), wherein the spectral clustering performed on the centroid values (265) determined for the target number of pre-clustered clusters is subject to the pairwise constraints (226).
22. The system (100) according to any one of claims 16 to 21, It is characterized in that in: Each speaker turn marker in the sequence of speaker turn markers (224) has a corresponding timestamp; and Segmenting the input audio signal (122) into the plurality of N speaker segments (225) based on the sequence of speaker turn markers (224) includes segmenting the input audio signal (122) into initial speaker segments (225), each initial speaker segment being delimited by the corresponding timestamps (223) of a pair of corresponding adjacent speaker turn markers (224) in the sequence of speaker turn markers (224).
23. The system (100) of claim 22, It is characterized in that The operations also include: for each initial speaker segment (225) having a corresponding duration exceeding a segment duration threshold, further segmenting the initial speaker segment (225) into two or more shortened duration speaker segments (225) having corresponding durations less than or equal to the segment duration threshold, The plurality of N speaker segments (225) obtained by segmenting the input audio signal (122) include: the initial speaker segments (225) having respective durations less than or equal to the segment duration threshold; and The shortened duration speaker segments (225) are further segmented from any of the initial speaker segments (225) having a corresponding duration exceeding the segment duration threshold.
24. The system (100) according to any one of claims 16 to 23, It is characterized in that Wherein extracting the corresponding speaker discriminative embedding from the speaker segment (225) comprises: receiving the speaker segments (225) as input to a speaker encoder model (230); and The corresponding speaker discriminative embedding is generated as an output from the speaker encoder model (230).
25. The system of claim 24, It is characterized in that The speaker encoder model (230) comprises a long short-term memory (LSTM) based speaker encoder model (230), wherein the LSTM based speaker encoder model is configured to extract the corresponding speaker discriminative embedding from each speaker segment (225).
26. The system (100) according to any one of claims 16 to 25, It is characterized in that The speech recognition model includes a streaming transducer-based speech recognition model, and the streaming transducer-based speech recognition model includes: An audio encoder (310), the audio encoder being configured to: receiving as input a sequence of acoustic frames; and generating, at each of a plurality of time steps, a high-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames; A tag encoder (320), wherein the tag encoder is configured to: receiving as input a sequence of non-null symbols output by the last softmax layer (340); and generating a dense representation at each of the plurality of time steps (322); and a joint network (330), the joint network being configured to: receiving as input the high-level feature representation generated by the audio encoder (310) at each of the plurality of time steps and the dense representation (322) generated by the label encoder (320) at each of the plurality of time steps; and At each of the plurality of time steps, a probability distribution of possible speech recognition hypotheses (342) at the corresponding time step is generated.
27. The system (100) of claim 26, It is characterized in that The audio encoder (310) comprises a neural network having multiple multi-head attention layers.
28. The system (100) according to claim 26 or 27, It is characterized in that Wherein the label encoder (320) comprises a bigram embedding lookup decoder model.
29. The system (100) according to any one of claims 1 to 28, It is characterized in that The speech recognition model is trained on training samples, each training sample comprising a pairing of training utterances (120) spoken by two or more different speakers and corresponding true value transcriptions (200) of the training utterances (120), each true value transcription (200) being injected with a true value speaker turn marker (224), the true value speaker turn marker indicating the position in the true value transcription (200) where a speaker turn occurs.
30. The system (100) of claim 29, It is characterized in that The corresponding true value transcription (200) of each training sample is not labeled with any timestamp information.
Citation Information
Cited By
Online speaker logging based on speaker conversion with constrained spectral clustering
CN117980991A