Detecting evaluation metrics based on evaluated speaker changes

By using the comparison of sequence transduction model and true value speaker change intervals in speaker change detection, the problem of annotation differences in the prior art affecting model performance is solved, and a more accurate and consistent speaker change detection is achieved.

CN120035859APending Publication Date: 2025-05-23GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072547.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-10-09
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, the performance of the speaker's change detection model is affected by significant differences in the annotated training data, resulting in subjective and inconsistent precise time to identify the speaker's transition points, which in turn affects the performance of the model.

Method used

By processing multi-discourse audio data using a sequence transduction model, the predicted speaker changes the lexical sequence and marks the correctness of the predicted lexical elements by comparing the change interval with the real value speaker to determine the accuracy index of the model.

Benefits of technology

This method can effectively reduce the impact of annotation differences on model performance, improve the accuracy and consistency of speaker change detection, and thus improve the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035859A_ABST
    Figure CN120035859A_ABST
Patent Text Reader

Abstract

A method (600) includes: obtaining a multi-utterance training sample (410), the multi-utterance training sample including audio data (412) characterizing utterances spoken by two or more different speakers (10); and obtaining a true value speaker change interval (414) indicative of a time interval in the audio data during which a speaker change occurs between two or more different speakers. The method further includes processing the audio data using a sequence transduction model (300) to generate a sequence of predicted speaker change lexical elements (302). For each corresponding predicted speaker change lexical element, the method includes marking the corresponding predicted speaker change lexical element as correct when the predicted speaker change lexical element overlaps with one of the true value speaker change intervals. The method further includes determining an accuracy indicator of the sequence transduction model based on the number of predicted speaker change lemmas marked as correct and the total number of predicted speaker change lemmas (442).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to evaluation based speaker change detection evaluation metrics. Background Art

[0002] Speaker change detection is a process that receives input audio data and outputs speaker turn tokens that identify speaker transition points during a conversation with multiple speakers (e.g., when one speaker stops speaking and another speaker starts speaking). Conventionally, speaker change detection maps input acoustic features to frame-level binary predictions that indicate whether a speaker change has occurred. However, models trained to perform speaker change detection can be affected by significant variance present in most annotated training data. That is, identifying the precise time when a speaker transition point occurs is highly subjective and depends on who is annotating the training data. Therefore, significant variance in annotated training data can lead to performance degradation of models performing speaker change detection. Summary of the invention

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for evaluating speaker change detection in a multi-speaker continuous conversation input audio stream. The operations include obtaining a multi-utterance sample, the multi-utterance sample including audio data representing utterances spoken by two or more different speakers. The operations also include obtaining a ground-truth speaker change interval, the ground-truth speaker change interval indicating a time interval in the audio data where a speaker change occurs between two or more different speakers. The operations also include processing the audio data using a sequence transduction model to generate a sequence of predicted speaker change tokens, each predicted speaker change token indicating a position of a corresponding speaker turn in the audio data. For each corresponding predicted speaker change token, the operations include: when the predicted speaker change token overlaps with one of the ground-truth speaker change intervals, marking the corresponding predicted speaker change token as correct. The operations also include determining an accuracy indicator for the sequence transduction model based on the number of predicted speaker change tokens marked as correct and a total number of predicted speaker change tokens in the sequence of predicted speaker change tokens.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, determining an accuracy metric for a sequence transduction model is based on a ratio between the number of predicted speaker change tokens marked as correct and the total number of predicted speaker change tokens in a sequence of predicted speaker change tokens. For each corresponding predicted speaker change token, the operation may further include: marking the corresponding predicted speaker change token as an incorrect prediction when the corresponding predicted speaker change token does not overlap with any of the true value speaker change intervals.

[0005] In some examples, the operation further includes: for each true value speaker change interval, when any of the predicted speaker change tokens overlaps with the corresponding true value speaker change interval, marking the corresponding true value speaker change interval as correctly matched, and determining a recall index of the sequence transduction model based on the duration of the true value speaker change interval marked as correctly matched and the total duration of all true value speaker change intervals. In these examples, the operation may further include determining a performance score of the sequence transduction model based on the precision index and the recall index. Here, determining the performance score includes calculating the performance score based on the following equation: 2*(precision index*recall score) / (precision index+recall score).

[0006] In some implementations, the multi-utterance training sample further includes true value speaker labels paired with the audio data, wherein the true value speaker labels each indicate a corresponding time-stamped segment in the audio data associated with a corresponding one of the utterances spoken by one of the two or more different speakers, and obtaining the true value speaker change interval includes: identifying each time interval in which two or more of the time-stamped segments overlap as a corresponding true value speaker change interval; and identifying each time gap indicating a pause between two adjacent time-stamped segments in the audio data associated with the corresponding one of the utterances spoken by the two different speakers. In these implementations, the operation may further include: determining a minimum start time and a maximum start time of the audio data based on the time-stamped segments indicated by the true value speaker labels, and omitting any predicted speaker change tokens having a time stamp earlier than the minimum start time or later than the maximum start time from determining the accuracy index of the sequence transduction model. In some examples, determining the accuracy index of the sequence transduction model is not based on any word-level speech recognition results output by the sequence transduction model. Determining an accuracy metric may not require performing a full speaker diarization on the audio data.

[0007] Another aspect of the present disclosure provides a system comprising data processing hardware and memory hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a multi-utterance sample, the multi-utterance sample including audio data representing utterances spoken by two or more different speakers. The operations also include obtaining a ground-truth speaker change interval, the ground-truth speaker change interval indicating a time interval in the audio data where a speaker change occurs between two or more different speakers. The operations also include processing the audio data using a sequence transduction model to generate a sequence of predicted speaker change tokens, each predicted speaker change token indicating a position of a corresponding speaker turn in the audio data. For each corresponding predicted speaker change token, the operations include: when the predicted speaker change token overlaps with one of the ground-truth speaker change intervals, marking the corresponding predicted speaker change token as correct. The operations also include determining an accuracy indicator for the sequence transduction model based on the number of predicted speaker change tokens marked as correct and a total number of predicted speaker change tokens in the sequence of predicted speaker change tokens.

[0008] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, determining an accuracy metric for a sequence transduction model is based on a ratio between the number of predicted speaker change tokens marked as correct and the total number of predicted speaker change tokens in a sequence of predicted speaker change tokens. For each corresponding predicted speaker change token, the operation may further include: marking the corresponding predicted speaker change token as an incorrect prediction when the corresponding predicted speaker change token does not overlap with any of the true value speaker change intervals.

[0009] In some examples, the operation further includes: for each true value speaker change interval, when any of the predicted speaker change tokens overlaps with the corresponding true value speaker change interval, marking the corresponding true value speaker change interval as correctly matched, and determining a recall index of the sequence transduction model based on the duration of the true value speaker change interval marked as correctly matched and the total duration of all true value speaker change intervals. In these examples, the operation may further include determining a performance score of the sequence transduction model based on the precision index and the recall index. Here, determining the performance score includes calculating the performance score based on the following equation: 2*(precision index*recall score) / (precision index+recall score).

[0010] In some implementations, the multi-utterance training sample further includes true value speaker labels paired with the audio data, wherein the true value speaker labels each indicate a corresponding time-stamped segment in the audio data associated with a corresponding one of the utterances spoken by one of the two or more different speakers, and obtaining the true value speaker change interval includes: identifying each time interval in which two or more of the time-stamped segments overlap as a corresponding true value speaker change interval; and identifying each time gap indicating a pause between two adjacent time-stamped segments in the audio data associated with the corresponding one of the utterances spoken by the two different speakers. In these implementations, the operation may further include: determining a minimum start time and a maximum start time of the audio data based on the time-stamped segments indicated by the true value speaker labels, and omitting any predicted speaker change tokens having a time stamp earlier than the minimum start time or later than the maximum start time from determining the accuracy index of the sequence transduction model. In some examples, determining the accuracy index of the sequence transduction model is not based on any word-level speech recognition results output by the sequence transduction model. Determining an accuracy metric may not require performing full speaker classification on the audio data.

[0011] The details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic diagram of an example automatic speech recognition system including an automatic speech recognition model and a sequence transduction model.

[0013] Figure 2 is a schematic diagram of an example automatic speech recognition model.

[0014] Figure 3 is a schematic diagram of an example sequence transduction model.

[0015] Figure 4 is a schematic diagram of an example two-stage training process for training a sequence transduction model.

[0016] Figure 5 is a graphical view of the multiple components used to calculate the precision and recall metrics.

[0017] Figure 6 is a flow chart of an example arrangement of operations of a computer-implemented method of evaluating a speaker change detection evaluation metric based on evaluation.

[0018] Figure 7 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.

[0019] Like reference numerals in the various drawings indicate like elements. DETAILED DESCRIPTION

[0020] In addition to converting an input sequence into an output sequence, a sequence transduction model is also constructed to detect special input conditions and generate special outputs (e.g., special output tokens or other types of indications) when the special input conditions are detected. That is, a sequence transduction model can be constructed and trained to process a sequence of input data to generate a predicted output sequence that includes, in addition to other normal / common predicted outputs (e.g., graphemes, word fragments, and / or words), special outputs when the sequence transduction model detects corresponding special input conditions in the input data. For example, a sequence transduction model can process input audio features and output a transcription representing the input audio features and a speaker change token "indicating a corresponding speaker turn in a multi-speaker conversation." <st>Here, speaker turn refers to a point in time during a conversation when one speaker stops speaking and / or another speaker starts speaking. Thus, a speaker change token indicates a point in time during a conversation where a speaker turn occurs.

[0021] However, training conventional sequence transduction models has several limitations. For example, conventional systems require accurate temporal information of speaker change points in the training data, which is difficult because deciding where to mark speaker change points is a very subjective process for human annotators. In addition, methods that use purely acoustic information ignore the rich semantic information in the audio signal to identify speaker change points.

[0022] Therefore, the implementation of the present invention relates to a method and system for performing an evaluation-based speaker change detection evaluation method. Specifically, a training process trains a sequence transduction model by: obtaining a multi-utterance training sample, the multi-utterance training sample including audio data representing utterances spoken by two or more different speakers; and obtaining a true value speaker change interval indicating a time interval in which a speaker change occurs in the audio data. The sequence transduction model processes the audio data to generate a predicted speaker change word sequence. Thereafter, when the predicted speaker change word overlaps with one of the true value speaker change intervals, the training process marks each predicted speaker change word in the predicted speaker change word sequence as correct. The training process determines an accuracy metric of the sequence transduction model based on the number of predicted speaker change words marked as correct and the total number of predicted speaker change words in the predicted speaker change word sequence, and uses the accuracy metric to train the sequence transduction model. As will become apparent, the training process can also use a recall metric to train the sequence transduction model in addition to or in lieu of the accuracy metric.

[0023] refer to Figure 1 , the example system 100 includes a user device 110 that captures speech utterances 106 spoken by two or more speakers 10, 10a-n during a conversation and communicates with a remote system 140 via a network 130. The remote system 140 can be a distributed system (e.g., a cloud computing environment) with scalable / elastic resources 142. The resources 142 include computing resources 144 (e.g., data processing hardware) and / or storage resources 146 (e.g., memory hardware). The user device 110 includes data processing hardware 112 and memory hardware 114. The user device 110 includes, but is not limited to, desktop computing devices and mobile computing devices such as laptops, tablet computers, smart phones, smart speakers / displays, smart appliances, Internet of Things (IoT) devices, and wearable computing devices (e.g., headphones and / or watches). The user device 110 may include an audio capture device (e.g., a microphone) for capturing speech utterances 106 from two or more speakers 10 and converting the speech utterances into an input audio stream (e.g., audio data or a sequence of acoustic frames) 108.

[0024] The user device 110 and / or the cloud computing environment 140 may execute an automatic speech recognition (ASR) system 118. In some implementations, the user device 110 is configured to execute a portion of the ASR system 118 locally (e.g., using the data processing hardware 112), while the remainder of the ASR system 118 is executed at the cloud computing environment 140 (e.g., using the data processing hardware 144). Alternatively, the ASR system 118 may be executed entirely at the user device 110 or the cloud computing environment 140.

[0025] The ASR system 118 is configured to receive an input audio stream 108 corresponding to a multi-speaker continuous conversation and generate a speaker change detection output 125. More specifically, the ASR system 118 includes an ASR model 200 configured to receive the input audio stream 108 and generate a transcription (e.g., speech recognition result / hypothesis) 120 as an output based on the input audio stream 108. In addition, the ASR system 118 includes a sequence transduction model (i.e., speaker change detection model) 300 configured to receive the input audio stream 108 and generate a sequence of predicted speaker change tokens 302 as an output, each token indicating a position of a corresponding speaker turn in the input audio stream 108. The ASR system 118 generates a speaker change detection output 125 that includes the transcription 120 generated by the ASR model 200 and the sequence of predicted speaker change tokens 302 generated by the sequence transduction model 300. In some examples, the speaker change detection output 125 includes a sequence of timestamps 122 corresponding to the transcription 120 and a sequence of predicted speaker change tokens 302. Thus, in these examples, the ASR system 118 can align the transcription 120 and the sequence of predicted speaker change tokens 302 based on the corresponding sequence of timestamps 122.

[0026] For example, in the example shown, a first speaker 10a speaks a first utterance 106a "How are you doing (how are you doing)", and a second speaker 10b speaks a second utterance 106b "I am good (I am good)". The ASR system 118 receives an input audio stream 108 corresponding to a multi-speaker continuous conversation input (e.g., the first utterance 106a and the second utterance 106b) spoken by the first speaker 10a and the second speaker 10b. In the example shown, the ASR model 200 generates a transcription 120 of "how areyou doing I am good (how are you doing I am good)", and the sequence transduction model 300 generates a predicted speaker change word-metaphor 302 sequence indicating speaker turns at a fifth timestamp 122 (e.g., T=5) and a ninth timestamp 122 (e.g., T=9). Notably, the predicted speaker change word-element 302 at the fifth timestamp indicates a transition point where the first speaker 10a stops speaking and the second speaker 10b starts speaking, and the predicted speaker change word-element 302 at the ninth timestamp indicates a transition point where the second speaker 10b stops speaking.

[0027] In some implementations, two or more speakers 10 and a user device 110 may be located within an environment (e.g., a room), wherein the user device 110 is configured to capture speech utterances 106 spoken by the two or more speakers 10 and convert the speech utterances into an input audio stream 108. For example, the two or more speakers 10 may correspond to colleagues having a conversation during a meeting, and the user device 110 may record the speech utterances 106 and convert the speech utterances into an input audio stream 108. The user device 110 may then provide the input audio stream 108 to the ASR system 118 to generate a speaker change detection output 125 including a speech recognition result 120 and a predicted sequence of speaker change tokens 302.

[0028] In some examples, at least a portion of the speech utterances 106 conveyed in the input audio stream 108 overlap, such that at a given instant in time, at least two speakers 10 are speaking simultaneously. Notably, when the sequence of acoustic frames 108 is provided to the ASR system 118 as input, the number N of the two or more speakers 10 may be unknown, whereby the ASR system 118 predicts the number N of the two or more speakers 10. In some implementations, the user device 110 is remote from one or more of the two or more speakers 10. For example, the user device 110 may include a remote device (e.g., a network server) that captures the speech utterances 106 from the two or more speakers 10 as participants in a phone call or video conference. In this case, each speaker 10 will speak to their own user device 110, which captures the speech utterances 106 and provides the speech utterances to the remote user device to convert the speech utterances 106 into the input audio stream 108. Of course, in this case, the speech utterance 106 may be processed at each of the user devices 110 and converted into a corresponding input audio stream 108, which is transmitted to the remote user device, which may further process the input audio stream 108 which is provided as input to the ASR system 118.

[0029] Reference now Figure 2 In some implementations, the example ASR model 200 includes a recurrent neural network transducer (RNN-T) model architecture that complies with latency constraints for interactive applications. The use of the RNN-T model architecture is exemplary only, as the ASR model 200 may include other architectures such as transformer-transducer and conformer-transducer model architectures. The RNN-T model 200 provides a small computational footprint and uses less memory requirements than conventional ASR architectures, making the RNN-T model architecture suitable for performing speech recognition entirely on the user device 110 (e.g., without the need to communicate with a remote server). The RNN-T model 200 includes an encoder network (e.g., an audio encoder) 210, a prediction network 220, and a joint network 230. The encoder network 210, which is generally similar to an acoustic model (AM) in a conventional ASR system, includes a stack of self-attention layers (e.g., Conformer layers or Transformer layers) or a recurrent network of stacked long short-term memory (LSTM) layers. For example, the audio encoder 210 reads a d-dimensional feature vector (eg, an acoustic frame 108 ( Figure 1 ))Sequence x = (x 1 ,x 2 ,...,x T ),in And at each output step a higher-order feature representation (e.g., audio encoding) is produced. This higher-order feature representation is represented as

[0030] Similarly, the prediction network 220 is also an LSTM network, which, like a language model (LM), converts the non-empty symbol sequence y currently output by the final Softmax layer 240 into 0 ,...,y ui-1 Processing as dense representation Finally, using the RNN-T model architecture, the representations produced by the encoder network 210 and the prediction / decoder network 220 are combined by the joint network 230. The prediction network 220 can be replaced by an embedding lookup table to improve latency by outputting the sparse embeddings looked up instead of processing dense representations. The joint network 230 then predicts It is the distribution about the next output symbol. In other words, the joint network 230 generates a probability distribution about possible speech recognition hypotheses at each output step (e.g., time step). Here, "possible speech recognition hypotheses" correspond to a set of output labels that each represent a symbol / character using a specified natural language. For example, when the natural language is English, the set of output labels may include twenty-seven (27) symbols, for example, one label for each of the 26 letters in the English alphabet, and one label specifies a space. Therefore, the joint network 230 can output a set of values, which indicates the possibility of each output label in a predetermined set of output labels to occur. This set of values ​​can be a vector and can indicate a probability distribution about a set of output labels. In some cases, the output label is a grapheme (e.g., a separate character, and possible punctuation and other symbols), but the set of output labels is not limited to this. For example, as an addition or replacement of a grapheme, the set of output labels may include a word piece, a phoneme and / or an entire word. The output distribution of the joint network 230 may include a posterior probability value for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output y of the joint network 230 is i 100 different probability values ​​may be included, one for each output label. The probability distribution may then be used (e.g., by a Softmax layer 240) to select candidate orthographic elements (e.g., graphemes, fragments, and / or words) and assign scores to them in a beam search process for use in determining speech recognition results (e.g., transcriptions) 120 ( Figure 1 ).

[0031] The Softmax layer 240 may employ any technique to select the output label / symbol with the highest probability in the distribution as the next output symbol predicted by the RNN-T model 200 at the corresponding output step. In this manner, the RNN-T model 200 does not make conditional independence assumptions, but rather the prediction of each symbol is conditioned not only on the acoustics, but also on the sequence of labels output so far. The RNN-T model 200 does assume that the output symbol is independent of future acoustic frames 108, which allows the RNN-T model to be employed in a streaming manner, a non-streaming manner, or some combination thereof.

[0032] In some examples, the audio encoder 210 of the RNN-T model includes multiple multi-head (e.g., 8-head) self-attention layers. For example, the multiple multi-head self-attention layers may include Conformer layers (e.g., Conformer encoders), Transformer layers, performer layers, convolutional layers (including lightweight convolutional layers), or any other type of multi-head self-attention layers. The multiple multi-head self-attention layers may include any number of layers, such as 16 layers. In addition, the audio encoder 210 may operate in a streaming manner (e.g., once an initial higher-order feature representation is generated, the audio encoder 210 outputs the initial higher-order feature representation), a non-streaming manner (e.g., the audio encoder 210 outputs subsequent higher-order feature representations by processing additional right contexts to improve the initial higher-order feature representations), or a combination of a streaming manner and a non-streaming manner.

[0033] Reference now Figure 3 In some implementations, the sequence transduction model 300 includes a Transformer-Transducer (TT) architecture. The use of the RNN-T model architecture is exemplary only, as the sequence transduction model 300 may include other architectures such as RNN-T and transformer-transducer model architectures. The sequence transduction model 300 includes an encoder network (e.g., an audio encoder) 310, a label encoder 320, a joint network 330, and a Softmax layer 340. The audio encoder 310 may include a stack of multi-head self-attention layers, such as 15 transformer layers, each with 32 left context frames and 0 right context frames. In addition, the audio encoder 310 includes a stacking layer after the second transformer layer to change the frame rate from 30 milliseconds to 90 milliseconds, and a destacking layer after the thirteenth transformer layer to change the frame rate from 90 milliseconds back to 30 milliseconds. For example, the audio encoder 310 reads a d-dimensional feature vector (e.g., an acoustic frame 108 ( Figure 1 ))Sequence x = (x 1 ,x 2 ,···,x T ),in And at each output step a higher-order feature representation (e.g., audio encoding) is produced. This higher-order feature representation is represented as That is, the sequence transduction model 300 extracts 128-dimensional logarithmic Mel-frequency cepstral coefficients, stacks them every 4 frames, and subsamples them every 3 frames to generate a 512-dimensional acoustic feature vector with a stride of 30 milliseconds as an input to the audio encoder 310 .

[0034] The label encoder 320 is, for example, a long short-term network (LSTM) network having a single 128-dimensional LSTM layer. The label encoder 320 receives the non-empty symbol sequence output by the final Softmax layer 340 and outputs a label code Ih u . The joint network 230 includes a stack of fully connected layers and a projection layer. The projection layer projects the audio encodings from the audio encoder 310 and the label encodings from the label encoder 320 to produce a probability distribution over possible speech recognition hypotheses. Here, "possible speech recognition hypotheses" correspond to a set of output labels that each represent a symbol / character using a specified natural language. For example, when the natural language is English, the set of output labels may include twenty-seven (75) symbols, such as one label for each of the 26 letters in the English alphabet, punctuation marks, special symbols (e.g., "$"), the predicted speaker change marker 302 ( Figure 1 ), and a label for a designated space. Thus, the joint network 330 may output a set of values ​​indicating the likelihood of each of a predetermined set of output labels occurring. This set of values ​​may be a vector and may indicate a probability distribution over a set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and possibly punctuation and other symbols), but the set of output labels is not limited thereto. For example, the set of output labels may include word fragments, phonemes, and / or entire words in addition to or in lieu of graphemes. The output distribution of the joint network 330 may include a posterior probability value for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output y of the joint network 330 may be y. i This can include 100 different probability values, one for each output label.

[0035] The probability distribution may then be used (eg, by a Softmax layer 340) to select candidate orthographic elements (eg, graphemes, fragments, and / or words) and assign scores thereto in a beam search process for use in determining a sequence output 345. Figure 4 Describing in more detail, the sequence output 345 generated by the sequence transduction model 300 includes a transcription 342 (or a probability distribution over possible speech recognition hypotheses) and a predicted sequence of speaker change tokens 302 during training. Specifically, each possible speech recognition hypothesis is associated with a corresponding sequence of predicted speaker change tokens 302. In some examples, each transcription 342 includes a predicted sequence of speaker change tokens 302 embedded directly in the transcription 342, e.g., "how are you doing <st>I am good <st>(Are you OK <st>I am fine <st>)”. On the other hand, Figure 1 As shown, the sequence transduction model 300 outputs only the predicted sequence of speaker-changed word-grams 302 , and not the transcription 342 .

[0036] Figure 4 An example two-stage training process 400 for training the sequence transduction model 300 is shown. The example two-stage training process 400 (also referred to simply as "training process 400") can be performed at the user device 110 and / or the cloud computing environment 140. The training process 400 obtains a plurality of multi-utterance training samples 410 for training the sequence transduction model 300. As will become apparent, the two-stage training process 400 trains the sequence transduction model 300 during a first stage by determining a negative log-likelihood loss term 422 for each of the plurality of multi-utterance training samples 410, and updates parameters of the sequence transduction model 300 based on the negative log-likelihood loss term 422. Thereafter, during the second stage, the two-stage training process 400 fine-tunes the sequence transduction model 300 trained during the first stage by determining the precision loss 442 and the recall loss 444 for each of the plurality of multi-utterance training samples 410 , and updates the parameters of the sequence transduction model 300 based on the precision loss 442 and the recall loss 444 .

[0037] Each multi-utterance training sample 410 includes audio data 412 representing utterances spoken by two or more different speakers. The audio data 412 of each multi-utterance training sample 410 can be paired with a corresponding true value speaker change interval 414, which indicates a time interval in the audio data 412 where a speaker change occurs between two or more different speakers. For example, in a conversation in which speaker A speaks from 0.1 to 10.5 seconds and speaker B speaks from 10.8 to 15.3 seconds, the true value speaker change interval 414 indicates that the time interval in which the speaker change occurs is 10.5 to 10.8 seconds. Therefore, in this example, the training process 400 will be correct for any predicted speaker change tokens during the time interval of 10.5 to 10.8. In addition, the training process 400 can apply a buffer (e.g., a buffer of 250 milliseconds) so that in the above example, the time interval is 10.75 to 11.05 seconds instead of 10.5 to 10.8 seconds. In addition, the audio data 412 of each multi-utterance training sample 410 can be paired with a ground-truth transcription 416, which indicates a textual representation of what was said in the audio data 412. That is, as described above, the sequence transduction model 300 can optionally output the transcription 342 during the training process 400.

[0038] In some implementations, each multi-utterance training sample 410 includes a true value speaker label 418 paired with the audio data 412, such that each true value speaker label 418 indicates a corresponding time-stamped segment in the audio data 412 associated with a corresponding one (or more) of the utterances spoken by one of the two or more different speakers. In these implementations, the training process 400 obtains the true value speaker change interval 414 by identifying each time interval in which two or more of the time-stamped segments overlap as a corresponding true value speaker change interval 414, and identifying each time gap indicating a pause between two adjacent time-stamped segments in the audio data associated with corresponding utterances in the utterances spoken by two different speakers as a corresponding true value speaker change interval.

[0039] For example, Figure 5 Depicts the training process 400( Figure 4 ), where the x-axis represents timestamps that increase from left to right (e.g., Tmin-Tmax). The graphical view 500 includes a first ground-truth speaker label 418, 418a indicating when speaker A is speaking, a second ground-truth speaker label 418, 418b indicating when speaker B is speaking, and a third ground-truth speaker label 418, 418c indicating when speaker C is speaking. In addition, the training process 400 ( Figure 4 ) can obtain the true value speaker change interval 414 by identifying the location where two or more true value speaker labels 418 overlap as the corresponding true value speaker change interval 414. For example, in the example shown, the first true value speaker label 418a and the third true value speaker label 418c overlap between timestamps T2 and T3, so that the training process 400 identifies the time interval between timestamps T2 and T3 as the corresponding true value speaker change interval 414. In some implementations, the training process 400 obtains the true value speaker change interval 414 by identifying where there is a time gap indicating a pause between two adjacent true value speaker labels 418. In the example shown, there is a time gap between timestamps T4 and T5 indicating a pause between the first true value speaker label 418a and the second true value speaker label 418b, so that the training process 400 identifies the time interval between timestamps T4 and T5 as the corresponding true value speaker change interval 414.

[0040] In some implementations, the training process 400 ( Figure 4 ) determines the minimum start time (Tmin) and the maximum start time (Tmax) of the audio data 412 based on the time-stamped segments indicated by the true value speaker labels 418. In these implementations, the training process 400 derives the accuracy metric 442 ( Figure 4 ) is omitted from the determination of . Figure 5 In the example shown, the training process 400 will omit any predicted speaker change tokens 302 that are to the left of Tmin or to the right of Tmax.

[0041] Return to reference Figure 4 During the first stage of the training process 400, the sequence transduction model 300 processes the audio data 412 of each multi-utterance training sample 410 to generate a corresponding first-stage sequence output 345, 345a, which includes a corresponding first-stage transcription 342, 342a and a corresponding first-stage predicted speaker change word 302, 302a sequence. The sequence transduction model 300 can determine whether to output the predicted speaker change word 302 based on the corresponding semantic information of the transcription 342. Thereafter, the log-likelihood loss function module 420 determines a negative log-likelihood loss term 422 based on comparing the corresponding first-stage output sequence 345a with the associated true value speaker change interval 414 and the associated true value transcription 416. That is, the log-likelihood loss function module 420 compares the corresponding first-stage transcription 342a with the corresponding true-value transcription 416, and compares the corresponding first-stage predicted sequence of speaker change tokens 302a with the corresponding true-value speaker change intervals 414 to determine a negative log-likelihood loss term 422. More specifically, the log-likelihood loss function module 420 determines whether the first-stage predicted speaker change tokens 302a overlap with any of the true-value speaker change intervals 414. Here, the overlap between the first-stage predicted speaker change tokens 302a and the true-value speaker change intervals 414 indicates a correct prediction by the sequence transduction model 300. The first stage of the training process 400 trains the sequence transduction model 300 by updating the parameters of the sequence transduction model 300 based on the negative log-likelihood loss term 422 determined for each multi-utterance training sample 410.

[0042] Thereafter, during the second phase, the training process 400 warm-starts the sequence transduction model 300 trained during the first phase of the training process 400. After warm-starting the sequence transduction model 300, the second phase of the training process 400 fine-tunes the model using the same plurality of multi-utterance training samples 410 used to train the sequence transduction model 300 during the first phase. Notably, during the second phase of the training process 400, the sequence transduction model 300 uses updated parameters generated by the first phase of the training process 400. That is, the sequence transduction model 300 processes the audio data 412 of each multi-utterance training sample 410 to generate a corresponding second phase sequence output 345, 345b, which includes a corresponding second phase transcription 342, 342b and a corresponding second phase predicted sequence of speaker change tokens 302, 302b.

[0043] The tagger 430 receives the second-stage sequence output 345b generated by the sequence transduction model 300 for each of the multi-utterance training samples 410 and performs a beam search to select the N best transcriptions from the second-stage transcriptions 342b and a sequence of second-stage predicted speaker change tokens 302b corresponding to the selected N best transcriptions. For example, the second-stage transcriptions 342b may include 10 candidate transcriptions based on the corresponding multi-utterance training samples 410, whereby the tagger 430 selects the top 3 candidate transcriptions with the largest confidence value scores and 3 corresponding second-stage predicted speaker change tokens 302b sequences. For each second-stage predicted speaker change token 302 (e.g., selected by the tagger 430), the tagger 430 generates a token label 432 that labels the corresponding second-stage predicted speaker change token 302b as correct, falsely accepted, or falsely rejected.

[0044] The correct word-gram label 432 indicates that the second-stage predicted speaker change word-gram 302b output by the sequence transduction model 300 correctly predicts the speaker change in the audio data 412. That is, when the corresponding second-stage predicted speaker change word-gram 302b overlaps with one of the true value speaker change intervals 414, the labeler 430 generates a word-gram label 432 that labels the corresponding second-stage predicted speaker change word-gram 302b as correct. For example, Figure 5 As shown, the sequence transduction model 300 outputs the corresponding second-stage predicted speaker change word-gram 302b, which indicates that the speaker change occurring between timestamps T8 and T9 overlaps with the true value speaker change interval 414 occurring between timestamps T8 and T9. Therefore, the word-gram label 432 between timestamps T8 and T9 indicates that the corresponding second-stage predicted speaker change word-gram 302b is correct (indicated by "X").

[0045] In some examples, the false acceptance word-unit label 432 indicates that the sequence transduction model 300 outputs the corresponding second-stage predicted speaker change word-unit 302b when no speaker change actually occurs in the audio data 412 (e.g., as shown by the true value speaker change interval 414). Here, when the corresponding second-stage predicted speaker change word-unit 302b does not overlap with any of the true value speaker change intervals 414, the marker 430 marks the corresponding second-stage predicted speaker change word-unit 302b as a false prediction. Figure 5 As shown, the sequence transduction model 300 outputs the corresponding second-stage predicted speaker change word-gram 302b, which indicates that the speaker change occurring between timestamps T1 and T2 does not overlap with any true value speaker change interval 414. Therefore, the word-gram label 432 between timestamps T1 and T2 indicates that the corresponding second-stage predicted speaker change word-gram 302b is a wrong prediction (indicated by "FA").

[0046] In yet other examples, the false acceptance word-unit label 432 indicates that the sequence transduction model 300 fails to output the corresponding second-stage predicted speaker change word-unit 302b when no speaker change actually occurs in the audio data 412 (e.g., as shown by the true value speaker change interval 414). Here, the tagger 430 generates a false rejection word-unit label 432 when the true value speaker change interval 414 occurs and no second-stage predicted speaker change word-unit 302b overlaps with the true value speaker change interval 414. Figure 5 As shown, there is a corresponding true value speaker change interval 414 between time stamps T4 and T5, and the sequence transduction model 300 does not output any speaker change word-gram 302b predicted by the second stage between time stamps T4 and T5. Therefore, the word-gram label 432 between time T4 and T5 indicates a false rejection prediction (indicated by "FR").

[0047] The word-level loss function module 440 receives the word-level label 432 generated by the tagger 430 for each multi-utterance training sample 410, and determines a precision index 442 and a recall index 444. That is, the word-level loss function module 440 determines the precision index 442 based on the number of speaker change word-levels 302b predicted by the second stage that are marked as correct and the total number of speaker change word-levels 302b predicted by the second stage in the sequence of speaker change word-levels 302b predicted by the second stage. In other words, the word-level loss function module 440 determines the precision index 442 based on the ratio between the number of speaker change word-levels 302b predicted by the second stage that are marked as correct and the total number of speaker change word-levels 302b predicted by the second stage in the sequence of speaker change word-levels 302b predicted by the second stage. Although Figure 4 Although not shown, the training process 400 may train the sequence transduction model 300 by updating the parameters of the sequence transduction model 300 directly based on the accuracy indicator 442 .

[0048] In some implementations, the tagger 430 tags each true value speaker change interval 414. That is, when any of the speaker change tokens 302b predicted by the second stage overlaps with the corresponding true value speaker change interval 434, the tagger 430 generates an interval token 434 indicating a correct match. On the other hand, when no speaker change token 302b predicted by the second stage overlaps with the corresponding true value speaker change interval 434, the tagger 430 generates an interval token 434 indicating an incorrect match. The token-level loss function module 440 receives the interval tokens 434 generated by the tagger 430 for each multi-utterance training sample 410. In some examples, the token-level loss function module 440 determines a recall index 444 of the sequence transduction model 300 based on the number of true value speaker change intervals 414 that are marked as being correctly matched and the total number of all true value speaker change intervals 414. In other examples, the word-level loss function module 440 determines the recall index 444 of the sequence transduction model 300 based on the duration of the true value speaker change interval marked as correctly matched and the total duration of all true value speaker change intervals. Determining the recall index based on duration is beneficial to the multi-utterance training samples 410 with longer speaker change intervals. Figure 4 Not shown, the training process 400 may train the sequence transduction model 300 by updating parameters of the sequence transduction model 300 directly based on the recall metric 444 .

[0049] In some implementations, M may represent the number of multi-utterance training samples 410, and N represents the number of hypotheses per training sample, such that H ij is the jth hypothesis of the i-th multi-utterance training sample 410. In addition, i is between [1, M], j is between [1, N], and P ij represents the sequence transduction model 300 generated by H ij The associated probability score, and R ij is the reference transcription. Therefore, the word-level loss function module 440 determines all H according to the following ij and R ij The minimum edit distance alignment (i.e., loss) between :

[0050]

[0051] In equations 1 and 2, r and h represent R ij and H ij In addition, k ≥ 1 controls the prediction <st>Tolerance for the offset at that time, such that if k = 1, the training process 400 expects the reference and the predicted <st>If k>1, the training process 400 allows a pair of reference and predicted <st>The maximum offset of k tokens between tokens for them to be considered correctly aligned.

[0052] In some examples, the loss combination module 450 receives the precision metric 442 and the recall metric 444, and determines a performance score 452 of the sequence transduction model 300 based on the precision metric 442 and the recall metric 444. For example, determining the performance score 452 can include calculating the performance score based on the following equation: "2*(precision metric*recall score) / (precision metric+recall score)". In some implementations, the loss combination module 450 determines the precision metric 442 based on the precision metric 442, the recall metric 444, and the negative log-likelihood loss term 422. The training process 400 trains the sequence transduction model 300 by updating the parameters of the sequence transduction model 300 based on the performance score 452. For example, the loss combination module 450 determines the word-level loss according to:

[0053]

[0054] In Equation 3, FA ij represents the number of false acceptance errors, FR ij represents the number of false rejection errors, W ij Indicates the number of spoken word errors, Q ij Represents R ij The total number of tokens in , and α, β, and γ control the influence (i.e., weight) of each subcomponent. In some examples, β and γ are significantly larger than γ to reduce the speaker-changing insertion and deletion rates. Therefore, the final per-batch training loss is expressed as:

[0055]

[0056] In Equation 4, -logP(y|x) is the negative log-likelihood of the true-valued transcription Y conditioned on the input acoustic frame X, thus acting as a regularization term. In addition, λ controls the weight of the negative likelihood loss.

[0057] Advantageously, in some examples, the training process 400 determines the accuracy metric 442 of the sequence transduction model 300 by determining the accuracy metric 442 that is not based on any word-level speech recognition results (e.g., transcription 342) output by the sequence transduction model 300. In other examples, determining the accuracy metric 442 does not require the training process 400 to perform a full speaker classification on the audio data 412. Instead, the training process 400 can simply output the predicted speaker change tokens 302 and the transcription 342 without assigning a speaker label to each frame or token of the transcription 342.

[0058] Figure 6 A flowchart including an example arrangement of operations of a computer-implemented method for evaluating speaker change detection in a multi-speaker continuous conversation input audio stream. The method 600 may use a computer program stored in memory hardware 720 ( Figure 7 ) on the data processing hardware 710 ( Figure 7 ) is executed on the computing device 700 ( Figure 7 ) corresponding to Figure 1 on the user device 110 and / or the cloud computing environment 140 .

[0059] At operation 602, the method 600 includes obtaining a multi-utterance training sample 410, which includes audio data 412 representing utterances spoken by two or more different speakers 10. At operation 604, the method 600 includes obtaining a true value speaker change interval 414, which indicates a time interval in the audio data 412 where a speaker change occurs between two or more different speakers. At operation 606, the method 600 includes processing the audio data 412 using the sequence transduction model 300 to generate a sequence of predicted speaker change tokens 302, each predicted speaker change token indicating the position of a corresponding speaker turn in the audio data. At operation 608, for each corresponding predicted speaker change token 302, the method 600 includes: when the predicted speaker change token 302 overlaps with one of the true value speaker change intervals 414, marking the corresponding predicted speaker change token 302 as correct. At operation 610, the method 600 includes determining a precision metric 442 of the sequence transduction model 300 based on the number of predicted speaker change tokens 302 marked as correct and the total number of predicted speaker change tokens 302 in the sequence of predicted speaker change tokens 302. In some examples, the method 600 includes determining a recall metric 444 of the sequence transduction model 300 based on the duration of the true value speaker change intervals 414 marked as correctly matched and the total duration of all true value speaker change intervals 414 (in addition to or in lieu of the precision metric 442).

[0060] Figure 7 is a schematic diagram of an example computing device 700 that can be used to implement the systems and methods described in this document. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit implementations of the inventions described and / or claimed in this document.

[0061] The computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and the storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 is interconnected using various buses and can be installed on a common motherboard or installed in other ways as appropriate. The processor 710 can process instructions for execution within the computing device 700, including instructions stored in the memory 720 or on the storage device 730, to display graphical information of a graphical user interface (GUI) on an external input / output device (such as a display 780 coupled to the high-speed interface 740). In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memories can be used as appropriate. In addition, multiple computing devices 700 can be connected (for example, as a server bank, a blade server group, or a multi-processor system), where each device provides part of the necessary operations.

[0062] The memory 720 stores information non-temporarily within the computing device 700. The memory 720 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-temporary memory 720 may be a physical device for storing programs (e.g., instruction sequences) or data (e.g., program state information) on a temporary or permanent basis for use by the computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware, such as bootloaders). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0063] The storage device 730 is capable of providing mass storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various implementations, the storage device 730 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices (including devices in a storage area network or other configuration). In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as the memory 720, the storage device 730, or a memory on the processor 710.

[0064] The high-speed controller 740 manages bandwidth-intensive operations of the computing device 700, while the low-speed controller 760 manages less bandwidth-intensive operations. This division of responsibilities is exemplary only. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., through a graphics processor or accelerator), and the high-speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and the low-speed expansion port 790. The low-speed expansion port 790, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner; or to a networking device, such as a switch or a router, for example, through a network adapter.

[0065] As shown, computing device 700 can be implemented in a variety of different forms. For example, the computing device can be implemented as a standard server 700a or multiple implementations in a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.

[0066] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuit systems, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor that can be either special purpose or general purpose and can be coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.

[0067] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or in assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0068] The processes and logic flows described in this specification can be performed by one or more programmable processors (also referred to as data processing hardware), which execute one or more computer programs to perform functions by operating on input data and generating outputs. The processes and logic flows can also be performed by a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). For example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as a magnetic disk, a magneto-optical disk, or an optical disk, or be operably coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer does not have to have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0069] To provide interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device for displaying information to the user—e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen—and may have a keyboard and pointing device—e.g., a mouse or trackball—through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.

[0070] A variety of implementations have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Therefore, other implementations are within the scope of the appended claims.< / st> < / st> < / st> < / st> < / st> < / st> < / st> < / st>

Claims

1. A computer-implemented method (600), It is characterized in that The method, when executed on data processing hardware (710), causes the data processing hardware (710) to perform operations for evaluating speaker change detection in a multi-speaker continuous conversation input audio stream, the operations comprising: Obtaining a multi-utterance training sample (410), the multi-utterance training sample comprising audio data (412) representing utterances spoken by two or more different speakers (10); obtaining a true value speaker change interval (414), the true value speaker change interval indicating a time interval in the audio data (412) where a speaker change occurs between the two or more different speakers (10); Processing the audio data (412) using a sequence transduction model (300) to generate a sequence of predicted speaker change tokens (302), each predicted speaker change token indicating a position of a corresponding speaker turn in the audio data (412); for each corresponding predicted speaker change word-gram (302), marking the corresponding predicted speaker change word-gram (302) as correct when the predicted speaker change word-gram (302) overlaps with one of the true value speaker change intervals (414); and An accuracy indicator (442) of the sequence transduction model (300) is determined based on the number of the predicted speaker change tokens (302) marked as correct and the total number of the predicted speaker change tokens (302) in the sequence of predicted speaker change tokens (302).

2. The computer-implemented method (600) of claim 1, It is characterized in that Determining the accuracy metric (442) of the sequence transduction model (332) is based on a ratio between the number of the predicted speaker change word-grams (302) marked as correct and the total number of the predicted speaker change word-grams (302) in the sequence of predicted speaker change word-grams (302).

3. The computer-implemented method (600) of any preceding claim, It is characterized in that The operation further includes: for each corresponding predicted speaker change word (302), when the corresponding predicted speaker change word (302) does not overlap with any of the true value speaker change intervals (414), marking the corresponding predicted speaker change word (302) as an incorrect prediction.

4. The computer-implemented method (600) of any preceding claim, It is characterized in that The operations further include: For each true value speaker change interval (414), when any of the predicted speaker change tokens (302) overlaps with a corresponding true value speaker change interval (414), marking the corresponding true value speaker change interval (414) as correctly matched; and A recall metric (444) of the sequence transduction model (300) is determined based on the duration of the true value speaker change intervals (414) marked as correctly matched and the total duration of all the true value speaker change intervals (414).

5. The computer-implemented method (600) of claim 4, It is characterized in that The operations further include determining a performance score (452) of the sequence transduction model (300) based on the precision metric (442) and the recall metric (444).

6. The computer-implemented method (600) of claim 5, It is characterized in that Determining the performance score (452) includes calculating the performance score (452) based on the following equation: 2*(precision indicator (442)*recall indicator (444)) / (precision indicator (442)+recall indicator (444)).

7. The computer-implemented method (600) of any preceding claim, It is characterized in that The operations further include: For each true value speaker change interval (414), when any of the predicted speaker change tokens (302) overlaps with the corresponding true value speaker change interval (414), marking the corresponding true value speaker change interval (414) as correctly matched; and A recall metric (444) of the sequence transduction model (300) is determined based on the number of the true value speaker change intervals (414) marked as correctly matched and the total number of all the true value speaker change intervals (414).

8. The computer-implemented method (600) of claim 7, It is characterized in that The operations further include determining a performance score (452) of the sequence transduction model (300) based on the precision metric (442) and the recall metric (452).

9. The computer-implemented method (600) of any preceding claim, It is characterized in that Determining the performance score (452) includes calculating the performance score (452) based on the following equation: 2*(precision indicator (442)*recall indicator (444)) / (precision indicator (442)+recall indicator (444)).

10. The computer-implemented method (600) of any preceding claim, Features: The multi-utterance training sample (410) further includes true value speaker labels (418) paired with the audio data (412), the true value speaker labels (418) each indicating a corresponding time-stamped segment in the audio data (412) associated with a respective one of the utterances spoken by one of the two or more different speakers (10); and Obtaining the true value speaker change interval (414) includes: identifying each time interval where two or more of the time-stamped segments overlap as a corresponding true-value speaker change interval (414); as well as Each time gap is identified, the time gap indicating a pause between two adjacent time-stamped segments in the audio data (412) associated with respective ones of the utterances spoken by two different speakers (10).

11. The computer-implemented method (600) of claim 10, It is characterized in that The operations further include: determining a minimum start time and a maximum start time of the audio data (412) based on the time-stamped segments indicated by the true value speaker labels (418); and Any predicted speaker change tokens (300) having a timestamp earlier than the minimum start time or later than the maximum start time are omitted from the determination of the accuracy metric (442) for the sequence transduction model (300).

12. The computer-implemented method (600) of any preceding claim, It is characterized in that Determining the accuracy indicator (442) of the sequence transduction model (300) is not based on any word-level speech recognition results (342) output by the sequence transduction model (300).

13. The computer-implemented method (600) of any preceding claim, It is characterized in that Determining the accuracy indicator (442) does not require performing a full speaker classification on the audio data (412).

14. A system (100), It is characterized in that include: Data processing hardware (710); as well as Memory hardware (720) in communication with the data processing hardware (710), the memory hardware (720) storing instructions that, when executed on the data processing hardware (710), cause the data processing hardware (710) to perform operations, the operations comprising: Obtaining a multi-utterance training sample (410), the multi-utterance training sample comprising audio data (412) representing utterances spoken by two or more different speakers (10); obtaining a true value speaker change interval (414), the true value speaker change interval indicating a time interval in the audio data (412) where a speaker change occurs between the two or more different speakers (10); Processing the audio data (412) using a sequence transduction model (300) to generate a sequence of predicted speaker change tokens (302), each predicted speaker change token indicating a position of a corresponding speaker turn in the audio data (412); For each corresponding predicted speaker change word-gram (302), when the predicted speaker change word-gram (302) overlaps with one of the true value speaker change intervals (414), marking the corresponding predicted speaker change word-gram (302) as correct; as well as An accuracy indicator (442) of the sequence transduction model (300) is determined based on the number of the predicted speaker change tokens (302) marked as correct and the total number of the predicted speaker change tokens (302) in the sequence of predicted speaker change tokens (302).

15. The system (100) of claim 14, It is characterized in that Determining the accuracy metric (442) of the sequence transduction model (332) is based on a ratio between the number of the predicted speaker change word-grams (302) marked as correct and the total number of the predicted speaker change word-grams (302) in the sequence of predicted speaker change word-grams (302).

16. The system (100) according to any one of claims 14 to 15, It is characterized in that The operation further includes: for each corresponding predicted speaker change word (302), when the corresponding predicted speaker change word (302) does not overlap with any of the true value speaker change intervals (414), marking the corresponding predicted speaker change word (302) as an incorrect prediction.

17. The system (100) according to any one of claims 14 to 16, It is characterized in that The operations further include: For each true value speaker change interval (414), when any of the predicted speaker change tokens (302) overlaps with a corresponding true value speaker change interval (414), marking the corresponding true value speaker change interval (414) as correctly matched; and A recall metric (444) of the sequence transduction model (300) is determined based on the duration of the true value speaker change intervals (414) marked as correctly matched and the total duration of all the true value speaker change intervals (414).

18. The system (100) of claim 17, It is characterized in that The operations further include determining a performance score (452) of the sequence transduction model (300) based on the precision metric (442) and the recall metric (444).

19. The system (100) of claim 18, It is characterized in that Determining the performance score (452) includes calculating the performance score (452) based on the following equation: 2*(precision indicator (442)*recall indicator (444)) / (precision indicator (442)+recall indicator (444)).

20. The system (100) according to any one of claims 14 to 19, It is characterized in that The operations further include: For each true value speaker change interval (414), when any of the predicted speaker change tokens (302) overlaps with the corresponding true value speaker change interval (414), marking the corresponding true value speaker change interval (414) as correctly matched; and A recall metric (444) of the sequence transduction model (300) is determined based on the number of the true value speaker change intervals (414) marked as correctly matched and the total number of all the true value speaker change intervals (414).

21. The system (100) of claim 20, It is characterized in that The operations further include determining a performance score (452) of the sequence transduction model (300) based on the precision metric (442) and the recall metric (452).

22. The system (100) according to any one of claims 14 to 21, It is characterized in that Determining the performance score (452) includes calculating the performance score (452) based on the following equation: 2*(precision indicator (442)*recall indicator (444)) / (precision indicator (442)+recall indicator (444)).

23. The system (100) according to any one of claims 14 to 22, Features: The multi-utterance training sample (410) further includes true value speaker labels (418) paired with the audio data (412), the true value speaker labels (418) each indicating a corresponding time-stamped segment in the audio data (412) associated with a respective one of the utterances spoken by one of the two or more different speakers (10); and Obtaining the true value speaker change interval (414) includes: identifying each time interval where two or more of the time-stamped segments overlap as a corresponding true-value speaker change interval (414); as well as Each time gap is identified, the time gap indicating a pause between two adjacent time-stamped segments in the audio data (412) associated with respective ones of the utterances spoken by two different speakers (10).

24. The system (100) of claim 23, It is characterized in that The operations further include: determining a minimum start time and a maximum start time of the audio data (412) based on the time-stamped segments indicated by the true value speaker labels (418); and Any predicted speaker change tokens (300) having a timestamp earlier than the minimum start time or later than the maximum start time are omitted from the determination of the accuracy metric (442) for the sequence transduction model (300).

25. The system (100) according to any one of the preceding claims 14 to 24, It is characterized in that Determining the accuracy indicator (442) of the sequence transduction model (300) is not based on any word-level speech recognition results (342) output by the sequence transduction model (300).

26. The system (100) according to any one of claims 14 to 25, It is characterized in that Determining the accuracy indicator (442) does not require performing a full speaker classification on the audio data (412).