Speaker diarization to support episodic content

JP2024516815A5Pending Publication Date: 2025-06-19DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023565384
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-04-30
Filing Date
2022-04-27
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing speaker diarization technologies struggle to effectively divide input audio streams containing multiple speakers' utterances without prior knowledge of speaker fingerprints and handle overlapping utterances, leading to inefficiencies and inaccuracies.

Method used

A method involving spatial transformation, block-based embedding extraction, clustering, and post-processing to identify and label speaker segments using machine learning models, optimizing embeddings for separability and employing multi-head attention and double clustering to enhance accuracy.

Benefits of technology

Improves speaker diarization by maximizing embedding separation, reducing memory usage, and enabling diarization across files of arbitrary length, with enhanced reliability and visualization tools for assessing success.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An embodiment for speaker diarization supporting episodic content is disclosed. In an embodiment, the method includes receiving media data including one or more utterances; dividing the media data into a plurality of blocks; identifying segments of each of the plurality of blocks associated with a single speaker; extracting embeddings for the identified segments according to a machine learning model, the extracting embeddings for the identified segments includes statistically combining the extracted embeddings for the identified segments corresponding to each successive utterance associated with a single speaker; clustering the embeddings for the identified segments into clusters; and assigning a speaker label to each of the embeddings for the identified segments according to the clustering results. In some embodiments, voiceprints are used to identify the speaker and speaker identity for the speaker labels.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 182,338, filed April 30, 2021.

[0002] Technical Field The present disclosure relates generally to audio signal processing, and more particularly to speaker diarization. [Background technology]

[0003] Speaker diarization is the process of dividing an input audio stream containing the speech of multiple individuals into homogenous segments associated with each speaker. Speaker diarization is used in many applications, such as understanding recorded conversations and video caption generation. Speaker diarization differs from speaker identification or speaker separation, because speaker diarization does not require a priori knowledge of the speaker's voice "fingerprint" or the number of speakers present in the input audio stream. Furthermore, speaker diarization differs from audio source separation, because speaker diarization is typically not applied to overlapping speech. Summary of the Invention [Problem to be solved by the invention]

[0004] SUMMARY OF THE DISCLOSURE Embodiments are disclosed for speaker diarization to support episodic content. [Means for solving the problem]

[0005] In some embodiments, the method includes: receiving, with at least one processor, media data including one or more utterances; dividing, with the at least one processor, the media data into a plurality of blocks; identifying, with the at least one processor, a segment of each of the plurality of blocks associated with a single speaker; extracting, with the at least one processor, embeddings for the identified segments according to a machine learning model, wherein extracting embeddings for the identified segments includes statistically combining extracted embeddings for the identified segments corresponding to respective consecutive utterances associated with a single speaker; clustering, with the at least one processor, embeddings for the identified segments into clusters; assigning, with the at least one processor, a speaker label to each of the embeddings for the identified segments according to results of the clustering; and outputting, with the at least one processor, speaker diarization information associated with the media data based in part on the speaker labels.

[0006] In some embodiments, prior to dividing the media data into blocks, a spatial transformation is performed on the media data.

[0007] In some embodiments, performing a spatial transformation on the media data includes: transforming a first plurality of channels of the media data into a second plurality of channels different from the first plurality of channels; and dividing the media data into a plurality of blocks includes: dividing each of the second plurality of channels independently into blocks.

[0008] In some embodiments, in response to a determination that the media data corresponds to a first media type, a machine learning model is generated from a first set of training data; and in response to a determination that the first media data corresponds to a second media type that is different from the first media type, a machine learning model is generated from a second set of training data that is different from the first set of training data.

[0009] In some embodiments, prior to clustering, the extracted embeddings for the identified segments are further optimized in response to determining that the optimization criterion is met.

[0010] In some embodiments, prior to clustering, further optimization of the extracted embeddings for the identified segments is foreclosed in response to a determination that the optimization criterion is not satisfied.

[0011] In some embodiments, optimizing the extracted embeddings for the identified segments includes performing at least one of a dimensionality reduction of the extracted embeddings or an embedding optimization of the extracted embeddings.

[0012] In some embodiments, the embedding optimization includes training a machine learning model to maximize separability between extracted embeddings for the identified segments; and updating the extracted embeddings by applying the machine learning model to maximize separability between extracted embeddings for the identified segments to the extracted embeddings for the identified segments.

[0013] In some embodiments, the clustering includes: for each identified segment, determining a respective length of the segment; in response to a determination that the respective length of the segment is greater than a threshold length, assigning an embedding associated with each identified segment according to a first clustering process; in response to a determination that the respective length of the segment is not greater than the threshold length, assigning an embedding associated with each identified segment according to a second clustering process different from the first clustering process.

[0014] In some embodiments, any of the aforementioned methods further include selecting a first clustering process from the plurality of clustering processes based in part on the determination of the number of distinct speakers associated with the media data.

[0015] In some embodiments, the first clustering process comprises spectral clustering.

[0016] In some embodiments, the media data includes multiple associated files.

[0017] In some embodiments, the method includes selecting a plurality of the associated files as the media data, wherein selecting the plurality of associated files is based in part on at least one of content similarity associated with the plurality of associated files, metadata similarity associated with the plurality of associated files, or received data corresponding to a request to process a particular set of files.

[0018] In some embodiments, a machine learning model is selected from a plurality of machine learning models according to one or more attributes shared by each of the plurality of related audio files.

[0019] In some embodiments, the method further includes the steps of: computing a voiceprint distance metric between the voiceprint embedding and the centroid of each cluster; computing the distance from each centroid to each embedding that belongs to that cluster; for each cluster, computing a probability distribution of the distances of the embeddings from the centroid of that cluster; for each probability distribution, computing the probability that the voiceprint distance belongs to that probability distribution; ranking the probabilities; assigning the voiceprint to one of the clusters based on the ranking; and combining speaker identification information associated with the voiceprint with the speaker diarization information.

[0020] In some embodiments, the probability distribution is modeled as a folded Gaussian distribution.

[0021] In some embodiments, the method further includes comparing each probability to a confidence threshold; and determining, based on the comparison, whether a speaker associated with the probability spoke.

[0022] In some embodiments, any of the aforementioned methods further include generating one or more analysis files or visualizations associated with the media data based in part on the assigned speaker labels.

[0023] In some embodiments, a non-transitory computer-readable storage medium stores at least one program for execution by at least one processor of an electronic device, the at least one program including instructions for performing any of the methods described above.

[0024] In some embodiments, a system comprises at least one processor and a memory coupled to the at least one processor storing at least one program for execution by the at least one processor, the at least one program including instructions for performing any of the methods described above.

[0025] Specific embodiments disclosed herein provide at least one or more of the following advantages: 1) an optimized architecture for speaker diarization that improves upon the standard diarization structure of existing architectures; 2) introduction of a pre-processing step to exploit spatial information present in stereo files before conversion to mono; 3) introduction of an embedding optimization step to maximize embedding separation and improve clustering based on a multi-head attention architecture or VBx clustering; 4) introduction of spectral clustering as an improved component in the pipeline; 5) introduction of a bi-clustering step to improve clustering and misclassification reliability of short speaker segments; 6) ability to perform diarization on files of any length, resulting in low memory occupation and processing load; 7) ability to perform diarization across different files, thus allowing diarization on episodic content; 8) statistics generation, error quantification and visualization to easily assess diarization success; and 9) a diarization pipeline that uses input voiceprints to determine if an audio file contains speech from that person. [Brief description of the drawings]

[0026] In the drawings, a particular arrangement or ordering of schematic elements, such as those representing devices, units, instruction blocks, and data elements, is shown for ease of explanation. However, it should be understood by those skilled in the art that the particular order or arrangement of the schematic elements in the drawings is not intended to imply that a particular order or sequence of processing, or separation of processes, is required. Furthermore, the inclusion of schematic elements in the drawings is not intended to imply that such elements are required in all embodiments, or that features represented by such elements may not be included in some embodiments or may be combined with other elements.

[0027] Furthermore, in the drawings, when a connecting element, such as a solid or dashed line or arrow, is used to indicate a connection, relationship, or association between two or more other schematic elements, the absence of such a connecting element is not intended to imply that the connection, relationship, or association may not exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the present disclosure. Furthermore, for ease of explanation, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents communication of a signal, data, or instruction, it should be understood by those skilled in the art that such element represents one or more signal paths, as necessary, to affect the communication.

[0028] [Figure 1] FIG. 2 is a block diagram of a speaker diarization processing pipeline according to an embodiment.

[0029] [Figure 2A] 2 is a flow diagram illustrating how different components of the pipeline shown in FIG. 1 act on an audio file, according to one embodiment. [Figure 2B]2 is a flow diagram illustrating how different components of the pipeline shown in FIG. 1 act on an audio file, according to one embodiment.

[0030] [Diagram 3] 1 is an embedding representation of a segment of audio containing four speakers after dimensionality reduction to two dimensions, according to one embodiment.

[0031] [Figure 4A] FIG. 2 is a flow diagram of a process in the pipeline shown in FIG. 1 that provides speaker change detection and overlap utterance detection using cross-entropy loss and a binary classification model, according to one embodiment.

[0032] [Figure 4B] 4 is a flow diagram 400b in a pipeline 100 that can use triplet loss to solve the embedding extraction process.

[0033] [Figure 5A] FIG. 1 is a flow diagram for processing stereo files in a diarization pipeline, according to one embodiment.

[0034] [Figure 5B] FIG. 13 is a flow diagram for alternative processing of stereo files in a diarization pipeline, according to an embodiment.

[0035] [Figure 6] 1 illustrates block segmentation and episodic content diarization according to one embodiment.

[0036] [Figure 7] 1 illustrates a multi-head attention plus GE2E loss function according to one embodiment.

[0037] [Figure 8]FIG. 1 is a flow diagram of a biclustering process, according to one embodiment.

[0038] [Figure 9] 13 illustrates the determination of the distance between the voiceprint embeddings and the centroids of the clusters identified by the diarization pipeline, according to one embodiment.

[0039] [Figure 10] 14 shows approximating the distance of centroids from embeddings belonging to clusters as a folded Gaussian distribution, according to one embodiment.

[0040] [Figure 11] 1 illustrates a probability ranking to identify which cluster has the highest probability of containing a voiceprint, according to one embodiment.

[0041] [Figure 12] 1 illustrates associating speakers with spoken voiceprints by combining speaker identification (ID) results with diarization pipeline output, according to one embodiment.

[0042] [Figure 13] 13 shows diarization performance for speech segments for two speakers using the ground truth as a baseline according to one embodiment.

[0043] [Figure 14] 14 illustrates speech segments for each speaker depicted in FIG. 13 according to one embodiment.

[0044] [Figure 15] FIG. 1 is a flow diagram of a process of speaker diarization to support episodic content, according to an embodiment.

[0045] [Figure 16]FIG. 16 is a block diagram of an exemplary device architecture for implementing the features and processes described with reference to FIGS. 1-15, according to one embodiment.

[0046] The same reference symbols used in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0047] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. It will be apparent to one of ordinary skill in the art that the various embodiments described may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described below that can be used independently of each other or in any combination with other features.

[0048] nomenclature As used herein, the term "comprises" and variations thereof should be read as open-ended terms meaning "including, but not limited to." The term "or" should be read as "and / or" unless the context clearly indicates otherwise. The terms "one exemplary embodiment" and "an exemplary embodiment" should be read as "at least one exemplary embodiment." The term "another embodiment" should be read as "at least one other embodiment." The terms "determined," "determine," or "determining" should be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. Additionally, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0049] Exemplary System 1 is a block diagram of a diarization processing pipeline 100 according to one embodiment. The diarization pipeline 100 is characterized by five main components: an audio pre-processing component 101, a block-based embedding extraction component 102, an embedding processing component 103, a clustering component 104, and a post-processing component 105.

[0050] The diarization pipeline 100 receives as input media data (e.g., mono or stereo audio files 106) containing utterances (e.g., speech) and returns as output corresponding segments (e.g., speech segments) with associated tags for each speaker detected in the audio file. In some embodiments, the diarization pipeline 100 also outputs analysis 121 and visualization 120 to provide a clear way of presenting the results of the diarization to a user. For example, the diarization pipeline 100 may identify the presence of multiple speakers in an audio file (e.g., three separate speakers in an audio file) and define the start and end times of each utterance. In some embodiments, the analysis 121 includes, but is not limited to, the number of speakers in the audio file, the confidence in the speaker identification, the time and percentage of participation in the conversation for each of the speakers, the dominant speaker in the conversation, overlapping speech sections, and conversation turns.

[0051] In some embodiments, the first component of the diarization pipeline 100 is an audio pre-processing component 101 that imports media data including speech. In some embodiments, the media data is an audio file 106. In some embodiments, the audio file 106 is a mono file. In some embodiments, the audio file 106 is a stereo file. In the case of a stereo file, a converter 107 converts the left (L) and right (R) stereo channels of the audio file 106 into two channels, namely channel 1 and channel 2.

[0052] In some embodiments, if speakers are spatialized and panned across different channels, the audio file 106 is transformed to maximize the spatial information present in channels 1 and 2. For example, if it is desirable to preserve the spatial information of the speaker localization, independent component analysis (ICA) can be used to generate two channels in which the spatial information is preserved and the speakers are separated into the two channels. In some embodiments, principal component analysis (PCA) is used to perform this task. In some embodiments, deep learning or machine learning is used to perform this task. In some embodiments, the audio file 106 is a stereo file that is converted to a mono file by submixing the L and R channels or by selecting one of the two channels (e.g., discarding one of the channels to improve processing efficiency).

[0053] After transforming the L and R channels into channels 1 and 2 to maximize separation and preservation of spatial information about the panned talker, each channel is processed by an audio block segmentation component 108 that divides audio channels 1 and 2 into blocks, which are sent to the block-based embedding extraction component 102. As used herein, a "block" is a unit of processing. Each block may contain one or more speech segments. A "speech segment" is a period of speech during which a unique speaker is speaking.

[0054] It should be noted that the diarization pipeline 100 is capable of processing audio files 106 of any duration and can also process and perform diarization across multiple files where the same speaker may be present. For example, in some embodiments, the diarization pipeline 100 can load (e.g., receive, retrieve) a stereo audio file and apply ICA to the two channels of the audio file 106 to maximize and exploit spatial information associated with panned speech to generate channel 1 and channel 2. The diarization pipeline 100 then divides the audio file into blocks of equal length (e.g., 1 second blocks), each of which is processed by downstream components of the pipeline 100. In some embodiments, the blocks include blocks of multiple lengths (e.g., 0.5 seconds, 0.1 seconds, or different durations).

[0055] Blocks of each audio channel are processed by a block-based embedding extraction component 102, which performs feature extraction 109 on the blocks and applies voice activity detection (VAD) 113 to those features. In some embodiments, VAD 113 detects speech from multiple speakers and performs overlapping speech detection 110. Speech detection is used to isolate label speech segments, i.e., portions of the blocks that contain speech. Speech detection is based on a combination of results from VAD 113, speaker change detection 112, and overlapping speech detection 110.

[0056] The isolated speech segments are input to an embedding extraction component 111 along with data for overlapping speech detection 110 and speaker change detection 112. The embedding extraction component 111 computes an embedding for each segment identified by speaker change detection 112, and overlapping speech is discarded so that no embeddings are extracted from the overlapping speech. The embedding extraction component 111 and embedding generation model extract embeddings (e.g., multi-dimensional vector representations of the speech segments) from the isolated speech segments. In some embodiments, the embedding extraction component 111 performs dimensionality reduction 114 on the embeddings (see FIG. 3) to generate improved embeddings 115 that are clustered by the clustering component 104.

[0057] The clustering component 104 receives as input the improved embeddings 115. A short segment classifier 116 identifies short and long speech segments which are clustered separately using long segment clustering 117 and short segment clustering 118, respectively. The embedding clusters are input to a post-processing component 105 which performs post-processing 119 and segmentation 120 on the embedding clusters which are used to generate analysis 121 and visualization 122.

[0058] FIG. 2 shows in more detail how the different components of the diarization pipeline 100 described above act on an audio file 106 according to an embodiment. In speech analysis, embeddings are multidimensional vector representations of segments of speech. Embeddings are generated to maximize the distinction of different speakers in a multidimensional feature space. For example, if two speakers speak in a conversation, their generated embeddings are multidimensional vectors that are clustered into two different regions of the multidimensional space. In some embodiments, extracting embeddings includes representing single-speaker speech segments or parts of single-speaker speech segments in a multidimensional subspace. In some embodiments, extracting embeddings includes converting each single-speaker speech segment into a multidimensional vector representation.

[0059] The motivation behind generating embeddings is to facilitate clustering of different speakers and thus perform speaker diarization successfully. A model, such as a generative embedding model, is used to generate embeddings by mapping or transforming speech segments to a representation in a multidimensional space. The model can be trained for embedding generation using different datasets and by using a loss function that maximizes the distance between embeddings of different speakers and minimizes the distance of embeddings generated from the same speaker. In some embodiments, the embedding model is a deep neural network (DNN) trained to have a speaker discriminatory embedding layer.

[0060] Referring to FIG. 2, in the example shown, the original audio includes speech from three different speakers and ambient sounds in the form of music and noise. Audio block segmentation generates a first block (Block 1) that captures the speech of Speaker 1 (SPK1), followed by music, followed by speech from Speaker 2 (SPK2). Block 1 is input to VAD 113, which detects the start and end of the SPK1 speech, resulting in Speech Segment 1 (SEGM1). Multiple embeddings of Speech Segment 1 are generated, features are extracted from each embedding of SEGM1, and the embeddings are statistically combined (e.g., by computing an average embedding from the multiple embeddings). An audio block is a uniform segmentation of the audio file (e.g., 1s, 2s, 3s, etc.).

[0061] The above process is repeated for blocks 2 and 3. For example, block 2 generates speech segment 4 (SEGM4) containing utterances from SPK3. Block 3 generates speech segment 7 (SEGM7) containing utterances from SPK3. Average embeddings of features are extracted from SEGM4 and SEGM7. The average embeddings of blocks 1, 2, and 3 are then input into a clustering process (e.g., short segment clustering and long segment clustering). Each cluster represents an utterance from one of three speakers, e.g., there are three clusters in three different regions of the multidimensional feature space.

[0062] 3 is an embedding representation of a segment of audio containing four speakers after dimensionality reduction to two dimensions using t-Distributed stochastic neighbor embedding (t-SNE) according to one embodiment. A 512-D embedding using the t-SNE algorithm described, for example, in [1], is reduced to 2D for visualization purposes. Each dot in FIG. 3 represents an embedding of a segment of speech. Embeddings belonging to the same speaker tend to cluster in space, and four separate clusters can be identified in FIG. 3, each cluster representing a different one of the four speakers. [Non-Patent Document 1] Van der Maaten, Laurens, and Geoffrey Hinton, “Visualizing data using t-SNE.” Journal of machine learning research 9.11 (2008)

[0063] Referring again to FIG. 1, the pipeline 100 includes three speech detections performed prior to embedding extraction: 1) VAD 113, 2) overlapping speech detection 110, and 3) speaker change detection 112. VAD 113 isolates speech segments from music or noise or other audio sources. In some embodiments, speech activity detection is performed using VAD as described in the AC-4 standard. In some embodiments, VAD 113 is based on the architecture described in FIG. 4A, 4B below. Overlapping speech detection 110 identifies regions where overlapping speech exists. Overlapping speech is typically discarded from diarization tasks. Process 400a may be used to detect overlapping speech. Other architectures, including different CNN and / or RNN architectures, may also be used for overlapping speech detection. Speaker change detection 112 detects conversation turns. A conversation turn is a transition from a first speaker to a second, different speaker. Speaker change detection 112 is performed on the output of the VAD 113. The process 400a may also be used for speaker change detection. Other architectures may also be used for speaker change detection, such as different CNN and / or RNN architectures.

[0064] 4A is a flow diagram of a process 400a in the pipeline 100 that provides speaker change detection 112 and overlapping speech detection 110 using cross-entropy loss and a binary classification model, according to some embodiments. In some embodiments, the process 400a includes a convolutional neural network (CNN) architecture 401a, followed by one or more convolutional layers 402a, followed by one or more recurrent neural network (RNN) layers 403a, followed by one or more feedforward neural network (FNN) layers 404a. In some embodiments, these layers are trained for a binary classification problem (e.g., speech present or absent) using a suitable training technique, such as backpropagation.

[0065] 4B is a flow diagram of a process 400b in the diarization pipeline 100 that can use triplet loss to solve the embedding extraction process. In some embodiments, the process 400b includes a convolutional neural network (CNN) layer 401b, followed by one or more convolutional layers 402b, followed by one or more recurrent layers 403b, followed by temporal pooling 404b, followed by one or more feedforward layers 404a. The output of the process 400b is a speech segment.

[0066] In the processes 400a and 400b, the CNNs 401a and 401b may be implemented using the SincCONV architecture, for example, as described in Non-Patent Document 2. In some embodiments, the SincCONV architecture may be replaced with hand-crafted features (e.g., Mel-frequency cepstral coefficients (MFCCs)). [Non-Patent Document 2] Ravanelli, M., & Bengio, Y. (2019), Speaker Recognition from Raw Waveform with SincNet. 2018 IEEE Spoken Language Technology Workshop, SLT 2018 - Proceedings, 1021-1028. https: / / doi.org / 10.1109 / SLT.2018.8639585

[0067] After segmenting the utterance into segments where only one speaker is present (e.g., each segment does not contain utterances from more than one speaker, as shown in FIG. 2) and isolating the segments of different speakers, embeddings from the utterance are extracted by the block-based embedding extraction component 102. In some embodiments, x-vector, d-vector, i-vector, or other strategies can be used for extraction of embeddings from the utterance. The architecture of FIG. 4B can be used for the same purpose after training a model on the utterance data to maximize the separability between different speakers. Different cost functions can be used to minimize the loss during training, such as triplet loss, cross entropy loss, probabilistic linear discriminant analysis (PLDA) based loss, etc.

[0068] In some embodiments, embeddings are extracted over a specified window length, as shown in FIG. 2. In some embodiments, embeddings are extracted over a fixed time window length (e.g., 20 ms). After extraction, embeddings belonging to consecutive segments of speech from the same speaker are averaged to reduce noise and improve accuracy, as shown in FIG. 2. The threshold defining the number of embeddings to be averaged and the maximum time window for averaging may be modified according to the length of the input audio file 106. For long input audio files, it is acceptable to average all embeddings that are part of an entire speech segment from the same speaker. In contrast, for short input audio files, the maximum acceptable time window for embedding averaging may be reduced to generate an overall larger number of embeddings that are fed to the clustering component 104 to improve clustering accuracy.

[0069] Different embedding models may be trained for different content types using different training datasets (e.g., training data corresponding to each content or media type). In some embodiments, the embedding models are optimized to extract (infer) embeddings from content associated with an environment (e.g., sports venue, concert venue, car cabin, recording studio), use case (e.g., podcast, lecture, interview, sportscast, music), subject or topic (e.g., chemistry), activity (e.g., commuting, walking), etc. In some embodiments, the type of content that needs to be diarized (e.g., podcast, phone call, education, user-generated content, etc.) is specified by the application or a user of the application, and the diarization pipeline 100 selects the best model for the type of content that needs to be processed according to the type of content specified. In some embodiments, the best model is selected according to various methodologies (e.g., lookup tables, content analysis).

[0070] After extraction of the embeddings 111, the embeddings are further processed as shown in Figure 1. In one embodiment, the processing of the embeddings is performed by two components: 1) an embedding dimensionality reduction component 114, and 2) an embedding optimization component 115. In some embodiments, one or both of components 114, 115 are skipped, for example, in response to determining that the extracted embeddings are above a quality threshold.

[0071] A dimensionality reduction component 114 is implemented to improve the ability to distinguish between different speakers and maximize the success of clustering. The embedding dimension can be reduced using PCA, t-SNE, uniform manifold approximation and projection (UMAP), PLDA, or other dimensionality reduction strategies.

[0072] The embedding optimization component 115 is described in more detail in the following section. The intent is to use a data-driven approach to train a model to maximize the separability between embeddings, further facilitating the clustering process. In some embodiments, this can be achieved using a multi-head attention plus generalized end-to-end (GE2E) loss architecture for embedding optimization. Other architectures can also be used for this purpose.

[0073] After embedding extraction and processing, the embeddings are clustered for different speakers. The clustering component 105 performs the clustering by distinguishing between short speech segments 117 and long speech segments 118, as shown in FIG. 1 and described in more detail below. Different clustering strategies can be used for clustering short or long segments. In some embodiments, an improved spectral clustering is described in [3]. [Non-Patent Document 3] Wang, Q., Downey, C., Wan, L., Mansfield, PA, & Moreno, IL (2018).c Speaker diarization with LSTM. ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2018-April, 5239-5243

[0074] Other clustering strategies, such as hierarchical clustering, affinity propagation, agglomerative clustering, or VBx clustering, may also be used. In some embodiments, the number of speakers in the audio file 106 is known or determined. For example, data indicative of the number of speakers associated with the audio file may be received (e.g., via an application programming interface (API) call) or may be derived and obtained based on an analysis of the audio file using, for example, k-means clustering or other suitable clustering algorithms. In one embodiment, when the number of speakers is not known in advance, an elbow method may be used in conjunction with the k-means clustering algorithm. The elbow method may include plotting a curve of the explained variation as a function of the number of clusters and then choosing the elbow of that curve as the number of clusters to use.

[0075] After clustering, post-processing 119 of the clustered embeddings is performed. Post-processing 119 is performed to remove false positives and improve the clustering accuracy. In some embodiments, methods such as Hidden Markov Models (HMMs) with predefined transition probabilities between different speakers are used for this purpose. The resulting improved post-processed embeddings are used by the segmentation component 120 to segment the audio file into segments with associated speaker labels / identifiers.

[0076] Analytics 121 are generated from the clustered embeddings as an output of the diarization pipeline 100. The analytics 121 may include information including, but not limited to, the number of speakers, the confidence of the speaker detection, the percentage and duration of speaker participation in the conversation, the most relevant speakers, and other analytics are generated. The diarization pipeline 100 may also generate a summary of the detection accuracy in the speaker diarization task.

[0077] Exemplary Pre-Processing Blocks As mentioned above, the audio pre-processing component 101 allows the diarization pipeline 100 to generate diarization outputs for stereo and mono input files. In the case of mono files, the audio file is treated as a mono file and processed by the audio block segmentation block 108 as described below. In the case of stereo files, processes 500a, 500b can be applied.

[0078] FIG. 5A is a flow diagram for processing a stereo file in the diarization pipeline 100 according to an embodiment. Process 500a (shown in FIG. 5A) illustrates an embodiment utilized when the left channel 501a and the right channel 502a are different and the speaker's voice is panned and not symmetric between the right and left channels. The panned speech has spatial information that may be useful for the diarization task. For each channel 501a, 502a, a spatial transform block 503a separates the spatial information into two different channels 504a (channel 1, channel 2) to improve the success of the diarization task. In some embodiments, ICA or PCA is implemented by the spatial transform block 503a. Deep learning strategies can also be used. The remaining blocks of the diarization pipeline 100 are applied separately to channel 1 and channel 2. After block processing 505a and embedding extraction stage 506a, the embeddings are processed together 507a (e.g., embeddings from each channel are combined for subsequent processing) to generate a processed version of the embeddings that exploits information associated with the speaker's spatial localization. In some embodiments, process 500a is applied to multi-channel audio (e.g., Dolby® Atmos®) instead of binaural audio.

[0079] Figure 5B is a flow diagram of an alternative processing of a stereo file in the diarization pipeline 100, according to an embodiment. In some embodiments, when the left and right channels 501a, 501b are identical (Figure 4B), the stereo file is converted to mono 503b and the remaining blocks of the diarization pipeline 100 are applied to the single mono file as described with reference to Figure 5A. The mono conversion is achieved by selecting one of the two channels, or by downmixing.

[0080] Exemplary Audio Block Segmentation Audio block segmentation 108 is another step in the pipeline 100 (see Figures 1, 2 and 6). The embeddings extracted for each speech segment (a segment is a time window of speech in which a unique speaker is speaking) are stored and clustered after all embeddings have been extracted from each segment across multiple audio blocks. This process introduces the possibility to process files of any length and limits the processing and memory usage required to process an entire file without impacting performance. This process can be applied to a list of input files, thus enabling the possibility to perform diarization on episodic content.

[0081] Referring to Figure 6, a collection of input audio files 601 from episodic content (e.g. a podcast season) are segmented by audio block segmentation 602 and fed through an extract embeddings component 603, as described above with reference to Figure 1. The extracted embeddings 604 are then combined together in an embedding process 605, clustered 606 and input to post-processing / segmentation 607. These steps allow speaker diarization to be performed across multiple audio files 106, thus enabling the possibility to compute analytics across episodic content.

[0082] Exemplary Embedding Embeddings play an important role in speaker diarization systems since embeddings are multi-dimensional vector representations of segments of speech. Existing solutions utilized Gaussian Mixture Modeling (GMM) based embedding extraction, i-vector, x-vector, d-vector, etc. Although these existing methods are robust and may show good performance in solving the speaker diarization problem, they do not fully utilize the temporal information. Therefore, a new module is introduced into the diarization pipeline 100 to further improve the effectiveness of embeddings by utilizing temporal information, thereby enhancing the accuracy of clustering in the next step.

[0083] Several related ideas have been proposed in the literature that utilize the temporal information of utterances to generate improved embeddings. For example, a long short-term memory (LSTM)-based vector-to-sequence scoring model has been proposed that utilizes adjacent temporal information to generate a similarity score from the embeddings. However, the LSTM structure has two limitations: 1) it focuses more on local information and may fail in tasks with long-term dependencies, and 2) it is both time and space consuming.

[0084] To address the limitations of the LSTM structure, architectures based entirely on attention mechanisms have shown promising value in sequence-to-sequence learning. In addition to providing significantly faster training, attention networks demonstrate efficient modeling of long-term dependencies. For example, a positional multi-head attention model structure and triplet loss function can be used. The positional encoding is a multi-dimensional vector that contains information about a specific position in the utterance. Furthermore, a triplet ranking loss is utilized to learn a similarity metric using the output from the multi-head attention model.

[0085] In some embodiments, a new loss function called the "end-to-end loss function" can be used to solve the speaker verification problem. GE2E training is based on processing many utterances at once, in the form of batches containing N speakers and, on average, M utterances from each speaker. The similarity matrix S ji,k is the set of all centroids c k For each embedding vector e ji It is defined as the scaled cosine similarity between

number

[0086] Here, e jirepresents the embedding vector of the i-th utterance of the j-th speaker (1 ≤ i ≤ M, j ≤ N). k represents the center of gravity of the embedding vector (k≦N).

[0087] Since the speaker diarization problem also tries to use the extracted embeddings to find a similarity matrix and then run a clustering algorithm based on the similarity matrix, the constraint on M utterances in equation [1] can be removed and a similar formula can be used to calculate the loss, since in the speaker diarization problem a speaker may speak an unlimited number of times. As in equation [1], e ji represents the embedding vector of the i-th utterance of the j-th speaker (i ≥ 1, j ≤ N), and c k represents the center of gravity of the embedding vector (k≦N). Then, when the softmax function is used for equation [1], each embedding vector e ji The loss on

number

[0088] FIG. 7 illustrates a multi-head attention plus GE2E loss function according to one embodiment. As shown in FIG. 7, a global embedding optimization module 700 sends input embeddings 701 to a multi-head attention model. A positional encoding 702 is applied at every time step, i.e., the training process uses temporal pooling 706 to take temporal information into account in the training process. The training process minimizes a loss function 707 (e.g., the GE2E loss function shown in Equation [2]). The loss function 707 moves each embedding vector closer to its centroid and away from all other centroids. During testing, the weights and biases trained within the multi-head attention model are utilized to predict an optimized embedding for each input and form a similarity matrix 708 that is input to clustering 709. Since we take temporal information into account, we should be able to build a better similarity matrix compared to the original embeddings.

[0089] Exemplary Clustering Techniques Spectral clustering is a standard clustering algorithm that has been used in several tasks, including diarization. In the diarization pipeline 100, a modified version of spectral clustering is used. FIG. 1 illustrates an embodiment that combines a modified version of spectral clustering with the diarization pipeline 100 described with reference to FIG. 1. The modified spectral clustering is one exemplary clustering strategy supported by the diarization pipeline 100. In some embodiments, other clustering strategies can be used, including but not limited to hierarchical clustering, k-means, affinity propagation, or agglomerative clustering.

[0090] As explained before, in order to remove noise and improve the estimation accuracy of the embedding, in some embodiments, the embeddings generated over speech segments with a single speaker are averaged. For this reason, speech segments of short duration may not provide enough information for speaker characterization. This may be due to the inaccuracy of the embeddings averaged over speech segments of long duration. Therefore, a double-step clustering approach is proposed, as described with reference to FIG. 8.

[0091] In some embodiments, Bayesian HMM clustering of x-vector sequences (VBx), as described in [4], can be used in the diarization pipeline 100. In the diarization pipeline 100, VBx is modified to take the x-vectors and other embeddings. [Non-Patent Document 4] Landini, Federico, Jan Profant, Mireia Diez, and Lukas Burget, “Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks”, Computer Speech & Language71 (2022): 101254

[0092] FIG. 8 is a flow diagram of a double clustering process according to an embodiment. As shown in FIG. 8, an embedding is extracted for each speech segment 801. The speech segments are separated into long duration speech segments 802 and short duration speech segments 803 based on a duration threshold. The long duration speech segments 802 are first clustered (803) and the centers of each cluster are calculated (804). The cosine distance (e.g., Eisen cosine correlation distance) between each of the short duration speech segments and each cluster is calculated (805) and the short duration speech segments 803 are assigned to the closest cluster. Other distance metrics (e.g., Euclidean, Manhattan, Pearson, Spearman, Kendall correlation distance) that assign each short duration speech segment 803 to the best cluster can be used or other strategies can be used. The overall clustering result is a combination of the clustering results of the short duration speech segments 803 and the long duration speech segments 802.

[0093] Exemplary Post-Processing Another component in the processing pipeline 100 is the post-processing component 105. The post-processing component 105 is introduced in the pipeline 100 to reduce errors in segmentation and clustering. The post-processing component 105 analyzes the annotations generated by the clustering and identifies and corrects possible errors. Several post-processing strategies can be used in the pipeline 100. For example, an HMM with associated probabilities of speaker changes can be used. Other algorithms that exploit temporal relationships between annotations can be used for the same purpose.

[0094] Exemplary Speaker Identification Speaker identification (speakerID) is the ability to understand who a speaker is based on an input voiceprint. In some embodiments, the diarization pipeline 100 uses an input voiceprint (e.g., 10-20 seconds of a person's speech) to determine if an audio file contains speech from that person. Additionally, based on knowledge of multiple speakers' voiceprints, a speakerID label can be assigned to the output of the diarization pipeline 100. As described above, the diarization pipeline 100 computes embeddings for each speech segment and then clusters the embeddings into a limited number of clusters.

[0095] 9 and 10 show how the distance of a reference point from an embedding belonging to a cluster is approximated as a folded Gaussian distribution, according to an embodiment. The reference point may be the centroid or any other suitable reference point, such as the medoid in partition around medois (PAM) clustering. As shown in FIG. 9, the first step in the speakerID process is to calculate the voiceprint distance metric, which in this example is the cosine distance metric d_vp between the voiceprint vp (905) and each reference point of each cluster (e.g., the centroid ci (for clusters 901-904)). i where i is an index that identifies a cluster from 1 to N, and N is the number of clusters. In the illustrated example, N=4 clusters represent the four main speakers.

[0096] For each cluster, the cosine distance from the cluster's reference point to each embedding that belongs to that cluster is calculated. For each cluster 901-904, the distribution of embedding distances from reference points that belong to that cluster is also calculated. In some embodiments, the distribution is a folded Gaussian distribution (p i) The probability that the voiceprint distance belongs to the corresponding distribution is then calculated. As shown in FIG. 11, the probabilities can be ranked to assign the voiceprint to the corresponding cluster. In this example, Spk1 has the highest probability. By defining a confidence threshold, one can better understand the probability that the person of the voiceprint spoke in the audio file at all. Understanding when the speaker associated with the voiceprint spoke can be determined by combining the speakerID result with the output of the diarization pipeline 100, as shown in FIG. 12. In the example shown in FIG. 12, the speakerID result is used to identify the speech segments 1201-1 to 1201-4 spoken by Spk1.

[0097] Exemplary Analysis / Visualization The diarization pipeline 100 generates analytics, metrics, and visualizations as output. The output provides the user with an objective and unambiguous quantification of the output of the diarization pipeline 100. In some embodiments, the analytics and visualizations are generated for a single file, multiple episodic files, and / or the entire dataset. Tables 1 and 2 below summarize exemplary analytics that may be generated by the diarization pipeline 100. In some embodiments, the visualization includes causing the display of a first object representing a first speaker, a second object representing each speaker, and a connector object connecting the first and second objects that varies in appearance (length, size, color, shape, style, etc.) according to one or more statistics derived from the clustering embedding. [Table 1] [Table 2]

[0098] Additionally, the performance of the diarization pipeline 100 can be evaluated. Using the diarization solutions, the evaluation metrics in Table 3 can be used to generate a report that allows a user to better evaluate the performance of the diarization pipeline 100. [Table 3]

[0099] In some embodiments, the diarization pipeline 100 includes a visualization module to help users understand the diarization performance in a more efficient manner. For example, FIG. 13 shows the percentage of time each speaker is speaking, which helps understand the relative duration of each speaker and how much of the utterance was correctly detected. FIG. 14 shows the speech segments of each speaker with the correct answer shown as a baseline, which helps to more intuitively understand the performance of the diarization pipeline 100.

[0100] Example Process 15 is a flow diagram of a process 1500 for context-aware audio processing, according to an embodiment. The process 1500 can be implemented, for example, using the device architecture 1500 described with reference to FIG.

[0101] The process 1500 includes receiving (1501) media data including one or more utterances, dividing (1502) the media data into a plurality of blocks, identifying (1503) segments of each of the plurality of blocks associated with a single speaker, extracting (1504) embeddings for the identified segments according to a machine learning model, clustering (1505) the embeddings of the identified segments into clusters, and assigning (1506) a speaker label to each of the embeddings for the identified segments according to the clustering results, each of which has been described in detail above with reference to Figures 1-14.

[0102] Exemplary System Architecture FIG. 16 shows a block diagram of an exemplary system 1600 suitable for implementing the exemplary embodiment described with reference to FIGS. 1 to 15. The system 1600 includes a central processing unit (CPU) 1601 capable of executing various processes according to a program stored, for example, in a read-only memory (ROM) 1602 or loaded, for example, from a storage unit 1608 into a random access memory (RAM) 1603. The RAM 1603 also stores data required by the CPU 1601 to execute various processes, as necessary. The CPU 1601, the ROM 1602, and the RAM 1603 are connected to each other via a bus 1604. An input / output (I / O) interface 1605 is also connected to the bus 1604.

[0103] The following components are connected to the I / O interface 1605: an input section 1606, which may include a keyboard, a mouse, etc.; an output section 1607, which may include a display, such as a liquid crystal display (LCD), and one or more speakers; a memory section 1608, which may include a hard disk or other suitable storage device; and a communication section 1609, which may include a network interface card, such as a network card (wired or wireless).

[0104] In some embodiments, the input 1606 includes one or more microphones in various positions (depending on the host device) that enable capture of audio signals in a variety of formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0105] In some embodiments, the output 1607 includes a system having a variety of speakers. The output 1607 can render audio signals in a variety of formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0106] The communication unit 1609 is configured to communicate with other devices (e.g., via a network). The drive 1610 is also connected to the I / O interface 1605 as needed. The drive 1610 mounts a removable medium 1611, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or other suitable removable media, and a computer program read from the removable medium 1611 is installed in the storage unit 1608 as needed. Those skilled in the art will understand that although the system 1600 is described as including the above-mentioned components, in actual applications, some of these components can be added, removed, and / or replaced, and all of these modifications or changes fall within the scope of the present disclosure.

[0107] According to an exemplary embodiment of the present disclosure, the processes described above may be implemented as a computer software program or on a computer readable storage medium. For example, an embodiment of the present disclosure includes a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing a method. In such an embodiment, the computer program may be downloaded and mounted from a network via the communication unit 709 and / or installed from a removable medium 1611 as shown in FIG. 16.

[0108] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuits (e.g., control circuits), software, logic, or any combination thereof. For example, the units described above may be executed by a control circuit (e.g., a CPU in combination with other components of FIG. 16), and thus the control circuit may perform the actions described in the present disclosure. Some aspects may be implemented in hardware, and other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuit). Although various aspects of the exemplary embodiments of the present disclosure have been illustrated and described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers, or other computing devices, or any combination thereof, as non-limiting examples.

[0109] Additionally, the various blocks illustrated in the flowcharts may be viewed as method steps and / or operations resulting from computer program code operations and / or as multiple combined logic circuit elements configured to perform the associated functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the methods described above.

[0110] In the context of this disclosure, a machine-readable medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0111] Computer program codes for carrying out the methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus having control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The program codes may be executed entirely on the computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0112] Although this document contains many specific embodiment details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as acting in certain combinations and may initially be claimed as such, one or more features from a claimed combination may, in some cases, be cut out of the combination, and the claimed combination may be directed to a subcombination or variation of the subcombination. The logic flows depicted in the figures do not require the particular order or sequential order depicted to achieve the desired results. In addition, other steps may be provided or steps may be removed from the described flows, and other components may be added to or removed from the described systems. Thus, other embodiments are within the scope of the following claims.

Claims

1. receiving, using at least one processor, media data including one or more utterances; dividing, using the at least one processor, the media data into a plurality of blocks; identifying, using the at least one processor, segments of each of the plurality of blocks associated with a single speaker; extracting, using the at least one processor, an embedding for the identified segments according to a machine learning model, wherein the extraction of the embedding for the identified segments includes statistically combining the extracted embeddings for the identified segments corresponding to each consecutive utterance associated with a single speaker; clustering, using the at least one processor, the embeddings for the identified segments into clusters; assigning, using the at least one processor, a speaker label to at least one of the embeddings for the identified segments according to the result of the clustering; and outputting, using the at least one processor, speaker diarization information associated with the media data based at least in part on the speaker label. A method.

2. Further comprising, before dividing the media data into a plurality of blocks, performing, using the at least one processor, a spatial transformation on the media data. The method according to claim 1.

3. Performing a spatial transformation on the media data comprises: converting a first plurality of channels of the media data into a second plurality of channels different from the first plurality of channels; Dividing the media data into a plurality of blocks includes: independently dividing each of the second plurality of channels into blocks. The method according to claim 2.

4. In response to determining that the media data corresponds to a first media type, the machine learning model is generated from a first set of training data; In response to determining that the first media data corresponds to a second media type different from the first media type, the machine learning model is generated from a second set of training data different from the first set of training data. The method according to claim 1.

5. Further comprising the step of further optimizing the extracted embedding for the identified segment in response to determining that an optimization criterion is satisfied before clustering. The method according to claim 1.

6. Further comprising deferring further optimizing the extracted embedding for the identified segment in response to determining that an optimization criterion is not satisfied before clustering. The method according to claim 5.

7. Optimizing the extracted embedding for the identified segment includes performing at least one of dimensionality reduction of the extracted embedding or embedding optimization of the extracted embedding, according to the method of claim 5.

8. Embedding optimization includes: Training the machine learning model to maximize the separability between the extracted embeddings for the identified segments; Updating the extracted embeddings by applying the machine learning model for maximizing the separability between the extracted embeddings for the identified segments to the extracted embeddings for the identified segments, The method according to claim 7. **Claim 9** The clustering is: For each identified segment, Determine the respective length of the segment; In response to a determination that the respective length of the segment is greater than a threshold length, assign the embeddings associated with the respective identified segments according to a first clustering process; In response to a determination that the respective length of the segment is not greater than the threshold length, assign the embeddings associated with the respective identified segments according to a second clustering process different from the first clustering process, The method according to claim 1. **Claim 10** Further comprising the step of selecting a first clustering process from a plurality of clustering processes, at least in part based on determining the number of different speakers associated with the media data, The method according to claim 1. **Claim 11** The method according to claim 1, wherein the first clustering process includes spectral clustering. **Claim 12** The method according to claim 1, wherein the media data includes a plurality of related files. **Claim 13** Further comprising the step of selecting a plurality of the related files as the media data, and selecting the plurality of related files is: Content similarity related to the plurality of related files; Metadata similarity related to the plurality of related files; or Received data corresponding to a request to process a specific set of files at least in part based on at least one of The method according to claim 12. **Claim 14** The method according to claim 12, wherein the machine learning model is selected from a plurality of machine learning models according to one or more attributes shared by each of the plurality of related audio files. **Claim 15** Calculating a voiceprint distance metric between the voiceprint embedding and a reference point of each cluster; Calculating the distance from each reference point to each embedding belonging to its cluster; For each cluster, calculating a probability distribution of the distances of the embeddings from the reference point for that cluster; For each probability distribution, calculating the probability that the voiceprint distance belongs to that probability distribution; Ranking the probabilities; Assigning the voiceprint to one of the clusters based on the ranking; Further comprising combining the speaker characteristics associated with the voiceprint with the speaker diarization information. The method according to claim 1. **Claim 16** The method according to claim 15, wherein the probability distribution is modeled as a folded Gaussian distribution. **Claim 17** Comparing each probability with a confidence threshold; Further comprising determining whether the speaker associated with the probability spoke based on the comparison. The method according to claim 15. **Claim 18** Further comprising generating one or more analysis files or visualizations associated with the media data based at least in part on the assigned speaker label. The method according to claim 1.

19. Receiving, using at least one processor, media data including one or more utterances; Dividing, using the at least one processor, the media data into a plurality of blocks; Identifying, using the at least one processor, segments of each of the plurality of blocks associated with a single speaker; Extracting, using the at least one processor, embeddings for the identified segments according to a machine learning model; Clustering, using the at least one processor, the embeddings for the identified segments into clusters; Assigning, using the at least one processor, a speaker label to at least one of the embeddings for the identified segments according to the result of the clustering; Outputting, using the at least one processor, speaker diarization information associated with the media data based at least in part on the speaker label. A method.

20. Further comprising, before dividing the media data into a plurality of blocks, performing, using the at least one processor, a spatial transformation on the media data. The method according to claim 19.

21. Performing a spatial transformation on the media data comprises: Converting a first plurality of channels of the media data to a second plurality of channels different from the first plurality of channels; Dividing the media data into a plurality of blocks comprises: independently dividing each of the second plurality of channels into blocks. The method according to claim 20.

22. A non-transitory computer-readable storage medium storing at least one program for execution by at least one processor of an electronic device, the at least one program including instructions for performing the method according to any one of claims 1 to 21. **Claim 23** At least one processor; A memory coupled to the at least one processor, storing at least one program for execution by the at least one processor, the at least one program including instructions for performing the method according to any one of claims 1 to 21. A system.