Method and system for real-time active speaker detection
The active speaker detection system uses an audio-visual encoder and classifier to process audiovisual data efficiently, addressing the accuracy-speed trade-off in existing algorithms and enabling real-time speaker identification.
Patent Information
- Application Number
- JP2025061359
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2025-04-02
- Publication Date
- 2025-10-15
AI Technical Summary
Existing active speaker detection algorithms face a trade-off between accuracy and computational speed, with longer sequences of audiovisual data improving accuracy but requiring excessive processing time, making real-time detection challenging.
An active speaker detection system using a computer system with an audio-visual encoder and classifier processes audiovisual data to generate embeddings and determine ASD scores, aggregating them to identify and focus on the active speaker in real-time without significant computational overhead.
The system achieves high-accuracy real-time active speaker detection by incorporating greater temporal context without a substantial increase in computational time, effectively identifying speakers in visual scenes.
Smart Images

Figure 2025157189000001_ABST
Abstract
Description
[Background technology]
[0001] A visual scene including one or more speakers, such as a video captured by one or more cameras, can be enhanced by identifying the active speaker and altering the display of the visual scene accordingly. For example, a display can be adapted to frame or depict only the active speaker once identified among one or more people captured in the visual scene. Algorithms generally falling under the category of machine learning models have been developed to detect active speakers using one or more of audio data and visual data, and any combination of audio data and visual data can be collectively referred to as audiovisual data. However, the accuracy of such methods is inversely proportional to the time required to computationally process the audiovisual data. Summary of the Invention [Problem to be solved by the invention]
[0002] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in limiting the scope of the claimed subject matter. [Means for solving the problem]
[0003] In general, in one aspect, an embodiment relates to an active speaker detection system including a visual sensor capturing a visual scene including a first person and a computer system. The computer system includes one or more computer processors and a detection model. The detection model includes an audio-visual encoder and a classifier. The computer system is communicatively coupled to the visual sensor and configured to acquire a first set of frames and a second set of frames from the visual sensor and generate a first embedding and a second embedding from the first set of frames and the second set of frames, respectively, using the audio-visual encoder. The computer system is further configured to generate one or more composite embeddings from the first embedding and the second embedding, and determine an active speaker detection (ASD) score for each of the one or more composite embeddings using the classifier. The computer system is further configured to aggregate the one or more ASD scores forming a detection result, determine whether the first person is speaking based on the detection result, and, if determining that the first person is speaking, adjust a display of the visual scene to focus on the first person.
[0004] In general, in one aspect, an embodiment relates to a method for determining whether a person in a visual scene including a first person is speaking. The method includes obtaining a first set of frames and a second set of frames from a visual sensor capturing the visual scene, and generating, using an audiovisual encoder, a first embedding and a second embedding from the first set of frames and the second set of frames, respectively. The method further includes generating one or more composite embeddings from the first embedding and the second embedding, and determining, using a classifier, an active speaker detection (ASD) score for each of the one or more composite embeddings. The method further includes aggregating the one or more ASD scores to form a detection result, determining whether the first person is speaking based on the detection result, and adjusting a display of the visual scene to focus on the first person in response to determining that the first person is speaking.
[0005] In general, in one aspect, an embodiment relates to a non-transitory computer-readable medium having computer-executable instructions that, when executed on a computer processor, cause the computer processor to perform various steps. The steps include acquiring a first set of frames and a second set of frames from a visual sensor that captures a visual scene, the visual scene including a first person. The steps further include generating, with an audiovisual encoder, a first embedding and a second embedding from the first set of frames and the second set of frames, respectively. The steps further include generating one or more composite embeddings from the first embedding and the second embedding, and determining, with a classifier, an active speaker detection (ASD) score for each of the one or more composite embeddings. The steps further include aggregating the one or more ASD scores to form a detection result, determining whether the first person is speaking based on the detection result, and adjusting a display of the visual scene to focus on the first person in response to determining that the first person is speaking. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 illustrates an active speaker detection system according to one or more embodiments of the present disclosure. [Figure 2] FIG. 1 illustrates a neural network in accordance with one or more embodiments of the present disclosure. [Figure 3A] FIG. 1 illustrates a recurrent neural network in accordance with one or more embodiments of the present disclosure. [Figure 3B] FIG. 1 illustrates a deployed recurrent neural network in accordance with one or more embodiments of the present disclosure. [Figure 4] FIG. 1 illustrates a speaker detection model in accordance with one or more embodiments of the present disclosure. [Figure 5]FIG. 1 illustrates a concatenation operation and a classifier according to one or more embodiments of the present disclosure. [Figure 6] FIG. 10 is a graphical illustration of a set of configuration parameters in accordance with one or more embodiments of the present disclosure. [Figure 7] FIG. 1 illustrates an exemplary use of the system of the present disclosure applied to an incoming video stream according to one or more embodiments. [Figure 8] FIG. 1 illustrates a flowchart in accordance with one or more embodiments of the present disclosure. [Figure 9] FIG. 1 illustrates a method for detecting an active speaker according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0007] Specific embodiments of the present disclosure will now be described in detail with reference to the drawings, in which like elements in the various figures are designated with like reference numerals for consistency.
[0008] In the following detailed description of embodiments of the present disclosure, numerous specific details are set forth in order to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without these specific details. In other instances, well-known features have not been described in detail to avoid unnecessarily complicating the description.
[0009] Throughout this application, ordinal numbers (e.g., first, second, third) may be used as adjectives for elements (e.g., any noun in this application). The use of ordinal numbers is not intended to imply or create a particular order of elements or limit any element to only a single element unless expressly disclosed, such as by using the terms "before," "after," "single," and other such terms. Rather, the use of ordinal numbers is to distinguish elements. As an example, a first element is different from a second element, and a first element may encompass two or more elements and may follow (or precede) a second element in the order of elements.
[0010]
[0003] Embodiments disclosed herein generally relate to an active speaker detection (ASD) system and its use. ASD is the task of detecting who is speaking in a visual scene of one or more people, where each person can be a candidate for an active speaker. Generally, at any moment in a visual scene, no one may be speaking, one person may be speaking, or multiple people may be speaking simultaneously. Detecting active speakers in a visual scene (e.g., video captured by at least one camera) is useful across applications such as content creation and video conferencing. Various ASD methods or algorithms have been proposed. These algorithms or methods detect active speakers using one or more of audio data and visual data. In this specification, the term "audiovisual data" is used to refer to audio data, visual data, or a combination of audio data and visual data. Thus, an ASD algorithm or method may be said to determine one or more active speakers in a visual scene by processing audiovisual data associated with the visual scene.
[0011] While determining the active speaker among a large number of possible speakers is typically an easy task for human participants, implementing this behavior algorithmically is challenging. Intuitively, at least from a human perspective, determining whether a given person is speaking is easier in a larger temporal context. For example, it is easier to determine whether a person is speaking when listening to or observing them for 30 seconds than when listening to or observing them for only one second. Advances in algorithmic ASD have therefore involved the use of longer data sequences (e.g., sequences of images or frames forming a video, and possibly associated audio) to contain greater temporal information. While processing longer sequences of audiovisual data improves accuracy, processing such long sequences is computationally expensive and challenging for real-time active speaker determination. For example, an ASD algorithm operating on a relatively long sequence of audiovisual data may require more than three minutes to process a video clip with a duration of one and a half minutes. In other words, with ASD, there is a trade-off between accuracy and speed; the use of longer sequences is associated with improved accuracy (consistent with intuition and human observation) but requires additional processing time.
[0012] Generally, audiovisual data consists of an ordered sequence of frames (e.g., images) and possibly associated audio, with the audio being segmented to associate audio segments with frames or images (usually one-to-one). The frame and / or audio sequences are ordered according to time. Thus, a sequence of audiovisual data can have a defined length depending on the number of frames in the sequence or the duration the sequence requires to play when the frames are displayed at a given frame rate. A sequence of frames forming audiovisual data can be simply referred to as audiovisual data or audiovisual data having a specified length (e.g., 50 frames). Furthermore, audiovisual data can be segmented, or spliced, or split into different sequences of various lengths. For example, audiovisual data of length 50, in which each frame is indexed using a number between 1 and 50, can be segmented into two different sequences of audiovisual data, each 25 frames long (e.g., a first audiovisual data consisting of frames 1 through 25 and a second audiovisual data consisting of frames 26 through 50, relative to the frame indexing according to the original audiovisual data having a length of 50 frames).
[0013] It is not uncommon for ASD algorithms to report computational time complexity proportional to the number of frames in the audiovisual data on which they operate. That is, using Big-O notation, an ASD algorithm may have a time complexity of O(n), where n is the number of frames in the provided audiovisual data. This consideration clearly demonstrates the trade-off between accuracy and required computational time (or speed): increasing the length of the audiovisual data to provide greater temporal context, and therefore greater accuracy in detecting active speakers, is associated with an increase in computational time proportional to the length of the audiovisual data. As demonstrated herein, an ASD system according to one or more embodiments enables the incorporation of greater temporal context without a significant increase (if any) in computational time. Thus, the embodiments disclosed herein can detect active speakers in a visual scene from audiovisual data with high accuracy in real time.
[0014] FIG. 1 illustrates an ASD system (100) according to one or more embodiments of the present disclosure. The ASD system (100) includes a visual sensor (102), e.g., one or more cameras. The visual sensor (102) may have a field of view (FOV). Field of view is used herein as a general term intended to indicate the extent of the observable world seen by the visual sensor (102) (e.g., a camera). The visual sensor (102) may capture a visual scene including one or more people, any of whom may or may not be actively speaking at any given time. The visual sensor (102) captures the visual scene as visual data organized into a sequence of frames, each frame representing the visual scene at a point in time. The ASD system (100) includes an audio sensor (104), e.g., an array of microphones or vibration sensors. Typically, the audio sensor (104) converts sound waves into electrical signals, i.e., a time series of discretized amplitudes representing the sound waves. The audio sensor (104) can capture sounds of its environment, including the sounds of the visual scene captured by the visual sensor (102). The captured audio or electrical signals generated by the audio sensor (104) may be referred to as audio data. Audiovisual data includes at least one of audio data and visual data. The visual sensor (102) and audio sensor (104) are packaged by a single device or subsystem and are considered the same sensor. For example, a camera may be configured to act as both a visual sensor (102) and an audio sensor (104) to capture both images and sounds associated with a visual scene.
[0015] Continuing with reference to FIG. 1 , the ASD system (100) includes a computer system (110) communicatively coupled to the visual sensor (102) and the audio sensor (104) and configured to receive and process audiovisual data. The computer system (110) includes one or more computer processors (112) and a data storage device (114) such as one or more of non-persistent storage devices (e.g., volatile memory such as random access memory (RAM) or cache memory) and persistent storage devices (e.g., hard disks, optical drives such as compact disc (CD) drives or digital versatile disc (DVD) drives, flash memory, etc.). The processor (112) may be part or all of an integrated circuit for processing instructions. For example, the processor (112) may be or include one or more cores or micro-cores. The computer system (110) further includes a communication interface (not shown), which may include integrated circuits for connecting to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, a mobile network, or any other type of network) and / or to another device, such as the visual sensor (102) and the audio sensor (104).
[0016] Software instructions configured to execute one or more embodiments of the present disclosure may be stored on a non-transitory computer-readable storage medium. The software instructions may include instructions to implement embodiments of the present disclosure, such as processing audiovisual data using a detection model (120) described below to detect or determine an active speaker in a visual scene. The non-transitory computer-readable medium may be, for example, a CD, DVD, flash memory, or storage device. In one or more embodiments, the non-transitory computer-readable medium is the data storage device (114) of the computer system (110). In other embodiments, the computer system (110) is configured to access, read, and execute the software instructions stored on the non-transitory computer-readable medium. Furthermore, the software instructions may be stored, in whole or in part, on a temporary or permanent basis on the non-transitory computer-readable medium.
[0017] The ASD system (100) includes a detection model (120). The detection model (120) may be stored, for example, in the data storage device (114) or as part of a non-transitory computer-readable medium. The detection model (120) accepts audiovisual data, or segments of audiovisual data, and returns detection results (130), which identify one or more active speakers (or a lack of active speakers) in a visual scene at various points in time. For example, a visual scene may include a dialogue between four characters A, B, C, and D over a time period T, where zero or more of characters A, B, C, and D may be actively speaking at any instant within the time period. The visual scene over time period T is represented as audiovisual data acquired using one or more of the vision sensor (102) and the audio sensor (104). The audiovisual data may consist of N frames over time period T. The detection results include a score for each character in each frame, where the score indicates whether the associated character is actively speaking at the associated frame instance. The score may be a categorical variable or a continuous variable. In one example, the score is a binary classification with classes being "actively speaking" and "not actively speaking." In another example, the score is a non-binary classification with classes, for example, "actively speaking," "not actively speaking," and "indeterminate." In yet another example, the score is a continuous variable in the range [0, 1] indicating the probability that a person is actively speaking. The detection model (120) processes the audiovisual data and returns a detection result (130), which indicates the speech status of a person in the visual scene represented by the audiovisual data.
[0018] The detection model (120) may be implemented using the computer system (110) and may include one or more machine learning models and further encompass various pre-processing and post-processing steps. Machine learning (ML) is broadly defined as the extraction of patterns and insights from data. Thus, in some implementations, the machine learning model determines an outcome, such as a prediction or detection, based on perceived patterns in the received data.
[0019] One type of machine-learned model is a neural network. A neural network (200), such as that shown in FIG. 2, may be used as a subcomponent of a larger machine-learned model, such as the detection model (120). The neural network (200) is shown here as a graph composed of nodes (202) and edges (204). The nodes (202) are represented as solid circles; to avoid cluttering the diagram, not all nodes are given numerical labels. Similarly, the edges (204) are shown as solid lines. Generally, the edges (204) of the neural network (200) are "directed," and, borrowing from graph terminology, the neural network (200) can be classified as a directed acyclic graph (DAG). Therefore, the edges (204) in FIG. 2 are more specifically depicted as directed lines. Again, to avoid cluttering the diagram, not all depicted edges (204) are given numerical labels.
[0020] Nodes (202) may be grouped to form layers. Figure 2 shows four layers (208, 210, 212, 214) of nodes (202), each consisting of a columnar group of nodes (202). In general, the grouping of nodes (202) and the formation of layers need not be as shown in Figure 2. For example, edges (204) may or may not connect to any node (202), regardless of which layer the node (202) is in. That is, edges (204) may form sparse and residual connections (e.g., so-called "skip" connections) between nodes (202). A layer and its adjacent layers are said to be fully, or densely, connected if all nodes (202) in the layer are connected to all nodes in the adjacent layers. In the neural network (200) of Figure 2, all of its layers are densely connected to their adjacent layers, if applicable. In this way, the neural network (200) of FIG. 2 may be said to be a fully connected or densely connected neural network (200).
[0021] The neural network (200) has at least two layers: an "input layer" (208) and an "output layer" (214). Zero or intermediate layers (210, 212) may exist between the input layer (208) and the output layer (214). The intermediate layers (210, 212) are generally referred to as "hidden layers." Furthermore, a neural network (200) having at least one hidden layer (210, 212) may be described as a "deep" neural network or a "deep learning method." In some embodiments, as described below, the detection model (120) includes a deep neural network. The output layer (214) of the neural network (200) may have multiple nodes (202). When the output layer (214) of the neural network (200) has multiple nodes (202), the neural network (200) may be referred to as a "multi-target" or "multi-output" network.
[0022] Additionally, each edge (204) in the neural network (200) has a numerical value associated with it. The numerical value of an edge (204), or even the edge (204) itself, is often referred to as a "weight" or "parameter." Thus, the neural network (200) may be said to include or be parameterized by a set of weights or parameters. The neural network (200) is "trained" by assigning a numerical value to each trainable edge of the neural network (200) through evaluation of a set of data, commonly referred to as training data (described below). The distinction "trainable edge" is introduced here if a trainable edge is an edge whose numerical value can be adjusted during a training routine. Generally, non-trainable edges have numerical values, but their values are determined using a process separate from the training process, such as direct assignment by a user.
[0023] Similarly, nodes (202) carry, pass, or temporarily store numerical values and are further associated with activation functions. Activation functions are not limited to any class of functions, but conventionally apply a function to the dot product of an array of values of nodes ("input nodes") connected to or directed at the node to which the activation function is to be applied ("activation node") and an array of weights or parameters of the edges connecting the input nodes to the activation node. An input node (202), when viewed as a graph (similar to Figure 2), is a node with a directed arrow pointing to the activation node whose numerical value is being calculated. Some commonly used activation functions are linear functions
number
number
number
[0024] When a neural network (200) receives an input, the input is propagated through the network according to the activation functions of the neural network's (200) nodes (202) and the neural network's (200) edge (204) values. Thus, the numerical values of the nodes (202) may change for each received input. In some cases, nodes (202) are assigned a fixed numerical value that is unaffected by the input, such as a value of 1. Nodes (202) with fixed numerical values (invariant to inputs) are often referred to as "biases" or "bias nodes" (206), shown in FIG. 2 by dashed circles.
[0025] In some implementations, the neural network (200) may include specialized layers, such as normalization layers, dropout layers, and concatenation layers. For the sake of brevity, such layers are not described herein, but those skilled in the art will recognize that the inclusion and use of such layers in the neural network (200) is within the scope of this disclosure.
[0026] As described above, the process of training a neural network (200) consists of, at a minimum, assigning values to the edges (204) of the neural network (200). Training begins with the neural network (200) having edge values initially provided by some initialization mechanism or procedure. The edge values may be assigned randomly, according to a prescribed distribution, manually, or by some other assignment procedure. Using the initial edge values, the neural network (200) may be said to act as a function that receives inputs and generates outputs. Thus, one or more inputs may be propagated through the neural network (200) to generate one or more associated outputs. During training, a training set, or training data, is provided to the neural network (200). The training set consists of inputs and associated targets, where the targets represent desired outputs, often observations or "ground truth" associated with the observed inputs. During training, the neural network (200) processes the inputs to generate outputs, which are compared to the associated targets. The comparison of the generated neural network output to a target is performed using a “loss function,” such as a mean squared error function, a mean absolute error function, or a logarithmic loss function (or a binary cross-entropy function). Generally, a loss function provides a numerical assessment of the similarity between the neural network (200) output and a given target. In some implementations, a loss function may be composed of multiple loss functions applied to different parts of the output-target comparison. A loss function may also be constructed to impose additional constraints on the values taken by the edges (204). For example, a loss function may include a regularization or penalty term, which may be physics-based, that influences or constrains the values of the edges (204). Overall, the goal of the training process is to modify the edge (204) values so that the output of the neural network (200) when processing a given input is similar to the target associated with the given input.In other words, the intent of training is to promote similarity between the neural network (200) outputs and associated targets across a dataset provided for training (e.g., training data). Changes in the values of the edges (204) are guided by a loss function, typically through a process called "backpropagation."
[0027] Backpropagation consists of calculating the gradient of a loss function with respect to the values of trainable edges (204). The gradient indicates the change in edge (204) values that, when applied to the edges (204), produces the largest change in the loss function with respect to the training data provided when calculating the gradient. The edge (204) values are typically updated by "steps" in the direction that follows the gradient. The step size, often referred to as the "learning rate," need not remain fixed during the training process. Furthermore, the step size updates to the edge (204) values may be informed by previously seen edge (204) values or previously calculated gradients.
[0028] Updates to the edge values of the neural network (200) are applied iteratively. In other words, the training process consists of repeatedly calculating the gradient of the loss function with respect to the edge (204) values and updating the edge (204) values in steps guided by the gradient. This process continues until a termination criterion is reached. For example, the termination criterion may consist of one or more of the following: reaching a fixed number of edge (204) updates, otherwise known as an iteration counter; noting that there is no noticeable change in the loss function between iterations (or that the change to the edge values between updates is less than a predetermined threshold); or reaching a specified performance metric, as evaluated on the training data or a separate holdout data set. Once the termination criterion is met and the edge (204) values are no longer intended to be updated, the neural network (200) is said to be "trained." The loss function can be constructed such that increasing the loss function increases the similarity between the output and the target, so that the training process can be viewed as maximizing the loss function. Similarly, a loss function can be constructed such that the similarity between the output and the target decreases as the loss function decreases, so that the training process can be viewed as minimizing the loss function. The tasks of maximization and minimization can be made equivalent through techniques such as sign reversal.
[0029] A machine-learned model architecture defines the "structure" of a machine-learned model. For example, in the case of a neural network (200), the structure is specified by, among other things, the number of hidden layers in the network, the type of activation function used, and the number of outputs, including the use and location of dedicated layers (e.g., batch normalization layers). The architecture of a machine-learned model is specified by a set of "hyperparameters." For example, in the case of a neural network (200), the number of hidden layers and the number of nodes (202) in each layer are hyperparameters of the neural network (200).
[0030] Another type of machine-learned model is the convolutional neural network (CNN). Similar to neural networks (200), CNNs can be thought of or depicted as consisting of a series of nodes connected by edges.
[0031] Yet another type of machine learning model is a recurrent neural network (RNN), as shown in Figure 3A. As shown, the RNN includes an RNN block (310) and a recurrent connection (350). The RNN block (310) can accept inputs (320) and states (330) and generate outputs (340). That is, the RNN block (310) applies one or more operations (e.g., matrix multiplication) to the inputs (320) and states (330) to generate the outputs (340). Additionally, as described below, the RNN block (310) can modify the states (330) of the inputs or between elements of the inputs.
[0032] The RNN block (310) typically includes one or more data structures (e.g., matrices, arrays, tensors, etc.) containing the weights or parameters of the RNN. These weights or parameters are similar to the weights or parameters of the neural network (200) (e.g., edge values) or filters of a CNN. A distinction can be made between data structures containing weights associated with biases (e.g., bias nodes in neural network terminology) and data structures containing weights applied to node values that vary based on received input. For simplicity, the data structure for weights associated with bias values is hereafter referred to as a bias vector, while the data structure for other weights is referred to as a matrix. Furthermore, the given example considers the input (320) as a vector. However, RNNs operating on high-dimensional inputs (e.g., inputs with a tensor rank of 2 or greater) can structure their weights in a higher-dimensional data structure such as a tensor rather than a matrix or vector.
[0033] In one or more implementations, the RNN block (310) has two weight matrices and a single bias vector. A commonly used naming convention is to refer to one weight matrix as
number
number
number
[0034] Continuing with reference to Figure 3A, the input (320) to the RNN block (310) is an element of a sequence. For example,
number
number
number
number
number
number
[0035] To process a sequence, the RNN receives the first ordered input (320) of the sequence along with the state (330).
number
number
number
number
number
number
number
number
[0036] More specifically, with reference to the matrix and vector labels discussed above, the process of the RNN block (310) can generally be described as follows:
number
number
number
number
number
[0037] Figure 3B shows an "unrolled" version of the RNN of Figure 3A.
number
[0038] As previously described with respect to the neural network (200), training a machine learning model such as an RNN requires that a pair of inputs and one or more targets (i.e., a training dataset) be provided to the machine learning model along with the implementation of the training process. In the context of an RNN, the RNN receives a sequence of one or more elements (input (320)) and processes the sequence to form an output, which may also be a sequence. The overall output of the RNN is referred to herein as the RNN result. In other words, the RNN receives a sequence of one or more elements and generates an RNN result. Thus, the training procedure for an RNN consists of determining values for the weight matrix and bias vector of the RNN block (310) by comparing the RNN result generated by processing the input sequence against the associated target using a comparison function (e.g., a loss function). As with the neural network (200) and CNN described above, the comparison function is typically used to guide changes made to the RNN weights through a process called "backpropagation through time," similar to the backpropagation process described above.
[0039] Various adaptations of RNNs can be made that result in specific, or at least well-specified, machine learning models. For example, gated recurrent networks (GRUs) and long short-term memory (LSTM) networks can each be considered instances or specific types of RNNs. The regression blocks of these machine learning models typically add additional weights (e.g., additional weight matrices, additional bias vectors, etc.) to process received inputs and states, as well as other quantities, in more complex ways. For example, LSTMs determine a separate "state-like" data structure, commonly called a "carry."
[0040] The ASD system (100) includes a detection model (120) that is configured with additional machine learning models. For example, the detection model (120) includes a neural network (200), a CNN, and an RNN. While at least some of the machine learning models described herein are neural networks (200), CNNs, and RNNs, other machine learning models can be used by or included in the detection model (120), including, but not limited to, the following: a neural network (200), a CNN, and an RNN. For example, the detection model (120) may include a Transformer (i.e., another type of machine learning model).
[0041] Figure 4 illustrates the flow of audiovisual data (401) through an audiovisual encoder (410) and a classifier (430) according to one or more embodiments. As will be described in more detail below, the detection model (120) can include both an audiovisual encoder (410) and a classifier (430), although the connections need not be as shown in Figure 4. The audiovisual data (401) is received as input by the audiovisual encoder (410) and ultimately converted into an ASD score (440). In other words, Figure 4 illustrates the generation of an ASD score (440) given audiovisual input. The ASD score (440) can be similar to the detection result (130) described above. That is, the ASD score (440) can include one or more continuous or categorical values indicating the speech status (e.g., "actively speaking" or "not actively speaking") of each person in the visual scene. However, as explained below, a distinction is made between the ASD scores (440) and the detection results (130) in that the detection results (130) can be formed from a collection of ASD scores (440).
[0042] FIG. 4 also depicts the audiovisual data (401) as consisting of visual data (402) (e.g., a sequence of frames or images) and audio data (404) (e.g., a time series of vibration amplitudes). In practice, the audiovisual data (401) need only include at least one of the visual data (402) and audio data (404). Furthermore, no distinction is made between the visual data (402) and the audio data (404). For example, if a single device includes or functions as both a visual sensor (102) and an audio sensor (104), the output of the device may simply be referred to as the audiovisual data (401). In other examples, the audiovisual data (401) may be separated or otherwise divided into the visual data (402) and the audio data (404). As shown in FIG. 4, the audiovisual data (401) is received as input by an audiovisual encoder (410). The audiovisual encoder (410) processes the visual data (402) and the audio data (404) separately using a visual encoder (412) and an audio encoder (414), respectively. That is, the audiovisual encoder (410) may include a visual encoder (412) and an audio encoder (414). The visual encoder (412) and the audio encoder (414) may each be a CNN. If the audiovisual data (404) includes only one of the visual data (402) or the audio data (404), the audiovisual encoder (410) may only include or be considered to be the visual encoder (412) or the audio encoder (414), respectively.
[0043] In Figure 4, which illustrates the audiovisual encoder (410) as including a visual encoder (412) and an audio encoder (414), the visual encoder (412) receives the visual data (402) and generates a visual encoding, and the audio encoder (414) receives the audio data (404) and generates an audio encoding. The visual encoding may be a vector of visual features (i.e., a visual feature vector), one for each frame of the audiovisual data (401). The visual encoding may be a two-dimensional array (or matrix) of values. For example, if the audiovisual data (401) includes N frames and the visual encoder (412) generates M v When determining features, the size of the visual encoding is N × M. v (or M v ×N). The audio encoding may be a vector of audio features (i.e., an audio feature vector), one for each frame of the audiovisual data (401). The audio encoding may also be a two-dimensional array (or matrix) of values. For example, if the audiovisual data (401) contains N frames, and the audio encoder (414) generates M a When determining features, the size of the visual encoding is N × M. a (or M a ×N). The visual and audio encodings are concatenated to form the embedding (420). Continuing with the previous example, the embedding (420) may also be N×(M v +M a ) can be a two-dimensional array (or matrix) with size
[0044] In other embodiments, the audiovisual encoder (410) cannot be divided or otherwise represented as a visual encoder (412) and an audio encoder (414). In such cases, the audiovisual encoder (410) accepts audiovisual data (401) and generates an embedding (420). The embedding (420) is a two-dimensional array (or matrix) with the length of one dimension corresponding to the number of frames in the audiovisual data (401) and the length of the other dimension corresponding to the number of features determined from each by the audiovisual encoder (410), where the features do not necessarily correspond exclusively to visual or exclusively to audio aspects of the audiovisual data (401).
[0045] Referring to FIG. 4, the embeddings (420) are passed to a classifier (430). The classifier (430) can be an RNN that operates on the embeddings as an ordered sequence of feature vectors, ordered according to their corresponding frames. The classifier returns an ASD score (440). The ASD score (440) provides an indication of whether each person in the visual scene is actively speaking. In some embodiments, the ASD score is applicable across the duration of the visual scene represented by the audiovisual data (401). That is, in some embodiments, the indication of whether a person is speaking is not determined on a frame-by-frame basis, but rather, frame-to-frame information (i.e., temporal context) is used to indicate each person's speech class (e.g., "actively speaking" or "not actively speaking") and / or speech probability (e.g., a value between 0 and 1), which indication is associated with the entirety of the audiovisual data (401).
[0046] According to one or more embodiments, Figure 5 illustrates in more detail the concatenation of the visual encoding (502) and audio encoding (504) returned by the audiovisual encoder (410) to form the embedding (420). As shown in Figure 5, given audiovisual data (401) having a length of N frames, the audiovisual encoder (410) (or, more specifically, the visual encoder (412)) can be used to determine the visual encoding (502). The visual encoding (502) is represented in Figure 5 as a feature vector
number
number
number
number
number
number
[0047] The visual feature vectors from the visual encoding (502) and the audio feature vectors from the audio encoding (504) are concatenated frame by frame to form the embedding (420). In other words, the embedding (420) is represented by the feature vectors
number
number
number
[0048] FIG. 5 further illustrates the internal operation of the classifier (430). The classifier (430) includes a forward RNN (510). As shown, the forward RNN (510) accepts a sequence of audiovisual feature vectors, starting with a "first" (e.g., of the first frame) audiovisual feature vector and progressing to a "last" (e.g., of the last frame) audiovisual feature vector. The forward RNN (510) is a gated recurrent network (GRU). FIG. 5 shows the forward RNN (510) "unfolded" using a GRU block. As shown, the classifier (430) further includes a backward RNN (520). As shown, the backward RNN (520) accepts a sequence of outputs generated each time the RNN block processes an input in the forward RNN (510). However, the backward RNN (520) accepts the sequence of forward RNN (510) outputs in reverse order. The inverse RNN (520) is a gated recurrent network (GRU). Figure 5 further illustrates the inverse RNN (520) "deployed" using a GRU block. The classifier (430) further includes a neural network (530). The neural network (530) receives as input a sequence of outputs generated each time the RNN block in the inverse RNN (520) processes an input. In response, the neural network (530) returns an ASD score (440) for the audiovisual data (401).
[0049] The detection model (120) of the ASD system (100) described herein includes an audiovisual encoder (410) and a classifier (430) similar to those depicted in Figures 4 and 5. Furthermore, the audiovisual encoder (410) has a processing time proportional to the length or number of frames in the received audiovisual data (401), and the classifier (430) is processing time invariant or nearly invariant to the length of the audiovisual data (401) (and thus the length of the embedding (420)). Using big-O notation, the time complexity of the audiovisual encoder (410) is O(n), where n is the number of frames provided in the audiovisual data (401), and the time complexity of the classifier is O(1) (or a constant).
[0050] When processed by the audiovisual encoder (410) of the detection model (120), the audiovisual data (401) is limited to a predetermined number of frames (e.g., 10 frames). For example, the detection model (120) can operate on an incoming (e.g., live) video stream, and the audiovisual encoder (410) can be applied to process the latest segment of N frames (e.g., 10 frames) of the video stream. Furthermore, a predetermined number of embeddings resulting from application of the audiovisual encoder (410) to previously encountered (and processed using the audiovisual encoder (410)) segmented audiovisual data (one segment per N frames) are concatenated and organized (e.g., ordered, spliced, etc.) according to a predetermined set of rules to form one or more composite embeddings. The one or more composite embeddings are then each processed by a classifier (430), which returns one or more ASD scores (440). The one or more ASD scores (440) are aggregated according to a predetermined aggregation function to form the detection result (130). Thus, the detection model (120) of the active speech detection (ASD) system (100) is configured such that the audiovisual encoder (410) processes periodic segments of the audiovisual data (401) constrained to a fixed size, while the classifier (430) operates on one or more combinations, concatenations, and / or splices of embeddings determined using the audiovisual encoder (410). As described above in accordance with one or more embodiments, an advantage of the detection model (120) separating the audiovisual encoder (410) and the classifier (430) is that, due to the limited number of frames processed by the audiovisual encoder (410), the detection model (120) can process segments of audiovisual data quickly (i.e., in real time) and simultaneously achieve high accuracy by taking into account pre-determined embeddings when applying the classifier (430). Thus, the ASD system (100) described herein can accurately determine active speakers in a visual scene in real time.The process of using an audiovisual encoder (410) to process audiovisual data constrained to a predetermined number of frames, followed by using a classifier (430) to process one or more composite embeddings, and calculating a detection result (130) based on an aggregation of one or more ASD scores returned by the classifier (430) is shown in Figure 6.
[0051] Referring to Figure 6, which illustrates an incoming (e.g., live) video stream (602) according to one or more embodiments, the video stream (602) may be acquired using one or more of the visual sensor (102) and the audio sensor (104). The video stream (602) is also comprised of a sequence of frames (audio data, if present, may also be segmented into frames). The frames of the video stream are numbered, with the first frame in the sequence being F1, the second frame being F2, and so on, until there are the same number of frames in the video stream (602). If the video stream (602) is live, new frames may be added to the acquired rendered video stream (602).
[0052] The video stream (602) is segmented into consecutive sets of frames according to a predetermined number indicating the number of frames in the set. That is, the detection model (120) can be configured according to the number of frames in the set. For example, a set is configured for every 10 frames in the video stream (602). Thus, each set forms a sequence of audiovisual data having a fixed predetermined length. If the video stream (602) is live, a set of frames (or a segment or sequence of audiovisual data) is formed in real time as frames are received according to the predetermined number of frames in the set. In the example of FIG. 6, the predetermined number of frames in the set is set to 10 frames. As can be seen in the example of FIG. 6, the video stream (602) currently has 50 frames, and frame F 41 ~F 50are the most recent frames in time, forming a set. Each set of frames (forming the audiovisual data) is processed by the audiovisual encoder (410). Figure 6 shows a set of the most recent 10 frames (F) of the video stream (602) that form the embedding (420). 41 ~F 50 ) is shown. The embedding is performed on the frame F 41 ~F 50 Since it is specific to
number
number
[0053] The previously determined embedding can be stored as a previous embedding (604). For example, FIG. 6 shows the previous embedding for frame F 11 ~F 20 , F 21 ~F 30 , and F 31 ~F 40 three other embeddings corresponding to the output of the audiovisual encoder (410) with processed audiovisual data constructed from a set of frames each including
number
number
number
number
number
number
number
[0054] Continuing with the example of FIG. 6, the previous embedding (604) and the last determined embedding (e.g.,
number
[0055] The concatenated embeddings (650) are processed according to one or more rules, with application of each rule resulting in one or more composite embeddings that can be passed as inputs to the classifier (430) of the detection model (120). In some embodiments, one or more rules can be applied to one or more of the previous embeddings (604) and the latest embedding without the need to form the concatenated embeddings (650). That is, in fact, the step of forming the concatenated embeddings (650) can be skipped, as shown in FIG. 6 for ease of understanding. In general, the detection model (120) can be configured with K rules, where K is an integer greater than or equal to 1. Each rule specifies a process for forming at least one composite embedding that is passed to the classifier to obtain an ASD score. For example, a rule can specify a window and stride by which a composite embedding is formed from audiovisual feature vectors bounded by the window applied to the concatenated embedding (650), where the window is moved through or slid over the audiovisual feature vectors of the concatenated embedding. For example, the first rule (i.e., rule 1) specifies a window size of 10 frames and a stride of 10 frames. Using the concatenated embedding (650) of FIG. 6 as an example, frames 11 through 50 (i.e., embedding
number
number
number
number
number
number
number
number
number
number
number
[0056] As shown in FIG. 6, the resulting compound embeddings, or embedding representations, generated by applying rules 1 through K (K≧1) to the concatenated embedding (650) are each independently processed by a classifier (430). Processing a given compound embedding through the classifier (430) results in an ASD score (440). Thus, the detection model (120) generates at least one ASD score (440). As shown in FIG. 6, one or more ASD scores are aggregated according to an aggregation function (660) to determine the detection result (130). Consider a vector V of length K, where the kth element in vector V indicates the number of compound embeddings returned through the application of the kth rule to the concatenated embedding (650). For example, in the example shown in FIG. 6, applying rule 1 returns four compound embeddings, and applying rule K returns one compound embedding. Thus, in this case, vector V is: [4,...,1]. Also, denoting the elements of vector V as V k Furthermore, the k-th rule is used to generate (V k The ASD score determined by processing the th composite embedding was
number
number
[0057] In one or more embodiments, the aggregation function is the average of the ASD scores after averaging over each rule, or mathematically,
number
[0058] The example in Figure 6 illustrates a case where the number of sets of frames processed by the audiovisual encoder (410) exceeds the number of listed embeddings already retained. If the number of sets of frames processed by the audiovisual encoder (410) is not yet equal to or greater than the number of frames specified to be retained, one or more of the rules for operating the concatenated embedding (450) may not be applicable. For example, if only the first 10 frames of the video stream (602) are available, then the previous embedding (604) does not contain any embeddings. Furthermore, applying the exemplary first rule, as described above, would simply return the audiovisual feature vectors of the first 10 frames (i.e.,
number
[0059] Figure 6 shows the results of a previously processed frame set (e.g., set F 11 ~F 20 , set F 21 ~F 30 , set F 31 ~F 40 ) embedding, the detection result (130) is calculated for a given set of frames (F 41 ~F 50 ), it is emphasized that the detection model (120) can be applied to any set of frames, taking into account previously determined embeddings, if any, for the frame set. This allows a detection result to be determined for any set of frames. In general, the detection model (120) can be configured to operate on a set of N frames, taking into account J previously obtained embeddings, if available. For example, for frame F i ~F i+N For a set of frames F i ~F i+N is processed by an audiovisual encoder (410) and embedded
number
number
[0060] The configuration of the detection model (120), or more generally, the ASD system (100), is defined through a set of configuration parameters. A set of configuration parameters (702) is shown in Figure 7. The set of configuration parameters defines, among other things, the number of frames (704) in a set of audiovisual data that are processed at a given time by the audiovisual encoder (410), the number of embeddings (706) that are retained and included in the previous embeddings (604), the definition of all embedding operation rules (708), and the definition of an aggregation function (710). Other configuration parameters that govern the behavior of the ASD system (100) can be included in the set of configuration parameters (702).
[0061] The ASD system (100) further includes speaker detection logic that modifies or enhances the detection result based on one or more previously determined or past detection results. That is, when determining whether a person is actively speaking, the ASD system (100) can further make the determination by taking into account the person's past active speech history. The detection result includes a speech metric for each person in the visual scene. The speech metric is a quantitative indication of whether a person in the visual scene is speaking. For example, the speech metric may be a probability that the person is speaking. Thus, the speech metric for each person obtained from the detection result can be compared to a threshold to determine whether the person is speaking. The threshold can be predefined by a user (e.g., set by an operator, administrator, or system administrator providing or using the ASD system (100)), customized on the fly (e.g., selected or adapted by a user while using the ASD system (100)), or selected or determined based on a given context (e.g., dependent on the audiovisual data (401) or embeddings (420), dependent on the number of people in the visual scene, etc.). For example, the threshold may be assigned a predetermined value based on a determined or user-provided number of people in the visual scene.
[0062] Each person is associated with a status. The status of each person can be one of "active" and "inactive," which indicate whether the person is speaking (or actively speaking) and whether the person is not speaking (or not actively speaking), respectively. The speaker detection logic of FIG. 8 determines the status of each speaker based on both the detection result (i.e., a quantitative indication of whether each person is speaking) and each person's last known status. Initially, the status of all people in the visual scene is set to "inactive."
[0063] Referring to Figure 8, in block 802, a person's speech metrics and state are obtained, where the state is the person's last known state and is initially set to "inactive." In block 804, the person's speech metrics are compared to a threshold. If the speech metric (e.g., the probability that the person is speaking) is greater than the detection threshold, the speaker detection logic of Figure 8 proceeds to block 810. In block 810, the given person's state is set to "active," thus detecting that the person is actively speaking. The threshold can be predefined by the user, customized on the fly, or selected or determined based on a given context.
[0064] However, if at block 804 the person's speech metric is less than or equal to the threshold, the speaker detection logic of FIG. 8 proceeds to block 806. At block 806, it is checked whether the given person's current state is "active." If the person's current state is not active, the speaker detection logic of FIG. 8 proceeds to block 812, where the person's state remains "inactive." If the person's current state is "active," the speaker detection logic of FIG. 8 proceeds to block 808. At block 808, the given speaker's speech metric is compared to the speech metrics of all other people in the visual scene. If the given person's speech metric is greater than the speech metrics of all other people, the person's state is kept as "active," as indicated by block 810. However, if another person in the visual scene has a speech metric greater than the given person's speech metric, then at block 812 the given person's state is set to "inactive."
[0065] 9 illustrates a method according to one or more embodiments of the present disclosure. As shown, in block 902, a first set of frames and a second set of frames are acquired from a visual sensor (e.g., a camera). The visual sensor captures a visual scene including at least a first person. The second set of frames are captured by the visual sensor after the first set of frames in time. For example, the visual sensor may generate a video stream. The first set of frames may include frames F1-F2. 10 and the second set of frames may include frames F 11 ~F 20 may include:
[0066] At block 904, each of the first and second sets of frames is processed by the audiovisual encoder to generate a first embedding and a second embedding, respectively. The first set of frames is processed by the audiovisual encoder as it is acquired, even though the second set of frames has not yet been acquired. Similarly, the second set of frames can be processed by the audiovisual encoder as it is received. That is, blocks 902 and 904 do not need to be performed in strict order, but can be performed in parallel.
[0067] At block 906, one or more embedding manipulation rules are applied to the first embedding and the second embedding (or a concatenated embedding formed from at least the first embedding and the second embedding). The application of the embedding manipulation rules forms at least one compound embedding. Thus, the application of the one or more embedding manipulation rules to the first and second embedding forms one or more compound embeddings. The embedding manipulation rules specify window sizes and strides for extracting splices (or windows) from the concatenated embedding formed from at least the first embedding and the second embedding.
[0068] At block 908, each of the one or more composite embeddings is processed by a classifier, which upon processing by the classifier, each results in an active speech detection (ASD) score, which, for each person in the visual scene that includes the first person, includes a probability that the given person is speaking.
[0069] At block 910, one or more ASD scores (resulting from each of the one or more composite embeddings) are aggregated according to an aggregation function to form a detection result. The aggregation function averages, for each person, the probability that the person is speaking across all ASD scores. The detection result includes, for each person in the visual scene that includes the first person, a speech metric that is a quantitative measure of whether the person is speaking.
[0070] At block 912, a determination is made based on the detection result as to whether the first person is speaking. The determination is made by comparing the speech metrics of the first person to a threshold. In some embodiments, the determination is further based on a current speech state of the first person and a comparison of the speech metrics of the first person to the speech metrics of all other people in the visual scene, if any.
[0071] In block 914, in response to determining that the first person is speaking, the display of the visual scene is adjusted to focus on the first person. For example, the display of the visual scene may frame or zoom in on only the first person.
[0072] Embodiments of the present disclosure have one or more of the following advantages. Embodiments of the present disclosure may provide high-accuracy ASD in real time. According to one or more embodiments, high-accuracy ASD in real time is achieved by: 1) using an audiovisual encoder to process a constrained number of frames (e.g., 10 frames as they are received) in real time (i.e., to form a current embedding); 2) retaining several sets of previously processed and embedded frames (i.e., previous embeddings); 3) applying a classifier to one or more combinations, splices, or windows of the current embedding and one or more previous embeddings; and 4) aggregating the results of the classifier. In other words, because the audiovisual encoder processes a constrained number of frames, it can generate embeddings for a given set of frames in real time. However, the classifier has access to and processes previous embeddings and can therefore integrate a larger temporal context when determining an ASD score.
[0073] While only a few exemplary embodiments have been described in detail above, those skilled in the art will readily appreciate that many modifications are possible in the exemplary embodiments without substantially departing from the invention, and all such modifications are therefore intended to be included within the scope of the present disclosure as defined in the appended claims. [Explanation of symbols]
[0074] 100 ASD System 102 Visual Sensor 104 Audio Sensor 110 Computer Systems 112 Computer Processors 114 Data storage device 120 detection models 130 Detection Results 200 Neural Networks 202 nodes 204 Edge 206 Bias Node 208 Input Layer 210, 212 hidden layer 214 Output Layer 310 RNN blocks 320 Input 330 Status 340 output 350 Recursive Bonds 401 Audiovisual Data 402 Visual Data 404 Audio data 410 Audiovisual Encoder 412 Visual Encoder 414 Audio Encoder 420 Embed 430 Classifier 440 ASD score 502 Visual Encoding 504 Audio Encoding 510 Forward RNN 520 Reverse RNN 530 Neural Networks 602 Video Streams Embedding before 604 650 Concatenated Embedding 660 Aggregate Functions 702 Set of configuration parameters 704 Number of frames in the set 706 Embeddings Retained 708 Embedding Operation Rules 710 Aggregate Functions
Claims
1. a visual sensor for capturing a visual scene including a first person; Computer systems and 1. An active speaker detection system comprising: The computer system includes: one or more computer processors; A detection model with an audiovisual encoder and a classifier Equipped with the computer system is communicatively coupled to the visual sensor; acquiring a first set of frames and a second set of frames from the visual sensor; generating a first embedding and a second embedding from the first set of frames and the second set of frames, respectively, using the audiovisual encoder; generating one or more composite embeddings from the first embedding and the second embedding; using the classifier to determine an active speaker detection (ASD) score for each of the one or more composite embeddings; aggregating the one or more ASD scores to form a detection result; determining whether the first person is speaking based on the detection result; Upon determining that the first person is speaking, adjusting the display of the visual scene to focus on the first person. It is configured as follows: Active speaker detection system.
2. The active speaker detection system of claim 1 , wherein the determination of whether the first person is speaking corresponds to the second set of frames.
3. The active speaker detection system of claim 1 , wherein the second set of frames is temporally subsequent to the first set of frames.
4. the first embedding and the second embedding each include a plurality of audiovisual feature vectors; the number of said audiovisual feature vectors is equal to the number of frames in said first set or said second set; The active speaker detection system of claim 1 .
5. The active speaker detection system of claim 1 , wherein the audiovisual encoder comprises a neural network and the classifier comprises a recurrent neural network.
6. the detection result includes a first speech metric for the first person; determining whether the first person is speaking includes comparing the first speech metric to a threshold; determining that the first person is speaking in response to the first speech metric being greater than the threshold; The active speaker detection system of claim 1 .
7. the visual scene further includes a second person, and the detection result further includes a second speech metric for the second person; The determining whether the first person is speaking includes: obtaining a state for the first person in response to the first speech metric of the first person being less than or equal to the threshold; determining whether the first speech metric is greater than the second speech metric and whether the state of the first person is active; further comprising determining that the first person is speaking in response to the state of the first person being active and the first speech metric being greater than the second speech metric; 7. The active speaker detection system of claim 6.
8. 1. A method for determining whether a person in a visual scene including a first person is speaking, comprising: acquiring a first set of frames and a second set of frames from a visual sensor capturing the visual scene; generating, using an audiovisual encoder, a first embedding and a second embedding from the first set of frames and the second set of frames, respectively; generating one or more composite embeddings from the first embedding and the second embedding; determining an active speaker detection (ASD) score for each of the one or more composite embeddings using a classifier; and aggregating the one or more ASD scores to form a detection result; and determining whether the first person is speaking based on the detection result; and adjusting a display of the visual scene to focus on the first person in response to the determination of whether the first person is speaking; A method comprising:
9. The method of claim 8 , wherein the determination of whether the first person is speaking corresponds to the second set of frames.
10. The method of claim 8 , wherein the second set of frames is temporally subsequent to the first set of frames.
11. the first embedding and the second embedding each include a plurality of audiovisual feature vectors; the number of said audiovisual feature vectors is equal to the number of frames in said first set or said second set; The method of claim 8.
12. The method of claim 8 , wherein the audiovisual encoder comprises a neural network and the classifier comprises a recurrent neural network.
13. the detection result includes a first speech metric for the first person; determining whether the first person is speaking includes comparing the first speech metric to a threshold; determining that the first person is speaking in response to the first speech metric being greater than the threshold; The method of claim 8.
14. the visual scene further includes a second person, and the detection result further includes a second speech metric for the second person; The determining whether the first person is speaking includes: obtaining a state for the first person in response to the first speech metric of the first person being less than or equal to the threshold; determining whether the first speech metric is greater than the second speech metric and whether the state of the first person is active; further comprising determining that the first person is speaking in response to the state of the first person being active and the first speech metric being greater than the second speech metric; The method of claim 13.
15. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed on a computer processor, cause the computer processor to: acquiring a first set of frames and a second set of frames from a visual sensor capturing a visual scene including a first person; generating, using an audiovisual encoder, a first embedding and a second embedding from the first set of frames and the second set of frames, respectively; generating one or more composite embeddings from the first embedding and the second embedding; determining an active speaker detection (ASD) score for each of the one or more composite embeddings using a classifier; and aggregating the one or more ASD scores to form a detection result; and determining whether the first person is speaking based on the detection result; and adjusting a display of the visual scene to focus on the first person in response to the determination that the first person is speaking; and A non-transitory computer-readable medium for implementing the above.
16. 16. The non-transitory computer-readable medium of claim 15, wherein the determination of whether the first person is speaking corresponds to the second set of frames.
17. 16. The non-transitory computer-readable medium of claim 15, wherein the second set of frames is temporally subsequent to the first set of frames.
18. the first embedding and the second embedding each include a plurality of audiovisual feature vectors; the number of said audiovisual feature vectors is equal to the number of frames in said first set or said second set; 16. The non-transitory computer-readable medium of claim 15.
19. the detection result includes a first speech metric for the first person; determining whether the first person is speaking includes comparing the first speech metric to a threshold; determining that the first person is speaking in response to the first speech metric being greater than the threshold; 16. The non-transitory computer-readable medium of claim 15.
20. the visual scene further includes a second person, and the detection result further includes a second speech metric for the second person; The determining whether the first person is speaking includes: obtaining a state for the first person in response to the first speech metric of the first person being less than or equal to the threshold; determining whether the first speech metric is greater than the second speech metric and whether the state of the first person is active; further comprising determining that the first person is speaking in response to the state of the first person being active and the first speech metric being greater than the second speech metric; 20. The non-transitory computer-readable medium of claim 19.
Citation Information
Patent Citations
Systems and methods for providing awareness of remote people in a room during video conferencing
JP2005510144A
Speech production section detector and computer program for speech production section detection
JP2013182150A
Detection program, detection method, and detection device
JP2021021749A
Speech recognition device, speech recognition method, and program
JP2024032655A
Speaking classification using audio-visual data
US20200117887A1