Active Speaker Detection via Contrastive Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional manual techniques for dubbing video content are time- and labor-intensive, making it impractical to achieve high-quality dubbing across multiple languages without significant expense.

Innovation Solution

The implementation of a machine learning model using contrastive learning to identify active speakers in videos by matching video and audio features, allowing for automated evaluation of dubbing quality and generation of metadata, which reduces labor and improves dubbing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual techniques are used for dubbing evaluation, then dubbing quality can be assessed, but the process is time-consuming and labor-intensive

Engineering Contradiction:
Improvedubbing quality assessment accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical evaluation with an automated machine learning system that uses contrastive learning to compare video and audio features. The system automatically detects active speakers, extracts visual and audio features, and evaluates dubbing quality without human intervention, thereby eliminating the time-consuming manual review process while maintaining assessment accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service evaluation by allowing the dubbing system to automatically assess its own output quality. The machine learning model independently evaluates whether dubbed audio matches visual cues without requiring external human reviewers, making the evaluation process autonomous and efficient.

Inventive Principle:
Principle #25Self-service

2Reliability

If manual dubbing evaluation is performed across multiple languages, then quality control is maintained, but extraordinary expense is incurred

Engineering Contradiction:
Improvedubbing quality controlVSAvoidproduction cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent replaces expensive manual evaluation across multiple languages with an automated machine learning system. The contrastive learning model processes video and audio features algorithmically, eliminating the need for human reviewers in each language and significantly reducing production costs while maintaining consistent quality control standards.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The machine learning system provides universal dubbing evaluation capability across multiple languages through a single automated platform. The contrastive learning model can assess any language's dubbed audio against visual cues without requiring separate manual evaluation processes for each language, achieving both cost efficiency and consistent quality control.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of manufacture

If traditional manual dubbing processes are used, then dubbing can be performed, but productivity is low due to repeated inspections

Engineering Contradiction:
Improvedubbing process feasibilityVSAvoiddubbing efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent implements continuous automated evaluation throughout the dubbing process. Instead of repeated discrete manual inspections, the machine learning system continuously monitors and evaluates dubbing quality in real-time, allowing the process to flow without interruption and significantly improving productivity while maintaining feasibility.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system replaces manual dubbing processes with automated machine learning-based evaluation. The contrastive learning model automatically compares audio and video features without human intervention, eliminating the need for repeated manual inspections and大幅提升 productivity while keeping the dubbing process feasible and controllable.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11983923B1Systems and methods for active speaker detection
Publication Date: 2024.05.14 NETFLIX INC
  • US11983923B1 patent drawing
  • US11983923B1 patent drawing
  • US11983923B1 patent drawing

AI summary

The disclosed computer-implemented method may include receiving, as input, an audio/video data object; isolating a video stream of a visible potential speaker over a plurality of frames of the audio/video data object; isolating an audio stream over the plurality of frames; providing the isolated video stream and the isolated audio stream to a machine learning model trained with contrastive learning, the contrastive learning using (i) a corpus of video segments of visible speakers with corresponding original audio for positive samples; and (ii) a corpus of video segments of visible speakers with corresponding dubbed audio for negative samples; and evaluating a match between the isolated audio stream and the isolated video stream based at least in part on an output of the machine learning model. Various other methods, systems, and computer-readable media are also disclosed.