Speaker Recognition for Media Content Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for searching and recommending electronic media content on the Internet are limited by reliance on metadata, which can be incomplete or unreliable, and visual analysis is computationally intensive and often inaccurate, leading to users finding undesirable content and decreased advertising revenue.

Innovation Solution

A computer-implemented method and system that generates speech and speaker models to detect and identify speech segments within electronic media content, calculating the probability of a speaker's involvement to improve content ranking and filtering, thereby enhancing the identification, ranking, and display of relevant media content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual analysis is used to identify people and objects in videos, then recall and precision are improved, but computational resource consumption increases and accuracy decreases

Engineering Contradiction:
ImproveprecisionVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent replaces visual analysis (mechanical/optical processing) with audio-based speaker recognition. Instead of analyzing visual frames to identify people and objects, the system extracts audio features, generates speaker models, and uses these models to recognize and identify speakers in videos. This substitution significantly reduces computational resource requirements while improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If visual analysis is used to identify people in videos, then identification capability is improved, but reliability decreases due to inaccurate recognition

Engineering Contradiction:
Improveidentification accuracyVSAvoidreliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces unreliable visual analysis with audio-based speaker recognition. The system extracts audio features from video tracks, generates speaker models based on these features, and uses the models to reliably identify speakers. This approach avoids the inaccuracies of visual recognition, especially in cases where visual identification fails or is ambiguous.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If metadata is used for video search, then search functionality is provided, but recall is limited due to incomplete and unreliable metadata

Engineering Contradiction:
Improvesearch functionalityVSAvoidrecall
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent performs preliminary action by extracting audio features and generating speaker models for all videos in the collection before search queries are received. This pre-processing creates a comprehensive index of speaker information that can be quickly queried during search operations, significantly improving recall without affecting search functionality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces speaker models as an intermediary between the video content and search queries. Instead of directly searching through incomplete metadata, the system uses speaker models as an intermediate representation that captures comprehensive speaker information, enabling accurate search results even when original metadata is insufficient.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11790933B2Systems and methods for manipulating electronic content based on speech recognition
Publication Date: 2023.10.17 VERIZON PATENT & LICENSING INC
  • US11790933B2 patent drawing
  • US11790933B2 patent drawing
  • US11790933B2 patent drawing

AI summary

Systems and methods are disclosed for displaying electronic multimedia content to a user. One computer-implemented method for manipulating electronic multimedia content includes generating, using a processor, a speech model and at least one speaker model of an individual speaker. The method further includes receiving electronic media content over a network; extracting an audio track from the electronic media content; and detecting speech segments within the electronic media content based on the speech model. The method further includes detecting a speaker segment within the electronic media content and calculating a probability of the detected speaker segment involving the individual speaker based on the at least one speaker model.