Emotion Detection Indexing for Video Content Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional search engines face challenges in finding and retrieving expressive images like GIFs and memes, as they rely on manual creation and web crawling, which limits their discoverability and often results in lost video context.

Innovation Solution

A method and apparatus for dividing videos into clips, extracting features, and building an index to detect emotional content, which allows proactive identification and enrichment of image indexes with similar web images, enhancing search engine capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional search engines crawl all web pages to find expressive images, then they can discover images that are already spread on the web, but they cannot find images with limited spreading scale and lose video context information

Engineering Contradiction:
Improvevideo context informationVSAvoidimage discovery efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system proactively processes videos before they are uploaded or shared, extracting emotional content and building indexes in advance. This preliminary action ensures that when videos are later searched or shared, the emotional content is already indexed and can be quickly retrieved without losing context information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The video processing system divides videos into segments or frames, analyzing each segment for emotional content independently. This segmentation allows the system to process large videos efficiently while maintaining the ability to retrieve specific emotional moments without processing the entire video each time.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If manual methods are used to create expressive images, then the quality and expressiveness of images can be ensured, but the quantity of available images is very limited

Engineering Contradiction:
Improvequantity of expressive imagesVSAvoidimage creation complexity
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The system automatically analyzes video content to identify and extract emotional expressions without requiring manual intervention. The automated emotional analysis and index building processes enable the system to self-generate expressive image candidates from video content, dramatically increasing quantity while maintaining quality through algorithmic emotional detection.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If the system extracts features from multiple clips to determine emotional content, then the accuracy of emotion detection is improved, but the processing time and computational complexity increase

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts features from a selected subset of clips rather than analyzing every single frame or clip uniformly. By focusing on key segments that are most likely to contain emotional content, the system achieves high detection accuracy while significantly reducing processing time compared to exhaustive analysis of all video content.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11328159B2Automatically detecting contents expressing emotions from a video and enriching an image index
Publication Date: 2022.05.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11328159B2 patent drawing
  • US11328159B2 patent drawing
  • US11328159B2 patent drawing

AI summary

The present disclosure provides method, apparatus and system for detecting contents expressing emotions from a video. The method may comprise: dividing the video into a plurality of clips; extracting, from a first clip and at least one second clip of the plurality of clips, features associated with the first clip; determining whether the first clip expresses emotions based on the features associated with the first clip; and building an index containing the first clip based on the features associated with the first clip if the first clip expresses emotions.