Video Content Moderation Model Using Salient Region Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual moderation of internet video content is time-consuming and inefficient, especially with the increasing volume of user-generated content, and existing machine learning approaches lack efficiency in identifying offensive content like terrorism, violence, and pornography.

Innovation Solution

A method and apparatus for training a content moderation model that extracts salient image regions from video frames containing offensive content, using deep neural networks to classify and moderate video content based on spatiotemporal positioning, reducing the need for manual annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual moderation is used to determine offensive content in videos, then accuracy of content classification is improved, but productivity and time efficiency deteriorate

Engineering Contradiction:
Improveaccuracy of content classificationVSAvoidproductivity of content moderation
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces an intermediary system consisting of a content moderation model trained on salient image regions. This intermediary automatically processes video content by extracting key frames, identifying salient regions, and classifying offensive content, thereby bridging the gap between manual accuracy requirements and automated productivity needs.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts only the most relevant parts of video content for moderation - specifically salient image regions from key frames that contain offensive content. By focusing computational resources on these extracted regions rather than processing entire videos, the system achieves both high accuracy and improved productivity.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If machine learning models process entire video content, then comprehensive content analysis is improved, but use of energy and computational resources deteriorate

Engineering Contradiction:
Improvecomprehensive content analysisVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only essential video frames (key frames) and their salient regions for model processing, rather than analyzing every frame of the video. This extraction approach maintains comprehensive content analysis capability while dramatically reducing computational energy consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by focusing computational analysis on specific salient regions within key frames rather than processing entire frames uniformly. This allows the model to concentrate computational resources on areas most likely to contain offensive content, improving energy efficiency while maintaining analysis comprehensiveness.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If salient image region extraction is implemented, then manufacturing precision of content identification is improved, but device complexity deteriorates

Engineering Contradiction:
Improveprecision of offensive content identificationVSAvoidcomplexity of processing system
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the video processing task into distinct stages: key frame extraction, salient region identification, and content classification. This segmentation of the processing pipeline improves precision at each stage while managing system complexity through modular design, where each component handles a specific function.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12513365B2Method for training content moderation model, method for moderating video content, computer device, and storage medium
Publication Date: 2025.12.30 BIGO TECH PTE LTD
  • US12513365B2 patent drawing
  • US12513365B2 patent drawing
  • US12513365B2 patent drawing

AI summary

Provided is a method for training a content moderation mode. The method includes extracting part of image data of a sample video file as sample image data; positioning a time point of the sample image data in the sample video file in the case that the sample image data contains offensive content; extracting salient image region data from the image data around the time point; and training the content moderation model based on the image region data and the sample image data.