Multimodal Video Detection System Using Message-Driven Condition Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting false information in videos rely heavily on manual processing, consuming significant labor and time, as they involve manual inspection and comparison of video frames, which is inefficient and resource-intensive.

Innovation Solution

A multimodal method and system that utilizes a processor to receive a message, generate detecting conditions, search a video database, compare multimodal data, and output target videos for display, leveraging multimodal association results to automate the detection process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual processing methods are used to detect false information in videos, then detection accuracy can be maintained through human judgment, but the consumption of labor and time resources increases significantly

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical inspection with an automated computer-based system that uses image processing and pattern recognition algorithms to detect false information in videos, thereby reducing time consumption while maintaining detection capability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary automated detection system that acts as a bridge between raw video data and human judgment, pre-processing and filtering video content to identify suspicious segments that require further manual review, thus reducing overall detection time

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual inspection of each video frame is performed to confirm alterations, then thorough detection can be achieved, but resource consumption including labor and time increases

Engineering Contradiction:
Improvedetection thoroughnessVSAvoiddetection efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent divides the video into discrete frames and further segments suspicious regions within frames, allowing the system to focus computational resources on specific areas of interest rather than processing entire frames uniformly, thereby improving detection efficiency without sacrificing thoroughness

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing automated detection on all frames but focusing detailed analysis only on identified suspicious segments, rather than applying full manual inspection to every frame, thus maintaining detection thoroughness while improving productivity

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If automated detection systems are implemented to reduce manual labor, then detection efficiency improves, but the complexity of the detection system increases

Engineering Contradiction:
Improvedetection efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent designs a detection system with universal components that can handle multiple detection tasks using the same core algorithms and processing pipeline, reducing overall system complexity while maintaining high detection efficiency across different video types and alteration methods

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12014546B2Multimodal method for detecting video, multimodal video detecting system and non-transitory computer readable medium
Publication Date: 2024.06.18 INSTITUTE FOR INFORMATION INDUSTRY
  • US12014546B2 patent drawing
  • US12014546B2 patent drawing
  • US12014546B2 patent drawing

AI summary

A multimodal method for detecting video includes following step of: receiving a message to be detected to obtain a multimodal association result, which message to be detected corresponds to a video to be detected; generating a plurality of detecting conditions according to multimodal association result; searching a plurality of videos in a video detection database to obtain a target video in videos according to detecting conditions, which each of videos includes a plurality of video paragraphs respectively, which each of video paragraphs includes a piece of multimodal related data respectively; comparing detecting conditions and piece of multimodal related data of video paragraphs to obtain a matching video paragraph and using video corresponding to matching video paragraph as the target video; and outputting the target video and the video to be detected to a display device for display.