Face Detection Video Concatenation System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video editing applications are time-consuming and require professional skills, making it difficult for laymen to seamlessly concatenate and synchronize videos from different sources, resulting in unprofessional and low-view videos.

Innovation Solution

A method that uses face detection to match and concatenate videos by processing client video frames to fit the size and shape of source video frames, matching frame rates and resolutions, and applying filtering techniques for a coherent and seamless output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual video editing is performed to concatenate and synchronize videos from different sources, then the quality and professionalism of the output video is improved, but the time consumption and complexity of the editing process increases significantly

Engineering Contradiction:
Improvevideo editing qualityVSAvoidediting time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs automatic video concatenation and synchronization by detecting faces in video frames and using them as reference points for alignment. The computer executes algorithms that automatically parse video frames, identify faces, match corresponding faces across different videos, and synchronize the videos without requiring manual intervention, thereby eliminating the need for professional editing skills and significantly reducing editing time while maintaining quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual editing process with an automated computer-based system that uses face detection algorithms and image processing techniques. Instead of manually aligning and synchronizing videos, the system automatically detects faces in video frames, extracts facial features, matches corresponding faces across different video sources, and performs synchronization based on these detected features, substituting manual mechanical editing with automated computational processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If professional editing skills are required to concatenate videos seamlessly, then the output video quality is improved, but the ease of operation deteriorates as laymen find it difficult to perform the editing

Engineering Contradiction:
Improvevideo concatenation qualityVSAvoidediting accessibility
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system performs automatic video concatenation and synchronization by detecting faces in video frames and using them as reference points for alignment. The computer executes algorithms that automatically parse video frames, identify faces, match corresponding faces across different videos, and synchronize the videos without requiring manual intervention, thereby eliminating the need for professional editing skills and significantly reducing editing time while maintaining quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual editing process with an automated computer-based system that uses face detection algorithms and image processing techniques. Instead of manually aligning and synchronizing videos, the system automatically detects faces in video frames, extracts facial features, matches corresponding faces across different video sources, and performs synchronization based on these detected features, substituting manual mechanical editing with automated computational processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If additional changes are made to the combined video, then the video can be improved, but the complexity increases as multiple corrections are required in multiple places

Engineering Contradiction:
Improvevideo modification flexibilityVSAvoidediting process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary face detection and synchronization by detecting faces in video frames and establishing correspondence between faces in different videos before final concatenation. This preliminary alignment creates a robust foundation that automatically adapts to modifications, reducing the need for multiple corrections across different places and simplifying the overall editing process

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10734027B2System and methods for concatenating video sequences using face detection
Publication Date: 2020.08.04 FUSIT INC
  • US10734027B2 patent drawing
  • US10734027B2 patent drawing
  • US10734027B2 patent drawing

AI summary

There are provided methods and devices for media processing, comprising: providing at least one media asset source selected from a media asset sources library, the at least one media asset source comprising at least one source video, via a network or client device; receiving via the network or the client device a media recording comprising a client video recorded by a user of the client device; parsing the client video and the source video, respectively, to a plurality of client video frames and a plurality of source video frames; identifying at least one face in at least one frame of the plurality of source video frames and at least another face in at least one frame of the plurality of client video frames by face detection; superposing one or more markers on the identified at least one face of the plurality of source video frames; processing said client video frames to fit the size or shape of said source video frames by using said one or more markers; concatenating said processed client video frames with said source video frames, wherein said concatenation comprises matching the frame rate and resolution of the processed client video frames to the frame rate and resolution of the plurality of client video frames to yield a mixed media asset.