Voice-Triggered Video Editing via Wake-Up Word Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals face inconvenience and high costs in editing videos to achieve high-quality broadcast content, as they need to manually edit each video to suit their personal style, which is time-consuming and requires specialized expertise.

Innovation Solution

A content editing apparatus and method that automatically edits videos based on a wake-up word and editing commands, determining the video category and applying corresponding templates to streamline the editing process, reducing manual effort and costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual video editing is performed to achieve high-quality broadcast content, then content quality is improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvecontent qualityVSAvoidediting time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system enables self-service editing by automatically analyzing video content, detecting key moments through audio and visual analysis, and generating edited videos without requiring manual intervention. The processor autonomously identifies wake-up words, determines video categories, selects appropriate templates, and performs editing operations, allowing the video content to edit itself based on predefined rules and algorithms.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual editing process with an automated electronic system. The processor substitutes human editors by performing audio analysis, visual analysis, template selection, and video rendering through software algorithms, eliminating the need for physical manual editing operations while maintaining high content quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual video editing is performed to achieve high-quality broadcast content, then content quality is improved, but cost burden increases

Engineering Contradiction:
Improvecontent qualityVSAvoidediting cost
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system enables self-service editing by automatically analyzing video content, detecting key moments through audio and visual analysis, and generating edited videos without requiring manual intervention. The processor autonomously identifies wake-up words, determines video categories, selects appropriate templates, and performs editing operations, allowing the video content to edit itself based on predefined rules and algorithms.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual editing process with an automated electronic system. The processor substitutes human editors by performing audio analysis, visual analysis, template selection, and video rendering through software algorithms, eliminating the need for physical manual editing operations while maintaining high content quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automatic editing based on wake-up words and commands is implemented, then editing efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveediting efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The editing system is segmented into distinct functional modules: audio analysis unit for wake-up word detection, visual analysis unit for moment detection, category determination unit for video classification, template selection unit for choosing editing patterns, and video rendering unit for final output. This segmentation allows each module to perform its specific function independently, improving editing efficiency while managing system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The processor acts as an intermediary that coordinates between various analysis units, template database, and rendering engine. It receives analysis results, determines video categories, selects appropriate templates, and orchestrates the editing process, thereby managing system complexity through centralized control while enabling automated high-efficiency editing.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of operation

If video category determination and template-based editing are applied, then ease of operation is improved, but adaptability to personal style may be reduced

Engineering Contradiction:
Improveediting convenienceVSAvoidpersonal style adaptation
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The template selection and editing process is made dynamic by allowing the system to adapt templates based on detected personal styles and preferences. The processor analyzes video content characteristics, determines appropriate categories, and selects or customizes templates that match the creator's style, enabling both ease of operation through automation and adaptability to personal preferences through dynamic adjustment of editing parameters.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11080531B2Editing multimedia contents based on voice recognition
Publication Date: 2021.08.03 LG ELECTRONICS INC
  • US11080531B2 patent drawing
  • US11080531B2 patent drawing
  • US11080531B2 patent drawing

AI summary

Disclosed are a content editing apparatus and method capable of editing a video filmed by a personal terminal in a 5G communication environment. The content editing apparatus of the present disclosure includes a processor, a memory operatively connected to the processor and which stores at least one code configured to be executed by the processor, and an interface for receiving a video. The memory stores codes that, when executed by the processor, cause the processor to recognize a set wake-up word from the video, and edit the video based on an editing command recognized within an interval of a preset time from a portion where the wake-up word of the video is located.