Video Scene Segmentation for Language Learning Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for language learning with video contents are inefficient, as learners find it difficult to combine images with spoken voices and subtitles, leading to boredom and loss of concentration due to repetitive playback of line-free parts.

Innovation Solution

A method that divides video contents into scenes without spoken lines, allowing users to select and play these scenes repeatedly in both native and target languages, with adjustable playback speed and order, to enhance language learning efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If video contents are played repeatedly for language learning, then memorization of lines is improved, but learners feel bored and lose concentration due to line-free parts

Engineering Contradiction:
Improvememorization effectivenessVSAvoidlearner concentration
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The video contents are segmented into multiple scenes based on subtitle display timing. Each scene corresponds to a specific subtitle display period, allowing the system to selectively replay only the relevant language-containing scenes rather than the entire video including line-free parts. This segmentation enables precise control over what is replayed, maintaining learner concentration while achieving memorization goals.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention extracts and removes line-free parts from the video contents before replay. By identifying scenes that contain spoken lines versus those that don't, the system extracts only the necessary language-containing scenes for replay. This extraction process eliminates the boring line-free portions that cause learners to lose concentration, while preserving the essential language material for memorization.

Inventive Principle:
Principle #2Taking out (Extraction)

2Area of stationary object

If subtitles are displayed within limited space, then display constraints are satisfied, but subtitles are partially omitted or fail to coincide in meaning with spoken voices

Engineering Contradiction:
Improvesubtitle display areaVSAvoidsubtitle completeness
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The system dynamically adjusts subtitle display by associating each scene with specific subtitle information corresponding to the spoken lines in that scene. Rather than displaying all subtitles simultaneously or statically, the invention enables dynamic presentation of subtitles that match the current scene's spoken content, ensuring completeness and accuracy while adapting to display space constraints.

Inventive Principle:
Principle #15Dynamics

3Productivity

If spoken lines are played in a non-native language only, then language learning is improved, but learners cannot understand the meaning in their native language

Engineering Contradiction:
Improvelanguage learning efficiencyVSAvoidmeaning understanding
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The invention makes the video playback system multi-functional by enabling it to provide both native language and non-native language audio tracks, along with corresponding subtitles in both languages. The system can switch between or combine these language modes, making it universally applicable to different learning stages and preferences while maintaining both understanding (through native language) and learning efficiency (through non-native language exposure).

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11138988B2Playing method for playing multi-language contents for education, and data structure and program therefor
Publication Date: 2021.10.05 KOYAMA OSAMU
  • US11138988B2 patent drawing
  • US11138988B2 patent drawing
  • US11138988B2 patent drawing

AI summary

Video contents having language information including spoken voices in a plurality of languages can be efficiently played to support language learning. After removing line-free parts from the video contents, the video contents are divided into divided scenes each corresponding to one or two consecutive displays of subtitles. Each of selected scenes of the divided scenes selected by a user is played (i) a first predetermined number of times selected by the user, in one of first and second languages selected beforehand by the user, together with images; and (ii) then a second predetermined number of times selected by the user, in the other of the first and second languages, together with the images.