Video Sound Importance Segmentation for Lecture Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video editing techniques that synchronize video and sound signals can lead to disparities between the content of video and sound signals, resulting in a loss of value in recorded lectures, especially when deleting sections based on importance, as it may cause the video and sound to become inconsistent, potentially leading to incomplete or failed recordings.

Innovation Solution

An information processing device generates reproduction assisting information by separately determining the degree of importance for video and sound sections within a target clip, allowing for independent editing of video and sound signals to maintain synchronization and consistency, ensuring that important content is preserved and inconsistencies are minimized.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If video and sound signals are deleted based on their respective degrees of importance independently, then editing efficiency is improved and more sections can be removed, but synchronization between video and sound content deteriorates causing inconsistency

Engineering Contradiction:
Improveediting efficiencyVSAvoidsynchronization between video and sound
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The recorded lecture is divided into multiple sections based on speaking times of persons, allowing independent evaluation of video and sound importance while maintaining overall structure. This segmentation enables efficient independent editing while preserving synchronization through structured divisions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system evaluates and adjusts the degree of importance as a parameter for both video and sound signals independently. By changing this importance parameter based on multiple factors (number of utterances, participants, discussion time, volume level, gestures, emotions), the system optimizes editing efficiency while maintaining content consistency through parameter-based control.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If the same time period is deleted from both video and sound signals according to degree of importance, then editing simplicity is improved, but content consistency deteriorates causing disparity between video and sound

Engineering Contradiction:
Improveediting simplicityVSAvoidcontent consistency
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system applies different editing treatments to different sections based on their local importance characteristics. Rather than uniform deletion, each section is evaluated independently with its own degree of importance calculation, allowing localized optimization of both video and sound while maintaining overall content consistency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses feedback from multiple evaluation factors (utterances, participants, discussion time, volume, gestures, emotions) to dynamically adjust the degree of importance for each section. This feedback mechanism ensures that editing decisions maintain content consistency by considering the interrelationship between video and sound characteristics.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240179381A1Information processing device, generation method, and program
Publication Date: 2024.05.30 SONY GROUP CORP
  • US20240179381A1 patent drawing
  • US20240179381A1 patent drawing
  • US20240179381A1 patent drawing

AI summary

The present technique relates to an information processing device, a generation method, and a program that can edit or reproduce a video in a proper form. The information processing device of the present technique includes a generation unit configured to generate reproduction assisting information for reproducing a video of contents and a sound related to the contents, the video and sound included in a target clip among a plurality of clips obtained by dividing data including the video and the sound, the reproduction assisting information generated according to a first degree of importance that is a degree of importance of a plurality of sections obtained by further dividing the video included in the target clip and a second degree of importance that is a degree of importance of a plurality of sections obtained by further dividing the sound included in the target clip. The present technique is applicable to, for example, a lecture capture system used for video-shooting a lecture.