Automated Videoconference Editing with Timed Speaker Transitions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing videoconference recording technologies lack efficient automated editing capabilities to seamlessly switch between multiple participants' views and excise irrelevant content, resulting in lengthy and unorganized recordings.

Innovation Solution

A method and system that access multiple video data streams, determine the views of participants making statements, and generate an edited video data stream with timed transitions between participants, allowing for automated or semi-autonomous editing based on various criteria, including content and time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual editing of videoconference recordings is used, then editing precision can be maintained, but productivity is reduced due to time-consuming manual processes

Engineering Contradiction:
Improveediting efficiencyVSAvoidediting precision
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system performs automated editing operations by itself without requiring continuous human intervention. The video conference recording system automatically identifies speakers, selects appropriate video streams, and generates edited recordings based on speaker attribution data, enabling the system to serve itself in the editing process while maintaining quality standards

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical editing operations with automated computational processes. Instead of manually reviewing and cutting video footage, the system uses automated speaker detection algorithms and video stream analysis to intelligently select and switch between different participant views based on who is speaking

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If multiple video data streams are processed to switch between speakers, then the representation accuracy is improved, but the device complexity increases

Engineering Contradiction:
Improverepresentation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The video conference recording system is designed to handle multiple video data streams simultaneously and perform multiple functions: recording, speaker detection, stream selection, and automated editing. This multi-functional capability allows the system to process diverse input streams and produce accurate speaker-focused recordings without requiring separate dedicated systems for each function

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces an intermediary processing layer that receives multiple video streams and speaker attribution data, then intelligently selects and switches between streams based on who is speaking. This intermediary layer manages the complexity of coordinating multiple streams while maintaining accurate representation of the conference

Inventive Principle:
Principle #24Intermediary (Mediator)

3Extent of automation

If automated speaker detection is implemented, then productivity is improved through automation, but the difficulty of detecting and measuring increases

Engineering Contradiction:
Improveautomation levelVSAvoidspeaker detection difficulty
Core Design Contradiction:
Extent of automationVSDifficulty of detecting and measuring

Solution Approach 1:

The system implements feedback mechanisms where speaker detection results are continuously refined based on audio signal analysis and video stream correlation. The automated editing process uses feedback from speaker attribution accuracy to improve stream selection, creating a self-correcting system that enhances detection capability while maintaining high automation levels

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9247205B2System and method for editing recorded videoconference data
Publication Date: 2016.01.26 FUJITSU LTD
  • US9247205B2 patent drawing
  • US9247205B2 patent drawing
  • US9247205B2 patent drawing

AI summary

In a method embodiment, a method for automatically editing data recorded during a videoconference includes accessing a plurality of video data streams. Each video data stream records a view of at least one of a plurality of human participants of the videoconference. The view recorded by each video data stream is different from the view recorded by each other video data stream. The method further includes determining, using one or more processors executing logic, that one of the plurality of video data streams recorded a view of a first one of the plurality of participants while the first one of the plurality of participants made a first statement. In addition the method includes determining, using one or more processors executing logic, that one of the plurality of video data streams recorded a view of a second one of the plurality of participants while the second one of the plurality of participants made a second statement after the first one of the plurality of participants made the first statement. An edited video data stream is generated using the plurality of video data streams. The edited video data stream comprises a transition that switches from the view of the first one of the plurality of participants to the view of the second one of the plurality of participants. The transition is timed such that when the edited video data stream is played the transition occurs before the commencement of the second statement.