Real-Time Audio-Video Compositing via Segmented Tracks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for social sharing of audio-video content lack the ability to seamlessly combine live video streams with overlays and audio, providing limited creative interaction and sharing options, especially in consumer devices.

Innovation Solution

The system records and composites a video track of an overlay alpha video with a live video stream and audio track, allowing real-time playback and mixing, with features like face detection, customizable containers, and interactive audio effects, enabling users to create and share enhanced video content across various devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If live video streams are combined with overlays and audio in real-time, then creative interaction and user engagement are improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improvecreative interactionVSAvoidprocessing requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system divides the video processing into separate tracks: a first video track for overlay alpha video and a second video track for the live video stream. This segmentation allows independent processing of each track, reducing the complexity of real-time compositing while maintaining creative flexibility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The overlay alpha video is prepared and processed in advance before being combined with the live video stream. This preliminary action allows complex overlay effects to be pre-rendered, reducing the real-time processing burden during live compositing operations.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If multiple video tracks and audio tracks are composited in real-time, then content creation capability is improved, but system resource consumption increases

Engineering Contradiction:
Improvecontent creation capabilityVSAvoidsystem resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system separates video and audio processing into distinct tracks, allowing efficient resource allocation. Video compositing operates on segmented video tracks while audio is processed separately, optimizing system resource usage during real-time content creation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses alpha channel copies to manage overlay transparency information without requiring full duplicate processing of the entire video data. This reduces memory bandwidth and processing requirements while maintaining the ability to create complex multi-track compositions.

Inventive Principle:
Principle #26Copying

3Ease of operation

If real-time playback and mixing is implemented, then user feedback and preview capability are improved, but processing time and computational load increase

Engineering Contradiction:
Improvepreview capabilityVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system pre-processes overlay alpha video and prepares compositing parameters before real-time playback. This preliminary preparation enables smooth real-time mixing and playback without computational delays, improving preview capability while minimizing processing time during actual content creation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10332560B2Audio-video compositing and effects
Publication Date: 2019.06.25 NOO INC
  • US10332560B2 patent drawing
  • US10332560B2 patent drawing
  • US10332560B2 patent drawing

AI summary

Systems, apparatuses, methods, and computer program products perform image and audio processing in a real-time environment, in which an overlay alpha-channel video is composited onto a camera stream received from a capture device, and in which an audio stream from a capture device is mixed with audio data are output to a storage file.