Edge Video-to-Text Conversion for Bandwidth Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cloud computing architectures face challenges in latency, availability, bandwidth usage, data privacy, and the capacity to process large volumes of data in real-time, particularly in remote edge locations with limited infrastructure, where data transmission and storage are cumbersome due to reliance on slower and more expensive wireless communication links.

Innovation Solution

The implementation of edge computing units that convert digital video data into natural language text descriptions, allowing for real-time processing and storage reduction by generating text-based video narratives, and providing question and answer sessions using trained machine learning models, which can operate in harsh environments with limited power or network connectivity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is transmitted from remote edge locations to centralized data centers, then data can be processed and stored, but transmission time and cost increase significantly due to limited bandwidth and reliance on wireless communication links

Engineering Contradiction:
Improvedata processing capabilityVSAvoiddata transmission time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the centralized processing function into distributed edge computing units deployed at remote locations. Each edge unit independently processes video data locally, converting it to text narratives and generating insights without requiring transmission to centralized data centers. This segmentation eliminates transmission delays while maintaining processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces edge computing units as intermediary devices between video cameras and centralized systems. These edge units perform local processing, converting video data to text narratives and generating insights locally before optionally transmitting only the processed results to centralized systems, thereby reducing transmission time and bandwidth requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If large volumes of raw video data are stored at edge locations, then data availability is maintained, but storage capacity requirements and costs increase significantly

Engineering Contradiction:
Improvedata availabilityVSAvoiddata storage volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts the essential information from raw video data by converting it to text narratives at the edge. Instead of storing entire video files, the system stores only the extracted text narratives and generated insights, which occupy minimal storage space while preserving the key information and maintaining data availability for querying and analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the data representation parameter from raw video format to compressed text narrative format. This transformation dramatically reduces the storage volume required while maintaining the essential information content, allowing the system to retain data availability with minimal storage requirements at edge locations.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If video data is transmitted over wireless communication links from remote locations, then data can be accessed centrally, but transmission cost and time increase due to limited bandwidth

Engineering Contradiction:
Improvedata access capabilityVSAvoiddata transmission time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of video data at the edge by converting it to text narratives and generating insights before transmission. This preliminary action ensures that only processed, high-value information needs to be transmitted to centralized systems, reducing transmission time and maintaining ease of central access to critical information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent inverts the traditional architecture by placing processing power at the edge rather than relying on centralized processing. Edge computing units locally convert video to text and generate insights, then transmit only these processed results to centralized systems, thereby reducing transmission time while maintaining central access capability.

Inventive Principle:
Principle #13The other way round (Inversion)

4Power

If centralized processing models are used, then data can be processed with high computational power, but latency increases and real-time processing becomes difficult

Engineering Contradiction:
Improvecomputational processing powerVSAvoidprocessing speed
Core Design Contradiction:
PowerVSSpeed

Solution Approach 1:

The patent segments the centralized processing model into distributed edge processing units. Each edge unit has sufficient computational power to perform local video-to-text conversion and insight generation independently, eliminating the need to transmit raw video data to centralized systems and thereby achieving real-time processing speed while maintaining adequate computational capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables edge computing units to self-process video data locally without requiring centralized processing. Each edge unit independently converts video to text narratives and generates insights autonomously, achieving real-time processing speeds while maintaining the computational power needed for complex video analysis tasks.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12141541B1Video to narration
Publication Date: 2024.11.12 ARMADA SYST INC
  • US12141541B1 patent drawing
  • US12141541B1 patent drawing
  • US12141541B1 patent drawing

AI summary

Disclosed are systems and methods that convert digital video data, such as two-dimensional digital video data, into a natural language text description describing the subject matter represented in the video. For example, the disclosed implementations may process video data in real-time, near real-time, or after the video data is created and generate a text-based video narrative describing the subject matter of the video. In addition, the disclosed implementations may also support a question and answer session in which a user may submit queries about the subject matter of one or more videos and the disclosed implementations will present natural language responses based on the subject matter of the video and any corresponding context.