Generative AI Media Enhancement System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio/video media platforms require manual input for supplemental information, limiting their ability to incorporate large amounts of content and providing language translation and transcript follow-along capabilities efficiently.

Innovation Solution

A system and method utilizing generative artificial intelligence, including a hardware-based processor, memory, and modules for transcoding media-to-text, summarizing, and chapterizing, which automatically generates text summaries and chapter headings, and translates content, enhancing media with neural networks and natural language processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual input method is used for supplemental information, then user can control content accuracy, but system cannot efficiently incorporate large amounts of supplemental content

Engineering Contradiction:
Improveefficiency of incorporating supplemental contentVSAvoidmanual input requirement
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system performs self-service by automatically generating summaries, transcripts, and chapter breaks through AI processing of the media content itself, eliminating the need for external manual input while maintaining high efficiency in incorporating supplemental content

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual typing system with an automated AI-based natural language processing system that can rapidly generate supplemental content from media files, dramatically increasing productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If manual annotation is used, then content quality can be ensured, but language translation and transcript follow-along capabilities cannot be readily provided

Engineering Contradiction:
Improvelanguage translation capabilityVSAvoidsystem complexity for multiple languages
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The AI processing system serves multiple functions simultaneously - generating summaries, creating transcripts, producing chapter breaks, and enabling language translation - all through the same automated pipeline, making the system versatile without proportionally increasing complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces AI language models as intermediaries that can translate between multiple languages and generate transcripts automatically, serving as a mediator between the original media content and diverse user language requirements without requiring separate manual processes for each language

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated AI processing is used, then efficiency and scalability are improved, but manual control over content accuracy is reduced

Engineering Contradiction:
Improveprocessing speedVSAvoidcontent accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system incorporates feedback mechanisms where generated content can be reviewed, corrected, and refined through iterative AI processing, allowing continuous improvement of accuracy while maintaining high processing speeds through automated loops

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12260883B1System and method to enhance audio and video media using generative artificial intelligence
Publication Date: 2025.03.25 MORGAN STANLEY SERVICES GROUP INC
  • US12260883B1 patent drawing
  • US12260883B1 patent drawing
  • US12260883B1 patent drawing

AI summary

A system and method enhance original media including a first audio using generative artificial intelligence, including large language models and media conversion modules. The system includes a graphic user interface (GUI) including a media player and a display region for outputting an enhanced media including the original media, a summary of the original media, and the plurality of chapter headings of the original media. The media player plays the original media, and the display region displays the summary and the plurality of chapter headings. The summary, the plurality of chapter headings, and each of a translation into a selected language and a second audio generated from the summary and the plurality of chapter headings are automatically generated from the original media. The method implements the system.