AI Media Enhancement for Automated Chapters and Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio/video media platforms struggle to efficiently incorporate large amounts of supplemental content and lack language translation and transcript follow-along capabilities due to reliance on manual annotation.

Innovation Solution

A system and method utilizing generative artificial intelligence, including hardware-based processors and large language models, automatically generates text, summaries, and chapter headings from audio, and translates them into different languages, enhancing media with a graphical user interface for playback and navigation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual annotation is used to supplement media with information, then the media can be enhanced with supplemental content, but the process becomes inefficient when incorporating large amounts of supplemental content

Engineering Contradiction:
Improveefficiency of incorporating supplemental contentVSAvoidtime required for manual annotation
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical annotation processes with automated artificial intelligence systems. The AI model automatically generates supplemental information including transcripts, summaries, and chapter headings from media content, eliminating the need for manual text input and significantly improving processing efficiency while reducing time loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The media enhancement system performs self-service by automatically generating all supplemental content without human intervention. The AI model processes media content autonomously to create transcripts, generate summaries, and produce chapter headings, allowing the system to serve itself rather than relying on manual annotation by users.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If manual annotation is used, then supplemental information can be added to media, but language translation and transcript follow-along capabilities cannot be readily provided

Engineering Contradiction:
Improvelanguage translation capabilityVSAvoidease of providing transcript follow-along
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements a universal AI-based media enhancement system that performs multiple functions simultaneously: generating transcripts, creating summaries, producing chapter headings, and providing language translation. This multi-functional approach enables transcript follow-along capabilities and language adaptation without requiring separate manual processes for each feature.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system replaces manual operations with automated AI processing for both transcript generation and language translation. The AI model automatically creates follow-along transcripts and provides real-time translation capabilities, making these features readily available without manual intervention and significantly improving ease of operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If large amounts of supplemental content are incorporated manually, then comprehensive media enhancement is achieved, but the process becomes inefficient and time-consuming

Engineering Contradiction:
Improveamount of supplemental contentVSAvoidefficiency of content incorporation
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent replaces manual content creation with automated AI generation. The system processes media content and automatically generates comprehensive supplemental material including full transcripts, detailed summaries, and chapter headings in large quantities without manual intervention, maintaining high productivity while incorporating substantial amounts of supplemental content.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The AI model performs preliminary action by pre-generating all supplemental content before user interaction. Transcripts, summaries, and chapter headings are created in advance through automated processing, allowing large amounts of content to be incorporated efficiently without manual effort during actual media consumption or editing processes.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260051336A1System and method to enhance audio and video media using generative artificial intelligence
Publication Date: 2026.02.19 MORGAN STANLEY SERVICES GROUP INC
  • US20260051336A1 patent drawing
  • US20260051336A1 patent drawing
  • US20260051336A1 patent drawing

AI summary

A system and method enhance original media including a first audio using generative artificial intelligence, including large language models and media conversion modules. The system includes a graphic user interface including a media player and a display region for outputting an enhanced media including the original media, a summary of the original media, the plurality of chapter headings of the original media, and generated text constituting chapters. The media player plays the original media, and the display region displays the summary in a first display region, and displays the plurality of chapter headings and generated text constituting chapters in a second display region. The summary, the plurality of chapter headings, the chapters, and each of a translation into a selected language and a second audio generated from the summary and the plurality of chapter headings are automatically generated from the original media. The method implements the system.