Video Content Generation Using Entity and Relation Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing systems lack the capability to generate comprehensive and coherent visual content in response to user inputs, limiting the richness and interactivity of human-computer interactions.

Innovation Solution

A system that performs automatic speech recognition, natural language understanding, coreference resolution, entity extraction, attribute and relation extraction, and composite image generation to create synchronized audio-visual outputs based on user inputs, using trained machine learning components to generate narratives and visual content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If natural language processing systems are used to process user inputs, then speech recognition and understanding capabilities are improved, but the ability to generate comprehensive visual content is insufficient

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidvisual content generation capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent merges speech recognition, natural language understanding, and visual content generation capabilities into a single integrated system. The processing system receives speech input, processes it through NLU to extract entities and attributes, then generates both audio responses and synchronized visual content (images or video clips) that correspond to the extracted information, creating a multi-modal response system.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to handle multiple types of user inputs (speech, text) and generate multiple types of outputs (audio responses, visual content, text transcripts). The same processing pipeline can accommodate different domains (weather forecasts, narratives, information delivery) and response modalities, making the system universally applicable across various interaction scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If existing NLP systems process user inputs, then language understanding is improved, but the richness and interactivity of human-computer interactions are limited

Engineering Contradiction:
Improvenatural language understanding accuracyVSAvoidinteraction richness
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system transitions from unimodal text-based interaction to multimodal interaction by adding visual and audio dimensions. Instead of only processing text input and generating text output, the system processes speech input and generates synchronized audio-visual responses, creating a more immersive and engaging interaction experience that mirrors natural human communication.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system incorporates feedback mechanisms where the generated visual content and audio responses are synchronized with the user's speech input in real-time. The system can adjust its responses based on the extracted entities and attributes from the user's input, creating an adaptive and interactive conversation flow that enhances engagement.

Inventive Principle:
Principle #23Feedback

3Reliability

If comprehensive visual content generation is implemented, then response coherence is improved, but system complexity increases

Engineering Contradiction:
Improveresponse coherenceVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system is divided into distinct processing modules: speech recognition module, natural language understanding module (with entity extraction and attribute extraction sub-components), and visual content generation module. Each module handles a specific aspect of processing independently, making the overall complex system more manageable and easier to implement through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The natural language understanding component acts as an intermediary between the speech recognition and visual content generation modules. It processes the raw speech input, extracts meaningful entities and attributes, and prepares structured data that guides the visual content generation, thereby coordinating the complex interactions between different system components.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of operation

If audio-visual content is generated from user inputs, then user engagement is improved, but processing time increases

Engineering Contradiction:
Improveuser engagementVSAvoidcontent generation time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system generates visual content selectively based on the extracted entities and attributes from the user input. Instead of always generating comprehensive audio-visual content, it determines the appropriate level of detail and type of visual content needed for each specific interaction, balancing engagement quality with processing time efficiency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12579721B2Generating video content from user input data
Publication Date: 2026.03.17 AMAZON TECH INC
  • US12579721B2 patent drawing
  • US12579721B2 patent drawing
  • US12579721B2 patent drawing

AI summary

Techniques for generating content associated with a user input/system generated response are described. Natural language data associated with a user input may be generated. For each portion of the natural language data, ambiguous references to entities in the portion may be replaced with the corresponding entity. Entities included in the portion may be extracted, and image data representing the entity may be determined. Background image data associated with the entities and the portion may be determined, and attributes which modify the entities in the natural language sentence may be extracted. Spatial relationships between two or more of the entities may further be extracted. Image data representing the natural language data may be generated based on the background image data, the entities, the attributes, and the spatial relationships. Video data may be generated based on the image data, where the video data includes animations of the entities moving.