Automated E-Commerce Video Generation from Images with Text Placement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

E-commerce sellers face burdens in creating professional-quality videos for their listings due to the need for manual video editing skills, leading to increased operational costs and suboptimal video presentation.

Innovation Solution

A system utilizing machine-learning models to automatically generate videos by sorting images, inserting text, and optimizing video data based on image analysis and user adjustments, ensuring coherent and effective video creation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual video editing tools are used by sellers, then video quality and professionalism are improved, but operational burden and costs increase

Engineering Contradiction:
Improvevideo qualityVSAvoidseller burden
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system enables automatic video generation where the computer system performs video editing tasks autonomously without requiring seller intervention. The machine learning model automatically sorts images, selects frames, inserts text, and optimizes video content based on item listing data, making the system self-sufficient in creating professional-quality videos.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical video editing operations with an automated machine learning-based system. Instead of sellers manually positioning text and adjusting video frames, the system uses computer vision and natural language processing to automatically generate video content from images and item descriptions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If automatic video generation is implemented, then seller burden is reduced, but video quality and professionalism deteriorate

Engineering Contradiction:
Improveseller burdenVSAvoidvideo quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system replaces manual video editing with advanced machine learning models including computer vision for image analysis, natural language processing for text generation, and automated video composition algorithms that replicate professional editing workflows automatically.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system incorporates feedback mechanisms where the machine learning model learns from item listing data, image metadata, and video performance metrics to continuously improve video generation quality. The system adjusts text placement, frame selection, and video timing based on analyzed patterns from successful e-commerce videos.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If manual text insertion and positioning is performed, then text placement accuracy is improved, but time consumption and operational complexity increase

Engineering Contradiction:
Improvetext placement accuracyVSAvoidediting time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically performs text insertion and positioning without seller intervention. The machine learning model analyzes video frames to identify optimal text placement regions, automatically inserts generated text descriptions, and adjusts positioning based on visual content analysis, completing tasks that would require manual editing in seconds.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediate machine learning processing layer between item listing data and final video output. This intermediary system generates text content, determines placement positions, and coordinates insertion timing, acting as an intelligent mediator that automates the entire text integration workflow.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If professional video editing tools are provided, then video effectiveness is improved, but device complexity and learning curve increase

Engineering Contradiction:
Improvevideo effectivenessVSAvoidtool complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system provides a universal automated video generation platform that handles multiple video creation tasks through a single machine learning model interface. The system can generate videos for different product categories, automatically adapt text styling, and handle various image formats without requiring sellers to learn different tools or settings.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The complex video editing functions are encapsulated within the automated system that performs all editing operations independently. Sellers interact only with simple item listing uploads, while the system self-manages complex tasks including frame selection, text generation, typography selection, and video composition without exposing underlying complexity to users.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12374075B2Method and system for automated video generation from images for e-commerce applications
Publication Date: 2025.07.29 EBAY INC
  • US12374075B2 patent drawing
  • US12374075B2 patent drawing
  • US12374075B2 patent drawing

AI summary

Systems and methods are provided for automatically generating a video associated with an item in the marketplace. An image receiver receives images associated with an item of an item listing. An image extractor generates visual descriptors for each image through computer vision analysis and extracts a unique set of images by removing redundant images. An image sorter sorts images in the unique set of images based on an item category and generates a sequence of images for generating a video. A text placer automatically identifies a region in an image and inserts text into the image using textual attributes as predicted by a model. A video data optimizes a generated video using another model trained based on manual adjustments previously made to other exemplary video data. The disclosed technology publishes the automatically generated video data for viewing by viewers in the marketplace.