Mobile Terminal Text Description Generation via Video Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The process of adding text descriptions to videos is complicated, hindering the richness of video content and user interaction on platforms.

Innovation Solution

A method and device for generating text descriptions using video content recognition, including image and voice recognition, to automatically provide text information, which can be edited and displayed on mobile terminals, enhancing user interaction and community engagement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual text input is used for video descriptions, then text accuracy can be ensured, but the operation complexity increases and time consumption increases

Engineering Contradiction:
Improvetext accuracyVSAvoidoperation complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs automatic text recognition on video content to generate descriptions without requiring manual user input. The video content itself serves as the source material, and the system extracts and generates text descriptions automatically, making the system self-sufficient and eliminating the need for manual text entry by users.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual text input process with automatic optical character recognition (OCR) technology. Instead of users manually typing or selecting text, the system uses image recognition algorithms to automatically extract and generate text descriptions from video frames, substituting human manual operation with automated computational processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual text input is used for video descriptions, then text accuracy can be ensured, but time consumption increases

Engineering Contradiction:
Improvetext accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs text recognition and generates video descriptions in advance during the video processing stage, before the user needs to post or share the video. By completing the text generation operation beforehand, the system eliminates the need for time-consuming manual text input at the moment of posting, significantly reducing user time consumption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The automatic text recognition system generates descriptions autonomously without requiring user time investment for manual input. The system serves itself by extracting text from video content and generating descriptions automatically, freeing users from time-consuming manual text entry tasks.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If automatic text recognition is used, then operation simplicity improves, but text recognition accuracy may deteriorate

Engineering Contradiction:
Improveoperation simplicityVSAvoidtext recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system provides feedback mechanisms that allow users to review, edit, and correct automatically generated text descriptions. Users can modify the AI-generated text to improve accuracy while maintaining the simplicity of automatic generation, combining automated efficiency with human oversight for quality assurance.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Instead of relying solely on automatic text recognition, the system combines partial automatic generation with optional manual supplementation. Users can add or modify text as needed, allowing the system to perform enough automatic work to simplify operation while permitting additional manual action to ensure accuracy when required.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11580290B2Text description generating method and device, mobile terminal and storage medium
Publication Date: 2023.02.14 BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
  • US11580290B2 patent drawing
  • US11580290B2 patent drawing
  • US11580290B2 patent drawing

AI summary

A text description generating method and device, a mobile terminal and a storage medium. The method of the embodiments of the present disclosure includes: obtaining the video content; performing text recognition according to the video content to obtain first text information, and displaying the first text information; and/or, generating second text information in response to user input operation on the video content, and displaying the second text information.