Mobile Terminal Text Description Generation via Video Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The process of adding text descriptions to videos is complicated, hindering the richness of video content and user interaction on platforms.
Innovation Solution
A method and device for generating text descriptions using video content recognition, including image and voice recognition, to automatically provide text information, which can be edited and displayed on mobile terminals, enhancing user interaction and community engagement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual text input is used for video descriptions, then text accuracy can be ensured, but the operation complexity increases and time consumption increases
Solution Approach 1:
The system performs automatic text recognition on video content to generate descriptions without requiring manual user input. The video content itself serves as the source material, and the system extracts and generates text descriptions automatically, making the system self-sufficient and eliminating the need for manual text entry by users.
Solution Approach 2:
The patent replaces the mechanical manual text input process with automatic optical character recognition (OCR) technology. Instead of users manually typing or selecting text, the system uses image recognition algorithms to automatically extract and generate text descriptions from video frames, substituting human manual operation with automated computational processing.
2Measurement precision
If manual text input is used for video descriptions, then text accuracy can be ensured, but time consumption increases
Solution Approach 1:
The system performs text recognition and generates video descriptions in advance during the video processing stage, before the user needs to post or share the video. By completing the text generation operation beforehand, the system eliminates the need for time-consuming manual text input at the moment of posting, significantly reducing user time consumption.
Solution Approach 2:
The automatic text recognition system generates descriptions autonomously without requiring user time investment for manual input. The system serves itself by extracting text from video content and generating descriptions automatically, freeing users from time-consuming manual text entry tasks.
3Ease of operation
If automatic text recognition is used, then operation simplicity improves, but text recognition accuracy may deteriorate
Solution Approach 1:
The system provides feedback mechanisms that allow users to review, edit, and correct automatically generated text descriptions. Users can modify the AI-generated text to improve accuracy while maintaining the simplicity of automatic generation, combining automated efficiency with human oversight for quality assurance.
Solution Approach 2:
Instead of relying solely on automatic text recognition, the system combines partial automatic generation with optional manual supplementation. Users can add or modify text as needed, allowing the system to perform enough automatic work to simplify operation while permitting additional manual action to ensure accuracy when required.
Data Source
AI summary
A text description generating method and device, a mobile terminal and a storage medium. The method of the embodiments of the present disclosure includes: obtaining the video content; performing text recognition according to the video content to obtain first text information, and displaying the first text information; and/or, generating second text information in response to user input operation on the video content, and displaying the second text information.


