Video Audio Text Interaction for Faster Information Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for obtaining valuable information from video sounds are inefficient and lack diversity.

Innovation Solution

An interaction method and apparatus that displays a list of text sentences corresponding to audio sentences in a video, allowing users to view and interact with the text sentences, including options for copying, sharing, and generating new media content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If users listen to video sounds to obtain information, then information can be acquired, but the efficiency of information acquisition is low

Engineering Contradiction:
Improveinformation acquisition efficiencyVSAvoidtime spent listening
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent converts audio information into visual text form by displaying text sentences corresponding to audio sentences on the screen. This allows users to acquire information through reading text rather than listening to audio, significantly improving information acquisition efficiency and reducing time consumption.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If only audio listening is used to obtain video information, then the method is simple, but the method lacks diversity

Engineering Contradiction:
Improvemethod diversityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent enables multiple information acquisition methods (listening to audio, reading text, clicking to play/pause) within the same system. The display module can show text sentences corresponding to audio content, providing both auditory and visual channels for information acquisition, thereby enhancing method diversity without significantly increasing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically switches between different interaction modes based on user actions. Text sentences are displayed in real-time corresponding to audio playback, and users can interact by clicking to pause or continue playback. This dynamic adaptation allows the system to respond to user needs flexibly, providing diverse interaction methods.

Inventive Principle:
Principle #15Dynamics

3Ease of operation

If text sentences are displayed for video content, then information acquisition efficiency is improved, but user interaction options are limited without additional features

Engineering Contradiction:
Improveuser interaction convenienceVSAvoidinteraction method variety
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent implements feedback mechanisms where user actions (clicking on text sentences) trigger specific responses (pausing or continuing video playback). This creates an interactive loop that enhances ease of operation while providing multiple interaction methods, allowing users to control video playback through text sentence selection.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The text sentences serve as an intermediary between the user and the video content. Users can interact with the video through the text interface by clicking on specific sentences to control playback, providing an additional layer of interaction that enhances both ease of operation and interaction variety.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260025559A1Interaction method and apparatus, electronic device, and storage medium
Publication Date: 2026.01.22 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260025559A1 patent drawing
  • US20260025559A1 patent drawing
  • US20260025559A1 patent drawing

AI summary

The present disclosure relates to an interaction method, an apparatus, an electronic device and a storage medium. The method comprises: receiving a text display operation for a first media content, wherein the first media content includes a video content; in response to the text display operation, displaying in a preset region a list of text sentences of the first media content, wherein the list of text sentences contains at least two text sentences, and the at least two text sentences each have a corresponding audio sentence in target audio data of the first media content.