Audio Video Text Outline Generation for Content Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio and video services lack an effective content overview function, making it difficult for users to quickly locate desired content, especially in long-duration videos, which can lead to decreased user engagement and poor experience due to the need for manual searching or complex video segmentation.

Innovation Solution

A method that extracts text information from audio or video data to generate a text outline with associated time periods, allowing for the creation of a display field that provides users with a quick overview, enabling rapid location of desired content within the data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If the duration of audio or video increases, then the amount of information transmitted increases, but user's attention and interest decrease

Engineering Contradiction:
Improveamount of informationVSAvoiduser attention and interest
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments long audio or video content into multiple chapters or sections with defined time ranges. Each chapter represents a logical division of the content, allowing users to navigate to specific segments of interest without having to watch or listen to the entire long duration content, thereby maintaining user attention while transmitting large amounts of information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a temporal dimension to content navigation by providing time range information for each chapter. This allows users to jump to specific time points within the audio or video content, transforming linear sequential consumption into non-linear selective access, thus maintaining engagement in long-duration content.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Quantity of substance

If the audio or video has a relatively long duration, then the content coverage increases, but the desired content cannot be located quickly and accurately

Engineering Contradiction:
Improvecontent coverageVSAvoidtime to locate desired content
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent performs preliminary organization of audio or video content into chapters with defined time ranges and descriptions before user consumption. This pre-structuring allows users to quickly locate desired content by browsing chapter titles and time ranges, eliminating the need to manually search through long-duration content and significantly reducing time to locate specific information.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If manual searching or complex video segmentation is used, then content location capability improves, but production costs increase

Engineering Contradiction:
Improvecontent location capabilityVSAvoidproduction costs
Core Design Contradiction:
Ease of operationVSEase of manufacture

Solution Approach 1:

The patent enables the system to automatically generate chapter divisions and time range information from the audio or video content itself, without requiring manual intervention. The processing device automatically analyzes the content structure, identifies logical breakpoints, and creates navigable chapters with time stamps, providing content location capability while eliminating the need for manual segmentation and reducing production costs.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12192598B2Method of processing audio or video data, device, and storage medium
Publication Date: 2025.01.07 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12192598B2 patent drawing
  • US12192598B2 patent drawing
  • US12192598B2 patent drawing

AI summary

A method of processing audio or video data is provided, which relates to a field of a natural language processing technology, and in particular to a semantic understanding of a natural language. The method includes: extracting a text information from the audio or video data; generating a text outline and a plurality of time periods according to the text information, the text outline includes multi-level outline entries, and the plurality of time periods are associated with the multi-level outline entries; generating a display field for the audio or video data according to the text outline and the plurality of time periods; adding the display field to the audio or video data, so as to obtain updated audio or video data. A device, and a storage medium are further provided.