Web Page Audio Generation via Server-Side Text Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face inconvenience in accessing web page information, as they must read web pages and individually access each page, with no efficient option to hear them or select alternative modes of information access.

Innovation Solution

A server process that generates a text version of a web page, filters out non-audible information, and creates an audio file, which is then transmitted to the client terminal, allowing users to hear the web page and providing options for audio file encoding based on connection speed and user preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If web pages are provided in traditional text form requiring user reading, then information delivery is straightforward, but user convenience and accessibility are reduced

Engineering Contradiction:
Improveuser convenienceVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary system comprising a server process and text-to-speech converter that mediates between the web page content and the user. The server process retrieves web page content, processes it through filtering to remove non-audible elements, converts it to speech audio, and delivers it to the user's client terminal. This intermediary architecture enables audio access to web pages without requiring complex client-side implementations, thus improving user convenience while maintaining manageable system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical interaction of reading text with the acoustic interaction of listening to speech. By substituting the visual-mechanical process of reading with an audio-based process, users gain the ability to access web page information through hearing, which improves accessibility and convenience particularly for users with visual impairments or those who prefer auditory learning.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If users must individually access each web page in a pre-selected list, then precise control over content access is maintained, but time efficiency is reduced

Engineering Contradiction:
Improvetime efficiencyVSAvoidoperational simplicity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent implements a multi-functional access system where a single interface provides multiple access modes for web page content. Users can select from text reading, audio playback, or a combination of both through a unified control mechanism. This universal access point eliminates the need to individually navigate and access each web page separately, as the system can process and deliver content in multiple formats through one interface, thereby improving time efficiency while maintaining operational simplicity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary processing of web page content by pre-converting text to speech audio and preparing multiple format versions before user request. This preliminary action allows users to immediately access content in their preferred format without waiting for conversion processes, significantly improving time efficiency when accessing pre-selected lists of web pages.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If text to speech converters are implemented on client terminals, then audio access to web pages is enabled, but client terminal requirements and user complexity increase

Engineering Contradiction:
Improveaccess flexibilityVSAvoidclient terminal requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent inverts the traditional client-server architecture for text-to-speech conversion. Instead of placing the text-to-speech converter on the client terminal as in conventional approaches, the invention moves the conversion functionality to the server side. The server process retrieves web page content, converts it to speech audio, and delivers the audio file to the client terminal for playback. This inversion eliminates the need for complex client-side text-to-speech software, reducing client terminal requirements while maintaining access flexibility.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent extracts the text-to-speech conversion functionality from the client terminal and relocates it to the server process. By separating the conversion function from the client device, the system reduces the computational burden and software requirements on client terminals. The client terminal only needs basic audio playback capability, while the server handles the complex text processing and speech synthesis, thus reducing device complexity while preserving adaptability.

Inventive Principle:
Principle #2Taking out (Extraction)

4Loss of information

If all web page content is converted to audio, then complete information access is provided, but audio file size and transmission time increase

Engineering Contradiction:
Improveinformation completenessVSAvoidtransmission time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts and removes non-audible information elements from web pages before converting the remaining content to speech. The server process filters out elements such as images, videos, advertisements, and other non-textual content that cannot be meaningfully converted to audio. This extraction process reduces the amount of content requiring text-to-speech conversion, thereby decreasing audio file sizes and transmission times while preserving all information that can be effectively communicated through audio.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing qualities to different portions of web page content based on their suitability for audio conversion. Textual information is fully converted to speech with high fidelity, while non-textual elements are either excluded or represented through descriptive audio cues. This local quality approach ensures that information completeness is maintained for audio-appropriate content while optimizing transmission efficiency by excluding or simplifying representation of non-audio content.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8838673B2Method and apparatus to generate audio versions of web pages
Publication Date: 2014.09.16 ASAPP INC
  • US8838673B2 patent drawing
  • US8838673B2 patent drawing
  • US8838673B2 patent drawing

AI summary

Audio files corresponding to a web page are generated by filtering a web page to remove characters that are non-audible. The audio files can be generated by a first server that receives a request for a web page or can be generated by a second server operating in cooperation with the first server. Additionally, web pages can be provided with a read me command button or other control object to allow audio versions of the web page to be selectively presented to a client terminal where the user desires to hear an audio version of the web page. Further, servers may maintain play lists of web pages, including audio versions thereof if desired. Some servers may maintain a preference list of web pages that users would like to hear audio versions of.