Web Page Audio Generation via Server-Side Text Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face inconvenience in accessing web page information, as they must read web pages and individually access each page, with no efficient option to hear them or select alternative modes of information access.
Innovation Solution
A server process that generates a text version of a web page, filters out non-audible information, and creates an audio file, which is then transmitted to the client terminal, allowing users to hear the web page and providing options for audio file encoding based on connection speed and user preferences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If web pages are provided in traditional text form requiring user reading, then information delivery is straightforward, but user convenience and accessibility are reduced
Solution Approach 1:
The patent introduces an intermediary system comprising a server process and text-to-speech converter that mediates between the web page content and the user. The server process retrieves web page content, processes it through filtering to remove non-audible elements, converts it to speech audio, and delivers it to the user's client terminal. This intermediary architecture enables audio access to web pages without requiring complex client-side implementations, thus improving user convenience while maintaining manageable system complexity.
Solution Approach 2:
The patent replaces the mechanical interaction of reading text with the acoustic interaction of listening to speech. By substituting the visual-mechanical process of reading with an audio-based process, users gain the ability to access web page information through hearing, which improves accessibility and convenience particularly for users with visual impairments or those who prefer auditory learning.
2Productivity
If users must individually access each web page in a pre-selected list, then precise control over content access is maintained, but time efficiency is reduced
Solution Approach 1:
The patent implements a multi-functional access system where a single interface provides multiple access modes for web page content. Users can select from text reading, audio playback, or a combination of both through a unified control mechanism. This universal access point eliminates the need to individually navigate and access each web page separately, as the system can process and deliver content in multiple formats through one interface, thereby improving time efficiency while maintaining operational simplicity.
Solution Approach 2:
The system performs preliminary processing of web page content by pre-converting text to speech audio and preparing multiple format versions before user request. This preliminary action allows users to immediately access content in their preferred format without waiting for conversion processes, significantly improving time efficiency when accessing pre-selected lists of web pages.
3Adaptability or versatility
If text to speech converters are implemented on client terminals, then audio access to web pages is enabled, but client terminal requirements and user complexity increase
Solution Approach 1:
The patent inverts the traditional client-server architecture for text-to-speech conversion. Instead of placing the text-to-speech converter on the client terminal as in conventional approaches, the invention moves the conversion functionality to the server side. The server process retrieves web page content, converts it to speech audio, and delivers the audio file to the client terminal for playback. This inversion eliminates the need for complex client-side text-to-speech software, reducing client terminal requirements while maintaining access flexibility.
Solution Approach 2:
The patent extracts the text-to-speech conversion functionality from the client terminal and relocates it to the server process. By separating the conversion function from the client device, the system reduces the computational burden and software requirements on client terminals. The client terminal only needs basic audio playback capability, while the server handles the complex text processing and speech synthesis, thus reducing device complexity while preserving adaptability.
4Loss of information
If all web page content is converted to audio, then complete information access is provided, but audio file size and transmission time increase
Solution Approach 1:
The patent extracts and removes non-audible information elements from web pages before converting the remaining content to speech. The server process filters out elements such as images, videos, advertisements, and other non-textual content that cannot be meaningfully converted to audio. This extraction process reduces the amount of content requiring text-to-speech conversion, thereby decreasing audio file sizes and transmission times while preserving all information that can be effectively communicated through audio.
Solution Approach 2:
The patent applies different processing qualities to different portions of web page content based on their suitability for audio conversion. Textual information is fully converted to speech with high fidelity, while non-textual elements are either excluded or represented through descriptive audio cues. This local quality approach ensures that information completeness is maintained for audio-appropriate content while optimizing transmission efficiency by excluding or simplifying representation of non-audio content.
Data Source
AI summary
Audio files corresponding to a web page are generated by filtering a web page to remove characters that are non-audible. The audio files can be generated by a first server that receives a request for a web page or can be generated by a second server operating in cooperation with the first server. Additionally, web pages can be provided with a read me command button or other control object to allow audio versions of the web page to be selectively presented to a client terminal where the user desires to hear an audio version of the web page. Further, servers may maintain play lists of web pages, including audio versions thereof if desired. Some servers may maintain a preference list of web pages that users would like to hear audio versions of.


