Speech-Enhanced Email Processing for Selective Playback and Pronunciation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice assistants struggle with efficiently converting HTML-based email content into speech, often resulting in lengthy and non-meaningful audio output, and lack the ability to optimize pronunciation of sender names and subject lines.

Innovation Solution

Incorporating speech-enhanced email clients that utilize speech markup and binary encoded audio within MIME parts to provide optimized audio playback, including pronunciation instructions and audio files for sender names and subject lines, while allowing user control over content selection through whitelists and authentication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional text to speech technology is used to convert HTML email content directly into speech, then the email content can be rendered as audio, but the audio output becomes lengthy and non-meaningful because the system cannot skim or selectively read content

Engineering Contradiction:
Improveemail to speech conversion capabilityVSAvoidaudio playback duration
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent segments email content into distinct parts with different speech enhancement levels. The HTML body is divided into sections that can be selectively converted to speech, while alternative speech content provides pre-processed audio versions. This allows the system to play only relevant segments rather than converting entire lengthy emails, directly reducing audio playback duration while maintaining conversion capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by pre-processing email content into alternative speech versions during email composition or pre-send. Speech-enhanced email clients prepare optimized audio content in advance using speech markers and synthesis, so when the email is delivered, the recipient receives ready-to-play audio segments rather than requiring real-time conversion of entire content, significantly reducing playback time.

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If HTML email content is converted directly to speech from top to bottom, then the conversion process is simple, but the output is not meaningful because HTML visual arrangement using Cascading Style Sheets does not reflect logical speech order

Engineering Contradiction:
Improveconversion process simplicityVSAvoidmeaningful content structure
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent introduces speech markers as an intermediary layer between HTML content and speech synthesis. These markers (such as <speech:start>, <speech:end>, <speech:skip>) act as mediators that guide the speech engine through the HTML structure, indicating which sections should be spoken and in what order. This preserves meaningful content structure without requiring complex HTML-to-speech mapping logic.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter of content representation by creating alternative speech versions of email content. Instead of directly synthesizing from HTML structure, the system transforms content into speech-optimized formats with explicit pronunciation and structure markers, changing how the content is parameterized for speech output while maintaining the simplicity of the conversion process.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If speech markers are embedded within email body parts to aid conversion, then certain elements like signature blocks can be excluded from speech, but the system still cannot optimize pronunciation or provide alternate speech versions

Engineering Contradiction:
Improveselective content renderingVSAvoidpronunciation optimization capability
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent makes the alternative speech content mechanism universal by applying it to multiple email components beyond just the body. Speech-enhanced subject lines, sender names, and other email parts can all include alternative speech versions with optimized pronunciation. This multi-functional approach allows selective rendering optimization across the entire email while maintaining the ability to exclude inappropriate content like signatures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent optimizes pronunciation by changing the parameter of text representation into phonetically annotated speech content. Alternative speech versions include pronunciation guides and phonetic transcriptions that enable accurate speech synthesis of proper nouns, company names, and specialized terms. This transforms plain text parameters into speech-optimized parameters with explicit pronunciation instructions.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If voice assistants read entire email content as speech, then all information is conveyed, but the user experience deteriorates due to lengthy audio playback and inability to skip to relevant sections

Engineering Contradiction:
Improvecompleteness of information deliveryVSAvoiduser experience quality
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments email content into speech-enhanced and non-speech-enhanced portions, allowing the voice assistant to selectively process different sections. Alternative speech content provides pre-processed audio segments that can be played independently or in sequence, enabling users to listen to specific sections without hearing the entire email, thus improving user experience while maintaining information completeness through selective playback.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial action by converting only the necessary portions of email content to speech rather than the entire content. Alternative speech versions provide optimized audio for key information sections, allowing the voice assistant to perform partial conversion that delivers essential information without the overhead of converting and playing back every section, balancing information completeness with user experience.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12363057B1System and method for processing of speech content in email messages
Publication Date: 2025.07.15 KHOO JUSTIN
  • US12363057B1 patent drawing
  • US12363057B1 patent drawing
  • US12363057B1 patent drawing

AI summary

A system and method to provide email content within a separate part of an email, where the content is optimized for speaking and audio playback. Embodiments of the invention are methods for a speech-enabled email client to identify and use substitute speech content when outputting audio or reading to a user instead of the regular HTML or text parts of an email.