Syllable-Based Speech Synthesis Using Rhythmic Beats

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition and synthesis technologies lack a systematic and accurate model to effectively convert spoken words into written text and vice versa, often resulting in unclear or artificial-sounding outputs due to challenges in segmenting and synchronizing sound patterns.

Innovation Solution

The approach models speech as cognitively-driven sensory-motor activity, focusing on syllable patterns and their rhythmic combinations to identify and generate natural-sounding speech, using syllable segmentation, gestural schemata, and prominence scales to correlate articulatory and acoustic patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech is segmented into sound segments (phones/phonemes) and arranged linearly to form words, then it should be possible to exhaustively pair sound segment arrangements with corresponding alphabetic letters, but the resulting speech is unclear and artificial-sounding

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidspeech naturalness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments speech into syllables rather than traditional phones/phonemes. Each syllable is defined by a specific acoustic pattern with a peak frequency and duration, creating a more natural segmentation unit that preserves the rhythmic and intonational characteristics of human speech while enabling systematic processing and pairing with text.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If speech is modeled as a series of ongoing and simultaneous gestures modified by neuro-muscular systems, then anatomical structures can be classified to produce gestures, but the challenge remains to systematically account for their synchronization

Engineering Contradiction:
Improvegesture modeling capabilityVSAvoidsynchronization complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent models speech as a sequence of periodic syllabic beats, where each syllable represents a rhythmic unit with characteristic duration and frequency patterns. This periodic structure naturally accounts for the timing and synchronization of multiple articulatory gestures without requiring complex ad hoc phase definitions, as the rhythmic framework provides a systematic temporal organization for all gesture coordination.

Inventive Principle:
Principle #19Periodic action

3Reliability

If constellations or molecules metaphors are used to bundle gestures together as a basis for speech synthesis, then gesture synchronization can be achieved, but a systematic and accurate model for speech synthesis and recognition is still lacking

Engineering Contradiction:
Improvegesture bundling effectivenessVSAvoidmodel systematicity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the abstract gesture bundling concept into concrete, measurable acoustic parameters. Each syllable is characterized by specific parameters including peak frequency, duration, amplitude envelope, and spectral features. This parameter-based approach provides a systematic and accurate model that bridges the gap between gesture metaphors and implementable speech synthesis and recognition systems, enabling precise control over speech production while maintaining naturalness.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9747892B1Method and apparatus for electronically sythesizing acoustic waveforms representing a series of words based on syllable-defining beats
Publication Date: 2017.08.29 FRIDMAN MINTZ BORIS
  • US9747892B1 patent drawing
  • US9747892B1 patent drawing
  • US9747892B1 patent drawing

AI summary

Speech is modeled as a cognitively-driven sensory-motor activity where the form of speech is the result of categorization processes that any given subject recreates by focusing on creating sound patterns that are represented by syllables. These syllables are then combined in characteristic patterns to form words, which are in turn, combined in characteristic patterns to form utterances. A speech recognition process first identifies syllables in an electronic waveform representing ongoing speech. The pattern of syllables is then deconstructed into a standard form that is used to identify words. The words are then concatenated to identify an utterance. Similarly, a speech synthesis process converts written words into patterns of syllables. The pattern of syllables is then processed to produce the characteristic rhythmic sound of naturally spoken words. The words are then assembled into an utterance which is also processed to produce a natural sounding speech.