Invertible Spectrogram Signal Reconstruction via Overlapping Windows

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current signal analysis methods, particularly in time-frequency representation, often result in discontinuous or irreversible spectrograms, limiting the ability to reconstruct original signals and being susceptible to background noise, which affects accuracy in applications like speech recognition and digital transmission.

Innovation Solution

A method involving sectioning digital signals into overlapping windows, generating energy pulses through a descriptive function, and integrating these pulses with an oscillating function across a band-pass frequency filter to create a time-dependent analytical signal that is invertible and capable of approximating the original signal, while mitigating noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the signal is windowed into sections and Fourier transform is applied to each temporal section, then a time-frequency representation or spectrogram is obtained, but the spectrogram becomes discontinuous and irreversible

Engineering Contradiction:
Improvetime-frequency representation accuracyVSAvoidsignal reconstruction capability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The signal is divided into overlapping windows (Hanning, Hamming, or Blackman windows) to create multiple temporal sections. Each window is processed independently through Fourier transform, allowing the system to capture time-frequency characteristics while maintaining continuity through the overlap between adjacent windows. This segmentation approach resolves the contradiction by enabling detailed time-frequency analysis without sacrificing reconstruction capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The spectral information from multiple windowed sections is merged through coherent accumulation, where the Fourier coefficients from adjacent windows are combined to form a continuous time-frequency representation. This merging process ensures that the spectrogram remains continuous and irreversible, allowing perfect reconstruction of the original signal while maintaining measurement precision.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If frequency filtering is applied to isolate and interpret information in a signal, then signal analysis is improved, but the theoretical upper limit to bitrates for digital transmission is reached

Engineering Contradiction:
Improvesignal information extraction accuracyVSAvoiddata transmission bitrate
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent transitions from traditional frequency-domain filtering to a time-frequency domain approach using spectrograms. By analyzing signals in the time-frequency plane rather than solely in the frequency domain, the system can extract information more efficiently and achieve higher bitrates. The spectrogram provides a two-dimensional representation that captures both temporal and spectral characteristics, enabling more compact and efficient data transmission.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system changes the analysis parameters from traditional frequency filtering to spectrogram-based time-frequency representation. This parameter transformation allows for more efficient information encoding by capturing signal characteristics in a compressed time-frequency domain, thereby increasing the theoretical upper limit of transmissible bitrates while maintaining information extraction accuracy.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If current speech recognition approaches are used, then patterns of phonetic sounds can be identified, but the accuracy is highly susceptible to background noise

Engineering Contradiction:
Improvephonetic sound identification accuracyVSAvoidbackground noise susceptibility
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts relevant speech information from the time-frequency domain by identifying and isolating energy pulses corresponding to phonetic sounds. By taking out only the significant spectral components and filtering out background noise in the spectrogram domain, the system achieves robust speech recognition. The extraction process separates speech signals from noise based on their distinct time-frequency signatures, maintaining high identification accuracy even in noisy environments.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The spectrogram serves as an intermediary representation that mediates between the raw audio signal and the final speech recognition decisions. By transforming the audio signal into a time-frequency representation, the system creates an intermediate domain where speech patterns can be clearly distinguished from background noise. This intermediary transformation enables more accurate phonetic identification by providing enhanced contrast between speech and noise components.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of manufacture

If pre-recorded phonetic sounds are concatenated to synthesize speech, then speech generation is achieved, but the output does not sound natural or human

Engineering Contradiction:
Improvespeech synthesis capabilityVSAvoidnaturalness of synthesized speech
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent applies dynamics by generating speech signals through dynamic modulation of energy pulses in the time-frequency domain. Rather than static concatenation of pre-recorded phonemes, the system dynamically synthesizes speech by varying the amplitude, frequency, and temporal characteristics of energy pulses according to linguistic rules. This dynamic generation process creates smooth transitions and natural prosody, resulting in synthesized speech that sounds human-like while maintaining ease of manufacture through algorithmic control.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11867733B2Systems and methods of signal analysis and data transfer using spectrogram construction and inversion
Publication Date: 2024.01.09 ENOUY ROBERT WILLIAM
  • US11867733B2 patent drawing
  • US11867733B2 patent drawing
  • US11867733B2 patent drawing

AI summary

A method of generating an analytical signal for signal analysis. The method includes obtaining a digital signal, sectioning the digital signal into a series of overlapping windows in time domain, generating a plurality of energy pulses by evaluating a function that describes energy information within each window or set of windows, and generating a time-dependent analytical signal by generating an oscillating signal by multiplying each of the plurality of energy pulses by an oscillating function; and integrating the oscillating function across a band pass frequency filter.