Autocorrelation Speech Recognition via ACF Factor Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition methods face challenges in handling wide frequency information, noise, and spatial information in acoustic fields, leading to prediction errors and difficulties in reflecting human auditory perception, especially in complex environments like the 'cocktail party effect'.

Innovation Solution

The method calculates running autocorrelation functions from speech signals, extracts specific ACF factors such as Wφ(0), τ1, φ1, and Δφ1/Δt, and uses these to identify syllables by comparing them with templates, allowing for segmentation and recognition without spectral analysis, effectively handling noise and environmental complexities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional spectral analysis methods are used to extract speech features, then speech signal spectra can be estimated, but complex parameters are required and prediction errors increase

Engineering Contradiction:
Improvespeech signal spectra estimation accuracyVSAvoidparameter complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential autocorrelation factors (peak position, peak amplitude, width at half maximum) from the autocorrelation function, discarding the complex spectral parameters. This extraction approach obtains the necessary speech features without requiring complex spectral analysis parameters, thereby reducing parameter complexity while maintaining recognition accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses autocorrelation function peaks as a simplified copy or representation of speech spectral characteristics. Instead of analyzing the full complex spectrum, the method captures the essential speech features through autocorrelation peak parameters, creating a simplified model that maintains recognition capability while reducing complexity.

Inventive Principle:
Principle #26Copying

2Measurement precision

If conventional spectral analysis methods are used, then speech features can be obtained, but noise handling capability is poor

Engineering Contradiction:
Improvespeech feature extraction accuracyVSAvoidnoise sensitivity
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent converts the potential harm of noise by using autocorrelation analysis which inherently provides noise robustness. The autocorrelation function and its peak parameters remain stable even in noisy environments, transforming the challenging noise condition into a manageable situation where the essential speech features can still be extracted accurately.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Productivity

If conventional analytical methods are used, then speech recognition can be performed, but spatial information and human auditory perception are not reflected

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidspatial information handling
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent changes the parameter representation from complex spectral parameters to autocorrelation function parameters (peak position, amplitude, width). These autocorrelation parameters better reflect human auditory perception characteristics and can capture spatial information in acoustic fields, thereby improving adaptability while maintaining speech recognition capability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9514738B2Method and device for recognizing speech
Publication Date: 2016.12.06 YOSHIMASA ELECTRONICS
  • US9514738B2 patent drawing
  • US9514738B2 patent drawing
  • US9514738B2 patent drawing

AI summary

A speech is recognized using ACF factors extracted from running autocorrelation functions calculated from the speech. The extracted ACF factors are a Wφ(0) (width of ACF amplitude around zero-delay origin), a Wφ(0)max (maximum value of the Wφ(0)), a τ1 (pitch period), a φ1 (pitch strength), and a Δφ1/Δt (rate of the pitch strength change). Syllables in the speech are identified by comparing the ACF factors with templates stored in a database.