Prosody Feature Extraction for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems fail to effectively utilize prosody features, leading to a loss of emotional and intentional aspects of human communication, which are crucial for unique identification and improved recognition accuracy.

Innovation Solution

A system and method that integrate prosody feature extraction and classification into voice pattern recognition, using adaptive signal processing techniques to enhance speech feature extraction and pattern matching, specifically employing dynamic time warping with adaptive search band limits and multi-layer recognition algorithms to improve speech recognition for speech-disabled populations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional speech recognition systems are used, then the system is simple and easy to operate, but the emotional and intentional aspects of human communication are lost

Engineering Contradiction:
Improveemotional and intentional aspectsVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the speech signal into multiple frames and extracts different types of features (acoustic features, prosody features, emotional features) from each frame. This segmentation allows the system to process complex information systematically while maintaining computational manageability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested feature extraction architecture where prosody features are extracted from acoustic features, which are in turn extracted from the speech signal. The emotional recognition module then processes these nested features to identify emotional states, creating a hierarchical structure that manages complexity through organization.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Measurement precision

If prosody features are added to speech recognition, then recognition accuracy is improved, but the device complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidfeature extraction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system changes the parameter space by adding prosodic parameters (pitch, duration, intensity) to the traditional acoustic parameters (spectral features). This parameter expansion enables the system to capture emotional and intentional information while maintaining a systematic approach to feature processing that manages complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system uses feedback mechanisms where extracted prosody features are fed back into the recognition process to refine and adjust the speech recognition results. This feedback loop allows the system to continuously improve accuracy while managing complexity through iterative processing.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If dynamic time warping with adaptive search band limits is used, then speech event transition detection is improved, but computational requirements increase

Engineering Contradiction:
Improvespeech event transition detection precisionVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system implements dynamic time warping with adaptive search band limits, where the search parameters are dynamically adjusted based on the speech signal characteristics and processing stage. This dynamic adaptation allows the system to optimize computational energy usage by focusing calculations on the most relevant temporal regions rather than processing the entire signal uniformly.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality principles by differentiating the processing intensity across different temporal regions of the speech signal. The adaptive search band limits concentrate computational resources on critical transition regions while using lighter processing in stable regions, thereby improving detection precision where needed while managing overall computational energy consumption.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9754580B2System and method for extracting and using prosody features
Publication Date: 2017.09.05 TECH FOR VOICE INTERFACE
  • US9754580B2 patent drawing
  • US9754580B2 patent drawing
  • US9754580B2 patent drawing

AI summary

A system for carrying out voice pattern recognition and a method for achieving same. The system includes an arrangement for acquiring an input voice, a signal processing library for extracting acoustic and prosodic features of the acquired voice, a database for storing a recognition dictionary, at least one instance of a prosody detector for carrying out a prosody detection process on extracted respective prosodic features, communicating with an end user application for applying control thereto.