Automatic Speech Fluency Measurement via Prosodic Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computerized speech recognition systems require manual transcription and human intervention, limiting their ability to analyze spontaneous speech and measure fluency effectively, especially for individuals with speech impairments such as aphasia or autism spectrum disorders.
Innovation Solution
An automated speech analysis system that collects and analyzes prosodic characteristics of a patient's speech, identifying phonemes, disfluencies, and intonation to measure fluency without manual transcription, using a speech analyzer with recognition engines and repetition detectors to produce objective and automatic assessments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual transcription and human intervention are used for speech analysis, then measurement precision of speech fluency can be achieved, but productivity and ease of operation deteriorate due to time-consuming manual processes
Solution Approach 1:
The system enables automatic speech fluency analysis by having the computer system perform transcription and measurement tasks autonomously without requiring human operators to manually transcribe speech. The speech-to-text program automatically processes audio samples, identifies disfluencies, and generates fluency measurements, allowing the system to serve itself in the analysis process.
Solution Approach 2:
The patent replaces the mechanical human transcription process with an automated computer-based speech recognition system. Instead of human operators manually transcribing speech and measuring fluency, the system uses statistical language models and automatic disfluency detection algorithms to perform these tasks, substituting mechanical automation for human manual operations.
2Ease of operation
If speech recognition programs discard disfluencies to create clean text documents, then ease of operation improves for text processing, but loss of information occurs regarding speech fluency characteristics
Solution Approach 1:
The system extracts and separately analyzes disfluencies from the speech signal rather than discarding them. The speech-to-text program identifies and isolates disfluency events (such as filled pauses, repetitions, and false starts) and uses these extracted features to calculate speech fluency measurements, preserving the information that would otherwise be lost in traditional text processing.
Solution Approach 2:
The patent introduces an intermediary measurement layer that operates between the raw speech signal and the final text output. The speech fluency measurement system acts as an intermediary that processes the speech signal to extract fluency characteristics, allowing both the cleaned text document and the fluency information to coexist without one eliminating the other.
3Productivity
If automated speech analysis is implemented, then productivity and ease of operation improve, but device complexity increases due to additional analysis components
Solution Approach 1:
The speech-to-text program is designed to perform multiple functions: it not only transcribes speech to text but also simultaneously performs disfluency detection and speech fluency measurement. By making the system universal and multi-functional, the patent avoids the need for separate dedicated devices for each task, thereby managing complexity while maintaining high productivity.
Solution Approach 2:
The patent merges the speech transcription function and the speech fluency measurement function into a single integrated system. The speech-to-text program and the speech fluency measurement components work together as a unified system, combining multiple functions into one device rather than requiring separate systems for each task.
Data Source
AI summary
Techniques are described for automatically measuring fluency of a patient's speech based on prosodic characteristics thereof. The prosodic characteristics may include statistics regarding silent pauses, filled pauses, repetitions, or fundamental frequency of the patient's speech. The statistics may include a count, average number of occurrences, duration, average duration, frequency of occurrence, standard deviation, or other statistics. In one embodiment, a method includes receiving an audio sample that includes speech of a patient, analyzing the audio sample to identify prosodic characteristics of the speech of the patient, and automatically measuring fluency of the speech of the patient based on the prosodic characteristics. These techniques may present several advantages, such as objectively measuring fluency of a patient's speech without requiring a manual transcription or other manual intervention in the analysis process.


