Speech Disorder Assistive Tool with FSM-Guided ASR Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional assistive technologies for individuals with speech disorders, such as aphasia and dysarthria, suffer from high error rates and inability to accurately interpret impaired speech, leading to frustration and discontinued use.

Innovation Solution

A system utilizing a finite state machine (FSM) model in conjunction with automatic speech recognition (ASR) to dynamically improve speech recognition accuracy by learning user patterns, presenting pictograms for selection, and updating the model based on user choices, thereby enhancing fluency and accuracy of communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional automatic speech recognition (ASR) systems are used for individuals with speech disorders, then the system can process speech input, but the recognition accuracy is low leading to high error rates

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system presents multiple candidate interpretations to the user and receives feedback through user selection. This feedback loop allows the system to learn from user corrections and improve recognition accuracy over time, directly addressing the low accuracy problem while maintaining reliability through iterative refinement

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary processing by generating multiple candidate interpretations before final selection. This preliminary action of creating several possible meanings allows the system to prepare multiple options for user selection, improving the chances of capturing the correct meaning while maintaining system reliability

Inventive Principle:
Principle #10Preliminary action

2Productivity

If conventional ASR systems are used, then speech input can be processed, but the system produces frequent errors causing user frustration and discontinued use

Engineering Contradiction:
Improvecommunication fluencyVSAvoidcommunication reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system dynamically adapts to individual users by learning from their speech patterns and selection feedback. This dynamic adaptation allows the system to improve communication fluency for each user over time while maintaining reliability through personalized modeling, directly countering the frustration caused by static error-prone systems

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

User selections provide continuous feedback that improves both fluency and reliability. The system learns from each interaction to better predict user intent, thereby improving communication fluency while simultaneously enhancing reliability through accumulated learning

Inventive Principle:
Principle #23Feedback

3Measurement precision

If a finite state machine (FSM) model with user learning is implemented, then speech recognition accuracy improves through adaptation, but the device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex learning task into manageable components: speech input processing, candidate generation, user feedback collection, and model updating. This segmentation makes the overall system more tractable and maintainable while achieving high recognition accuracy through coordinated operation of these modular components

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The FSM model acts as an intermediary between raw speech input and final interpretation. This intermediary layer processes and structures information, making the system more manageable despite increased complexity, while enabling accurate recognition through systematic pattern learning

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If multiple candidate words and pictograms are presented for user selection, then communication accuracy improves, but the time required for communication increases

Engineering Contradiction:
Improvecommunication accuracyVSAvoidcommunication time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system presents a limited set of top candidate interpretations rather than all possible meanings. This partial action approach focuses computational resources on the most likely candidates, reducing the time burden on users while maintaining high accuracy by presenting only the most relevant options for selection

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12361951B2Tool for assisting people with speech disorder
Publication Date: 2025.07.15 CERNER INNOVATION INC
  • US12361951B2 patent drawing
  • US12361951B2 patent drawing
  • US12361951B2 patent drawing

AI summary

Various tools are disclosed for providing assistive or augmentative means to enhance the fluency and accuracy of persons having speech disabilities. These technologies may automatically ascertain and dynamically improve the accuracy with which automatic speech recognition (ASR) systems recognize utterances of persons having impaired speech conditions. In an embodiment, digitized audio information about a speaker's utterance is processed to determine a set of candidate words matching the utterance. From these candidate words, a set of concepts is determined using a finite state machine model. A pictogram representing each concept is identified and presented to the speaker so that the speaker may select the pictogram corresponding to the best match of his or her intended meaning associated with the utterance. An action corresponding to speaker's selection then may be performed. For example, displaying or synthesizing speech from textual information describing the selected concept.