Speech Recognition Engine for Real-Time Reading Fluency Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for oral reading fluency training require heavy preprocessing and extensive installation, making them inconvenient and difficult to maintain, and often lack the availability of human instructors for providing timely feedback.

Innovation Solution

A system that uses a processor-readable medium to initiate a speech recognition engine on a client device, allowing for real-time visual highlighting and fluency feedback without human intervention, by downloading and executing the speech recognition engine from a server, and synchronizing audio with text for visual highlighting and feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If known systems use computer speech recognition feedback for oral reading exercises, then instructional feedback is provided without human instructors, but heavy preprocessing and extensive installation are required

Engineering Contradiction:
Improveautomated speech recognition feedbackVSAvoidpreprocessing and installation complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent extracts the speech recognition engine from the complex preprocessing system and applies it directly to raw audio streams from reading exercises. This separation allows the system to provide automated feedback without requiring heavy manual crafting and alignment of audio and text content, resolving the contradiction between automation and complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system enables self-service by allowing users to directly upload audio files for automatic analysis without requiring manual preprocessing or intervention from human editors. The speech recognition engine autonomously processes the uploaded audio and provides feedback, eliminating the need for extensive installation and preprocessing steps

Inventive Principle:
Principle #25Self-service

2Reliability

If human instructors provide feedback for oral reading practice, then instructional quality is high, but availability of instructors is limited

Engineering Contradiction:
Improveinstructional feedback qualityVSAvoidavailability of instructors
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent replaces the mechanical system of human instructors with an automated speech recognition engine that processes audio uploads and provides feedback. This substitution maintains instructional quality through algorithmic analysis while dramatically improving availability, as the system can serve unlimited users simultaneously without human resource constraints

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system implements automated feedback loops where speech recognition technology analyzes uploaded audio recordings and generates instructional feedback without human intervention. This feedback mechanism maintains reliability by using established speech recognition algorithms while solving the availability problem by operating autonomously around the clock

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10210769B2Method and system for reading fluency training
Publication Date: 2019.02.19 ROSETTA STONE LTD
  • US10210769B2 patent drawing
  • US10210769B2 patent drawing
  • US10210769B2 patent drawing

AI summary

A non-transitory processor-readable medium stores code representing instructions to be executed by a processor. The code causes the processor to receive a request from a user of a client device to initiate a speech recognition engine for a web page displayed at the client device. In response to the request, the code causes the processor to (1) download, from a server associated with a first party, the speech recognition engine into the client device; and then (2) analyze, using the speech recognition engine, content of the web page including text in an identified language to produce analyzed content based on the identified language, where the content of the web page is received from a server associated with a second party. The code further causes the processor to send a signal to cause the client device to present the analyzed content to the user at the client device.