Grapheme-Phoneme Learning With Augmented Spectrogram Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing educational methods lack effective automated systems for real-time feedback and modeling in grapheme-phoneme correspondence learning, making it difficult for students to accurately learn letter-sound correspondence.

Innovation Solution

A computer-implemented system with a deep neural network provides real-time feedback and modeling by recognizing individual letter sounds, correcting mistakes, and allowing students to repeat correct pronunciations, using a grapheme-phoneme model trained with augmented spectrogram data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If automated systems are implemented for real-time feedback in grapheme-phoneme learning, then feedback speed and accuracy are improved, but system complexity increases

Engineering Contradiction:
Improvefeedback accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

A spectrogram generator serves as an intermediary component that transforms audio data into visual spectrogram representations. This mediator enables the neural network to process speech sounds more effectively by converting them into a format that highlights phonetic features, thereby improving feedback accuracy without requiring direct complex audio analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Traditional manual phonics instruction methods are replaced with an automated neural network system. The mechanical process of human teacher-student interaction is substituted with an automated system comprising audio recording, spectrogram generation, and neural network-based phoneme recognition, which provides consistent real-time feedback without human intervention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If data augmentation is applied to training spectrogram data, then model generalization is improved, but training time increases

Engineering Contradiction:
Improvemodel generalizationVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Data augmentation techniques are applied in advance during the training phase to create diverse spectrogram variations. By pre-processing and augmenting the training data before model training, the system prepares robust training examples that improve model generalization to unseen speech patterns, making the initial investment of time worthwhile for long-term performance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Various parameters of the spectrogram data are modified during augmentation, including frequency scaling, time stretching, and noise addition. These parameter changes create diverse training examples from limited data, improving the model's ability to generalize across different speech conditions without requiring additional recording sessions

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12387619B2Systems and methods for grapheme-phoneme correspondence learning
Publication Date: 2025.08.12 617 EDUCATION INC
  • US12387619B2 patent drawing
  • US12387619B2 patent drawing
  • US12387619B2 patent drawing

AI summary

Systems and methods are described for grapheme-phoneme correspondence learning. In an example, a display of a device is caused to output a grapheme graphical user interface (GUI) that includes a grapheme. Audio data representative of a sound made by the human user is received based on the grapheme shown on the display. A grapheme-phoneme model can determine whether the sound made by the human corresponds to a phoneme for the displayed grapheme based on the audio data. The grapheme-phoneme model is trained based on augmented spectrogram data. A speaker is caused to output a sound representative of the phoneme for the grapheme to provide the human with a correct pronunciation of the grapheme in response to the grapheme-phoneme model determining that the sound made by the human does not correspond to the phoneme for the grapheme.