Multisyllabic Speech Assessment via Polygonal Line Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current language learning devices and speech recognition systems are inadequate for teaching and assessing pronunciation of multisyllabic words, particularly for hearing-impaired individuals, as they primarily operate on isolated characters rather than continuous speech, failing to effectively address the unique challenges of Chinese tone perception and consecutive word pronunciation.
Innovation Solution
A speech assessment device and method for a multisyllabic-word learning machine, which includes a standard speech database, a speech playing unit, a central processing system, and an assessment unit, capable of comparing and visualizing continuous audio files by converting spoken words into polygonal lines for evaluation, allowing users to select and play standard or learner audio files, and providing assessment results based on slope, turning time, and slope deviation analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If language learning devices operate on isolated characters, then character-by-character pronunciation can be assessed, but continuous speech flow and multisyllabic word pronunciation cannot be effectively evaluated
Solution Approach 1:
The patent segments continuous speech into individual character units while maintaining their sequential arrangement. The system divides a multisyllabic word into constituent characters, extracts pronunciation features from each character, and evaluates them individually while preserving the continuous speech context. This allows both character-level precision and continuous speech flow assessment.
Solution Approach 2:
The patent introduces a temporal dimension by arranging character pronunciation assessments in sequential order along a time axis. The system evaluates pronunciation features of each character in the order they appear in continuous speech, creating a time-ordered sequence of assessment results that reflects the natural flow of speech while maintaining individual character evaluation precision.
2Use of energy by moving object
If hearing aids are used to restore hearing, then auditory input is improved, but perception of unique Chinese tone frequency remains insufficient
Solution Approach 1:
The patent introduces visual aids as an intermediary between the hearing-impaired user and the speech content. The system converts audio speech into visual waveforms and pronunciation feature representations, allowing users to perceive tone frequency and pronunciation characteristics through visual channels rather than relying solely on auditory input, thereby compensating for the limitations of hearing aids in capturing subtle tone variations.
Solution Approach 2:
The patent creates a multi-functional assessment system that simultaneously evaluates multiple pronunciation features (pitch, duration, intensity) across multiple dimensions. The system provides both auditory feedback and visual representation, serving both hearing-impaired and non-hearing-impaired users, and can assess both isolated characters and continuous speech, making it universally applicable to diverse learning and rehabilitation needs.
3Measurement precision
If speech recognition systems are designed for isolated characters, then character recognition accuracy is improved, but actual spoken communication which is consecutive and pausable is not accurately assessed
Solution Approach 1:
The patent creates a dynamic assessment system that adapts to the user's speech pace and pauses. The system can evaluate pronunciation features whether the user speaks continuously or with pauses between characters, adjusting the assessment methodology to match the natural rhythm of speech. This allows accurate evaluation of both isolated characters and continuous speech without requiring the user to speak mechanically without pauses.
Data Source
AI summary
A speech assessment device and method for a multisyllabic-word learning machine, and a method for visualizing continuous audio are provided. By performing the step of starting the assessment mode, the step of selecting words to be assessed, the step of choosing to play or record, the step of recording, the step of visualization (including the step of picking out fundamental frequency, the step of defining analysis point, the step of transforming polygonal lines, and the step of simplifying the polygonal lines), the step of repeating, and the step of assessment, the speech assessment device and method for a multisyllabic-word learning machine are capable of providing assistance in oral language learning, and capable of rehabilitating patients with hearing impairment through visual aids.


