Conversation-Based Foreign Language Learning Method Using Speech Recognition and TTS
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional foreign language learning methods using speech recognition and TTS technologies do not effectively facilitate continuous conversation between learners and terminals, limiting opportunities for learners to practice speaking foreign languages naturally, as they often require button presses and lack accurate determination of correct pronunciation, leading to inefficient learning processes.
Innovation Solution
A conversation-based foreign language learning method that utilizes speech recognition and TTS functions to enable interactive conversation between learners and terminals through speech transmission, where the terminal enters a speech waiting state in response to voice commands, allows for repeated practice of learning target sentences, and provides feedback on pronunciation accuracy, minimizing screen interaction and addressing pronunciation differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition and TTS technologies are used for foreign language learning, then interactive conversation capability is improved, but the ability to accurately determine correct pronunciation and ensure continuous conversation is insufficient
Solution Approach 1:
The terminal provides feedback by comparing the learner's speech with reference audio and determining whether the speech corresponds to the reference. This feedback mechanism enables the terminal to accurately determine pronunciation correctness and guide the learner through iterative practice until the desired accuracy is achieved.
Solution Approach 2:
The terminal pre-stores reference audio data for learning target sentences before the learning process begins. This preliminary preparation allows the terminal to accurately compare and evaluate the learner's pronunciation in real-time without requiring complex real-time processing.
2Adaptability or versatility
If conventional speech recognition programs are used to execute functions with voice commands, then voice input capability is improved, but the program is not suitable for foreign language learning due to limitations in conversation topics and expression accuracy
Solution Approach 1:
The learning content is segmented into multiple learning target sentences with associated reference audio. The terminal processes and evaluates each sentence separately, allowing for precise control over conversation topics and guaranteed accuracy for each individual expression before moving to the next topic.
Solution Approach 2:
The terminal uses feedback mechanisms to verify that the learner's speech matches the reference audio for each learning target sentence. This ensures reliability of expression accuracy while maintaining versatility across multiple conversation topics through structured learning modules.
3Ease of operation
If the terminal requires button presses to enter speech waiting state, then operational control is improved, but screen interaction increases and natural conversation flow is interrupted
Solution Approach 1:
The terminal automatically enters speech waiting state based on audio detection without requiring manual button presses or screen interactions. The system serves itself by monitoring audio input and initiating the appropriate learning mode autonomously, eliminating the need for additional user actions.
Solution Approach 2:
The patent replaces mechanical button presses with acoustic field detection. Instead of requiring physical interaction with the device interface, the terminal uses audio sensors to detect the learner's voice and automatically initiates the speech waiting state, substituting mechanical input with acoustic input.
4Productivity
If the terminal does not provide repeated practice opportunities, then learning efficiency is improved, but pronunciation accuracy cannot be ensured
Solution Approach 1:
The terminal implements periodic repeated practice by continuously asking the learner to speak the learning target sentence until the speech corresponds to the reference audio. This periodic repetition maintains learning momentum while ensuring pronunciation accuracy through iterative feedback.
Solution Approach 2:
The terminal uses feedback to determine whether the learner's speech matches the reference and controls the repetition accordingly. When the speech does not correspond to the reference, the terminal repeats the learning target sentence for the learner to speak again, creating a feedback-driven practice loop that ensures accuracy.
Data Source
AI summary
A method for foreign language learning between a learner and a terminal, based on video or audio containing foreign language, particularly, to a conversation-based foreign language learning method using a speech recognition function and a TTS function of a terminal, a learner learns a foreign language in a way that: the terminal reads a current learning target sentence to the learner to allow the learner to speak the current learning target sentence after the terminal, when speech input by the learner in a speech waiting state of the terminal is the same as the current learning target sentence or belongs to the same category as the current learning target sentence; and the terminal and the learner alternately speak sentences one-by-one when the speech input by the learner is the same as the next sentence of the current learning target sentence or belongs to the same category as the next sentence.


