Age-Adaptive ASR Module Selection for Language Level Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems for electronic devices are limited in accurately determining a target's language level, particularly for children, as they rely solely on voice input and use adult-oriented automatic speech recognition (ASR) modules, failing to consider the target's age and behavior.

Innovation Solution

An electronic device and method that utilize both voice and image data to identify a target's language level by selectively applying ASR modules appropriate for the target's age, converting the data into text, and generating outputs based on the identified language level.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If adult-oriented ASR modules are used for speech recognition, then the system can process voice input, but the accuracy of language level identification deteriorates for children

Engineering Contradiction:
Improvevoice input processingVSAvoidlanguage level identification accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent segments the ASR module into multiple versions tailored for different age groups (children, adolescents, adults). Each ASR module is specifically designed and trained for its target demographic, allowing the system to select the appropriate module based on the user's age. This segmentation resolves the contradiction by ensuring that children's speech is processed by a child-optimized ASR module, thereby improving language level identification accuracy while maintaining automated voice processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by making the ASR module's characteristics match the specific needs of different age groups. The child-oriented ASR module has different performance characteristics (vocabulary, grammar handling, speech patterns) compared to adult-oriented modules. This localized optimization ensures that each demographic receives processing tailored to their specific language development stage, resolving the accuracy deterioration issue for children.

Inventive Principle:
Principle #3Local quality

2Device complexity

If only voice input data is used for language level identification, then the process is simple, but the accuracy of language acquisition assessment deteriorates

Engineering Contradiction:
Improvedata processing complexityVSAvoidlanguage acquisition assessment accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges multiple data sources (voice input data and image data) to comprehensively assess language level. The system captures both auditory information from speech and visual information from images, then integrates these data streams to perform a more accurate language level identification. This combination resolves the contradiction by improving assessment accuracy through multi-modal data fusion while keeping the overall process integrated and manageable.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal identification system that handles multiple data types (voice and image) and serves multiple functions (language level identification, language acquisition assessment, disability detection). This multi-functional approach resolves the contradiction by demonstrating that increased data processing capability can be achieved within a unified framework that maintains reasonable complexity while significantly improving assessment accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If a single ASR module is used for all targets, then the device complexity is low, but the adaptability to different age groups deteriorates

Engineering Contradiction:
ImproveASR module configurationVSAvoidage group adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic ASR module selection mechanism that automatically chooses the appropriate ASR module based on the target user's age group. Rather than using a static single module for all users, the system dynamically adapts by selecting from multiple specialized modules (child-oriented, adolescent-oriented, adult-oriented). This dynamic approach resolves the contradiction by maintaining low device complexity through automated selection while achieving high adaptability to different age groups.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11961505B2Electronic device and method for identifying language level of target
Publication Date: 2024.04.16 SAMSUNG ELECTRONICS CO LTD
  • US11961505B2 patent drawing
  • US11961505B2 patent drawing
  • US11961505B2 patent drawing

AI summary

Methods and devices for identifying language level are provided. A first automatic speech recognition (ASR) module is identified, from among a plurality of ASR modules, based on information on a target received at the electronic device. First voice data and first image data for the target are received. The first voice data and the first image data are converted to first text data using the first ASR module. A first language level of the target is identified based on the first text data. Data including at least one of a voice output and an image output is output based on the first language level satisfying a condition.