Personalized Spoken Language Identification with Multimodal Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing spoken language identification systems often rely solely on audio data, leading to inaccurate language identification, especially in cases of short-duration speech or limited speech data, which can result in improper processing.

Innovation Solution

A method that combines a trained spoken language identification model with characteristics of the person or electronic device, such as location, language history, and device settings, to determine a weighted sum of probabilities for improved language identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spoken language identification relies solely on audio data, then the system is simple and fast, but language identification accuracy deteriorates, especially for short-duration speech or limited speech data

Engineering Contradiction:
Improvelanguage identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple data sources (audio data, location data, language history) into a unified language identification system. The audio module provides acoustic features, the location module provides geographic context, and the language history module provides temporal patterns, all of which are combined to improve identification accuracy without requiring a completely new system architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates a universal language identification approach that can handle multiple scenarios (short-duration speech, limited speech data, different languages) by leveraging multiple data sources. The same core architecture can process various input types and adapt to different user contexts, making the system universally applicable across diverse language identification challenges.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple data sources are combined for language identification, then language identification accuracy improves, but processing time and computational complexity increase

Engineering Contradiction:
Improvelanguage identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing and storing location data and language history data before they are needed for language identification. Location information is cached and language patterns are pre-analyzed, so when actual language identification is required, the system only needs to process the audio data against these pre-prepared references, significantly reducing real-time processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality by tailoring the language identification approach to specific local contexts. Instead of using a single uniform method for all languages and situations, the system adjusts its behavior based on local factors such as geographic location and user-specific language history, optimizing accuracy for each local context while managing computational resources efficiently.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12367872B2Personalized multi-modal spoken language identification
Publication Date: 2025.07.22 SAMSUNG ELECTRONICS CO LTD
  • US12367872B2 patent drawing
  • US12367872B2 patent drawing
  • US12367872B2 patent drawing

AI summary

A method includes obtaining an audio input of a person speaking, where the audio input is captured by an electronic device. The method also includes, for each of multiple language types, (i) determining a first probability that the person is speaking in the language type by applying a trained spoken language identification model to the audio input, (ii) determining at least one second probability that the person is speaking in the language type based on at least one characteristic of the person or the electronic device, and (iii) determining a score for the language type based on a weighted sum of the first and second probabilities. The method further includes identifying the language type associated with a highest score as a spoken language of the person in the audio input.