Voice-Controlled Media Selection Using Phonetic Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled multimedia systems face challenges in recognizing speech commands when media file names are in different languages, as the language of the entry is often unknown, leading to low recognition rates and difficulties in processing variable vocabulary.

Innovation Solution

Incorporating phonetic data into file identification metadata, which includes different pronunciations of file names across languages, allowing the speech recognition unit to compare user inputs more accurately and improve selection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition is used to select media files with names in different languages, then voice-controlled operation is enabled, but recognition rate decreases when the language of the entry is unknown

Engineering Contradiction:
Improvevoice-controlled operationVSAvoidrecognition rate
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary action by storing multiple phonetic transcriptions of media file names in different languages before the speech recognition process. When a user speaks a media file name in their native language, the system can match it against pre-stored transcriptions in various languages, enabling successful recognition even when the user's language differs from the original media file language.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system achieves universality by implementing a multi-language phonetic transcription system that can handle speech recognition across multiple languages. The phonetic data structure stores transcriptions in various languages for each media file, allowing the speech recognition unit to universally process user inputs regardless of the user's native language or the original media file language.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If phonetic transcriptions are generated automatically, then processing speed is improved, but recognition accuracy decreases

Engineering Contradiction:
Improveprocessing speedVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary action by pre-storing multiple phonetic transcriptions of media file names in different languages before the speech recognition process. When a user speaks a media file name in their native language, the system can match it against pre-stored transcriptions in various languages, enabling successful recognition even when the user's language differs from the original media file language.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system achieves universality by implementing a multi-language phonetic transcription system that can handle speech recognition across multiple languages. The phonetic data structure stores transcriptions in various languages for each media file, allowing the speech recognition unit to universally process user inputs regardless of the user's native language or the original media file language.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9153233B2Voice-controlled selection of media files utilizing phonetic data
Publication Date: 2015.10.06 HARMAN BECKER AUTOMOTIVE SYST GMBH
  • US9153233B2 patent drawing
  • US9153233B2 patent drawing
  • US9153233B2 patent drawing

AI summary

A voice-controlled data system is providing that has a storage medium for storing media files, the media files having associated file identification data for allowing the identification of the media files, the file identification data including phonetic data having phonetic information corresponding to the file identification data. The phonetic data is supplied to a speech recognition unit that compares the phonetic data to a speech command input into the speech recognition unit. The data system further includes a file selecting unit that selects one of the media files based on the comparison result.