Phonetic Refrain Extraction for Speech Audio Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-driven selection of audio files is hindered by language barriers, unknown character encodings, and the lack of phonetic or orthographic information in metadata, leading to difficulties in accurately identifying and selecting audio files.

Innovation Solution

A method and system that detect the refrain in an audio file by generating a phonetic transcription and identifying repeated vocal segments, which are then used to create a phonetic or acoustic representation for speech recognition, allowing users to select audio files based on the best matching voice command.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If speech-driven selection is implemented without phonetic information, then the system can operate with simple metadata, but the accuracy of audio file identification deteriorates due to language barriers and encoding issues

Engineering Contradiction:
Improvesimplicity of metadata storageVSAvoidaccuracy of audio file identification
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by generating phonetic transcriptions of audio file contents and refrains during the audio file processing stage, before the speech recognition stage. This pre-computed phonetic information is stored alongside the audio file, enabling accurate speech-driven selection without requiring complex real-time transcription during playback. The phonetic representation is created in advance using speech recognition technology, resolving the contradiction between simple metadata storage and accurate file identification.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If traditional metadata tags are used for audio file identification, then the system structure remains simple, but the selection accuracy deteriorates due to unknown character encodings, spelling mistakes, and language variations

Engineering Contradiction:
Improvesystem structureVSAvoidfile selection accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces phonetic transcriptions as an intermediary layer between the audio file content and the speech recognition system. Instead of directly comparing user speech against potentially inaccurate metadata tags (titles, artists, albums), the system mediates through phonetic representations that capture the actual spoken content of the audio file, particularly the refrain. This intermediary phonetic layer resolves encoding issues, spelling variations, and language barriers, improving reliability while maintaining relatively simple system structure.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If the entire audio file is transcribed for speech recognition, then complete information is available for selection, but the processing time and computational resources increase significantly

Engineering Contradiction:
Improvecompleteness of audio informationVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent applies the extraction principle by isolating and transcribing only the refrain portion of each audio file rather than the entire content. The system identifies and extracts the refrain - the repeated, characteristic vocal segment that users are most likely to sing along to or recognize. This selective extraction provides sufficient information for accurate speech-driven file selection while dramatically reducing processing time and computational resources compared to transcribing complete audio files.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If phonetic transcriptions of refrains are generated and stored, then speech-driven selection accuracy improves, but the data storage requirements and processing complexity increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs phonetic transcription and refrain identification in advance during audio file processing, storing the results for later speech recognition operations. This preliminary processing converts audio content into phonetic representations that can be efficiently compared against user speech inputs. By doing this work beforehand, the system achieves high speech recognition accuracy without requiring complex real-time processing during the actual file selection operation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8106285B2Speech-driven selection of an audio file
Publication Date: 2012.01.31 HARMAN BECKER AUTOMOTIVE SYST GMBH
  • US8106285B2 patent drawing
  • US8106285B2 patent drawing
  • US8106285B2 patent drawing

AI summary

A system and method for detecting a refrain in an audio file having vocal components. The method and system includes generating a phonetic transcription of a portion of the audio file, analyzing the phonetic transcription and identifying a vocal segment in the generated phonetic transcription that is repeated frequently. The method and system further relate to the speech-driven selection based on similarity of detected refrain and user input.