Speaker-Independent Speech Indexing for Multi-Language File Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-activated devices require pre-programming for voice input association, are limited by English language dominance, and lack multi-language recognition capabilities, making them inconvenient and costly to manufacture for diverse markets.

Innovation Solution

A method and apparatus that generate an index of digital file information entries, receive speaker-independent speech input in multiple languages, determine the input language, and compare it with the index to access digital files, allowing for seamless access without pre-recording and accommodating various character code sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pre-programming is used to associate voice input with particular entries, then speech recognition accuracy is improved, but device complexity and ease of operation deteriorate due to tedious setup process

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsetup convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs preliminary indexing of all digital files during initialization, creating a searchable database of file information including names, metadata, and content keywords. This preliminary action eliminates the need for users to pre-program voice associations, as the system is already prepared to recognize and match speech inputs against the indexed data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically indexes digital files and manages the speech recognition database without requiring user intervention. The device serves itself by continuously maintaining an updated index of files, allowing users to simply speak their requests without any setup or programming effort.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If multiple language recognition capabilities are added, then adaptability is improved, but manufacturing cost increases due to dedicated production lines

Engineering Contradiction:
Improvelanguage recognition capabilityVSAvoidmanufacturing cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system uses a universal speech recognition engine that can handle multiple languages through software configuration rather than hardware differentiation. The same device firmware and hardware platform support multiple languages, eliminating the need for dedicated production lines for each language version.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the language parameter dynamically based on user input or system detection. By modifying software parameters (language models, character encodings) rather than hardware configuration, the same device can adapt to different languages without requiring different manufacturing processes.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If character code set management is improved for multiple languages, then adaptability is improved, but device complexity increases

Engineering Contradiction:
Improvemulti-language supportVSAvoidcode set management
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system introduces an intermediary layer (software abstraction layer) that handles character code set conversions between different encodings (ASCII, Big-5, GB, JIS, etc.). This intermediary manages the complexity of multiple code sets internally, presenting a unified interface to users while handling encoding conversions transparently.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8015013B2Method and apparatus for accessing a digital file from a collection of digital files
Publication Date: 2011.09.06 CREATIVE TECHNOLOGY LTD
  • US8015013B2 patent drawing
  • US8015013B2 patent drawing
  • US8015013B2 patent drawing

AI summary

There is provided a method for accessing at least one digital file from a collection comprising more than one digital file in an electronic device, including: generating one index comprising of information entries obtained from each of the more than one digital file in the collection, with each digital file in the collection information being linked to at least one information entry; receiving a speaker independent speech input in at least one language during a speech reception mode; determining a language of the speech input; and setting the speech reception mode to the language of the speech input; comparing the speech input received during the speech reception mode with the entries in the index. The file may advantageously be accessed when the speech input coincides with at least one of the information entries in the index. The digital files may be stored in the electronic device, any device functionally connected to the electronic device or a combination of the aforementioned. The at least one digital file may be received from a source selected from: a memory device, a wired computer network or a wireless computer network. An apparatus that is able to carry out the aforementioned method is also disclosed.