Speech Recognition Using Human Position Sensing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems require significant memory and time to recognize utterances, making them inefficient and frustrating to use for interfacing with electronic devices, as they rely heavily on visual cues and are not intuitive enough for users.
Innovation Solution
A method that uses sensory inputs of human position, such as changes in orientation, direction, or motion, to select a recognition set and recognize speech input signals, allowing for faster and more accurate speech recognition through the use of sensors like tactile and gyroscope sensors, which can sense body positions and convert them into digital data for processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional speech recognition methods are used, then speech recognition capability is achieved, but memory requirements and processing time increase significantly
Solution Approach 1:
The patent segments the speech recognition process by dividing recognition sets into multiple subsets based on sensory inputs. Instead of using one large recognition set, the system creates multiple smaller subsets (e.g., first recognition set, second recognition set) that are selected based on detected sensory conditions, thereby reducing memory requirements while maintaining recognition capability.
Solution Approach 2:
The patent implements dynamic selection of recognition sets based on real-time sensory inputs. The system dynamically adjusts which recognition set is used depending on detected conditions (such as tactile sensor inputs or gyroscope readings), making the speech recognition system adaptive and context-aware rather than static.
2Reliability
If traditional speech recognition methods are used, then speech recognition capability is achieved, but processing time increases considerably
Solution Approach 1:
By segmenting the recognition process into multiple context-specific subsets, the system reduces the search space for speech recognition. Each subset contains only relevant speech patterns for a particular context, significantly reducing processing time compared to searching through a single large recognition set.
Solution Approach 2:
The patent performs preliminary actions by detecting sensory inputs and pre-selecting the appropriate recognition set before speech recognition begins. This preliminary context establishment narrows down the recognition scope in advance, reducing the time required for actual speech processing.
3Ease of operation
If visual cues are used for interfacing, then device control is achieved, but user learning time and complexity increase
Solution Approach 1:
The patent makes the interface universal by accepting multiple input modalities (tactile, motion, and speech) rather than relying solely on visual cues. This multi-functional approach allows users to interact with the device through their most comfortable modality, reducing learning time and improving ease of operation.
Solution Approach 2:
The system provides self-service by automatically detecting the user's context through sensory inputs and adjusting the recognition set accordingly, without requiring the user to manually configure or learn complex visual interfaces. The system adapts to user needs automatically.
Data Source
AI summary
Embodiments of the present invention improve methods of performing speech recognition using sensory inputs of human position. In one embodiment, the present invention includes a speech recognition method comprising sensing a change in position of at least one part of a human body, selecting a recognition set based on the change of position, receiving a speech input signal, and recognizing the speech input signal in the context of the first recognition set.


