Adaptive Speech Compression for Navigation POI Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional navigation devices require high-performance equipment to provide enhanced searching functions for geographical names and POIs, making them complex and difficult to implement effectively.
Innovation Solution
An information terminal system that includes an audio input unit, a communication unit, a POI specifying unit, and a route searching unit, which compresses and transmits speech information to a server for noise removal and recognition, allowing for efficient POI identification and route searching, even with lower processing capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech information is transmitted at high quality without compression, then speech recognition accuracy is improved, but communication bandwidth consumption increases
Solution Approach 1:
The patent applies parameter changes by dynamically adjusting the compression rate of speech information based on noise level conditions. When noise exceeds a threshold, a lower compression rate is used to preserve speech quality for accurate recognition, while in low-noise conditions, higher compression rates reduce bandwidth consumption. This resolves the contradiction by making compression parameters adaptive rather than fixed.
Solution Approach 2:
The system dynamically changes the compression rate based on real-time noise level detection. The noise detection unit continuously monitors the speech signal, and the compression rate is adjusted accordingly - using lower compression (higher quality) when noise is high, and higher compression (lower quality) when noise is low. This dynamic adaptation resolves the static contradiction between quality and bandwidth.
2Measurement precision
If noise removal processing is applied to speech information, then speech recognition accuracy is improved, but processing time increases
Solution Approach 1:
The noise removal processing is performed preliminarily on the compressed speech information before transmission and recognition. By removing noise in advance during the compression stage rather than during recognition, the system reduces the computational burden on the recognition engine and overall processing time, while still achieving high accuracy.
Solution Approach 2:
The system takes preliminary anti-action by detecting noise levels and applying appropriate noise removal and compression strategies before the speech recognition process begins. This preemptive handling of noise prevents degradation of recognition accuracy and reduces the need for complex post-processing, thereby minimizing total processing time.
3Productivity
If compression rate is increased to reduce data size, then communication efficiency is improved, but speech quality deteriorates
Solution Approach 1:
The compression rate parameter is changed dynamically based on noise conditions. When ambient noise is high, the system uses lower compression rates to maintain speech quality for accurate recognition. When noise is low, higher compression rates are applied to maximize communication efficiency. This parameter adaptation resolves the contradiction by making compression intensity conditional rather than constant.
Solution Approach 2:
The system transitions from static compression to dynamic compression where the compression rate adjusts in real-time based on noise level detection. This dynamic approach allows the system to optimize the balance between communication efficiency and speech quality according to actual environmental conditions, resolving the fixed contradiction.
4Productivity
If high-performance devices are used for speech recognition and POI searching, then searching function performance is improved, but device complexity and cost increase
Solution Approach 1:
The patent extracts the complex speech recognition and noise removal processing from the local navigation device and relocates it to a server environment. The terminal device only needs to perform simple audio input, compression based on noise levels, and transmission, while the server handles the computationally intensive recognition and POI searching. This extraction resolves the contradiction by moving complexity from the client device to the server.
Solution Approach 2:
The system introduces a server as an intermediary between the terminal device and the speech recognition process. The terminal communicates speech data to the server, which performs noise removal, recognition, and POI searching, then returns results to the terminal. This intermediary handles the complex processing, allowing the terminal to maintain simplicity while achieving high-performance searching functions.
Data Source
AI summary
An object of the present invention is to provide a technique of an information terminal that allows more efficient utilization of high-level searching functions. The information terminal is provided with an audio input accepting unit to accept an input of speech information, a communication unit to establish communication with a predetermined server device via a network, an output unit, a POI specifying unit to transmit the speech information accepted by the audio input accepting unit to the server device and receive information specifying a candidate of a POI (Point Of Interest) associated with the speech information, a POI candidate output unit to output to the output unit, the information specifying the candidate of the POI received by the POI specifying unit, and a route searching unit to accept a selective input of the information specifying the candidate of the POI, and search for a route directed to the POI.


