Local Audio Fingerprint Recognition on Portable Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio recognition systems on portable devices require constant remote server connectivity, leading to high computational and bandwidth loads, and are limited by data plan usage and availability of Wi-Fi networks, making them inefficient and user-unfriendly.
Innovation Solution
A system that performs audio content recognition locally on portable devices by calculating a fingerprint of the recorded audio and comparing it with a resident fingerprint database, eliminating the need for remote server connectivity during recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio recognition is performed by sending samples to a remote server, then identification accuracy is improved, but network bandwidth consumption increases and recognition speed decreases
Solution Approach 1:
The system segments the audio recognition process into two parts: fingerprint extraction performed locally on the portable device, and database matching performed remotely on the server. This segmentation allows the computationally intensive fingerprint extraction to be done locally with minimal network transmission, while the database matching is handled by the server, thus reducing overall network bandwidth consumption while maintaining identification accuracy.
Solution Approach 2:
The system performs preliminary action by pre-computing and storing audio fingerprints in a database before actual recognition is needed. When recognition is required, the system only needs to extract a fingerprint from the query audio and match it against pre-computed fingerprints, rather than comparing entire audio files. This preliminary preparation significantly reduces the data transmission and processing requirements during actual recognition operations.
2Adaptability or versatility
If audio samples are transmitted via cell phone link, then remote server access is achieved, but data plan allotment is consumed
Solution Approach 1:
The system extracts only the essential fingerprint data from audio samples for transmission to the remote server, rather than transmitting entire audio files. This extraction principle reduces data plan consumption from potentially megabytes of audio data to just a compact fingerprint representation, while still enabling complete remote server access and recognition functionality.
3Ease of operation
If Wi-Fi is not available and cell network is unavailable, then portable device operates independently, but audio recognition cannot be performed
Solution Approach 1:
The system performs preliminary action by downloading and storing a local copy of the audio fingerprint database on the portable device before offline operation is needed. This allows the device to perform audio recognition independently without network connectivity. When online connectivity is restored, the local database can be updated with new fingerprints, thus maintaining both independent operation capability and continuous audio recognition functionality.
4Adaptability or versatility
If multiple users simultaneously access the remote server for audio recognition, then service coverage is improved, but network bottlenecks increase
Solution Approach 1:
The system segments the recognition workload by performing fingerprint extraction locally on each user's portable device, eliminating the need for multiple users to simultaneously transmit entire audio files to the server. Only compact fingerprint data needs to be transmitted and matched, significantly reducing network bottlenecks while maintaining service coverage for multiple simultaneous users.
5Reliability
If remote server hardware is over-provisioned to handle peak loads, then service reliability is improved, but hardware cost increases
Solution Approach 1:
The system applies partial action by having the remote server perform only the database matching function rather than complete audio analysis. This reduces the computational power and hardware resources required on the server side, as the majority of processing (fingerprint extraction) is performed locally on user devices. The server hardware can be right-sized for matching operations only, reducing over-provisioning while maintaining service reliability.
Data Source
AI summary
According to a preferred aspect of the instant invention, there is provided a system and method for content recognition in portable devices. Content, preferably audio content is recorded by the instant invention, preferably a sample with a length between 1 and 10 seconds. A fingerprint will be generated from the recorded sample and automatically, and preferably without further user interaction, prompting, notification, etc. (e.g. invisible to the user), compared with the fingerprints in a fingerprint database that is stored locally in the portable device and the result thereafter presented to the user.


