Sound Recognition via Time-Frequency Map Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice recognition systems face information loss due to compression of time-frequency maps during image processing, which affects the accuracy of sound recognition.
Innovation Solution
The method involves converting original sounds into time-frequency maps without compression, segmenting and statistically sorting sound intensity information, and using a convolutional neural network-based image recognition method to enhance and recognize sound images, ensuring accurate matching with preset database sounds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If time-frequency map is compressed according to aspect ratio of image processing model, then image processing can be performed, but sound information is lost
Solution Approach 1:
The time-frequency map is divided into multiple sub-maps, each maintaining the original aspect ratio and sound information integrity. This segmentation allows the system to process audio data without compression while still being compatible with image processing models that require specific aspect ratios.
Solution Approach 2:
The patent introduces a new dimension by creating a one-to-many mapping relationship between time-frequency maps and sound images. Instead of directly compressing the time-frequency map, the system generates multiple sound images from segmented time-frequency sub-maps, adding a dimensional transformation that preserves information while enabling image processing compatibility.
2Productivity
If time-frequency map is compressed to fit image processing model, then processing efficiency is improved, but recognition accuracy decreases due to information loss
Solution Approach 1:
By segmenting the time-frequency map into multiple sub-maps, the system maintains processing efficiency while preventing information loss. Each sub-map can be processed independently, allowing the system to achieve both speed and accuracy.
Solution Approach 2:
The system creates multiple sound images as copies from the segmented time-frequency sub-maps. These copies preserve the original sound information integrity while being suitable for image processing, thereby maintaining both processing efficiency and recognition accuracy.
3Measurement precision
If original sound data integrity is maintained, then sound recognition accuracy is enhanced, but data processing complexity increases
Solution Approach 1:
Segmenting the time-frequency map into sub-maps reduces the complexity of processing large datasets while maintaining data integrity. Each sub-map can be processed with simpler algorithms, reducing overall computational complexity.
Solution Approach 2:
The system changes the parameter representation from a single compressed time-frequency map to multiple sound images derived from segmented sub-maps. This parameter transformation maintains data integrity while simplifying the processing requirements through the one-to-many mapping relationship.
Data Source
AI summary
A voice recognition system includes a computing device and at least one mobile terminal communicatively coupled to the computing device through a network. The computing device obtains an original sound from the at least one mobile terminal and converts the original sound into a digitized time-frequency map, performs compression segmentation on the time-frequency map to obtain a sound image corresponding to the time-frequency map, and uses an image recognition method to recognize the sound image, obtain an enhanced sound image, and search a preset database for sound information corresponding to the enhanced sound image.


