Sound Recognition via Time-Frequency Map Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice recognition systems face information loss due to compression of time-frequency maps during image processing, which affects the accuracy of sound recognition.

Innovation Solution

The method involves converting original sounds into time-frequency maps without compression, segmenting and statistically sorting sound intensity information, and using a convolutional neural network-based image recognition method to enhance and recognize sound images, ensuring accurate matching with preset database sounds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If time-frequency map is compressed according to aspect ratio of image processing model, then image processing can be performed, but sound information is lost

Engineering Contradiction:
Improveimage processing compatibilityVSAvoidsound information loss
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The time-frequency map is divided into multiple sub-maps, each maintaining the original aspect ratio and sound information integrity. This segmentation allows the system to process audio data without compression while still being compatible with image processing models that require specific aspect ratios.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension by creating a one-to-many mapping relationship between time-frequency maps and sound images. Instead of directly compressing the time-frequency map, the system generates multiple sound images from segmented time-frequency sub-maps, adding a dimensional transformation that preserves information while enabling image processing compatibility.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If time-frequency map is compressed to fit image processing model, then processing efficiency is improved, but recognition accuracy decreases due to information loss

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsound recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

By segmenting the time-frequency map into multiple sub-maps, the system maintains processing efficiency while preventing information loss. Each sub-map can be processed independently, allowing the system to achieve both speed and accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system creates multiple sound images as copies from the segmented time-frequency sub-maps. These copies preserve the original sound information integrity while being suitable for image processing, thereby maintaining both processing efficiency and recognition accuracy.

Inventive Principle:
Principle #26Copying

3Measurement precision

If original sound data integrity is maintained, then sound recognition accuracy is enhanced, but data processing complexity increases

Engineering Contradiction:
Improvesound recognition accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Segmenting the time-frequency map into sub-maps reduces the complexity of processing large datasets while maintaining data integrity. Each sub-map can be processed with simpler algorithms, reducing overall computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter representation from a single compressed time-frequency map to multiple sound images derived from segmented sub-maps. This parameter transformation maintains data integrity while simplifying the processing requirements through the one-to-many mapping relationship.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11183204B2Sound recognition system and method
Publication Date: 2021.11.23 HON HAI PRECISION INDUSTRY CO LTD
  • US11183204B2 patent drawing
  • US11183204B2 patent drawing
  • US11183204B2 patent drawing

AI summary

A voice recognition system includes a computing device and at least one mobile terminal communicatively coupled to the computing device through a network. The computing device obtains an original sound from the at least one mobile terminal and converts the original sound into a digitized time-frequency map, performs compression segmentation on the time-frequency map to obtain a sound image corresponding to the time-frequency map, and uses an image recognition method to recognize the sound image, obtain an enhanced sound image, and search a preset database for sound information corresponding to the enhanced sound image.