Sound Identification Using Banach Space Vector Centroids

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound identification systems in closed captioning applications suffer from a high miss rate, leading to inaccurate text labeling and reduced recall, as they struggle to balance precision and recall in sound recognition tasks.

Innovation Solution

The method involves comparing features from sound documents to acoustic files, generating sound vectors, creating centroid vectors, and redesignating false negatives as true positives using Banach spaces to optimize precision and recall, thereby improving the accuracy of sound identification and text labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional sound identification methods are used to maintain precision in sound recognition, then false positives are reduced, but the miss rate increases and recall decreases

Engineering Contradiction:
Improvesound identification precisionVSAvoidsound identification recall
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces Banach spaces as an intermediary mathematical framework between the acoustic feature extraction and sound identification processes. This framework provides a rigorous metric space for comparing sound vectors, enabling the system to accurately measure distances and similarities while maintaining both precision and recall in sound recognition

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter space by representing sounds as vectors in Banach spaces with specific norms (such as L1, L2, or L-infinity norms). By adjusting the norm parameter and threshold values, the system can optimize the balance between precision and recall based on specific application requirements

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the sound identification threshold is lowered to reduce miss rate, then recall improves, but precision decreases and false positives increase

Engineering Contradiction:
Improvesound identification recallVSAvoidsound identification precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where the Banach space distance calculations provide continuous information about the similarity between query sounds and database sounds. This feedback enables dynamic threshold adjustment and iterative refinement of identification results, allowing the system to maintain high recall while filtering false positives through multiple comparison rounds

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If more acoustic files are added to the database to improve coverage, then the ability to identify diverse sounds improves, but the complexity of the system increases

Engineering Contradiction:
Improvesound database coverageVSAvoidacoustic library management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses vector representations (copies) of acoustic features in Banach spaces instead of storing and processing raw acoustic files. This allows the system to work with compact mathematical representations that capture essential sound characteristics, reducing storage requirements and processing complexity while maintaining the ability to identify diverse sounds

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11176924B2Reduced miss rate in sound to text conversion using banach spaces
Publication Date: 2021.11.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11176924B2 patent drawing
  • US11176924B2 patent drawing
  • US11176924B2 patent drawing

AI summary

A computer-implemented method includes: comparing features extracted from a first document that include a sound to features extracted from acoustic files related to the sound; designating the sound in a document of the plurality of documents as a true; designating the sound in the first document as a false negative; generating a first sound vector for the sound in the first document in response to the sound in the first document being designated a false negative; generating a sound vector for each of the documents designated as a true positive; creating a centroid vector for the sound vectors of the documents designated as a true positive; and redesignating the sound in the first document from a false negative to a true positive in response to the first sound vector and the centroid vector being a Banach space.