Sound Recognition Confidence Adjustment via Learned Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound recognition systems face challenges in accurately identifying non-verbal sound events and scenes from audio signals, particularly in dynamic environments where prior knowledge and device-specific properties are not adequately considered, leading to inefficiencies and inaccuracies.

Innovation Solution

A method that adjusts sound class scores in audio signals by using a learned model to process confidence values associated with properties like energy and environment, allowing for dynamic adaptation without retraining the machine learning model, and incorporating external knowledge to improve recognition accuracy across various devices and conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sound class scores are adjusted using a learned model that incorporates device-specific properties and environmental factors, then recognition accuracy is improved, but system complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing audio frames to extract device-specific properties and environmental factors before sound class classification. This preliminary processing enables the learned model to receive enriched input data that captures contextual information, thereby improving recognition accuracy without requiring complex modifications to the core classification algorithm.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary learned model that acts as a mediator between raw audio data and sound class scores. This intermediary component processes device-specific properties and environmental factors, transforming them into adjusted sound class scores that reflect contextual conditions. The learned model serves as a bridge that reconciles raw observations with contextual knowledge, improving accuracy while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system dynamically adapts to changing environments and devices, then recognition robustness is improved, but computational resources increase

Engineering Contradiction:
Improvedynamic adaptation capabilityVSAvoidcomputational resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by selectively processing only the most relevant device-specific properties and environmental factors for each audio frame. Rather than comprehensively analyzing all possible parameters, the learned model focuses on key features that have the greatest impact on sound class classification. This selective processing enables dynamic adaptation to changing environments while constraining computational resource consumption to essential operations only.

Inventive Principle:
Principle #16Partial or excessive action

3Stability of the object's composition

If sound class scores are adjusted based on device-specific properties, then consistency across devices is improved, but processing time increases

Engineering Contradiction:
Improverecognition consistencyVSAvoidprocessing time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent implements parameter changes by modifying sound class scores based on device-specific properties extracted from audio frames. The learned model adjusts classification parameters (sound class scores) according to device characteristics such as microphone sensitivity, frequency response, and environmental conditions. This parameter adjustment mechanism ensures that recognition results remain consistent across different devices by compensating for device-specific variations in the scoring parameters.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10878840B1Method of recognising a sound event
Publication Date: 2020.12.29 META PLATFORMS TECHNOLOGIES LLC
  • US10878840B1 patent drawing
  • US10878840B1 patent drawing
  • US10878840B1 patent drawing

AI summary

A method for recognising at least one of a non-verbal sound event and a scene in an audio signal comprising a sequence of frames of audio data, the method comprising: for each frame of the sequence: receiving at least one sound class score, wherein each sound class score is representative of a degree of affiliation of the frame with a sound class of a plurality of sound classes; for a sound class score of the at least one sound class scores: determining a confidence that the sound class score is representative of a degree of affiliation of the frame with the sound class by processing a value for a property associated with the frame, wherein the value is processed using a learned model for the property; adjusting the sound class score for the frame based at least on the determined confidence.