Audio Event Detection Using Context-Aware Confidence Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing complexity of classifying captured audio due to various overlapping sounds in monitoring sensors leads to incorrect identification of audio events, such as false detections, especially when environmental noise is present, making it challenging to accurately notify users of relevant events.

Innovation Solution

A system that uses context information and location data from sensors to adjust the confidence level of audio event classification, employing audio and video analysis modules with machine learning algorithms to differentiate between relevant and irrelevant sounds, thereby reducing false positives and improving notification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the monitoring sensor monitors more types of sounds to increase detection coverage, then the quantity of detected audio events increases, but the complexity of classifying captured audio at a high confidence level increases and false detections increase

Engineering Contradiction:
Improvedetection coverageVSAvoidclassification confidence
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces context information as an intermediary element that mediates between the raw audio detection and the final classification decision. This context information (including location data, time of day, and environmental factors) acts as a mediator to adjust confidence levels and resolve ambiguities in audio event classification, thereby maintaining high detection coverage while improving classification reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically changes the confidence level parameter based on context information. Instead of using a fixed confidence threshold, the system adjusts the confidence level by incorporating contextual parameters such as sensor location, time of day, and environmental conditions. This allows the system to maintain versatility in detecting various sound types while improving the reliability of classifications through context-aware confidence adjustment.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the monitoring sensor monitors more types of sounds to increase detection coverage, then the quantity of detected audio events increases, but false detections increase due to environmental noise

Engineering Contradiction:
Improvedetection coverageVSAvoidfalse detections
Core Design Contradiction:
Adaptability or versatilityVSObject-generated harmful factors

Solution Approach 1:

Context information serves as an intermediary that filters out false detections caused by environmental noise. By introducing location data, time of day, and environmental context as mediating factors, the system can distinguish between relevant audio events and noise, thereby maintaining broad detection coverage while reducing false positives.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent converts potentially harmful environmental noise into beneficial context information. By analyzing the context in which sounds occur (location, time, environmental conditions), the system transforms what would be interfering noise into useful data for distinguishing false detections from genuine events, thereby converting a harmful factor into a beneficial filtering mechanism.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Reliability

If context information and location data are used to adjust confidence levels, then false positives are reduced and notification accuracy improves, but the device complexity increases

Engineering Contradiction:
Improvenotification accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements a multi-functional context analysis module that handles multiple tasks: location tracking, time monitoring, environmental sensing, and confidence level adjustment. By creating a universal context processing system that performs all these functions through a unified approach, the patent reduces the need for separate specialized components, thereby improving notification accuracy while limiting the increase in overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If machine learning algorithms are used to differentiate between relevant and irrelevant sounds, then classification accuracy improves, but the device complexity and computational requirements increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary filtering of audio events using context information before applying machine learning algorithms. By pre-processing the audio data and filtering out obviously irrelevant sounds based on location, time, and environmental context, the system reduces the computational burden on the machine learning model. This preliminary action maintains high classification accuracy while limiting the increase in computational complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220067090A1Systems, methods, and apparatuses for intelligent audio event detection
Publication Date: 2022.03.03 COMCAST CABLE COMM LLC
  • US20220067090A1 patent drawing
  • US20220067090A1 patent drawing
  • US20220067090A1 patent drawing

AI summary

Methods, systems, and apparatuses for intelligent audio event detection are described herein. Audio data and video data from a sensor is analyzed. The audio data may include an audio event of interest that is associated with a confidence level. The confidence level may be adjusted based on a location of the sensor and context data associated with the audio event. Notifications may be sent based on the adjusted confidence level and the context data.