Sound Detection with Image Capture System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound recognition systems in computer-monitored environments lack the ability to effectively integrate sound event information with images and automate camera operations based on synchronized audio and visual data, leading to suboptimal image processing and user experience.

Innovation Solution

A computing device configured to receive image and audio metadata, detect target sounds or scenes, and output camera control commands to process images accordingly, enhancing image quality and user interaction by incorporating sound-related tags at the point of capture rather than relying on cloud-based processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sound recognition systems process audio data in the cloud, then comprehensive sound analysis can be achieved, but computational load and latency increase

Engineering Contradiction:
Improvesound analysis accuracyVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The system segments the sound processing workflow into edge-based preliminary detection (using lightweight sound event detection models on the device) and cloud-based comprehensive analysis. This segmentation allows immediate local responses while reducing the volume and complexity of data transmitted to the cloud, thereby lowering computational load and energy consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary sound event detection and filtering at the edge device before transmitting data to the cloud. By pre-processing audio data locally and only sending relevant segments or metadata to the cloud, the system reduces the computational burden on cloud servers and minimizes energy consumption during data transmission and processing.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If camera operations are controlled manually, then user flexibility is maintained, but automation and response speed to sound events are reduced

Engineering Contradiction:
Improveuser flexibilityVSAvoidcamera operation automation
Core Design Contradiction:
Ease of operationVSExtent of automation

Solution Approach 1:

The system implements dynamic camera control where automation level adjusts based on sound event characteristics. For critical sound events (e.g., baby crying, glass breaking), the system automatically triggers appropriate camera actions. For less critical events, the system provides suggestions or waits for user confirmation, thereby maintaining user flexibility while achieving necessary automation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system provides feedback to users about detected sound events and suggested camera actions. Users can review the sound event details, suggested actions, and captured images before final confirmation or modification of camera operations. This feedback loop maintains user control and flexibility while enabling automated assistance for rapid response.

Inventive Principle:
Principle #23Feedback

3Loss of information

If all audio data is captured during image capture, then complete audio context is preserved, but data redundancy and storage requirements increase

Engineering Contradiction:
Improveaudio context completenessVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only the relevant sound events and their temporal boundaries from the continuous audio stream during image capture. Instead of storing all audio data, the system identifies and extracts specific sound event segments (e.g., laughter, applause, speech) that co-occur with images, thereby preserving essential audio context while dramatically reducing data volume for storage and processing.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different data retention strategies to different time segments of the audio stream. For segments containing significant sound events co-occurring with images, the system preserves complete audio data with high quality. For segments without significant events, the system reduces storage requirements by storing only metadata or skipping storage entirely, thereby optimizing the balance between audio context completeness and data volume.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11468904B2Computer apparatus and method implementing sound detection with an image capture system
Publication Date: 2022.10.11 META PLATFORMS TECHNOLOGIES LLC
  • US11468904B2 patent drawing
  • US11468904B2 patent drawing
  • US11468904B2 patent drawing

AI summary

A computing device comprising a processor, the processor configured to: receive, from an image capture system, an image captured in an environment and image metadata associated with the image, the image metadata comprising an image capture time; receive a sound recognition message from a sound recognition module, the sound recognition message comprising (i) a sound recognition identifier indicating a target sound or scene that has been recognised based on captured audio data captured in the environment, and (ii) time information associated with the sound recognition identifier; detect that the target sound or scene occurred at a time that the image was captured based on the image metadata and the time information in the sound recognition message; and output a camera control command to said image capture system based on said detection.