Surveillance Camera Audio Zooming via Beam-Forming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Surveillance camera systems face challenges in maintaining sound quality due to environmental noise and the need to selectively amplify audio signals from specific positions, which existing systems fail to address effectively.

Innovation Solution

A surveillance camera system that uses a microphone array and camera to receive video and audio signals, generate a sound-based heatmap, and perform beam-forming to selectively amplify human voice signals by correcting the audio zooming point based on video and audio data, allowing for precise audio zooming in user-designated areas.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a surveillance camera system collects sounds and videos simultaneously, then comprehensive surveillance data is obtained, but sound quality is degraded due to environmental noise and diffraction phenomena

Engineering Contradiction:
Improvesurveillance dataVSAvoidsound quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments the audio signal processing by dividing the surveillance area into multiple zones and applying different beam-forming parameters to each zone. The audio zooming function separates specific areas of interest from the general surveillance area, allowing independent processing of audio signals from different spatial regions. This segmentation enables the system to extract clear audio from specific zones while filtering out environmental noise from other areas.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality enhancement by applying audio zooming and beam-forming selectively to specific areas within the surveillance field. Different audio processing parameters are applied to different spatial zones, with enhanced audio quality focused on designated areas of interest while maintaining standard processing elsewhere. This allows the system to prioritize sound quality in critical surveillance zones.

Inventive Principle:
Principle #3Local quality

2Ease of manufacture

If the camera and microphone array are located in different planes or surfaces, then the device structure is simplified and easier to manufacture, but audio zooming precision is reduced

Engineering Contradiction:
Improvedevice structureVSAvoidaudio zooming precision
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent replaces the need for precise mechanical alignment between camera and microphone array with a software-based solution. The beam-forming algorithm and audio zooming processing compensate for the spatial misalignment between the camera plane and microphone array plane. By using signal processing techniques rather than mechanical precision, the system achieves accurate audio localization despite the simplified physical structure.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a composite solution by combining data from two different spatial perspectives - the camera's visual field and the microphone array's audio field. The system integrates video signals and audio signals from different planes, using the visual information to guide audio processing and vice versa. This composite approach allows the system to achieve precise audio zooming even when the sensor planes are not co-located.

Inventive Principle:
Principle #40Composite materials

3Quantity of substance

If beam-forming is performed on all audio signals, then comprehensive audio coverage is achieved, but processing time and computational resources increase

Engineering Contradiction:
Improveaudio coverageVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent applies partial action by performing beam-forming selectively rather than uniformly across all audio signals. The audio zooming function identifies specific areas of interest and applies intensive beam-forming processing only to those regions. For other areas, the system uses standard audio processing, reducing overall computational load while maintaining comprehensive surveillance coverage.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent segments the audio processing workload by dividing the surveillance area into multiple zones with different processing priorities. High-priority zones requiring audio zooming receive full beam-forming processing, while other zones receive lighter processing. This segmentation allows the system to manage computational resources efficiently while maintaining comprehensive audio coverage across the entire surveillance area.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system effectively enhances sound quality by selectively amplifying human voice signals in specific areas, improving audio zooming precision and reducing noise interference, thereby enhancing surveillance capabilities.

Implementation Method 1

a processor configured to change an audio zooming point of the camera device from a first point, in an image of the area captured by the camera device, to a second point based on the information about the area, and perform a beam-forming on an audio signal corresponding to the second point

Methodology Applied
Scientific EffectBeam-forming:

Data Source

PatentUS11462235B2Surveillance camera system for extracting sound of specific area from visualized object and operating method thereof
Publication Date: 2022.10.04 HANWHA VISION CO LTD
  • US11462235B2 patent drawing
  • US11462235B2 patent drawing
  • US11462235B2 patent drawing

AI summary

A camera system for extracting a sound of a specific area includes: a camera device configured to receive video signals and audio signals from an area; at least one memory configured to store information about the area including data corresponding to the video signals and the audio signals from the area; and a processor configured to change an audio zooming point of the camera device from a first point, in an image of the area captured by the camera device, to a second point based on the information about the area, and perform a beam-forming on an audio signal corresponding to the second point.