Sound Source Visualization on Video Feed
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In surveillance systems, users cannot easily and visually recognize sound sources generating abnormal sounds due to the lack of integration of sound information with video data, limiting the efficiency of monitoring operations.
Innovation Solution
A monitoring system and method that utilize an omnidirectional camera and microphone array to image and collect sound data, processing sound parameters to visually represent sound sources on the video feed, allowing users to identify sound sources through a heat map overlay on the video data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If sound information is not integrated with video data on the monitor device, then the system structure remains simple, but users cannot visually recognize sound sources generating abnormal sounds
Solution Approach 1:
The patent combines sound information and video data into a single integrated display on the monitor device. The sound source position detection unit identifies the location of sound sources, and the image generation unit superimposes this sound information onto the video feed, creating a unified visual output that displays both video and sound source positions together.
Solution Approach 2:
The patent introduces an image generation unit as an intermediary component that processes and integrates sound source position data with video data. This intermediary unit creates the composite image by superimposing sound information onto the video feed, enabling visual recognition of sound sources without directly modifying the core video or sound collection systems.
2Ease of operation
If sound information is displayed on the monitor device, then users can visually recognize sound sources, but the device complexity increases
Solution Approach 1:
The patent divides the monitoring system into distinct functional modules: a sound source position detection unit that processes sound data, an image generation unit that creates the visual overlay, and a display unit that shows the composite image. This segmentation allows each component to perform its specific function independently, improving monitoring efficiency while managing system complexity through modular design.
Solution Approach 2:
The patent adds a visual dimension to sound information by converting audio data into spatial visual representations on the video feed. The sound source position detection determines the location of sound sources, and this positional information is displayed as visual markers on the video image, enabling users to intuitively understand sound source locations through spatial visualization.
3Loss of information
If multiple processing units are added to integrate sound and video, then sound source visualization is achieved, but processing time increases
Solution Approach 1:
The patent performs preliminary processing of sound data to detect sound source positions before generating the final composite image. The sound source position detection unit continuously monitors and identifies sound source locations in advance, preparing this information for rapid integration with video data when needed, thereby reducing real-time processing delays.
Solution Approach 2:
The patent replaces complex mechanical or computational processing with a streamlined image superimposition method. Instead of performing extensive real-time audio analysis and video processing separately, the system uses the detected sound source positions to directly overlay visual markers on the video feed, substituting complex processing with a more efficient graphical composition approach.
Data Source
AI summary
In a sound source display system, an omnidirectional camera captures an image of a monitoring area. A microphone array collects a voice in the monitoring area. A monitoring monitor displays the image of an imaging area captured by the omnidirectional camera. A sound pressure calculator in a directivity control device calculates a sound pressure indicating a source of a sound in the image of the imaging area using voice data of the voice collected by the microphone array. An output controller in the directivity control device compares the sound pressure and threshold values (first threshold value, second threshold value), and causes sound image information in which the sound pressure is converted into visual information according to the result of comparison, to be displayed on the monitoring monitor so as to be superimposed on the image of the imaging area.


