Multi-Modal Monitoring Device Combining Voice and Video Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing monitoring methods face challenges in detecting unexpected abnormal situations solely based on video features, as defining pre-registered video features for various abnormal situations is impractical due to the diversity of physiognomic features, behaviors, and crimes, and sound analysis alone cannot determine the necessity of a response.
Innovation Solution
A monitoring system that combines voice acquisition, person identification, video analysis, and abnormal situation evaluation, where a voice acquisition unit captures abnormal voices, a person identification unit identifies the speaker, and an analysis unit searches for and evaluates the person's facial expression and motion in video footage to assess the situation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If video analysis with pre-registered features is used to detect abnormal situations, then detection capability for known abnormalities is improved, but adaptability to unexpected abnormal situations deteriorates
Solution Approach 1:
The patent combines video analysis with audio analysis (specifically voice recognition and acoustic analysis) to create a multi-modal monitoring system. This merging allows the system to detect abnormal situations through multiple channels: video features for known abnormalities and audio features (screams, shouts, glass breaking sounds) for unexpected situations, thereby resolving the contradiction between detection precision for known cases and adaptability to unknown cases.
Solution Approach 2:
The monitoring system is designed to perform multiple functions: it can detect pre-registered video features, recognize abnormal voices, analyze acoustic events, and estimate sound source positions. This multi-functionality enables the system to handle both expected abnormal situations (through video feature matching) and unexpected situations (through universal audio detection), improving both detection precision and adaptability simultaneously.
2Measurement precision
If multiple monitoring sensors are deployed to improve detection coverage, then detection accuracy is improved, but device complexity increases
Solution Approach 1:
The monitoring system is segmented into distinct functional modules: video analysis module, voice recognition module, acoustic analysis module, and sound source position estimation module. Each module processes specific types of data independently, then results are integrated for comprehensive abnormal situation detection. This segmentation allows high detection accuracy through multiple sensors while managing complexity through modular architecture.
Solution Approach 2:
The patent introduces an intermediary integration layer that coordinates between multiple sensors and analysis modules. This intermediary manages data flow, synchronizes processing, and integrates results from video, voice, and acoustic analyses, thereby enabling high detection accuracy through multiple sensors while controlling overall system complexity through centralized coordination.
Data Source
AI summary
Provided is a novel technology with which the occurrence of an abnormal situation can be detected and the abnormal situation can be appropriately ascertained. A monitoring device (1) comprises: a voice acquisition unit (2) that acquires prescribed speech spoken by a person due to the occurrence of an abnormal situation in a monitoring target area; a person identification unit (3) that identifies the person who spoke the prescribed speech, on the basis of a feature obtained from the prescribed speech; an analysis unit (4) that searches for the identified person in the images from a camera which images the monitoring target area, and that analyzes an expression or action of the person; and an abnormal situation evaluation unit (5) that evaluates the abnormal situation in the monitoring target area, on the basis of the analysis results.


