Sound Activity Detection Using P-Norm Histogram Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing elderly watch services face challenges in accurately determining the active or inactive state of a person using sound information, as they often fail to distinguish between human activity sounds and background noises, leading to incorrect classifications, especially with sounds like rain being mistakenly identified as active states.
Innovation Solution
The system employs a method that calculates the variety of sounds within a certain time frame as an index to determine the active state, using the p-order norm of a normalized histogram to differentiate between activity and background sounds, thereby enhancing the robustness of activity detection and reducing false positives.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If sound information is used to determine active state, then coverage area is improved, but measurement precision deteriorates due to inability to distinguish human activity sounds from background noises
Solution Approach 1:
The patent segments sound data into multiple feature dimensions (frequency spectrum, time domain, cepstral coefficients) and further segments different sound environments into separate classification models. This allows precise identification of human activity sounds by analyzing multiple segmented features rather than treating all sounds uniformly, thereby improving measurement precision while maintaining broad coverage.
Solution Approach 2:
The patent changes parameters by extracting multiple acoustic features (frequency spectrum, zero-crossing rate, cepstral coefficients) from sound data and adapting classification parameters based on different sound environments. This enables the system to distinguish human activity sounds from background noises by analyzing changes in these parameters, resolving the contradiction between coverage and precision.
2Area of stationary object
If multiple sensors are used to expand detection range, then coverage area is improved, but device complexity increases
Solution Approach 1:
The patent makes a single sensor perform multiple functions by extracting various acoustic features (frequency spectrum, time domain characteristics, cepstral coefficients) from the same sound data. This multi-functional approach allows one sensor to achieve detection capabilities that would otherwise require multiple sensors, reducing device complexity while maintaining expanded detection range through advanced signal processing.
Solution Approach 2:
The patent replaces the mechanical approach of deploying multiple physical sensors with an information-processing approach using a single sensor combined with sophisticated acoustic feature extraction and classification algorithms. This substitution reduces installation complexity while maintaining or expanding detection capabilities through computational methods.
3Ease of operation
If traditional sound classification is used, then ease of operation is improved, but reliability deteriorates due to false positives from background sounds
Solution Approach 1:
The patent changes parameters by utilizing multiple acoustic features (frequency spectrum, zero-crossing rate, cepstral coefficients) and adapting classification thresholds based on different sound environments. This enables reliable distinction between human activity sounds and background noises while maintaining ease of operation through automated multi-parameter analysis.
Solution Approach 2:
The patent implements feedback by using learned classification models that adapt to different sound environments and provide feedback on classification confidence. This feedback mechanism improves reliability by continuously refining the distinction between active and inactive states based on accumulated data, reducing false positives while maintaining operational simplicity.
Data Source
AI summary
An information processing apparatus including: a memory, and a processor coupled to the memory and the processor configured to: detect a plurality of sounds in sound data captured in a space within a specified period, classify the plurality of sounds into a plurality of kinds of sound based on similarities of the plurality of sounds respectively, and determine a state of a person in the space within the specified period based on counts of the plurality of kinds of sound.


