ART Network Background Model for Dynamic Video Scenes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video surveillance systems face difficulties in distinguishing between scene background and foreground, especially in dynamic or complex backgrounds with noise, compression artifacts, or varying illumination, leading to incorrect classification of pixels.
Innovation Solution
The use of an adaptive resonance theory (ART) network to generate a background model by mapping pixel appearance values to clusters, allowing for classification of pixels as background or foreground, and adapting to different background states over time without supervision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional background modeling methods are used in dynamic scenes, then the system is simple to implement, but the accuracy of pixel classification deteriorates due to noise, compression artifacts, and varying illumination
Solution Approach 1:
The patent segments the background model into multiple independent state representations (e.g., different lighting conditions, camera positions, or temporal states). Each state is modeled separately using probability distributions, allowing the system to handle complex dynamic scenes by combining multiple simplified models rather than using a single complex model.
Solution Approach 2:
The patent employs adaptive parameter adjustment where the background model parameters (mean, variance, state probabilities) are dynamically updated based on observed pixel values over time. This allows the model to adapt to changing conditions such as varying illumination, noise patterns, and compression artifacts without requiring manual reconfiguration.
2Reliability
If simple background modeling is used, then the device complexity is low, but the reliability of distinguishing background from foreground deteriorates in complex scenes
Solution Approach 1:
The patent implements a dynamic background model that evolves over time by incorporating temporal information. The model adapts to changing scene conditions by updating state probabilities and parameters based on sequences of video frames, allowing it to reliably distinguish between dynamic background elements and foreground objects even in complex environments.
Solution Approach 2:
The patent uses feedback mechanisms where the classification results are used to update the background model parameters. By continuously comparing observed pixel values with the modeled background states and adjusting the model accordingly, the system improves its reliability in distinguishing background from foreground over time without requiring increasing structural complexity.
3Measurement precision
If dynamic background states are modeled explicitly, then the accuracy of foreground detection improves, but the computational complexity increases
Solution Approach 1:
The patent uses efficient parameter updates for dynamic background states by leveraging temporal correlations and using compact probabilistic representations. Instead of full re-computation, the model parameters are incrementally adjusted using recursive formulas that require minimal computational resources while maintaining high foreground detection accuracy.
Solution Approach 2:
The patent applies different modeling complexities to different regions of the image based on their characteristics. Areas with high dynamic content use more sophisticated state modeling, while static regions use simpler models, optimizing the balance between detection accuracy and computational power consumption across the entire scene.
Data Source
AI summary
Techniques are disclosed for learning and modeling a background for a complex and/or dynamic scene over a period of observations without supervision. A background/foreground component of a computer vision engine may be configured to model a scene using an array of ART networks. The ART networks learn the regularity and periodicity of the scene by observing the scene over a period of time. Thus, the ART networks allow the computer vision engine to model complex and dynamic scene backgrounds in video.


