Human Detection via Spatiotemporal Volume Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional human detection methods using still images or single differential images struggle with predicting shape characteristic changes, leading to false or missed detections, especially when the walking direction of humans is not consistent or when the ankle position is unknown, and require pre-defined slit positions, limiting detection areas.
Innovation Solution
A human detection device that generates a three-dimensional spatiotemporal image from a moving picture, extracts and verifies spatiotemporal fragments using a human movement model, allowing for detection of human presence and gait direction without pre-defined detection areas or known walking directions, using a spatiotemporal volume generation unit, extraction unit, output unit, verification unit, and attribute output unit.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional methods use still images or single differential images for human detection, then the detection process is simple, but false detection and non-detection occur due to inability to predict shape characteristic changes
Solution Approach 1:
The patent transitions from two-dimensional still images to three-dimensional spatiotemporal images by adding the time dimension. Multiple frame images are arranged along the temporal axis to form a 3D spatiotemporal volume, enabling detection of temporal changes in human shape characteristics and reducing false detections
2Measurement precision
If the first conventional art extracts spatiotemporal fragments based on known ankle position, then human detection is possible, but detection is limited to left-right walking directions throughout the image
Solution Approach 1:
The patent creates a universal detection method that does not depend on pre-known anatomical positions or walking directions. The spatiotemporal fragment extraction and verification process works for humans walking in any direction, making the detection system adaptable to various scenarios without requiring prior knowledge of gait patterns
Solution Approach 2:
The patent dynamically determines fragment extraction lines based on the spatiotemporal volume data rather than using fixed anatomical references. The system adapts to different walking directions by analyzing temporal pixel value changes and generating verification patterns that match observed motion, enabling detection of humans walking in any direction
3Area of stationary object
If the second conventional art places a plurality of slits throughout the image for human detection, then detection area is expanded, but slit positions must be decided in advance by designer
Solution Approach 1:
The patent enables the system to automatically determine fragment extraction lines and verification patterns from the spatiotemporal volume data itself, without requiring pre-configured slit positions or manual design input. The system self-adapts to detect humans in any region of the image by analyzing temporal pixel value changes and generating appropriate verification patterns
4Measurement precision
If initial detection of human is required to detect ankle position, then detection can be performed, but it becomes difficult to detect humans walking in various directions within the image
Solution Approach 1:
The patent inverts the conventional approach by not starting with anatomical landmark detection (ankle position) and then determining walking direction. Instead, it analyzes temporal pixel value changes across the entire spatiotemporal volume to directly identify human presence and gait characteristics, eliminating the need for initial ankle position detection and enabling detection of humans walking in any direction
Data Source
AI summary
The present invention provides a human detection device which detects a human contained in a moving picture, and includes the following: a spatiotemporal volume generation unit which generates a three-dimensional spatiotemporal image in which frame images that make up the moving picture in which a human has been filmed are arranged along a temporal axis; a spatiotemporal fragment extraction unit which extracts a real image spatiotemporal fragment, which is an image appearing in a cut plane or cut fragment when the three-dimensional spatiotemporal image is cut, from the generated three-dimensional spatiotemporal image; a human body region movement model spatiotemporal fragment output unit which generates and outputs, based on a human movement model which defines a characteristic of the movement of a human, a human body region movement spatiotemporal fragment, which is a spatiotemporal fragment obtained from a movement by the human movement model; a spatiotemporal fragment verification unit which verifies between a real image spatiotemporal fragment and a human body region movement model spatiotemporal fragment; and an attribute output unit which outputs a human attribute which includes the presence/absence of a human in the moving picture, based on that verification result.


