Human Parsing Using Spatial Distribution Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing human parsing methods using fully convolutional networks (FNCs) overlook spatial and statistical characteristics of the human body, leading to inefficiencies in processing time and memory utilization.
Innovation Solution
A method and device for human parsing using spatial distribution, which employs an attention mechanism based on spatial distribution according to spatial statistics. This involves generating height and width distribution maps, acquiring attention maps and scaled feature maps, calculating a distribution loss rate, and using this information to improve feature maps for accurate human parsing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If fully convolutional networks are used for human parsing, then the method can process human images, but processing time increases and memory utilization decreases
Solution Approach 1:
The patent segments the human body into distinct body parts (head, arms, legs, etc.) and compositional parts (clothes) for independent processing and classification. This segmentation allows the network to handle complex human images by breaking them down into manageable components, improving processing efficiency while maintaining parsing accuracy
Solution Approach 2:
The patent introduces spatial distribution characteristics as an additional dimension for analyzing human images. By incorporating spatial statistics and distribution patterns alongside traditional pixel-based analysis, the method achieves more efficient processing and better memory utilization while maintaining high parsing accuracy
2Measurement precision
If fully convolutional networks are used for human parsing, then the method can segment body parts, but spatial and statistical characteristics are overlooked
Solution Approach 1:
The patent incorporates feedback mechanisms that continuously compare predicted parsing results against spatial distribution statistics and ground truth data. This feedback loop allows the network to refine its predictions by incorporating spatial and statistical information, improving both parsing accuracy and information retention
Solution Approach 2:
The patent changes the parameters used for analysis by incorporating spatial distribution statistics, height maps, and width maps alongside traditional image data. These parameter changes enable the network to capture and utilize spatial and statistical characteristics that were previously overlooked, improving overall parsing precision
Data Source
AI summary
Provided are a device and method for human parsing. The method includes receiving at least one piece of image data and ground truth values for human parsing, generating a height distribution map and a width distribution map for the image data, acquiring attention maps and scaled feature maps for each of a height and a width of the image data using the distribution maps, calculating a distribution loss rate by concatenating the scaled feature maps, acquiring an improved feature map on the basis of the calculated distribution loss rate, and performing human parsing on an object included in the image data using the improved feature map.


