Human Parsing Using Spatial Distribution Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing human parsing methods using fully convolutional networks (FNCs) overlook spatial and statistical characteristics of the human body, leading to inefficiencies in processing time and memory utilization.

Innovation Solution

A method and device for human parsing using spatial distribution, which employs an attention mechanism based on spatial distribution according to spatial statistics. This involves generating height and width distribution maps, acquiring attention maps and scaled feature maps, calculating a distribution loss rate, and using this information to improve feature maps for accurate human parsing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If fully convolutional networks are used for human parsing, then the method can process human images, but processing time increases and memory utilization decreases

Engineering Contradiction:
Improveprocessing speedVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments the human body into distinct body parts (head, arms, legs, etc.) and compositional parts (clothes) for independent processing and classification. This segmentation allows the network to handle complex human images by breaking them down into manageable components, improving processing efficiency while maintaining parsing accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial distribution characteristics as an additional dimension for analyzing human images. By incorporating spatial statistics and distribution patterns alongside traditional pixel-based analysis, the method achieves more efficient processing and better memory utilization while maintaining high parsing accuracy

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If fully convolutional networks are used for human parsing, then the method can segment body parts, but spatial and statistical characteristics are overlooked

Engineering Contradiction:
Improveparsing accuracyVSAvoidspatial and statistical characteristics
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent incorporates feedback mechanisms that continuously compare predicted parsing results against spatial distribution statistics and ground truth data. This feedback loop allows the network to refine its predictions by incorporating spatial and statistical information, improving both parsing accuracy and information retention

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameters used for analysis by incorporating spatial distribution statistics, height maps, and width maps alongside traditional image data. These parameter changes enable the network to capture and utilize spatial and statistical characteristics that were previously overlooked, improving overall parsing precision

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12272172B2Method and device for human parsing
Publication Date: 2025.04.08 AJOU UNIV IND ACADEMIC COOP FOUND
  • US12272172B2 patent drawing
  • US12272172B2 patent drawing
  • US12272172B2 patent drawing

AI summary

Provided are a device and method for human parsing. The method includes receiving at least one piece of image data and ground truth values for human parsing, generating a height distribution map and a width distribution map for the image data, acquiring attention maps and scaled feature maps for each of a height and a width of the image data using the distribution maps, calculating a distribution loss rate by concatenating the scaled feature maps, acquiring an improved feature map on the basis of the calculated distribution loss rate, and performing human parsing on an object included in the image data using the improved feature map.