Crowd scene image captioning method based on multi-layer attribute guidance

The image captioning description method for crowd scenes guided by multi-level attributes solves the problem of insufficient dataset relevance in existing technologies, generates more vivid and detailed descriptions of complex crowd scenes, and improves the descriptive capabilities of image captioning datasets.

CN115294353BActive Publication Date: 2026-01-20UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210837834.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-16
Publication Date
2026-01-20
Estimated Expiration
2042-07-16

AI Technical Summary

Technical Problem

Existing image captioning datasets suffer from problems such as insufficient dataset relevance, limited descriptive perspectives, simplistic sentence structures, and simplistic backgrounds in crowd scene understanding research, making them unable to effectively address the challenges of complex crowd scenes.

Method used

A crowd scene image captioning description method based on multi-level attribute guidance is adopted. Through image feature extraction, visual feature embedding, multi-level dense crowd perception processing, feature fusion and dense crowd-oriented decoding steps, a more characteristic description of the crowd is generated.

Benefits of technology

It enables detailed descriptions of complex crowd scenes, generates more vivid and detailed captions, and improves the multimodal description capabilities of image caption datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294353B_ABST
    Figure CN115294353B_ABST
Patent Text Reader

Abstract

This invention proposes a method for describing crowd scene images based on multi-layer attribute guidance. It extracts regional visual features, corresponding location information, and human action features from the input image. A multi-layer perceptron (MLP) is used to obtain visual, location, and action features after feature embedding and mapping. Through a set feature processing layer and the MLP, global visual features, local features, object-level features, action-level features, and state-level features are obtained sequentially. A fusion feature is obtained using the global visual features, object-level features, action-level features, state-level features, and the hidden layer state from the previous time step. The semantic features for the current time step are obtained using the global visual features, the fusion feature, and the semantic features from the previous time step. Finally, the probability distribution of the current word is predicted based on the semantic features at the current time step and output. This invention extracts different levels of crowd attribute features, thereby generating more vivid and detailed descriptions specific to the crowd.
Need to check novelty before this filing date? Find Prior Art