Crowd scene image captioning method based on multi-layer attribute guidance
The image captioning description method for crowd scenes guided by multi-level attributes solves the problem of insufficient dataset relevance in existing technologies, generates more vivid and detailed descriptions of complex crowd scenes, and improves the descriptive capabilities of image captioning datasets.
Patent Information
- Application Number
- CN202210837834.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-16
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2042-07-16
AI Technical Summary
Existing image captioning datasets suffer from problems such as insufficient dataset relevance, limited descriptive perspectives, simplistic sentence structures, and simplistic backgrounds in crowd scene understanding research, making them unable to effectively address the challenges of complex crowd scenes.
A crowd scene image captioning description method based on multi-level attribute guidance is adopted. Through image feature extraction, visual feature embedding, multi-level dense crowd perception processing, feature fusion and dense crowd-oriented decoding steps, a more characteristic description of the crowd is generated.
It enables detailed descriptions of complex crowd scenes, generates more vivid and detailed captions, and improves the multimodal description capabilities of image caption datasets.
Smart Images

Figure CN115294353B_ABST