A method for city perception analysis combining street view images and eye movement data
By combining street view images with eye-tracking data and utilizing ViT, Mask2former, and CNN-LSTM models, an urban perception model is constructed, which solves the problem of insufficient perception prediction accuracy in existing technologies and achieves more accurate urban perception analysis.
CN122116129APending Publication Date: 2026-05-29BEIJING TECH & BUSINESS UNIV
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING TECH & BUSINESS UNIV
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-29
AI Technical Summary
Technical Problem
Existing urban perception research mainly relies on street view images, which cannot fully reflect human visual perception, resulting in insufficient accuracy in perception prediction.
Method used
By combining street view images and eye-tracking data, we use ViT, Mask2former, and CNN-LSTM models to extract image and eye-tracking features, construct an urban perception model, and predict perception scores using the RF model.
Benefits of technology
It improves the accuracy of urban perception prediction, reveals the relationship between the attention of visual elements and perception dimensions, and draws urban perception maps to intuitively present differences in perception distribution.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN122116129A_ABST
Abstract
The present application relates to the technical field of urban perception analysis, and particularly relates to a kind of urban perception analysis method combined with street view image and eye movement data, wherein the method comprises: obtaining street view image in Beijing six-ring area based on Baidu map to carry out eye movement experiment, and creates urban perception score dataset;A new urban perception model based on visual semantics is proposed, a deep learning model based on Transformer architecture is used as a semantic segmentation model, VIT is used as a classification model, LSMT+CNN is used as an eye movement feature extraction model, the category probability, semantics and eye movement features of street view image are extracted respectively, and the three feature results are fused into a random forest model, and finally the perception score is predicted.
Need to check novelty before this filing date? Find Prior Art