This invention provides a multimodal
crowd counting method based on multi-scale alignment and fusion. First, a multimodal crowd scene dataset is acquired and divided into training, validation, and test sets. Preprocessed RGB images and multimodal auxiliary images are input into a VGG16
backbone network to extract multi-stage high-level feature maps. These are then sequentially fed into a density-sharing local contrastive learning module, a local
feature fusion module, and an adaptive Mamba context-aware fusion module to achieve cross-
modal fine-grained alignment, local
feature fusion, and global context modeling. Subsequently, the global fused features are input into a dynamically upsampled multi-scale feature decoder to generate a high-resolution
crowd density map.
Supervised training is performed using a composite
loss function, and the optimal model is saved for testing. This invention adopts an "align-then-fuse" architecture, effectively mitigating cross-
modal heterogeneity and density fluctuation problems through multi-module
collaborative design, significantly improving the accuracy and robustness of
crowd counting in complex scenes.