This invention discloses a dynamically densely connected spatiotemporal feature decoupling network for recognizing cross-view
gait, relating to the field of
computer vision. It includes: an initial feature
processing module, a dynamically dense spatiotemporal decoupling
feature extraction module, and a feature enhancement
processing module. The module uses dense spatiotemporal feature decoupling blocks and
concatenation operations to achieve the sharing of shallow and deep network features, thereby addressing the problem of insufficient representation ability. Simultaneously, it employs an enhanced convolutional block attention mechanism to allow the network to focus on more important
gait features. Finally, the feature enhancement
processing module processes the five-dimensional
feature mapping into multiple lateral features and performs batch
standardization of the features to enhance representation ability and model generalization ability. This allows for the mining of correlations between shallow and deep features, alleviating the problem of insufficient
information representation ability.