The invention provides a dynamic audio-
visual navigation method based on intensity calibration geometric coding and cross-
modal scene aggregation. The dynamic audio-
visual navigation method is suitable for application scenes such as service robots, security inspection and
virtual reality interaction. According to the method, a depth visual image and a binaural
spectrogram serve as input, firstly, a
sound intensity map and depth geometric information are explicitly fused through an intensity calibration geometric
encoder, and physical-space correlation across sound categories is modeled; then, a cross-
modal scene aggregator is utilized to convert the multiple groups of spatial features into a Token sequence, fine-grained and content-aware multi-
modal dynamic fusion is realized through Transform, and global
state representation rich in reasoning ability is generated; and finally, training a navigation strategy of
time sequence perception under a near-end strategy optimization framework based on the representation, and outputting a tracking action sequence aiming at the mobile sound source. According to the method, the navigation success rate and the path efficiency are remarkably improved under the conditions of a mobile sound source and unknown sound, and the generalization ability and the decision robustness of the
system in a dynamic environment are enhanced.