The application discloses a
gaze spectrum
estimation method based on a first-view high dynamic long video and belongs to the field of
image processing. The
gaze spectrum prediction method of the application effectively solves the key problems existing in the prior art, such as
time sequence information loss in a long video, insufficient fusion of local and
global information, and difficulty of a static
mask in adapting to dynamic scene changes. An enhanced long-time memory
encoder effectively encodes a long video through multi-scale attention and hierarchical memory mechanism, solves the problems of
information loss and redundancy under long-time dependence, and ensures the integrity of
time sequence information. A high-pass global-local
information aggregation module combines a global background and local details in a multi-layer network through the design of a cross-layer dynamic
information transmission channel, and enhances the
gaze spectrum prediction capability in a complex dynamic scene. A dynamic
mask fusion module adopts an adaptive mechanism, can adjust a
mask weight in real time, solves the problem that a static attention cannot cope with a rapidly changing background, and improves the flexibility and accuracy of the model.