The weakly supervised temporal action localization method and device based on semantic and saliency knowledge collaborative propagation comprises the following steps: 1) extracting temporal features and saliency foreground features from an uncropped video; 2) constructing a basic
branch and a saliency
perception branch to respectively process the temporal features and the saliency target features to obtain a
basic class activation sequence, a motion, appearance representation
score and a saliency class activation sequence, and weighting and fusing the four sequences to obtain a fused action
score sequence; 3) utilizing
branch distillation and branch action consistency constraint to interact
semantic information and saliency information, and perfecting the fused action
score sequence; 4) extracting key segments and ambiguous segments of the basic branch and the saliency
perception branch, and utilizing the key segments and the ambiguous segments between branches and within branches for comparative learning to improve feature representation, combining the
distillation result to perfect the fused action score sequence and obtaining an action localization result. The present application can perceive subtle human actions and accurate temporal action boundaries in uncropped videos.