The application relates to a skeleton
action recognition method based on region
perception motion contrast learning, and relates to the field of
computer vision. The method solves the problems that existing contrast learning skeleton
action recognition methods focus on global feature contrast learning, ignore the importance of local action patterns for semantic discrimination, and a large amount of redundancy exists in the time and space dimensions of a skeleton sequence, and the method comprises the following steps: skeleton data preprocessing and
standardization; multi-
stream data conversion, generating three kinds of representations of joint flow,
motion flow and skeleton flow;
motion perception time sequence enhancement, adaptively retaining key frames based on interframe
motion intensity; spatial multi-
level data enhancement, implementing progressive three-
level data enhancement operation; spatiotemporal
feature extraction and region
perception mining, extracting features through an
encoder and unsupervisedly identifying significant motion regions by using a contrast
motion perception region mining method; contrast learning optimization, constructing a dynamic
negative sample queue to learn features; and multi-
stream model fusion, weightedly integrating recognition results of the data streams.