This invention provides a method and
system for identifying bipedal foot-to-
ground contact states based on
monocular video, relating to the field of
data processing technology. The method includes: acquiring a two-dimensional keypoint sequence of human posture from
monocular video; performing robust
outlier suppression and
smoothing differentiation on the two-dimensional keypoint sequence to obtain a smoothed keypoint coordinate and velocity sequence; extracting multiple kinematic features related to the bipedal foot-to-
ground contact state; determining the estimated direction of human movement through a viewpoint
adaptation mechanism, and projecting the foot velocity based on the direction of human movement to obtain viewpoint-adaptive velocity features; modeling the bipedal foot-to-
ground contact state using an unsupervised two-component
Gaussian mixture model and calculating the
posterior probability that the feet are in the support phase; performing joint temporal decoding of the bipedal foot-to-ground contact state using a four-state coupled
hidden Markov model and a Viterbi decoding
algorithm to obtain a temporally consistent bipedal foot-to-ground contact
state sequence; and completing the identification of the bipedal foot-to-ground contact state by extracting ground contact events and
ground lift events.