The application discloses a video
human behavior prediction method based on residual
diffusion theory and skeleton points, comprising the following steps: S10, acquiring an original group motion sequence, decomposing the multi-person sequence into individual trajectories, and then converting the time-domain motion sequence of each individual into frequency coefficients through
discrete cosine transform (DCT); S20, based on the
frequency domain representation, introducing a residual term to construct a forward
diffusion process; S30, defining a weighted
residual noise as a learning target, and realizing conversion from
noise to
residual noise; S40, determining an optimal step through transformation, realizing acceleration and unified training and reasoning; S50, applying physical constraints to single-person prediction, introducing a weighted combined biomechanical regularization
loss function to constrain the physical feasibility of the generated sequence; and S60, synthesizing a group sequence and outputting a predicted sequence. The application utilizes
deep learning technology, extracts
human skeleton point information, combines a residual
diffusion mechanism, predicts
human behavior in a
video sequence or a real-time scene, and realizes high-precision prediction of future behavior.