This invention discloses a text-driven
human motion editing method based on comprehensive positive and negative
supervised learning, belonging to the fields of
artificial intelligence and
computer vision. The method comprises two stages: model training and application. In the training stage, the original motion, target motion, and positive and negative text instructions are acquired. After
feature extraction and fusion, these are input into a
diffusion Transformer model for denoising and prediction. The core lies in simultaneously calculating three losses: retrospective feature supervision,
motion preservation, and triple
semantic alignment, and jointly optimizing them with the model's main loss to obtain a comprehensively supervised motion editing model. In the application stage, user instructions and the motion to be edited are input into the model, which outputs a target motion that accurately responds to
semantics while highly maintaining the coherence of non-edited regions. This invention significantly improves editing accuracy, preservation capability, and generation stability through a multi-layered, positive-negative combined comprehensive supervision mechanism.