This invention discloses a semi-automatic
annotation method and
system for student
facial expression data based on multimodal fusion, comprising: S1 generating
multidimensional expressions based on a semantic conditional
diffusion model, and simultaneously outputting discrete expression classifications and continuous VAD emotion dimension labels through text prompts in an educational setting; S2 obtaining synthetic expression images and their corresponding 68 facial key point coordinates based on S1, performing classroom scene-adaptive style transfer, and using a key point-constrained adversarial generative network to preserve expression features and adapt to classroom lighting and viewing angle; S3 identifying the classification probability of discrete expression categories and the three-dimensional predicted value of the continuous VAD emotion dimension, using
hybrid uncertainty-driven active learning, and combining classification entropy and regression variance to select high-value samples; S4 performing identity decoupling and privacy anonymization on the selected high-value samples, reconstructing the face through a 3D deformation model and specifically blurring the
eyebrow region to achieve
anonymity protection. This invention can ensure privacy compliance, reduce
annotation costs, and improve model accuracy and robustness.