The invention belongs to the technical field of
image restoration and
computer vision, and particularly relates to a multi-mode face restoration and expression
recognition system and method based on semantic guidance of a facial action unit. According to the
system, multi-scale features are extracted through visual coding, and an AU activation probability is detected by using a graph neural network; the semantic conversion module converts the numerical probability into an interpretable biomechanical
structured text; the multi-
modal reasoning module fuses vision and text information, introduces the common sense reasoning ability of a multi-
modal large
language model, and improves the student
network performance through knowledge
distillation; and finally, the conditional generation module realizes
image restoration by taking the semantic features as guidance. The facial action unit is used as a biomechanical medium, the multi-
modal reasoning ability is converted into restoration constraint, the defects that in the prior art, restoration of
physiology is distorted, recognition depends on
image quality, and two tasks are isolated are overcome, collaborative enhancement of face restoration and expression recognition is achieved, and it is ensured that the restoration result is clear in vision and conforms to the physiological law of facial muscles.