The invention relates to the technical field of intelligent education, in particular to an interactive English teaching method and
system based on AI vision, and the method comprises the steps: obtaining the voice and mouth dynamic
image frame segments of a student, extracting an offset frame segment to generate a semantic motion disjunction interval, analyzing the
time sequence consistency of a behavior
signal and a motion cause word triggering frame segment, and calculating the motion matching degree. Time
synchronism of the
speech output and the visual action is evaluated. According to the invention, through real-time analysis of the voice and mouth dynamic image of the student, the alignment offset of the pronunciation and the mouth action is identified, and in combination with behavior signals such as eye fixation and
head rotation of the student, the interactivity between the language action and the behavior of the learner is comprehensively evaluated, so that the efficient alignment of the language expression and the standard template is ensured; by analyzing the distribution frequency and the matching degree of the voice behaviors, a personalized teaching path is optimized. In addition,
delay frames and contour offset are deeply analyzed, the
synchronism of voice and visual actions is enhanced, and the real-time feedback and interaction effect in the learning process is improved.