The invention relates to the field of voice
signal processing, in particular to a multi-language
interactive learning system based on voice recognition, which is characterized in that sound
waves and
mouth shape images are synchronously acquired and discretized by the
system, and cross-
modal coding is formed after alignment; secondly, the code is injected into a micro-ring
photon reserve network through
phase modulation to unfold
time sequence characteristics, the code is mapped into a
quaternion graph to be embedded, and segment boundaries are extracted through a
diffusion-pulse
coupling method; then, according to a
fixed field sequence, encapsulating the
quaternion graph embedding and segmentation data into an object mark prompt, inputting the object mark prompt into a low-rank adaptive
language model, and generating semantic segmentation data and text transcription; and finally, the
adaptive learning module implements sparse gradient updating on the low-rank weight and pulse network by using an integer fractal hash exclusive-or difference
mask, and adopts support-query element learning for synchronous iteration after user clarification. According to the
system, low-power-consumption, high-precision and second-level accent self-adaptive multi-language voice interaction is realized on the end side.