Intelligent language learning method based on multi-modal fusion
By employing a multimodal fusion-based intelligent language learning method, the problem of insufficient cross-modal association in existing systems is solved, achieving effective alignment of speech, lip-sync video, and text data, thereby improving learning efficiency and personalized learning experience.
Patent Information
- Application Number
- CN202511253941.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-28
AI Technical Summary
Existing intelligent language learning systems have significant shortcomings in multimodal fusion and personalized learning. They fail to effectively establish cross-modal associations when processing speech, lip-sync video, and text data, resulting in large time alignment errors and affecting the learning experience.
This approach employs a multimodal fusion-based intelligent language learning method. By acquiring speech, lip-sync video, and text data, it utilizes short-time Fourier transform, 3D convolutional neural networks, and multi-layer text feature extraction. Combined with cross-modal attention mechanisms and conditional generation models, it generates personalized learning content and provides multimodal feedback.
It improves learning efficiency and memory retention by achieving effective alignment of speech features, text features, and visual features through multi-sensory memory encoding, providing a personalized learning experience and real-time feedback.
Smart Images

Figure CN121034288A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of intelligent language learning, and particularly relates to an intelligent language learning method based on multi-modal fusion. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, intelligent language learning systems have become an important research direction in the field of educational technology. Although the language learning solutions on the current market have been relatively mature in single-modal processing, there are still significant deficiencies in multi-modal fusion and personalized learning.
[0003] When processing speech, lip shape video and text data, existing systems usually adopt a simple feature splicing method and fail to establish effective cross-modal association. For example: the speech recognition module and the text analysis module run independently, resulting in pronunciation errors and spelling errors that cannot be associated and diagnosed; lip shape video is only used for auxiliary display and does not establish a quantitative association with speech spectrum features; the time alignment error between modalities generally exceeds 200 ms, affecting the learning experience. Lip feature extraction mostly uses 2D-CNN, which cannot effectively capture the timing dynamic features of pronunciation. Text feature modeling does not consider the particularity of the language learning scene.
[0004] In view of this, there is a need for an intelligent language learning method based on multi-modal fusion. SUMMARY
[0005] An intelligent language learning method based on multi-modal fusion, characterized by the following steps:
[0006] S1. Obtain the user's speech input signal, lip shape video data and text input data;
[0007] S2. Perform frame windowing processing on the speech input signal, calculate its short-time Fourier transform, extract mel frequency cepstral coefficient features through a mel filter bank, and obtain a speech feature vector;
[0008] S3. Perform face detection and position the lip region on the lip shape video data, and use a three-dimensional convolutional neural network to extract a lip shape motion feature vector;
[0009] S4. Input the text input data into a pre-trained language model to extract a text semantic feature vector;
[0010] S5. Input the speech feature vector, lip shape motion feature vector and text semantic feature vector into a multi-modal alignment module, calculate the alignment weight between the features of each modality through a cross-modal attention mechanism, and obtain an aligned multi-modal feature representation;
[0011] S6. Based on the aligned multimodal feature representation, combined with user language proficiency features and interest tags, personalized language learning content is dynamically generated through a conditional generation model;
[0012] S7. Generate multimodal feedback information including speech, vision, and text based on the user's response to the language learning content.
[0013] Preferably, the formula for calculating the short-time Fourier transform is:
[0014] Where STFT(k,t) represents the short-time Fourier transform, k represents the frequency index, t represents the time index, N represents the window length, s'(n) represents the windowed speech signal, and w(n) represents the window function.
[0015] Preferably, the formula for calculating the Mel frequency cepstral coefficients is:
[0016]
[0017] Where MFCC(m) represents the Mel frequency cepstral coefficients, m represents the Mel filter bank index, and H m (k) represents the m-th Mel filter, and K represents the number of frequency points.
[0018] Preferably, the text feature extraction adopts a three-level hierarchical architecture, achieves efficient computation through knowledge distillation, and employs a multi-granularity distillation strategy, including word-level distillation—KL divergence alignment of output distribution, sentence-level distillation—MSE matching of hidden states, and document-level distillation—graph topological similarity constraints.
[0019] Preferably, the calculation of the cross-modal attention mechanism includes:
[0020] Calculate query vector and key vector
[0021] Calculate attention weights:
[0022] Among them, Q i K represents the query vector. j Represents the key vector, v i and v j These represent feature vectors of different modalities. and Let d be a learnable parameter matrix, where d represents the feature dimension and α is the number of parameters. i-j This represents the attention weights from the i-th modality to the j-th modality.
[0023] Preferably, the cross-modal attention mechanism employs a multi-head attention mechanism, including:
[0024] Calculate the output of h attention heads:
[0025] The outputs of each attention point are concatenated and then subjected to a linear transformation:
[0026] MultiHead(Q,K,V)=concat(head1,...,head h W o ;
[0027] Among them, head i This represents the output of the i-th attention head. W represents the query, key, and value transformation matrices of the i-th attention head, respectively; O represents the output projection matrix; h represents the number of attention heads; Q represents the query matrix, which is composed of the query vectors of all modalities; K represents the key matrix, which is composed of the key vectors of all modalities; V represents the value matrix, which is composed of the feature vectors of all modalities.
[0028] Preferably, the conditional generation model employs a conditional variational autoencoder, and its objective function is:
[0029] Where, θ G θ represents the set of learnable parameters of the generator. E Let represent the set of learnable parameters of the encoder, φ represent the set of parameters of the inference network, x represent the multimodal feature vector of the input, z represent the latent representation of the latent variable space, β represent the hyperparameter that adjusts the weights of the KL divergence term, and D KL Let q represent the KL divergence. φ (z|x) represents the approximate posterior distribution defined by the encoder, p θ (x|z) represents the generator distribution defined by the decoder, p(z) represents the prior distribution, and E represents the mathematical expectation operator.
[0030] Preferably, the generation of the multimodal feedback information includes:
[0031] Calculate the difference between the user's response and the expected content: e = max(0, m + sim(r, c) - )-sim(r,c), where m represents the boundary value, c - sim(·) represents the content of negative samples, and sim(·) represents the similarity function.
[0032] Based on the difference e, corresponding voice, visual, and text feedback are generated.
[0033] Furthermore, to achieve the above objectives, the present invention also proposes an intelligent language learning system for implementing the method, characterized in that it includes:
[0034] A multimodal data acquisition device for acquiring user voice, video, and text input;
[0035] The feature extraction module is used to extract modal features from the input data;
[0036] The multimodal alignment module is used to achieve alignment between features of different modalities;
[0037] The content generation module is used to generate personalized language learning content;
[0038] The feedback generation module is used to generate multimodal feedback information;
[0039] An interactive interface is used to display learning content and feedback information.
[0040] Furthermore, to achieve the above objectives, the present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any step of the method.
[0041] Compared with the prior art, the advantages of this invention are:
[0042] A multimodal attention alignment mechanism is adopted to improve learning efficiency; multisensory memory encoding is used, with phonological features leading to auditory memory, textual features leading to semantic memory, and visual features leading to image memory, thereby improving memory retention rate. Attached Figure Description
[0043] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the description of the specific embodiments will be briefly introduced below.
[0044] Figure 1 This is a flowchart illustrating the steps of an intelligent language learning method based on multimodal fusion according to the present invention.
[0045] Figure 2 This is a flowchart of the three-level feature extraction pipeline constructed according to the present invention. Detailed Implementation
[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] Reference Figure 1 , Figure 1 This is a flowchart illustrating the steps of an intelligent language learning method based on multimodal fusion according to the present invention.
[0048] This invention constructs a complete closed-loop learning environment of "perception-analysis-generation-feedback". Its core technological breakthrough lies in achieving deep semantic alignment of multimodal data and dynamic content generation, solving the pain points of static content and singular feedback in traditional language learning. When the system is working, it first collects the user's raw learning data through multi-channel sensors, then performs feature extraction and cross-modal fusion to generate personalized content that matches the user's current cognitive level, and finally provides real-time feedback through multi-sensory channels, forming a continuously optimized learning closed loop.
[0049] This invention provides an intelligent language learning method based on multimodal fusion, comprising the following steps:
[0050] S1. Acquire the user's voice input signal, lip-sync video data, and text input data.
[0051] Audio was captured using a professional-grade USB microphone (Blue Yeti) at a 48kHz sampling rate and 24-bit depth. In the preprocessing stage, a pre-emphasis filter (H(z) = 1 - 0.97z) was first applied. -1 To compensate for high-frequency attenuation, automatic gain control (AGC) based on a noise threshold is performed, with the dynamic range controlled between -30 dBFS and -3 dBFS. To eliminate environmental noise, a real-time denoising algorithm based on an RNN is adopted, whose network structure consists of 3 layers of bidirectional LSTM (256 units per layer), achieving a 15 dB SNR improvement when trained on the DNS Challenge dataset.
[0052] Video streams were captured at 60fps using a Logitech Brio 4K camera. Face detection was performed using an improved RetinaFace model, which, after fine-tuning on the 300W-LP dataset, achieved a keypoint detection error of <3.5 pixels. For the lip region, an adaptive ROI cropping algorithm was employed: first, 68 facial feature points were located, and then a 160×160 pixel dynamic tracking window was established based on the geometric center of the lips, maintaining stable lip tracking even with head movement.
[0053] The system features a hybrid input channel that supports both traditional keyboard input and an integrated Conformer-Transformer-based ASR system. This ASR system was fine-tuned on the LibriSpeech dataset, reducing the word error rate (WER) to 5.3%. To handle multilingual input, an automatic switching processing model is employed with a Language Identification (LID) front-end, supporting real-time transcription of 12 languages, including Chinese, English, Japanese, and Korean.
[0054] S2. The speech input signal is processed by frame-by-frame windowing, and its short-time Fourier transform is calculated. Then, the Mel frequency cepstral coefficient features are extracted through the Mel filter bank to obtain the speech feature vector.
[0055] The formula for calculating the short-time Fourier transform is as follows:
[0056] Where STFT(k,t) represents the short-time Fourier transform, k represents the frequency index, t represents the time index, N represents the window length, s'(n) represents the windowed speech signal, and w(n) represents the window function.
[0057] An improved frequency cepstral coefficient (MFCC) extraction process is employed: first, a 512-point STFT is calculated, then 40 non-uniform Mel filters are applied, with their center frequencies distributed according to the following formula:
[0058] Among them, f mel f represents frequency on the Mel scale, a unit of perceived frequency; Hz The actual physical frequency is represented in Hertz; 2595 is the scaling constant that maps Hertz to the Mel scale; 700 is the offset that controls the nonlinearity of the curve.
[0059] The formula for calculating the Mel frequency cepstral coefficients is as follows:
[0060] Where MFCC(m) represents the Mel frequency cepstral coefficients, m represents the Mel filter bank index, and H m (k) represents the m-th Mel filter, and K represents the number of frequency points.
[0061] This formula simulates the nonlinear perception of frequency by the human ear. To enhance the representational ability of speech features, the following derived features are added to the traditional MFCC:
[0062] The first-order difference ΔMFCC characterizes the slope;
[0063] The second-order difference ΔΔMFCC characterizes the curvature;
[0064] PLP coefficients simulate the characteristics of human hearing.
[0065] RASTA filtering enhances robustness.
[0066] The final 78-dimensional hypervector is generated, which is then normalized by L2 and input into the downstream network.
[0067] The formulas for calculating the dynamic characteristics ΔMFCC and ΔΔMFCC are as follows:
[0068] Among them, c tdenoted as MFCC coefficients in frame t; N represents the difference window size, typically 2-3 frames; the second-order difference ΔΔMFCC is calculated twice.
[0069] S3. Perform face detection and localize the lip region on the lip shape video data, and use a three-dimensional convolutional neural network to extract lip shape motion feature vectors.
[0070] Lip movement analysis requires capturing both spatial morphology and temporal dynamics, such as the degree of lip opening and instantaneous changes during phoneme transitions. Traditional 2D CNNs cannot model temporal relationships, while pure 3D CNNs are computationally too expensive. This embodiment designs a dedicated 3D-ResNet18 architecture to process lip video sequences, and its innovations include:
[0071] Spatiotemporal separable convolution: decomposes 3D convolution into 2D spatial convolution + 1D temporal convolution.
[0072] The D convolution kernel is decomposed into: in, Processing spatial dimensions, In terms of processing the time dimension, ⊕ indicates tensor concatenation.
[0073] Residual connections: solve the vanishing gradient problem in deep networks.
[0074] Each basic module contains:
[0075]
[0076] During backpropagation, the gradient is transmitted through two paths:
[0077] Attention pooling: Automatically focusing on keyframes
[0078] The network was pre-trained on the LRW dataset and achieved an accuracy of 82.3%. After being transferred to this system, its performance was further improved through domain adaptation (DA) technology.
[0079] S4. Input the text input data into the pre-trained language model and extract the text semantic feature vector.
[0080] Traditional text processing models have two key limitations: first, single-granularity features are difficult to capture multi-level information of words, sentences, and paragraphs simultaneously; second, long text modeling suffers from context fragmentation, such as the resolution of references in dialogues.
[0081] like Figure 2 This is a flowchart illustrating the three-level feature extraction pipeline constructed in this invention.
[0082] This embodiment of text feature extraction adopts a three-level hierarchical architecture and achieves efficient computation through knowledge distillation. Its core technical details are as follows:
[0083] Hierarchical feature extraction:
[0084] Word-level processing employs improved BERT character embeddings and adds a character attention mechanism to the output layer:
[0085] Where, α i Let w be the attention weight for the i-th character, and h be the attention query vector. i Let be the character embedding vector, W be the learnable parameter matrix, tanh be the hyperbolic tangent activation function, and exp be the exponential function. This design significantly improves spell correction capabilities.
[0086] Sentence-level modeling, based on the Transformer-XL architecture, innovatively introduces:
[0087] Fragment loop mechanism: Among them, h τ h represents the hidden state of the current segment τ. τ-1 The hidden state of the previous segment τ-1, f is the word embedding matrix for the current segment, and f is the gated cyclic transformation function.
[0088] Dynamic masking strategy: Adjust the mask rate according to the CEFR level (15% for A1 level, 30% for C2 level).
[0089] Document-level analysis involves constructing a text semantic graph and then performing graph convolution operations. in, Adjacency matrix with self-connections Let H be the degree matrix, σ be the ReLU activation function, and H be the degree matrix. (l) Features of the l-th layer nodes, W (l) This is a trainable weight matrix. This technique improves the accuracy of logical reasoning.
[0090] Knowledge distillation optimization:
[0091] Employing a multi-particle size distillation strategy:
[0092] Word-level distillation: KL divergence aligned output distribution;
[0093] Sentence-level distillation: MSE matches hidden states;
[0094] Column-level distillation: Graph topological similarity constraints.
[0095] Combined with dynamic temperature control: Where T is the current temperature value, T max Let T be the initial temperature. min Let t be the final temperature, and t be the current training step number. total This represents the total number of training steps.
[0096] Ultimately, this achieves the beneficial effects of reducing the number of parameters, increasing inference speed, and maintaining high accuracy.
[0097] S5. Input the speech feature vector, lip movement feature vector and text semantic feature vector into the multimodal alignment module, calculate the alignment weights between each modality feature through the cross-modal attention mechanism, and obtain the aligned multimodal feature representation.
[0098] The calculation of the cross-modal attention mechanism includes:
[0099] Calculate query vector and key vector
[0100] Calculate attention weights:
[0101] Among them, Q i K represents the query vector. j Represents the key vector, v i and v j These represent feature vectors of different modalities. and Let d be a learnable parameter matrix, where d represents the feature dimension and α is the number of parameters. i-j This represents the attention weights from the i-th modality to the j-th modality.
[0102] The cross-modal attention mechanism described in this embodiment employs an improved multi-head attention formula: MultiHead(Q,K,V)=concat(head1,...,head) h W o ;
[0103] Among them, head i This represents the output of the i-th attention head; W represents the query, key, and value transformation matrices of the i-th attention head, respectively; O represents the output projection matrix; h represents the number of attention heads; Q represents the query matrix, which is composed of the query vectors of all modalities; K represents the key matrix, which is composed of the key vectors of all modalities; V represents the value matrix, which is composed of the feature vectors of all modalities.
[0104] The calculation for each attention head is as follows:
[0105] Thus, relative position encoding solves the sequence length extrapolation problem, sparse attention, and reduces O(n) time complexity. 2 ) Computational complexity.
[0106] The design incorporates both intermodal and intramodal attention mechanisms. Intermodal attention establishes connections between speech, text, and vision, while self-attention strengthens the organization of features within each modality.
[0107] Dynamically adjust attention weights through a gating mechanism: g = σ(W) g [v i ;v j ]+b g ), where σ is the sigmoid function, [;] denotes the concatenation operation; g is the gate vector; W g For trainable weight matrix; v i Feature vector of modality i, speech features; v j Feature vector of modality j, text features; b g Bias term; d is the feature dimension, default 512.
[0108] S6. Based on the aligned multimodal feature representation, combined with the user's language proficiency features and interest tags, personalized language learning content is dynamically generated through a conditional generation model.
[0109] The conditional generation model described in this embodiment is an improved conditional variational autoencoder, whose objective function is: L(θ) G ,θ E ) = E qφ(z|x) [logp θ (x|z)]-βD KL (q φ (z|x)||p(z)).
[0110] Where, θ G θ represents the set of learnable parameters of the generator. E Let represent the set of learnable parameters of the encoder, φ represent the set of parameters of the inference network, x represent the multimodal feature vector of the input, z represent the latent representation of the latent variable space, β represent the hyperparameter that adjusts the weights of the KL divergence term, and D KL Let q represent the KL divergence. φ (z|x) represents the approximate posterior distribution defined by the encoder, p θ (x|z) represents the generator distribution defined by the decoder, p(z) represents the prior distribution, and E represents the mathematical expectation operator.
[0111] Design layered potential space:
[0112] Each layer adopts a conditional normal distribution: z l ~N(μ) l (z <l ),σ l (z <l )).
[0113] Where L is the total number of potential spatial levels; z l z is the latent variable of the l-th level; <l For the set of all predecessor layer variables; μ l The l-th layer mean network; σ l The l-th layer variance network; N represents the normal distribution.
[0114] S7. Generate multimodal feedback information including speech, vision, and text based on the user's response to the language learning content.
[0115] The generation of the feedback information content is divided into three stages:
[0116] Conceptual planning: Generating a content outline based on latent variable z;
[0117] Detailed Filling: Adjust vocabulary complexity based on user skill level;
[0118] Style adaptation: Matching the expression style to the user's preferred style.
[0119] The course-based learning strategy is adopted, and the difficulty is dynamically adjusted as the user progresses.
Claims
1. An intelligent language learning method based on multimodal fusion, characterized in that, Includes the following steps: S1. Acquire the user's voice input signal, lip-sync video data, and text input data; S2. The speech input signal is processed by frame-by-frame windowing, and its short-time Fourier transform is calculated. Then, the Mel frequency cepstral coefficient features are extracted through the Mel filter bank to obtain the speech feature vector. S3. Perform face detection and localize the lip region on the lip shape video data, and use a three-dimensional convolutional neural network to extract lip shape motion feature vectors; S4. Input the text input data into a pre-trained language model and extract the text semantic feature vector; S5. Input the speech feature vector, lip movement feature vector and text semantic feature vector into the multimodal alignment module, calculate the alignment weight between each modality feature through the cross-modal attention mechanism, and obtain the aligned multimodal feature representation; S6. Based on the aligned multimodal feature representation, combined with user language proficiency features and interest tags, personalized language learning content is dynamically generated through a conditional generation model; S7. Generate multimodal feedback information including speech, vision, and text based on the user's response to the language learning content.
2. The method according to claim 1, characterized in that, The formula for calculating the short-time Fourier transform is as follows: Where STFT(k,t) represents the short-time Fourier transform, k represents the frequency index, t represents the time index, N represents the window length, s'(n) represents the windowed speech signal, and w(n) represents the window function.
3. The method according to claim 1, characterized in that, The formula for calculating the Mel frequency cepstral coefficients is as follows: Where MFCC(m) represents the Mel frequency cepstral coefficients, m represents the Mel filter bank index, and H m (k) represents the m-th Mel filter, and K represents the number of frequency points.
4. The method according to claim 1, characterized in that, The text feature extraction adopts a three-level hierarchical architecture, achieves efficient computation through knowledge distillation, and employs a multi-granularity distillation strategy, including word-level distillation—KL divergence alignment of output distribution, sentence-level distillation—MSE matching of hidden states, and document-level distillation—graph topological similarity constraints.
5. The method according to claim 1, characterized in that, The calculation of the cross-modal attention mechanism includes: Calculate query vector and key vector Calculate attention weights: Among them, Q i K represents the query vector. j Represents the key vector, v i and v j These represent feature vectors of different modalities. and Let d be a learnable parameter matrix, where d represents the feature dimension and α is the number of parameters. i-j This represents the attention weights from the i-th modality to the j-th modality.
6. The method according to claim 4, characterized in that, The cross-modal attention mechanism employs a multi-head attention mechanism, including: Calculate the output of h attention heads: The outputs of each attention point are concatenated and then subjected to a linear transformation: MultiHead(Q,K,V)=concat(head1,...,head h )W o ; Among them, head i This represents the output of the i-th attention head. W represents the query, key, and value transformation matrices of the i-th attention head, respectively; O represents the output projection matrix; h represents the number of attention heads; Q represents the query matrix, which is composed of the query vectors of all modalities; K represents the key matrix, which is composed of the key vectors of all modalities; V represents the value matrix, which is composed of the feature vectors of all modalities.
7. The method according to claim 1, characterized in that, The conditional generation model employs a conditional variational autoencoder, and its objective function is: Where, θ G θ represents the set of learnable parameters of the generator. E Let represent the set of learnable parameters of the encoder, φ represent the set of parameters of the inference network, x represent the multimodal feature vector of the input, z represent the latent representation of the latent variable space, β represent the hyperparameter that adjusts the weights of the KL divergence term, and D KL Let q represent the KL divergence. φ (z|x) represents the approximate posterior distribution defined by the encoder, p θ (x|z) represents the generator distribution defined by the decoder, p(z) represents the prior distribution, and E represents the mathematical expectation operator.
8. The method according to claim 1, characterized in that, The generation of the multimodal feedback information includes: Calculate the difference between the user's response and the expected content: e = max(0, m + sim(r, c) - )-sim(r,c); Where m represents the boundary value, c - sim(·) represents the content of negative samples, and sim(·) represents the similarity function. Based on the difference e, corresponding voice, visual, and text feedback are generated.
9. An intelligent language learning system implementing the method of any one of claims 1-7, characterized in that, include: A multimodal data acquisition device for acquiring user voice, video, and text input; The feature extraction module is used to extract modal features from the input data; The multimodal alignment module is used to achieve alignment between features of different modalities; The content generation module is used to generate personalized language learning content; The feedback generation module is used to generate multimodal feedback information; An interactive interface is used to display learning content and feedback information.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Cited By
Heterogeneous modal sequence-oriented alignment method, device and system and storage medium
CN121397286A