Korean pronunciation teaching auxiliary system based on AI speech recognition
Through the Korean pronunciation teaching assistance system based on AI speech recognition, the enhanced XLSR model and Mel filter group are used, combined with deep learning technology, the problem that traditional systems cannot provide real-time feedback and detailed analysis is solved, and the pronunciation improvement and learning effect of Korean learners are achieved.
Patent Information
- Application Number
- CN202510175959.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
AI Technical Summary
The traditional Korean pronunciation teaching auxiliary system cannot provide real-time feedback, and insufficient analysis of pronunciation details, resulting in difficulty in the process of correcting pronunciation and affecting the learning effect.
Using a Korean pronunciation teaching assistance system based on AI speech recognition, the introduction of an enhanced XLSR model and Mel filter group, combined with deep learning technology, real-time and accurate pronunciation scores and detailed pronunciation analysis are achieved.
It significantly improves the pronunciation accuracy and fluency of Korean language learners. Through real-time feedback and detailed analysis, learners can effectively improve pronunciation skills and improve learning results.
Smart Images

Figure CN119993211A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular to a Korean pronunciation teaching auxiliary system based on AI speech recognition. Background Art
[0002] With the continuous advancement of the globalization process, Korean, as an important foreign language, has received more and more attention and attention from learners; traditional Korean learning methods usually rely on teacher guidance and classroom teaching, however, Korean pronunciation rules are complex and varied, and pronunciation accuracy and fluency are the most challenging parts of the learning process; traditional Korean pronunciation teaching auxiliary systems have the following shortcomings: first, traditional Korean pronunciation teaching auxiliary systems usually rely on pre-recorded audio feedback to evaluate learners' pronunciation; the biggest problem of this type of system is that it cannot provide real-time feedback, and the detailed analysis of pronunciation is insufficient. The lack of dynamic interaction and timely feedback teaching mode often makes learners feel difficult in the process of pronunciation correction, which in turn affects the learning effect; secondly, the traditional system cannot deeply identify the subtle differences between syllables, especially in the complex Korean pronunciation structure, the system's judgment of high-order syllables often produces errors, resulting in inaccurate scoring results, which in turn affects learners' pronunciation improvement; therefore, there is an urgent need for a new Korean pronunciation teaching auxiliary system that can overcome the shortcomings of the traditional system, improve the accuracy of pronunciation scoring through more sophisticated pronunciation feature analysis and more efficient real-time feedback mechanism, and effectively support learners in improving their pronunciation skills. Summary of the invention
[0003] The present invention provides a Korean pronunciation teaching auxiliary system based on AI speech recognition, aiming to solve the problems of untimely pronunciation feedback and inaccurate analysis in traditional Korean pronunciation teaching, and realizes real-time and accurate pronunciation scoring and detailed pronunciation analysis by introducing advanced speech recognition and deep learning technologies, thereby improving the pronunciation accuracy and fluency of Korean learners; the system adopts an enhanced XLSR model to extract pronunciation features, and optimizes the feature data of audio signals by introducing Mel filter groups and diffusion processes, captures key frequency features and timing information in learners' pronunciation, and generates effective linear projection data; combined with a bidirectional Mamba multi-order timing model, the system conducts in-depth analysis of the extracted features, evaluates learners' syllable, tone and rhythm pronunciation features, generates pronunciation scores, and a detailed pronunciation analysis report, thereby formulating personalized learning paths and improvement suggestions for learners, and can significantly improve the pronunciation level and learning effect of Korean learners.
[0004] The present invention provides a Korean pronunciation teaching auxiliary system based on AI speech recognition, which includes an audio input and preprocessing module, a feature extraction module, a pronunciation scoring and analysis module, and a feedback module;
[0005] The audio input and preprocessing module uses a microphone to monitor and capture the Korean audio signal input by the student in real time to obtain the original Korean audio signal, and performs noise suppression, signal enhancement and removal of irrelevant background noise on the original Korean audio signal to ensure the audio quality and obtain the preprocessed Korean audio signal;
[0006] Feature extraction module, establish XLSR model, introduce Mel filter and diffusion process and inverse diffusion process to improve XLSR model, build enhanced XLSR model, pre-process Korean audio signal through enhanced XLSR model, generate linear projection feature data;
[0007] The pronunciation scoring and analysis module combines the self-attention mechanism, high-order nonlinear feature transformation and state space modeling to build a bidirectional Mamba multi-order time series model. The bidirectional Mamba multi-order time series model processes the linear projection feature data to generate K-PronScore feature data. Based on the K-PronScore feature data, the students' Korean pronunciation is scored, and the syllables, tones and rhythms of the pronunciation are analyzed to generate scoring results and analysis reports. The bidirectional Mamba multi-order time series model includes a forward Mamba layer, a reverse Mamba layer and a linear layer.
[0008] The feedback module develops personalized learning paths and suggestions based on the scoring results and analysis reports.
[0009] Furthermore, the feature extraction module generates linear projection feature data, which specifically includes the following steps:
[0010] Step S1: performing short-time Fourier transform on the preprocessed Korean audio signal to convert the time domain signal into a frequency domain representation; dividing the preprocessed Korean audio signal into a number of small time windows, calculating the spectrum of each time window to capture the frequency characteristics of the audio signal, and generating a frequency characteristic graph;
[0011] Step S2: Processing the frequency feature map through a Mel filter bank, converting the frequency feature map into a Mel frequency scale to simulate the perceptual characteristics of the human auditory system; the conversion enhances the frequency resolution of the low-frequency region in the frequency feature map, and downsamples the high-frequency region of the frequency feature map to obtain a Mel spectrum;
[0012] Step S3: Perform discrete cosine transform on the Mel spectrum, compress the redundant information of the Mel spectrum, and extract the Mel spectrum cepstrum coefficients to generate MFCC feature data; introduce a diffusion process to perturb the MFCC feature data by adding noise to generate perturbed MFCC feature data, and then restore the perturbed MFCC feature data through an inverse diffusion process to remove irrelevant noise information and generate optimized MFCC feature data. The formula used is as follows:
[0013] ;
[0014] in, Represents the MFCC dimension index, Indicates the Mel spectrum index, represents the total dimension of the Mel spectrum, Indicates Dimensional Mel frequency cepstrum coefficients; Indicates The energy of the Mel band, Indicates the Mel spectrum value after taking the logarithm, represents the cosine transform term;
[0015] ;
[0016] in, represents the time step index, Indicates that at time step The perturbed MFCC feature data under Represents the previous time step The MFCC feature data under represents the diffuse noise parameter, represents the normalization coefficient, represents standard Gaussian noise, represents a multivariate Gaussian distribution with mean 0 and unit covariance matrix;
[0017] ;
[0018] in, represents the trainable parameters, Indicated in the parameter Under the given disturbance Estimated original The probability distribution of represents the mean, represents the variance, Indicates the Gaussian distribution that the predicted MFCC obeys;
[0019] Step S4: further processing the optimized MFCC feature data through the XLSR model to generate a rich audio feature representation;
[0020] Step S5: linearly project the rich audio feature representation, standardize the rich audio feature representation, and obtain linear projection feature data.
[0021] Further, step S4 specifically includes the following steps:
[0022] Step S41: performing preliminary feature extraction on the optimized MFCC feature data through a convolutional neural network to generate convolutional MFCC feature data;
[0023] Step S42: Use the Transformer encoder to perform deeper temporal dependency modeling on the convolutional MFCC feature data, capture global context information, and generate rich audio feature representation.
[0024] Furthermore, the pronunciation scoring and analysis module generates K-PronScore feature data, which specifically includes the following steps:
[0025] Step C1: Forward feature processing: The linear projection feature data is processed layer by layer in time order through the forward Mamba layer, the forward features are analyzed, the forward time dependency is captured, and the forward Mamba feature data is generated;
[0026] Step C2: Backward feature processing: Arrange the linear projection feature data in reverse order to obtain reverse linear projection feature data, process the reverse linear projection feature data through the reverse Mamba layer, analyze the backward features, capture the backward time dependency, enhance the understanding of the global context, and generate reverse Mamba feature data;
[0027] Step C3: feature splicing: splicing the forward Mamba feature data and the reverse Mamba feature data to generate spliced Mamba feature data;
[0028] Step C4: The concatenated Mamba feature data is input into the linear layer and further processed by linear transformation to obtain K-PronScore feature data.
[0029] Furthermore, step C1 specifically includes the following steps:
[0030] Step C11: local feature extraction: the linear projection feature data is processed by one-dimensional convolution to extract local time features and generate forward convolution feature data;
[0031] Step C12: Capture global information: The forward convolution feature data is used to capture global context information through the self-attention mechanism to generate attention feature data;
[0032] Step C13: Perform nonlinear activation on the attention feature data to generate nonlinear feature data, introduce high-order polynomial feature transformation to capture the complex temporal relationship of the nonlinear feature data; introduce cross terms to perform power combination on the nonlinear feature data to capture the nonlinear interaction between the nonlinear feature data and generate optimized nonlinear feature data. The formula used is as follows:
[0033] ;
[0034] in, represents power, Indicates the maximum order of polynomial transformation, that is, the maximum power; represents the cross-term feature data, represents the weight matrix of the cross terms, represents the attention feature data, Express conduct Power transformation, Express conduct Power transformation;
[0035] ;
[0036] in, represents the optimization of nonlinear characteristic data, represents the parameters of the activation function, represents a learnable activation function, represents the bias term, express The weight of represents the polynomial transformation weight, represents a polynomial feature transformation;
[0037] Step C14: Modeling long-term dependencies: Establishing a state space model, passing the optimized nonlinear feature data to the state space model, updating the current state and capturing long-term dependencies, and generating forward state space feature data;
[0038] Step C15: Feature fusion: The forward state space feature data is fused with the optimized nonlinear feature data through residual connection to generate forward Mamba feature data.
[0039] Furthermore, step C2 specifically includes the following steps:
[0040] Step C21: Process the reverse linear projection feature data through one-dimensional convolution to extract local time features and generate reverse convolution feature data;
[0041] Step C22: pass the reverse convolution feature data to the state space model, update the current state and capture the long-term temporal dependency, and generate reverse state space feature data as reverse Mamba feature data.
[0042] By adopting the above scheme, the beneficial effects achieved by the present invention are as follows:
[0043] The present invention provides a Korean pronunciation teaching auxiliary system based on AI speech recognition. By introducing an enhanced XLSR model and combining the Mel filter and the diffusion process and the inverse diffusion process, the extraction process of Korean pronunciation features is effectively optimized; the Mel filter group simulates the perceptual characteristics of the human auditory system, and can more accurately capture the frequency characteristics of the learner's pronunciation, especially in the resolution processing of the low-frequency and high-frequency regions, thereby improving the adaptability of the system to complex audio signals; at the same time, the introduction of the diffusion process and the inverse diffusion process removes noise and irrelevant background information by disturbing the feature data and then restoring it, making the feature data cleaner and more accurate; this technology significantly improves the system's ability to capture the learner's pronunciation features, provides more reliable input data for subsequent pronunciation scoring and analysis, and ultimately improves the accuracy of pronunciation scoring and the system's sensitivity to the learner's pronunciation quality;
[0044] In addition, the present invention introduces a self-attention mechanism, high-order nonlinear feature transformation and state space modeling in the pronunciation scoring and analysis module, which effectively solves the problem that the traditional pronunciation scoring system cannot accurately capture pronunciation details and the analysis is insufficient; the self-attention mechanism can strengthen the focus on important information in the audio and capture global context information, so that the system can fully understand the pronunciation characteristics of the learner; the high-order nonlinear feature transformation captures the complex time series relationship in the pronunciation data by introducing high-order polynomial feature transformation, and improves the system's adaptability to complex pronunciation patterns; at the same time, the introduction of cross terms enhances the nonlinear interaction between features in the system, further improving the model's ability to analyze and score pronunciation features in detail;
[0045] Through the combination of the above-mentioned technical means, the system can capture the detailed changes in the learners' pronunciation in real time, and adjust the learning path in time according to the scoring results, helping learners to correct pronunciation errors more efficiently; the above-mentioned technical solution not only optimizes the pronunciation teaching process, but also significantly improves the learners' learning effects, enabling them to achieve more efficient pronunciation improvement in a short period of time. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram of a module of a Korean pronunciation teaching auxiliary system based on AI speech recognition proposed by the present invention;
[0047] Figure 2 The thermal map obtained by processing the linear projection feature data in Example 6;
[0048] Figure 3 This is the thermal map obtained by processing the linear projection feature data in Example 7. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0050] Embodiment 1, according to Figure 1 , the present invention provides a Korean pronunciation teaching auxiliary system based on AI speech recognition, the system includes an audio input and preprocessing module, a feature extraction module, a pronunciation scoring and analysis module and a feedback module;
[0051] The audio input and preprocessing module uses a microphone to monitor and capture the Korean audio signal input by the student in real time to obtain the original Korean audio signal, and performs noise suppression, signal enhancement and removal of irrelevant background noise on the original Korean audio signal to ensure the audio quality and obtain the preprocessed Korean audio signal;
[0052] Feature extraction module, establish XLSR model, introduce Mel filter and diffusion process and inverse diffusion process to improve XLSR model, build enhanced XLSR model, pre-process Korean audio signal through enhanced XLSR model, generate linear projection feature data;
[0053] The pronunciation scoring and analysis module combines the self-attention mechanism, high-order nonlinear feature transformation and state space modeling to build a bidirectional Mamba multi-order time series model. The bidirectional Mamba multi-order time series model processes the linear projection feature data to generate K-PronScore feature data. Based on the K-PronScore feature data, the students' Korean pronunciation is scored, and the syllables, tones and rhythms of the pronunciation are analyzed to generate scoring results and analysis reports. The bidirectional Mamba multi-order time series model includes a forward Mamba layer, a reverse Mamba layer and a linear layer.
[0054] The feedback module develops personalized learning paths and suggestions based on the scoring results and analysis reports.
[0055] Embodiment 2: This embodiment is based on embodiment 1. In this embodiment, the process of generating linear projection feature data by the feature extraction module specifically includes the following steps:
[0056] Step S1: performing short-time Fourier transform on the preprocessed Korean audio signal to convert the time domain signal into a frequency domain representation; dividing the preprocessed Korean audio signal into a number of small time windows, calculating the spectrum of each time window to capture the frequency characteristics of the audio signal, and generating a frequency characteristic graph;
[0057] Step S2: Processing the frequency feature map through a Mel filter bank, converting the frequency feature map into a Mel frequency scale to simulate the perceptual characteristics of the human auditory system; the conversion enhances the frequency resolution of the low-frequency region in the frequency feature map, and downsamples the high-frequency region of the frequency feature map to obtain a Mel spectrum;
[0058] Step S3: Perform discrete cosine transform on the Mel spectrum, compress the redundant information of the Mel spectrum, and extract the Mel spectrum cepstrum coefficients to generate MFCC feature data; introduce a diffusion process to perturb the MFCC feature data by adding noise to generate perturbed MFCC feature data, and then restore the perturbed MFCC feature data through an inverse diffusion process to remove irrelevant noise information and generate optimized MFCC feature data. The formula used is as follows:
[0059] ;
[0060] in, Represents the MFCC dimension index, Indicates the Mel spectrum index, represents the total dimension of the Mel spectrum, Indicates Dimensional Mel frequency cepstrum coefficients; Indicates The energy of the Mel band, Indicates the Mel spectrum value after taking the logarithm, represents the cosine transform term;
[0061] ;
[0062] in, represents the time step index, Indicates that at time step The perturbed MFCC feature data under Represents the previous time step The MFCC feature data under represents the diffuse noise parameter, represents the normalization coefficient, represents standard Gaussian noise, represents a multivariate Gaussian distribution with mean 0 and unit covariance matrix;
[0063] ;
[0064] in, represents the trainable parameters, Indicated in the parameter Under the given disturbance Estimated original The probability distribution of represents the mean, represents the variance, Indicates the Gaussian distribution that the predicted MFCC obeys;
[0065] Step S4: further processing the optimized MFCC feature data through the XLSR model to generate a rich audio feature representation;
[0066] Step S5: linearly project the rich audio feature representation, standardize the rich audio feature representation, and obtain linear projection feature data.
[0067] Embodiment 3: This embodiment is based on embodiment 1. In this embodiment, the process of generating linear projection feature data by the feature extraction module specifically includes the following steps:
[0068] Step E1: performing short-time Fourier transform on the preprocessed Korean audio signal to convert the time domain signal into a frequency domain representation; dividing the preprocessed Korean audio signal into a number of small time windows, calculating the spectrum of each time window to capture the frequency characteristics of the audio signal, and generating a frequency characteristic graph;
[0069] Step E2: further process the frequency feature map through the XLSR model to generate a rich audio feature representation;
[0070] Step E3: linearly project the rich audio feature representation, standardize the rich audio feature representation, and obtain linear projection feature data.
[0071] Embodiment 4: This embodiment is based on embodiment 2. In this embodiment, step S4 specifically includes the following steps:
[0072] Step S41: performing preliminary feature extraction on the optimized MFCC feature data through a convolutional neural network to generate convolutional MFCC feature data;
[0073] Step S42: Use the Transformer encoder to perform deeper temporal dependency modeling on the convolutional MFCC feature data, capture global context information, and generate rich audio feature representation.
[0074] Embodiment 5, this embodiment is based on embodiment 4. In this embodiment, the pronunciation scoring and analysis module generates the process of K-PronScore feature data, which specifically includes the following steps:
[0075] Step C1: Forward feature processing: The linear projection feature data is processed layer by layer in time order through the forward Mamba layer, the forward features are analyzed, the forward time dependency is captured, and the forward Mamba feature data is generated;
[0076] Step C2: Backward feature processing: Arrange the linear projection feature data in reverse order to obtain reverse linear projection feature data, process the reverse linear projection feature data through the reverse Mamba layer, analyze the backward features, capture the backward time dependency, enhance the understanding of the global context, and generate reverse Mamba feature data;
[0077] Step C3: feature splicing: splicing the forward Mamba feature data and the reverse Mamba feature data to generate spliced Mamba feature data;
[0078] Step C4: The concatenated Mamba feature data is input into the linear layer and further processed by linear transformation to obtain K-PronScore feature data.
[0079] Embodiment 6, according to Figure 2 This embodiment is based on the fifth embodiment. In this embodiment, step C1 specifically includes the following steps:
[0080] Step C11: local feature extraction: the linear projection feature data is processed by one-dimensional convolution to extract local time features and generate forward convolution feature data;
[0081] Step C12: Capture global information: The forward convolution feature data is used to capture global context information through the self-attention mechanism to generate attention feature data;
[0082] Step C13: Nonlinear activation and feature transformation: Perform nonlinear activation on the attention feature data to generate nonlinear feature data, introduce high-order polynomial feature transformation to capture the complex temporal relationship of the nonlinear feature data; introduce cross terms to perform power combination on the nonlinear feature data to capture the nonlinear interaction between the nonlinear feature data and generate optimized nonlinear feature data. The formula used is as follows:
[0083] ;
[0084] in, represents power, Indicates the maximum order of polynomial transformation, that is, the maximum power; represents the cross-term feature data, represents the weight matrix of the cross terms, represents the attention feature data, Express conduct Power transformation, Express conduct Power transformation;
[0085] ;
[0086] in, represents the optimization of nonlinear characteristic data, represents the parameters of the activation function, represents a learnable activation function, represents the bias term, express The weight of represents the polynomial transformation weight, represents a polynomial feature transformation;
[0087] Step C14: Modeling long-term dependencies: Establishing a state space model, passing the optimized nonlinear feature data to the state space model, updating the current state and capturing long-term dependencies, and generating forward state space feature data;
[0088] Step C15: Feature fusion: The forward state space feature data is fused with the optimized nonlinear feature data through residual connection to generate forward Mamba feature data.
[0089] Embodiment seven, according to Figure 3 This embodiment is based on the fifth embodiment. In this embodiment, step C1 specifically includes the following steps:
[0090] Step C11: Process the linear projection feature data through one-dimensional convolution to extract local time features and generate forward convolution feature data;
[0091] Step C12: The forward convolution feature data is used to capture global context information through a self-attention mechanism to generate attention feature data;
[0092] Step C13: performing nonlinear activation on the attention feature data to generate nonlinear feature data;
[0093] Step C14: Establish a state space model, transfer the nonlinear characteristic data to the state space model, update the current state and capture the long time series dependency, and generate forward state space characteristic data;
[0094] Step C15: The forward state space feature data is fused with the nonlinear feature data through residual connection to generate forward Mamba feature data.
[0095] Embodiment 8: This embodiment is based on embodiment 6. In this embodiment, step C2 specifically includes the following steps:
[0096] Step C21: Process the reverse linear projection feature data through one-dimensional convolution to extract local time features and generate reverse convolution feature data;
[0097] Step C22: pass the reverse convolution feature data to the state space model, update the current state and capture the long-term temporal dependency, and generate reverse state space feature data as reverse Mamba feature data.
[0098] The present invention and its embodiments are described above, which is not restrictive. What is shown in the accompanying drawings is only one of the embodiments of the present invention, and the actual structure is not limited to this. In short, if ordinary technicians in this field are inspired by it and do not deviate from the purpose of the invention, they can creatively design structural methods and embodiments similar to the technical solution, which should all fall within the scope of protection of the present invention.
Claims
1. A Korean pronunciation teaching auxiliary system based on AI speech recognition, comprising an audio input and preprocessing module; the audio input and preprocessing module collects original Korean audio signals, preprocesses the original Korean audio signals, and obtains preprocessed Korean audio signals; characterized in that: The system also includes a feature extraction module and a pronunciation scoring and analysis module; The feature extraction module establishes an XLSR model, introduces a Mel filter and a diffusion process and an inverse diffusion process to improve the XLSR model, constructs an enhanced XLSR model, and pre-processes the Korean audio signal through the enhanced XLSR model to generate linear projection feature data; The pronunciation scoring and analysis module combines the self-attention mechanism, high-order nonlinear feature transformation and state space modeling to construct a bidirectional Mamba multi-order time series model; the linear projection feature data is processed by the bidirectional Mamba multi-order time series model to generate K-PronScore feature data.
2. The Korean pronunciation teaching auxiliary system based on AI speech recognition according to claim 1, characterized in that: The bidirectional Mamba multi-order time series model includes a forward Mamba layer, a reverse Mamba layer and a linear layer.
3. The Korean pronunciation teaching auxiliary system based on AI speech recognition according to claim 1 is characterized in that: The process of generating linear projection feature data by the feature extraction module specifically includes the following steps: Step S1: performing short-time Fourier transform on the preprocessed Korean audio signal to generate a frequency characteristic graph; Step S2: Processing the frequency feature map through the Mel filter bank, converting the frequency feature map into a Mel frequency scale, simulating the perceptual characteristics of the human auditory system, and obtaining a Mel spectrum; Step S3: Perform discrete cosine transform on the Mel spectrum, compress the redundant information of the Mel spectrum, and extract the Mel spectrum cepstrum coefficients to generate MFCC feature data; introduce a diffusion process to disturb the MFCC feature data by adding noise to generate disturbed MFCC feature data, and then restore the disturbed MFCC feature data through an inverse diffusion process to remove irrelevant noise information and generate optimized MFCC feature data; Step S4: further processing the optimized MFCC feature data through the XLSR model to generate a rich audio feature representation; Step S5: linearly project the rich audio feature representation, standardize the rich audio feature representation, and obtain linear projection feature data.
4. The Korean pronunciation teaching auxiliary system based on AI speech recognition according to claim 3 is characterized in that: Step S4 specifically includes the following steps: Step S41: performing preliminary feature extraction on the optimized MFCC feature data through a convolutional neural network to generate convolutional MFCC feature data; Step S42: Use the Transformer encoder to perform deeper temporal dependency modeling on the convolutional MFCC feature data, capture global context information, and generate rich audio feature representation.
5. The Korean pronunciation teaching auxiliary system based on AI speech recognition according to claim 2, characterized in that: The process of generating K-PronScore feature data in the pronunciation scoring and analysis module specifically includes the following steps: Step C1: The linear projection feature data is processed layer by layer in time order through the forward Mamba layer, the forward features are analyzed, the forward time dependency is captured, and the forward Mamba feature data is generated; Step C2: Arrange the linear projection feature data in reverse order to obtain reverse linear projection feature data, process the reverse linear projection feature data through the reverse Mamba layer, analyze the backward features, capture the backward time dependency, enhance the understanding of the global context, and generate reverse Mamba feature data; Step C3: splicing the forward Mamba feature data and the reverse Mamba feature data to generate spliced Mamba feature data; Step C4: The concatenated Mamba feature data is input into the linear layer and further processed by linear transformation to obtain K-PronScore feature data.
6. The Korean pronunciation teaching auxiliary system based on AI speech recognition according to claim 5, characterized in that: Step C1 specifically includes the following steps: Step C11: Process the linear projection feature data through one-dimensional convolution to extract local time features and generate forward convolution feature data; Step C12: The forward convolution feature data is used to capture global context information through a self-attention mechanism to generate attention feature data; Step C13: nonlinearly activate the attention feature data to generate nonlinear feature data, introduce high-order polynomial feature transformation to capture the complex temporal relationship of the nonlinear feature data; introduce cross terms to perform power combination on the nonlinear feature data to capture the nonlinear interaction between the nonlinear feature data and generate optimized nonlinear feature data; Step C14: Establish a state space model, transfer the optimized nonlinear characteristic data to the state space model, and generate forward state space characteristic data; Step C15: The forward state space feature data is fused with the optimized nonlinear feature data through residual connection to generate forward Mamba feature data.