Teaching management method and system based on big data analysis
By combining the data collected by microphone and vibration sensors, deep learning technology is used to generate student pronunciation training defect reports, which solves the problem of difficulty in correcting students' pronunciation problems in the existing technology, and realizes timely identification and correction of defects in student pronunciation training, and improves the efficiency of English oral learning.
Patent Information
- Application Number
- CN202510337558.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing English teaching methods are difficult to detect and correct problems in students' pronunciation in a timely manner, which may lead to students developing wrong pronunciation habits, which will affect the difficulty of correcting pronunciation.
By obtaining the vibration signals of the student pronunciation training process collected by microphones and the vibration training process collected by vibration sensors, deep learning technology is used to extract the Mel spectrum feature map and vibration correlation feature map of the student pronunciation training logarithmic Mel spectrum to generate a report on student pronunciation training defects to help students correct pronunciation problems in a timely manner.
It realizes timely and effectively identify and correct the defects in students' pronunciation training, avoids students' wrong pronunciation habits, and improves the efficiency of English oral learning.
Smart Images

Figure CN120220723A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of teaching management, and more specifically, to a teaching management method and system based on big data analysis. Background Art
[0002] English teaching is to help learners master the ability of the English language through a series of teaching activities. It not only includes the cultivation of basic skills such as listening, speaking, reading, and writing, but also covers multiple aspects such as grammar, vocabulary, pronunciation, and the practical application of the language. The core goal of English teaching is to improve students' English communication ability, enabling them to use English fluently in different situations to meet the needs of academic, professional, and daily communication.
[0003] In modern English classrooms, students' pronunciation problems have received increasing attention. Existing teaching methods usually involve teachers explaining pronunciation techniques or having students follow along with recordings. Students improve their pronunciation through imitation and practice. However, since teachers cannot make individual and timely pronunciation adjustments for each student in the classroom, they often cannot discover and correct students' pronunciation problems in a timely manner. This may cause students to develop incorrect pronunciation habits during long-term practice, making it difficult to correct.
[0004] Therefore, there is a need for a teaching management method and system based on big data analysis. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a teaching management method and system based on big data analysis. First, it acquires the student pronunciation training audio collected by a microphone and the vibration signal during the student pronunciation training process collected by a vibration sensor. Then, using deep learning technology, it performs feature extraction and correlation analysis on the two. Finally, through a generator, it generates a student pronunciation training defect report, thereby helping students improve their pronunciation level in a timely and effective manner, make corrections, and further avoid developing incorrect pronunciation habits, improving the efficiency of oral English learning.
[0006] According to one aspect of this application, there is provided a teaching management method based on big data analysis, which includes:
[0007] Acquiring the student pronunciation training audio collected by a microphone and the vibration signal during the student pronunciation training process collected by a vibration sensor;
[0008] Extracting the student pronunciation training logarithmic mel spectrogram feature map and the student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal during the student pronunciation training process collected by the vibration sensor;
[0009] Generate a student pronunciation training defect report based on the student pronunciation training log Mel spectrogram feature map and the student pronunciation training vibration correlation feature map.
[0010] According to another aspect of the present application, there is provided a teaching management system based on big data analysis, which includes:
[0011] An English teaching management data acquisition module for acquiring the student pronunciation training audio collected by a microphone and the vibration signal during the student pronunciation training process collected by a vibration sensor;
[0012] An English teaching management data extraction module for extracting a student pronunciation training log Mel spectrogram feature map and a student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal during the student pronunciation training process collected by the vibration sensor;
[0013] A student pronunciation training defect report generation module for generating a student pronunciation training defect report based on the student pronunciation training log Mel spectrogram feature map and the student pronunciation training vibration correlation feature map.
[0014] Compared with the prior art, a teaching management method and system based on big data analysis provided by the present application first acquire the student pronunciation training audio collected by a microphone and the vibration signal during the student pronunciation training process collected by a vibration sensor, then use deep learning technology to perform feature extraction and correlation analysis on the two, and finally generate a student pronunciation training defect report through a generator, so as to help students improve their pronunciation level in a timely and effective manner, correct it, and further avoid developing wrong pronunciation habits and improve the learning efficiency of spoken English. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 It is a flowchart of a teaching management method based on big data analysis according to an embodiment of the present application.
[0017] Figure 2 It is a flowchart of extracting a student pronunciation training log Mel spectrogram feature map and a student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal during the student pronunciation training process in the teaching management method based on big data analysis according to an embodiment of the present application.
[0018] Figure 3 It is a flowchart for extracting features from the student pronunciation training audio collected by the microphone according to the embodiments of the present application to obtain the student pronunciation training logarithmic Mel spectrogram feature map in the teaching management method based on big data analysis.
[0019] Figure 4 It is a flowchart for extracting frequency domain features from the vibration signals of the student pronunciation training process collected by the vibration sensor according to the embodiments of the present application to obtain the student pronunciation training vibration frequency domain statistical feature map.
[0020] Figure 5 It is a block diagram of a teaching management system based on big data analysis according to the embodiments of the present application. Detailed implementation manners
[0021] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0022] The special term "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here is not necessarily to be construed as superior to or better than other embodiments.
[0023] In addition, for a better description of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present application.
[0024] Furthermore, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0025] Figure 1 It is a flowchart of a teaching management method based on big data analysis according to the embodiments of the present application. As Figure 1As shown, the teaching management method based on big data analysis according to an embodiment of the present application includes: S110, obtaining the student pronunciation training audio collected by a microphone and the vibration signal of the student pronunciation training process collected by a vibration sensor; S120, extracting the student pronunciation training logarithmic Mel spectrogram feature map and the student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process collected by the vibration sensor; S130, generating a student pronunciation training defect report based on the student pronunciation training logarithmic Mel spectrogram feature map and the student pronunciation training vibration correlation feature map.
[0026] In the above teaching management method 100 based on big data analysis, in step S110, the student pronunciation training audio collected by a microphone and the vibration signal of the student pronunciation training process collected by a vibration sensor are obtained. It should be understood that the microphone records the audio signal when the student pronounces by capturing the sound waves in the air. These audio signals can be used to analyze the characteristics of the voice such as frequency, volume, pitch, etc., to help evaluate the clarity and accuracy of pronunciation. The audio signals collected by the microphone usually need to go through digital signal processing (DSP) to extract features, such as the fundamental frequency (F0) of the voice, formants, etc., in order to objectively evaluate the student's pronunciation. At the same time, the vibration sensor captures the physiological vibration patterns related to the pronunciation process by monitoring the minute vibration signals of the student's mouth, throat or chest when pronouncing. These vibration signals can reflect the muscle activity, sound generation mechanism and air flow condition of the student when pronouncing. By analyzing the vibration signals, the pronunciation actions and physiological feedback of the student can be obtained, so as to evaluate the coordination, strength and control accuracy of pronunciation. Combining these two signals (audio signal and vibration signal), a more comprehensive analysis of the student pronunciation training process can be carried out, and then the pronunciation training method can be optimized to help the student correct pronunciation problems and improve pronunciation ability.
[0027] Specifically, current teaching methods usually rely on teachers to explain pronunciation skills, or let students imitate and follow the recording to help students improve their pronunciation level. However, since teachers cannot provide personalized and immediate pronunciation correction for each student in the classroom, many students' pronunciation problems cannot be discovered and corrected in time. In this case, students may develop incorrect pronunciation habits during repeated practice, making it more difficult to correct these problems in the future. Therefore, in the technical solution of the present application, by obtaining the student pronunciation training audio collected by a microphone and the vibration signal of the student pronunciation training process collected by a vibration sensor, and combining deep learning technology, a student pronunciation training defect report is generated, so as to help students improve their pronunciation level in a timely and effective manner, correct it, and then avoid developing incorrect pronunciation habits and improve the efficiency of English oral learning.
[0028] In the above teaching management method 100 based on big data analysis, in step S120, the student pronunciation training logarithmic mel spectrogram and the student pronunciation training vibration correlation spectrogram are extracted from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process collected by the vibration sensor. It should be understood that by combining the logarithmic mel spectrogram of the audio signal and the correlation features of the vibration signal, the physiological process and sound performance of the student's pronunciation can be more comprehensively understood, so as to provide accurate feedback and improvement suggestions for pronunciation training.
[0029] Figure 2 It is a flowchart for extracting the student pronunciation training logarithmic mel spectrogram and the student pronunciation training vibration correlation spectrogram from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process in the teaching management method based on big data analysis according to the embodiments of the present application. As Figure 2 shown, in a specific embodiment of the present application, step S120 of extracting the student pronunciation training logarithmic mel spectrogram and the student pronunciation training vibration correlation spectrogram from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process includes: S121, performing feature extraction on the student pronunciation training audio collected by the microphone to obtain the student pronunciation training logarithmic mel spectrogram; S122, performing frequency domain feature extraction on the vibration signal of the student pronunciation training process collected by the vibration sensor to obtain the student pronunciation training vibration frequency domain statistical spectrogram; S123, performing waveform feature extraction on the vibration signal of the student pronunciation training process collected by the vibration sensor to obtain the student pronunciation training vibration waveform spectrogram; S124, correlating the student pronunciation training vibration frequency domain statistical spectrogram and the student pronunciation training vibration waveform spectrogram to obtain the student pronunciation training vibration correlation spectrogram.
[0030] It should be understood that by extracting the logarithmic mel spectrogram from the audio signal during the student pronunciation training process, detailed pronunciation process analysis can be provided, reflecting the pronunciation deviation and progress of the student during training, and providing an important basis for further speech training and personalized feedback.
[0031] Furthermore, performing frequency domain feature extraction on the vibration signal of the student pronunciation training process collected by the vibration sensor can analyze the vibration characteristics of the body parts (such as the larynx, oral cavity, and chest) during the student's pronunciation, and reveal the physiological characteristics of the student's pronunciation from the perspective of the frequency domain. By transforming the vibration signal into the frequency domain, frequency components, vibration modes related to pronunciation and their changes can be extracted, so as to more accurately evaluate the pronunciation strength, coordination, and control accuracy of the student.
[0032] Furthermore, waveform feature extraction is performed on the vibration signals collected during the student pronunciation training process by the vibration sensor. The purpose is to more directly obtain time-series information about the vibration of body parts during pronunciation from the vibration signals, helping to evaluate the performance in aspects such as muscle control, coordination, and strength during the pronunciation process. Among them, waveform feature extraction can obtain key features from the time-domain waveform of the vibration signal, and can detail the change rules and performance of muscle vibration when students pronounce.
[0033] Specifically, correlating the vibration frequency-domain statistical feature map and the vibration waveform feature map of student pronunciation training can integrate the time-domain and frequency-domain information of the vibration signal, reveal details such as muscle control, pronunciation strength, and stability during the pronunciation process, and provide more accurate analysis and feedback. Among them, the vibration frequency-domain statistical feature map extracts information such as the energy distribution, frequency peak, and spectral bandwidth of the vibration signal at different frequencies. These features reflect the intensity and distribution of different frequency components during the student pronunciation process, helping to evaluate the frequency characteristics of pronunciation, coordination, and force output during the pronunciation process. The vibration waveform feature map, based on the characteristics of the time-domain signal, such as peaks, zero-crossing points, and the change rate of the waveform, reveals the time-varying characteristics of the vibration, reflecting the dynamic changes, vibration intensity, and its stability during the pronunciation process. By correlating these two feature maps, the frequency-domain and time-domain information can be combined to obtain a more comprehensive vibration correlation feature map of student pronunciation training. In specific operations, first, the frequency-domain statistical features are matched and fused with the time-domain waveform features. The statistical features in the frequency domain (such as frequency peak and bandwidth, etc.) can be compared with the amplitude changes and periodic features in the waveform to analyze their mutual relationship during the pronunciation process. At the same time, through multi-dimensional feature fusion, potential correlations between the time domain and the frequency domain are discovered, such as whether the changes in the frequency domain are synchronized with the instantaneous amplitude or the changes in the waveform amplitude in the waveform, thus forming a multi-level vibration correlation feature map. Finally, the vibration correlation feature map of student pronunciation training provides a comprehensive evaluation of muscle control and coordination during the pronunciation process, making the problems in training (such as unstable pronunciation, inaccurate force control, etc.) clearer, and providing a scientific basis for personalized training programs.
[0034] Figure 3 It is a flowchart for extracting features from the student pronunciation training audio collected by the microphone according to the embodiments of the present application to obtain the student pronunciation training logarithmic mel spectrogram feature map. As Figure 3As shown, in a specific embodiment of the present application, step S121 of extracting the logarithmic Mel spectrogram features of the student pronunciation training audio collected by the microphone to obtain a logarithmic Mel spectrogram feature map of the student pronunciation training includes: S1211, extracting the logarithmic Mel spectrogram of the student pronunciation training audio collected by the microphone; S1212, passing the logarithmic Mel spectrogram of the student pronunciation training through a logarithmic Mel spectrogram feature encoder of the student pronunciation training to obtain the logarithmic Mel spectrogram feature map of the student pronunciation training.
[0035] It should be understood that considering that the student pronunciation audio signal collected by the microphone is a time-domain signal containing rich audio information, in order to effectively extract the content related to the pronunciation features, it needs to be converted into a frequency-domain representation. This process is usually achieved through the short-time Fourier transform (STFT), which decomposes the audio signal into frequency components within several short-time windows, thereby revealing the spectral features of the audio signal at different time points. Next, the Mel filter bank is applied to the spectrum, processing the spectrum according to the sensitivity of the human ear to different frequencies, compressing the high-frequency information and enhancing the low-frequency signal to adapt to the human auditory characteristics. The use of the Mel frequency scale can make the spectrum more in line with the requirements of speech recognition and pronunciation analysis because the human ear is more sensitive to low frequencies and less sensitive to high frequencies. Then, in order to enhance the contrast of the spectrum and improve the adaptability to low signal-to-noise ratios, a logarithmic transformation is performed on the Mel spectrum. This step converts the amplitude information of the Mel spectrum into a logarithmic value, strengthening the weak components in the signal and improving the distinguishability of the features, especially being more sensitive to slight changes in pronunciation training. Finally, by converting the student pronunciation audio signal collected by the microphone into a logarithmic Mel spectrogram, various features such as the temporal variation, pitch, rhythm, and pronunciation clarity of the speech during the pronunciation process can be effectively revealed.
[0036] Furthermore, processing the log Mel spectrogram of the student pronunciation training through the log Mel spectrogram feature encoder for student pronunciation training can extract more efficient, concise, and discriminative features from the spectrogram, thereby providing more accurate data support for the analysis, evaluation, and feedback of student pronunciation training. It should be understood that considering that the log Mel spectrogram is the spectrogram of the audio signal processed by the short-time Fourier transform (STFT) and the Mel filter bank, it can effectively reflect the frequency characteristics during the pronunciation process. However, these features often have high dimensions and a large amount of redundant information. Directly using these spectrograms for analysis and modeling may reduce efficiency and increase computational complexity. Therefore, the log Mel spectrogram feature encoder is introduced to map the original log Mel spectrogram to a more compact and representative feature space. Among them, the feature encoder encodes the log Mel spectrogram through a series of deep learning models (such as convolutional neural networks, autoencoders, etc.) to extract the key information therein. In specific operations, the encoder first performs dimensionality reduction processing on the input log Mel spectrogram, removes the redundant information therein, and converts the high-dimensional spectrogram into a low-dimensional feature representation. These low-dimensional feature representations can retain the important speech features in the pronunciation training while eliminating the irrelevant or noisy parts. The feature encoder usually learns how to effectively extract the potential patterns related to the pronunciation quality through the training data, so that the encoder can identify the subtle changes in the student's pronunciation, such as the clarity, stability, and accuracy of the pronunciation. By performing feature encoding on the log Mel spectrogram of the student pronunciation training, not only can a simplified low-dimensional feature map be obtained, but also more discriminative pronunciation features can be extracted. Specifically, obtaining the log Mel spectrogram feature map for student pronunciation training by passing the log Mel spectrogram of the student pronunciation training through the log Mel spectrogram feature encoder for student pronunciation training includes: using each layer of the log Mel spectrogram feature encoder for student pronunciation training to perform convolution processing, mean pooling processing based on the local feature matrix, and non-linear activation processing on the input data respectively during the forward pass of the layer, and outputting the log Mel spectrogram feature map for student pronunciation training by the last layer of the log Mel spectrogram feature encoder for student pronunciation training, where the input of the log Mel spectrogram feature encoder for student pronunciation training is the log Mel spectrogram of the student pronunciation training.
[0037] Figure 4 A flowchart for extracting the frequency domain features of the vibration signal during the student pronunciation training collected by the vibration sensor according to the embodiments of the present application to obtain the vibration frequency domain statistical feature map for student pronunciation training. As Figure 4As shown, in a specific embodiment of the present application, step S122, which extracts the frequency-domain features of the vibration signals during the student pronunciation training collected by the vibration sensor to obtain the vibration frequency-domain statistical feature map of the student pronunciation training, includes: S1221, performing Fourier transform on the vibration signals during the student pronunciation training collected by the vibration sensor to obtain multiple vibration frequency-domain statistical feature values of the student pronunciation training; S1222, arranging the multiple vibration frequency-domain statistical feature values of the student pronunciation training and passing them through the vibration frequency-domain statistical convolutional neural network of the student pronunciation training as a feature encoder to obtain the vibration frequency-domain statistical feature map of the student pronunciation training.
[0038] It should be understood that the Fourier transform represents the time-domain signal as a combination of different frequency components, thereby revealing the energy distribution and vibration characteristics of the signal in each frequency band. During the pronunciation process, students are involved in vibrations of multiple aspects such as muscles and airflows, and the frequency characteristics of these vibrations can reflect the force application, stability, and muscle coordination during pronunciation. Through the Fourier transform, different vibration components from low frequency to high frequency can be effectively separated, and the key frequency information during the student pronunciation training process can be extracted. Specifically, when performing Fourier transform on the vibration signal, first, the collected vibration signal is preprocessed to remove noise and perform smoothing. Then, the signal is transformed into the frequency domain through the short-time Fourier transform (STFT) or the fast Fourier transform (FFT) to obtain the amplitude spectrum of the frequency components. In the frequency domain, the main characteristics of the vibration signal include frequency distribution, energy concentration, spectral width, etc., all of which can reflect the frequency characteristics during the pronunciation process. Then, multiple frequency-domain statistical feature values can be extracted from the spectrum, including frequency peak, bandwidth, average frequency, maximum frequency, etc. These feature values help analyze the vibration mode and pronunciation stability during the student pronunciation training process. For example, the frequency peak reflects the frequency position of the main energy in the signal, the bandwidth can describe the frequency distribution range of the vibration, and the average frequency helps evaluate the overall frequency characteristics of the pronunciation. Through these frequency-domain statistical feature values, the changes in high-frequency or low-frequency components during the student pronunciation process can be effectively revealed, helping teachers identify problems during the pronunciation process, such as muscle tension, unclear pronunciation, or unstable pronunciation. Finally, the multiple vibration frequency-domain statistical feature values obtained by the Fourier transform provide rich information for further analyzing the pronunciation performance of students.
[0039] Furthermore, the statistical eigenvalue of the vibration frequency domain of multiple students' pronunciation training is arranged and processed through a convolutional neural network (CNN) of the vibration frequency domain statistics of students' pronunciation training as a feature encoder, aiming to transform the frequency domain statistical features into more efficient and richly expressive feature maps, so as to provide accurate and intuitive data support for the analysis and evaluation of pronunciation training. In specific operations, first, multiple frequency domain statistical eigenvalues need to be arranged or sorted, and they are organized into a suitable format (such as a two-dimensional matrix or a multi-dimensional array) so that they can be input into the convolutional neural network. Usually, these eigenvalues are arranged in the order of time series or different frequency bandwidths to ensure that the time series of data and the spatiality of the spectrum are effectively retained. Then, these sorted feature data are processed through the convolutional neural network. The convolutional neural network extracts features from the input data through multiple convolutional layers and pooling layers, and can automatically learn the potential patterns and relationships between frequency domain features, and extract higher-level features with stronger expressive power. The convolutional operation in the convolutional neural network can capture the associated features in the local area, and reduce the computational amount through the pooling layer while retaining important information, and finally compress multiple frequency domain statistical eigenvalues into a more compact and low-dimensional feature map. This feature map can comprehensively reflect the vibration characteristics in the process of students' pronunciation, including information such as frequency distribution, pronunciation stability, and energy concentration. Through this method, teachers can more accurately understand the problems in students' pronunciation training, such as unstable pitch, pronunciation fatigue, or improper force control, etc., so as to adjust the training strategy in time and improve the students' pronunciation level. Specifically, arranging the statistical eigenvalue of the vibration frequency domain of multiple students' pronunciation training and passing it through the convolutional neural network of the vibration frequency domain statistics of students' pronunciation training as a feature encoder to obtain the statistical feature map of the vibration frequency domain of students' pronunciation training includes: arranging the statistical eigenvalue of the vibration frequency domain of multiple students' pronunciation training into an input matrix of the vibration frequency domain of students' pronunciation training; using each layer of the convolutional neural network of the vibration frequency domain statistics of students' pronunciation training as a feature encoder to respectively perform the following operations on the input data during the forward transmission of the layer: using the convolutional units of each layer of the convolutional neural network of the vibration frequency domain statistics of students' pronunciation training as a feature encoder to perform convolutional processing on the input data based on a two-dimensional convolutional kernel to obtain a convolutional feature map; using the pooling units of each layer of the convolutional neural network of the vibration frequency domain statistics of students' pronunciation training as a feature encoder to perform pooling processing on the convolutional feature map along the channel dimension to obtain a pooled feature map; and, using the activation units of each layer of the convolutional neural network of the vibration frequency domain statistics of students' pronunciation training as a feature encoder to perform non-linear activation on the eigenvalues at each position in the pooled feature map to obtain an activated feature map; wherein, the output of the last layer of the convolutional neural network of the vibration frequency domain statistics of students' pronunciation training as a feature encoder is the statistical feature map of the vibration frequency domain of students' pronunciation training.
[0040] In a specific embodiment of the present application, step S123 of extracting waveform features from the vibration signals collected during the student pronunciation training process to obtain a vibration waveform feature map of the student pronunciation training includes: extracting a vibration waveform map of the student pronunciation training from the vibration signals collected during the student pronunciation training process by the vibration sensor; and passing the vibration waveform map of the student pronunciation training through a student pronunciation training vibration waveform feature encoder based on a channel attention mechanism to obtain the vibration waveform feature map of the student pronunciation training.
[0041] It should be understood that vibration signals are mechanical vibrations generated by students during the pronunciation process, and these vibrations are closely related to aspects such as pronunciation strength, frequency, and stability. The vibration waveform map can visually display the time-varying characteristics of the signals, helping teachers or researchers to evaluate the pronunciation quality in real time. The process of extracting the vibration waveform map first requires collecting the vibration signals. The vibration sensor converts the minute vibrations generated when students speak during the pronunciation training process into electrical signals. At this time, the vibration signals themselves are time-domain signals, containing all the dynamic information during pronunciation. Once the signals are processed, they can be plotted into a vibration waveform map. The vibration waveform map is a graphical representation with time as the horizontal axis and vibration intensity as the vertical axis, capable of showing the instantaneous changes, amplitude fluctuations, and periodicity of the signals. By observing the vibration waveform map, one can intuitively understand the changes during the student's pronunciation, including the pronunciation duration, pitch fluctuations, and pronunciation stability. The periodic changes in the waveform map reflect the repeatability of pronunciation, while the amplitude changes show the pronunciation strength changes. Through a detailed analysis of these waveforms, factors such as pronunciation intensity, clarity, and stability during the student's pronunciation can be evaluated, thereby providing a basis for personalized pronunciation training programs.
[0042] Furthermore, processing the vibration waveform diagram of students' pronunciation training through a student pronunciation training vibration waveform feature encoder based on the channel attention mechanism can extract more meaningful and discriminative features from the original vibration waveform diagram, thus more effectively analyzing and evaluating the pronunciation quality of students. Among them, the core idea of the channel attention mechanism is to weight and adjust the importance of different feature channels, enabling the network to focus on the most relevant feature channels and thus ignoring those channels that contribute little to the task or have a large amount of noise. In the vibration waveform diagram, the vibration characteristics in certain frequency bands or certain time periods have a greater impact on the pronunciation quality, while the impact of other parts is relatively small. The channel attention mechanism assigns different importance to different channels (i.e., different vibration characteristics) by learning adaptive weights, so that the encoder can concentrate on those parts that are crucial for evaluating the pronunciation quality when processing the waveform diagram. Specifically, the feature encoder based on the channel attention mechanism first performs preliminary feature extraction on the waveform diagram through a convolutional layer, and then introduces the channel attention mechanism to calculate the importance weights of each feature channel. These weights are used to adjust the contributions of different channels, so that the useful information in the feature map is strengthened and the redundant or irrelevant information is suppressed. Finally, the encoder outputs a feature map that contains the most discriminative features extracted from the original vibration waveform diagram. These feature maps can reflect the key patterns in the pronunciation process, such as the stability of pronunciation, force control, frequency fluctuation, etc., providing accurate and effective data support for subsequent pronunciation quality evaluation, training strategy adjustment, and feedback. In this way, the vibration waveform feature map of students' pronunciation training not only simplifies the original data, reduces information redundancy, but also enhances the expression ability and classification ability of the feature map, helping to better identify problems in pronunciation and providing strong support for personalized pronunciation training and optimization. Among them, obtaining the vibration waveform feature map of students' pronunciation training by passing the vibration waveform diagram of students' pronunciation training through the student pronunciation training vibration waveform feature encoder based on the channel attention mechanism includes: using each layer of the student pronunciation training vibration waveform feature encoder based on the channel attention mechanism to respectively perform the following operations on the input data during the forward pass of the layer: performing convolutional processing on the input data based on a convolutional kernel to generate a convolutional feature map; performing pooling processing on the convolutional feature map to generate a pooling feature map; performing activation processing on the pooling feature map to generate an activation feature map; calculating the quotient of the mean value of the eigenvalues of the feature matrix corresponding to each channel in the activation feature map and the sum of the mean values of the eigenvalues of the feature matrix corresponding to all channels as the weighting coefficient of the feature matrix corresponding to each channel; and weighting the feature matrix of each channel with the weighting coefficient of each channel in the activation feature map to generate a channel attention feature map; wherein, the output of the last layer of the student pronunciation training vibration waveform feature encoder based on the channel attention mechanism is the vibration waveform feature map of students' pronunciation training.
[0043] In the above teaching management method 100 based on big data analysis, in step S130, a student pronunciation training defect report is generated based on the student pronunciation training logarithmic mel spectrogram and the student pronunciation training vibration correlation feature map. It should be understood that the logarithmic mel spectrogram can reflect the frequency components and their energy distribution in the student's pronunciation, while the vibration correlation feature map reflects the mechanical vibration and strength control during the pronunciation process. By combining these two feature maps, the quality, stability, and potential defects of the student's pronunciation can be comprehensively evaluated, and then a detailed pronunciation training defect report can be generated. Through precise feature analysis, potential problems in the student's pronunciation process, such as unstable pronunciation, insufficient pronunciation strength, or poor pronunciation coordination, can be revealed, thus providing a scientific basis for personalized training.
[0044] In a specific embodiment of the present application, in step S130, generating a student pronunciation training defect report based on the student pronunciation training logarithmic mel spectrogram and the student pronunciation training vibration correlation feature map includes: performing pooling on the student pronunciation training logarithmic mel spectrogram and the student pronunciation training vibration correlation feature map to obtain a student pronunciation training logarithmic mel spectrogram feature vector and a student pronunciation training vibration correlation feature vector; fusing the student pronunciation training logarithmic mel spectrogram feature vector and the student pronunciation training vibration correlation feature vector to obtain a pronunciation defect report generation feature vector; performing lightweight feature dynamic activation response optimization based on structural constraint perception on the pronunciation defect report generation feature vector to obtain an optimized pronunciation defect report generation feature vector; and passing the optimized pronunciation defect report generation feature vector through a generator to generate a student pronunciation training defect report.
[0045] It should be understood that by performing pooling operations on the student pronunciation training logarithmic Mel spectrogram feature map and the student pronunciation training vibration correlation feature map, more concise and representative feature vectors can be effectively extracted, thereby reducing the computational complexity to a certain extent and improving the efficiency of subsequent analysis. The pooling operation mainly extracts important information from the local regions in the feature map by statistical methods, while ignoring some irrelevant and redundant information. For the student pronunciation training logarithmic Mel spectrogram feature map, the pooling operation can be performed by pooling the feature map along the channel dimension, retaining the most representative features in each channel while removing unnecessary details. This approach can reduce the dimensionality of the data, enabling subsequent analysis to focus more on the core features without being interfered by noise or irrelevant factors. A similar operation also applies to the student pronunciation training vibration correlation feature map. Through pooling, the most important parts of each vibration feature can be extracted, reducing the risk of information loss and enhancing the model's sensitivity to key information. The student pronunciation training logarithmic Mel spectrogram feature vector and the student pronunciation training vibration correlation feature vector obtained through pooling will retain the most prominent features while improving the computational efficiency and performance of the model.
[0046] Furthermore, although the student pronunciation training logarithmic Mel spectrogram feature vector and the student pronunciation training vibration correlation feature vector each provide information about different aspects of pronunciation, there is a certain correlation between them. The pronunciation defects can be more accurately identified by fusing these two feature vectors. During the fusion process, various methods can be adopted, such as concatenation, weighted average, or feature fusion through a deep learning model. In these ways, the student pronunciation training logarithmic Mel spectrogram feature vector and the student pronunciation training vibration correlation feature vector can be integrated into a unified feature representation, capturing information in more dimensions. On this basis, models such as deep neural networks or convolutional neural networks can be used to further process the fused feature vector to generate a feature vector for generating pronunciation defect reports. This feature vector for generating pronunciation defect reports will comprehensively display various problems in the student's pronunciation process, such as unstable pitch, uneven intensity, unclear pronunciation, etc.
[0047] Furthermore, in the technical solution of the present application, considering that there may be cross-redundancy in some features when audio signals and vibration signals capture pronunciation features, that is, they reflect the same or similar pronunciation problems in a certain aspect. This redundancy will affect the efficiency of subsequent analysis. Especially when applying deep learning models for training, it not only wastes computing resources but also may cause the model to over-rely on certain redundant features, thereby reducing the accuracy of its prediction and analysis. Moreover, the information dimensions represented by different types of features (such as log Mel spectrogram features and vibration correlation feature maps) vary greatly. Some features may contain high-dimensional information, while other features may only have low-dimensional information, resulting in an imbalance in feature dimensions. When these features with different dimensions are combined into a feature vector, the dimension differences between them may cause the model to be overly sensitive to certain high-dimensional features during training while ignoring low-dimensional features, thus affecting the accuracy of the overall analysis result. The imbalance of feature dimensions will cause different types of features to lose their due balance during integration, leading to biases in the model learning process and reducing the effect of generating the final pronunciation defect report. Therefore, in the technical solution of the present application, the feature vector for generating the pronunciation defect report is optimized by lightweight feature dynamic activation response based on structure constraint perception to obtain an optimized feature vector for generating the pronunciation defect report.
[0048] Among them, optimizing the feature vector for generating the pronunciation defect report by lightweight feature dynamic activation response based on structure constraint perception to obtain an optimized feature vector for generating the pronunciation defect report includes: extracting a pronunciation defect structured constraint matrix; performing denoising filtering on the pronunciation defect structured constraint matrix to obtain a set of pronunciation defect model core prior information structure-aware encoding vectors; constructing a pronunciation defect model core prior information non-linear coupling matrix between the feature vector for generating the pronunciation defect report and each pronunciation defect model core prior information structure-aware encoding vector in the set of pronunciation defect model core prior information structure-aware encoding vectors to obtain a set of pronunciation defect model core prior information non-linear coupling matrices; calculating the pronunciation defect prior information response dynamic activation mode factor of each pronunciation defect model core prior information non-linear coupling matrix in the set of pronunciation defect model core prior information non-linear coupling matrices to obtain a set of pronunciation defect prior information response dynamic activation mode factors; based on the set of pronunciation defect prior information response dynamic activation mode factors, performing lightweight feature fusion on the set of pronunciation defect model core prior information non-linear coupling matrices to obtain a pronunciation defect model prior information response projection coding matrix; mapping the feature vector for generating the pronunciation defect report to the feature space of the pronunciation defect model prior information response projection coding matrix to obtain the optimized feature vector for generating the pronunciation defect report.
[0049] Among them, performing lightweight feature dynamic activation response optimization based on structural constraint perception on the pronunciation defect report generation feature vector to obtain an optimized pronunciation defect report generation feature vector includes:
[0050] First, extract the pronunciation defect structured constraint matrix. It should be understood that extracting the pronunciation defect structured constraint matrix is not just an information acquisition process, but actually a key operation for structuring and computabilizing prior knowledge. The principle is to condense domain expertise, model design concepts, or macroscopic laws inferred from data into a mathematical form of the pronunciation defect structured constraint matrix for subsequent effective utilization by algorithms.
[0051] Then, perform core prior knowledge denoising filtering on the pronunciation defect structured constraint matrix to obtain a set of pronunciation defect model core prior information structure-aware encoding vectors, which is represented by the core prior knowledge denoising filtering formula as:
[0052]
[0053] Among them, U represents the set of pronunciation defect model core prior information structure-aware encoding vectors, v1, v2, v m respectively represent the first, second, and m-th pronunciation defect model core prior information structure-aware encoding vectors, T represents the transpose operation, CoreExtraction represents core prior knowledge denoising filtering, M p represents the pronunciation defect structured constraint matrix, Λ represents the diagonal matrix, and λ1, λ m respectively represent the values at the first and m-th positions on the diagonal of the diagonal matrix. It should be understood that the pronunciation defect structured constraint matrix may contain redundant or high-dimensional information, and direct application may lead to an excessive computational burden and interference from non-critical information in the optimization process. Therefore, the principle of core prior knowledge denoising filtering is to denoise and refine prior knowledge, similar to the filtering process in signal processing, retaining the main components and filtering out redundancy.
[0054] Next, construct a pronunciation defect model core prior information non-linear coupling matrix between the pronunciation defect report generation feature vector and each pronunciation defect model core prior information structure-aware encoding vector in the set of pronunciation defect model core prior information structure-aware encoding vectors to obtain a set of pronunciation defect model core prior information non-linear coupling matrices, which is represented by the non-linear coupling formula as:
[0055]
[0056] Among them, x o represents the pronunciation defect report generation feature vector, l i (x o ) represents xo Perform a linear transformation. The eigenvector after the linear transformation has the same feature scale as the corresponding perceptual coding vector of the core prior information of the pronunciation defect model, v i represents the i-th perceptual coding vector of the core prior information of the pronunciation defect model, represents matrix multiplication, L represents the length of the perceptual coding vector of the core prior information of the pronunciation defect model, MR i represents the i-th non-linear coupling matrix of the core prior information of the pronunciation defect model. It should be understood that the core of this step lies in constructing the interaction relationship between the pronunciation defect report generation eigenvector and the refined prior knowledge. Specifically, through latent space mapping, the non-linear response pattern of the pronunciation defect report generation eigenvector to prior knowledge in different aspects is learned. Essentially, the non-linear coupling matrix of the core prior information of the pronunciation defect model is a re-encoding of the pronunciation defect report generation eigenvector from the perspective of prior knowledge, integrating the interpretation and processing of knowledge. Its function exceeds information association, realizing the directional enhancement of feature representation and the extraction of multi-perspective feature information.
[0057] Subsequently, calculate the dynamic activation mode factor of the prior information response of each non-linear coupling matrix of the core prior information of the pronunciation defect model in the set of non-linear coupling matrices of the core prior information of the pronunciation defect model to obtain a set of dynamic activation mode factors of the prior information response of the pronunciation defect, which is expressed by the dynamic activation mode formula as:
[0058] S i =‖MR i ‖ F
[0059] where, ‖·‖ F represents the F-norm of the matrix, S i represents the i-th dynamic activation mode factor of the prior information response of the pronunciation defect. It should be understood that the principle of this step is to extract key information and streamline the feature representation for each non-linear coupling matrix of the core prior information of the pronunciation defect model. Just like generating an information summary, the most representative summary or eigenvector of the core information is extracted. The dynamic activation mode factor of the prior information response of the pronunciation defect needs to be representative and discriminative. It contains the ideas of feature selection and feature aggregation, retaining the most informative parts and aggregating them into a concise dynamic activation mode factor to achieve deeper compression and extraction. The function of the dynamic activation mode factor of the prior information response of the pronunciation defect is not only information compression, but more importantly, to improve the efficiency and robustness of the subsequent fusion process, reduce the data dimension, and reduce the computational burden, especially in the case of high-dimensional matrices.
[0060] Subsequently, based on the prior information of the pronunciation defect, a set of dynamic activation mode factors is responded to, and lightweight feature fusion is performed on the set of non-linear coupling matrices of the core prior information of the pronunciation defect model to obtain a projection coding matrix of the prior information response of the pronunciation defect model, which is expressed by the lightweight feature fusion formula as:
[0061]
[0062] Among them, softmax represents the normalized exponential function, and P represents the projection coding matrix of the prior information response of the pronunciation defect model. It should be understood that the core principle of the lightweight feature fusion lies in emphasizing the adaptive and selective prior information integration strategy. Sparsity reflects that not all prior information responses are equally important, and the fusion should be selective, focusing on more important responses, weakening or ignoring unimportant responses, improving the feature selection ability and generalization ability, and avoiding overfitting. Dynamics means that the weights or methods of lightweight feature fusion are not fixed, but are adaptively adjusted according to the input data or model state, improving the flexibility and adaptability of the model. The essence of performing lightweight feature fusion on the set of non-linear coupling matrices of the core prior information of the pronunciation defect model is to optimally combine the response information from different prior knowledge perspectives to form a projection coding matrix of the prior information response of the pronunciation defect model that comprehensively reflects the prior response of the model.
[0063] Finally, the pronunciation defect report generation feature vector is mapped to the feature space of the projection coding matrix of the prior information response of the pronunciation defect model to obtain the optimized pronunciation defect report generation feature vector, which is expressed by the mapping formula as:
[0064]
[0065] where x optRepresents the optimized pronunciation defect report generation feature vector. It should be understood that the entire adaptation process is finally completed by mapping the pronunciation defect report generation feature vector to the feature space of the pronunciation defect model prior information response projection coding matrix. The principle of this step is to utilize the feature space defined by the pronunciation defect model prior information response projection coding matrix to project the pronunciation defect report generation feature vector into this space. Substantially, the pronunciation defect model prior information response projection coding matrix acts as a transformation matrix, performing a linear or non-linear mapping transformation on the pronunciation defect report generation feature vector, enabling it to be embedded in the feature space integrating model prior information. The optimized pronunciation defect report generation feature vector better conforms to the constraints and guidance of the model prior, achieving boundary adaptation on the feature manifold. Compared with the pronunciation defect report generation feature vector, the optimized pronunciation defect report generation feature vector is usually significantly improved in terms of expression ability, discriminability, and generalization, providing a better feature representation for subsequent machine learning tasks.
[0066] In particular, the optimized pronunciation defect report generation feature vector contains various key features in the student's pronunciation process, such as the frequency stability of pronunciation, volume control, pronunciation clarity, etc. Through the generator model, the optimized pronunciation defect report generation feature vector can be transformed into meaningful report content, helping teachers or researchers deeply understand the deficiencies in the student's pronunciation and formulate targeted improvement plans. Specifically, the generator can transform the abstract optimized pronunciation defect report generation feature vector into a specific and structured pronunciation defect report. This process can be achieved through a generative adversarial network (GAN) or other generative models. The generator takes the optimized pronunciation defect report generation feature vector as input, undergoes a series of neural network processes, and outputs an easily understandable text report. During the generation process, the generator model will first analyze the key information in the optimized pronunciation defect report generation feature vector, such as frequency fluctuations during pronunciation, abnormal vibration amplitudes, or instability of pronunciation, and generate a detailed report based on this information. The report content may include descriptions of pronunciation problems, the severity of defects, and suggestions for targeted improvement. Through the optimization process of the generator, the pronunciation defect report can not only have high precision but also make the report content more readable and practically guiding. Finally, teachers and students can formulate more personalized and precise pronunciation training plans based on the analysis and suggestions in these reports, thereby improving the effect of pronunciation training. The introduction of the generator model makes the report generation process automated and precise, and greatly improves the efficiency of training feedback.
[0067] In summary, in the embodiments of the present application, first, the student pronunciation training audio collected by the microphone and the vibration signal during the student pronunciation training process collected by the vibration sensor are obtained. Then, using deep learning technology, feature extraction and correlation analysis are performed on the two. Finally, through the generator, a student pronunciation training defect report is generated, so as to help students improve their pronunciation level in a timely and effective manner, correct it, and further avoid developing incorrect pronunciation habits and improve the efficiency of oral English learning.
[0068] As described above, the teaching management method 100 based on big data analysis according to the embodiments of the present application can be implemented in various terminal devices. In one example, the teaching management method 100 based on big data analysis can be integrated into the terminal device as a software module and / or a hardware module. For example, the teaching management method 100 based on big data analysis can be a software module in the operating system of the terminal device, or can be an application program developed for the terminal device; of course, the teaching management method 100 based on big data analysis can also be one of the many hardware modules of the terminal device.
[0069] Alternatively, in another example, the teaching management method 100 based on big data analysis and the terminal device can also be separate devices, and the teaching management method 100 based on big data analysis can be connected to the terminal device through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.
[0070] Figure 5 FIG. is a block diagram of a teaching management system based on big data analysis according to an embodiment of the present application. As Figure 5 shown, the teaching management system 100 based on big data analysis according to the embodiments of the present application includes: an English teaching management data acquisition module 110, configured to acquire the student pronunciation training audio collected by the microphone and the vibration signal during the student pronunciation training process collected by the vibration sensor; an English teaching management data extraction module 120, configured to extract the student pronunciation training log Mel spectrogram feature map and the student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal during the student pronunciation training process collected by the vibration sensor; a student pronunciation training defect report generation module 130, configured to generate a student pronunciation training defect report based on the student pronunciation training log Mel spectrogram feature map and the student pronunciation training vibration correlation feature map.
[0071] Here, those skilled in the art can understand that the specific operations of each step in the above teaching management system based on big data analysis have been described in detail in the description of the teaching management method based on big data analysis above with reference to Figures 1 to 4 and therefore, the repeated description thereof will be omitted.
Claims
1. A teaching management method based on big data analysis, characterized in that: include: Acquire the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process collected by the vibration sensor; Extracting a student pronunciation training logarithmic Mel frequency spectrum feature graph and a student pronunciation training vibration correlation feature graph from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process collected by the vibration sensor; A student pronunciation training defect report is generated based on the student pronunciation training logarithmic Mel frequency spectrum feature graph and the student pronunciation training vibration association feature graph.
2. The teaching management method based on big data analysis according to claim 1 is characterized in that: Extracting a student pronunciation training logarithmic Mel spectrum feature map and a student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process collected by the vibration sensor, including: Performing feature extraction on the student pronunciation training audio collected by the microphone to obtain a logarithmic Mel frequency spectrum feature graph of the student pronunciation training; Performing frequency domain feature extraction on the vibration signal of the student pronunciation training process collected by the vibration sensor to obtain a frequency domain statistical feature graph of the student pronunciation training vibration; Extracting waveform features of the vibration signal of the student pronunciation training process collected by the vibration sensor to obtain a student pronunciation training vibration waveform feature graph; The student pronunciation training vibration frequency domain statistical feature graph and the student pronunciation training vibration waveform feature graph are associated to obtain the student pronunciation training vibration association feature graph.
3. The teaching management method based on big data analysis according to claim 2 is characterized in that: The feature extraction is performed on the student pronunciation training audio collected by the microphone to obtain a logarithmic Mel spectrum feature graph of the student pronunciation training, including: Extracting a student pronunciation training logarithmic Mel-frequency spectrogram from the student pronunciation training audio collected by the microphone; The student pronunciation training logarithmic Mel spectrum graph is passed through the student pronunciation training logarithmic Mel spectrum feature encoder to obtain the student pronunciation training logarithmic Mel spectrum feature graph.
4. The teaching management method based on big data analysis according to claim 3 is characterized in that: The vibration signal of the student pronunciation training process collected by the vibration sensor is subjected to frequency domain feature extraction to obtain a frequency domain statistical feature graph of the student pronunciation training vibration, including: Performing Fourier transform on the vibration signal of the student pronunciation training process collected by the vibration sensor to obtain a plurality of student pronunciation training vibration frequency domain statistical characteristic values; The plurality of student pronunciation training vibration frequency domain statistical feature values are arranged and passed through a student pronunciation training vibration frequency domain statistical convolutional neural network as a feature encoder to obtain the student pronunciation training vibration frequency domain statistical feature graph.
5. The teaching management method based on big data analysis according to claim 4 is characterized in that: The vibration signal of the student pronunciation training process collected by the vibration sensor is subjected to waveform feature extraction to obtain a student pronunciation training vibration waveform feature graph, including: Extracting a student pronunciation training vibration waveform from the vibration signal of the student pronunciation training process collected by the vibration sensor; The student pronunciation training vibration waveform graph is passed through a student pronunciation training vibration waveform feature encoder based on a channel attention mechanism to obtain the student pronunciation training vibration waveform feature graph.
6. The teaching management method based on big data analysis according to claim 5 is characterized in that: The student pronunciation training vibration waveform graph is passed through a student pronunciation training vibration waveform feature encoder based on a channel attention mechanism to obtain the student pronunciation training vibration waveform feature graph, including: Each layer of the student pronunciation training vibration waveform feature encoder based on the channel attention mechanism performs the following operations on the input data in the forward pass of the layer: Performing convolution processing on the input data based on a convolution kernel to generate a convolution feature map; Performing pooling processing on the convolutional feature map to generate a pooled feature map; Performing activation processing on the pooled feature map to generate an activation feature map; Calculating the quotient of the eigenvalue mean of the feature matrix corresponding to each channel in the activation feature map and the sum of the eigenvalue means of the feature matrices corresponding to all channels as the weighting coefficient of the feature matrix corresponding to each channel; and Weighting the feature matrix of each channel by the weight coefficient of each channel in the activation feature map to generate a channel attention feature map; Among them, the output of the last layer of the student pronunciation training vibration waveform feature encoder based on the channel attention mechanism is the student pronunciation training vibration waveform feature map.
7. The teaching management method based on big data analysis according to claim 6 is characterized in that: Based on the student pronunciation training logarithmic Mel frequency spectrum feature graph and the student pronunciation training vibration association feature graph, a student pronunciation training defect report is generated, including: Pooling the student pronunciation training logarithmic Mel spectrum feature map and the student pronunciation training vibration association feature map to obtain a student pronunciation training logarithmic Mel spectrum feature vector and a student pronunciation training vibration association feature vector; Fusion of the student pronunciation training logarithmic Mel spectrum feature vector and the student pronunciation training vibration association feature vector to obtain a pronunciation defect report generation feature vector; Performing a lightweight feature dynamic activation response optimization based on structural constraint perception on the pronunciation defect report generation feature vector to obtain an optimized pronunciation defect report generation feature vector; The optimized pronunciation defect report generates a feature vector which is passed through a generator to generate a student pronunciation training defect report.
8. The teaching management method based on big data analysis according to claim 7 is characterized in that: The pronunciation defect report generation feature vector is subjected to a lightweight feature dynamic activation response optimization based on structural constraint perception to obtain an optimized pronunciation defect report generation feature vector, including: Extracting the pronunciation defect structured constraint matrix; Performing core prior knowledge denoising and filtering on the pronunciation defect structured constraint matrix to obtain a set of core prior information structure perception coding vectors of the pronunciation defect model; Constructing a pronunciation defect model core prior information nonlinear coupling matrix between the pronunciation defect report generation feature vector and each pronunciation defect model core prior information structure perception coding vector in the set of the pronunciation defect model core prior information structure perception coding vector to obtain a set of pronunciation defect model core prior information nonlinear coupling matrices; Calculating the pronunciation defect prior information response dynamic activation pattern factor of each pronunciation defect model core prior information nonlinear coupling matrix in the set of the pronunciation defect model core prior information nonlinear coupling matrix to obtain a set of pronunciation defect prior information response dynamic activation pattern factors; Based on the set of dynamic activation pattern factors of the pronunciation defect prior information response, a set of nonlinear coupling matrices of the core prior information of the pronunciation defect model is subjected to lightweight feature fusion to obtain a projection coding matrix of the pronunciation defect model prior information response; The pronunciation defect report generation feature vector is mapped to the feature space of the pronunciation defect model prior information response projection coding matrix to obtain the optimized pronunciation defect report generation feature vector.
9. A teaching management system based on big data analysis, characterized in that: include: An English teaching management data acquisition module is used to acquire the student pronunciation training audio collected by a microphone and the vibration signal of the student pronunciation training process collected by a vibration sensor; An English teaching management data extraction module, used for extracting a student pronunciation training logarithmic Mel spectrum feature map and a student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process collected by the vibration sensor; The student pronunciation training defect report generating module is used to generate a student pronunciation training defect report based on the student pronunciation training logarithmic Mel spectrum feature map and the student pronunciation training vibration association feature map.
10. The teaching management system based on big data analysis according to claim 9 is characterized in that: The English teaching management data extraction module includes: Extracting a student pronunciation training logarithmic Mel spectrum feature map and a student pronunciation training vibration correlation feature map from the student pronunciation training audio collected by the microphone and the vibration signal of the student pronunciation training process collected by the vibration sensor, including: Performing feature extraction on the student pronunciation training audio collected by the microphone to obtain a logarithmic Mel frequency spectrum feature graph of the student pronunciation training; Performing frequency domain feature extraction on the vibration signal of the student pronunciation training process collected by the vibration sensor to obtain a frequency domain statistical feature graph of the student pronunciation training vibration; Extracting waveform features of the vibration signal of the student pronunciation training process collected by the vibration sensor to obtain a student pronunciation training vibration waveform feature graph; The student pronunciation training vibration frequency domain statistical feature graph and the student pronunciation training vibration waveform feature graph are associated to obtain the student pronunciation training vibration association feature graph.