Intelligent evaluation method for real-time learning and teaching effect feedback

By collecting and analyzing multimodal teaching scenario data, using ViT and variational automatic encoder to extract and reduce dimensionality features, identifying teaching behavior characteristics, the problem of insufficient objective and real-time evaluation of teaching effect in the existing technology is solved, and efficient and accurate teaching quality evaluation is achieved.

CN119941053AInactive Publication Date: 2025-05-06BEIJING FUTURE GENE EDUCATION TECH CO LTD

Patent Information

Application Number
CN202510427292.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for existing technology to achieve large-scale, objective and real-time teaching effect evaluation, and the evaluation results are often limited to surface data, making it difficult to deeply analyze the essential characteristics of teaching effect.

Method used

By collecting multimodal teaching scenario data (video, audio, timestamp), using a visual basic model and a variational automatic encoder based on ViT architecture for feature extraction and dimensionality reduction processing, identifying teaching behavior characteristics, student behavior characteristics and teaching interactive behavior characteristics, and generating a real-time teaching quality assessment report.

Benefits of technology

It realizes a comprehensive capture of classroom dynamics, improves the accuracy of feature extraction and behavior recognition, improves the real-time performance of the system, provides objective teaching quality assessment support, and helps teachers adjust their teaching strategies in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941053A_ABST
    Figure CN119941053A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent evaluation method for real-time learning and teaching effect feedback, and the method comprises the steps: carrying out the preprocessing of collected multi-modal teaching scene data of a target classroom, and obtaining a multi-modal data stream which comprises the aligned target video data, target audio data and a timestamp; pre-training the visual basic model based on the ViT architecture by adopting a plurality of education scene data sets to obtain a first model, carrying out alignment fusion on the first model and a preset variational automatic encoder structure to obtain a preset target model, and carrying out dimension reduction processing on the extracted spatial-temporal features and feature vectors by utilizing the model to obtain target features; and determining a classroom atmosphere index, a student participation index and a teaching interaction effect index based on teaching behavior characteristics, student behavior characteristics and teaching interaction behavior characteristics identified from the target characteristics, and synchronously generating a real-time teaching quality evaluation report. According to the method, the classroom information can be accurately and comprehensively analyzed, and a real-time and objective evaluation result is synchronously generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent education technology, and in particular to an intelligent evaluation method for real-time learning and teaching effect feedback. Background Art

[0002] The field of intelligent education is undergoing a digital transformation, in which real-time monitoring and evaluation of classroom teaching quality has become a key challenge. Traditional teaching evaluation mainly relies on manual observation and subjective evaluation, which makes it difficult to achieve large-scale, objective and real-time teaching effectiveness evaluation.

[0003] Currently common teaching evaluation methods include manual class inspections and basic video surveillance. Among them, manual class inspections involve teaching supervision experts regularly entering the classroom to conduct on-site evaluations and collect teaching data. Basic video surveillance uses fixed cameras to record the classroom for later manual review and analysis.

[0004] The more advanced solution uses a single video analysis technology to identify classroom behaviors, and uses computer vision algorithms to perform basic analysis of teachers’ teaching behaviors and students’ reactions. However, this solution uses a single deep learning model to process video data and can only achieve face detection and simple behavior recognition.

[0005] Therefore, the following problems exist in the existing technology: first, a single video analysis model is difficult to accurately capture complex classroom teaching scenarios, especially in dynamic scenarios such as teacher-student interaction; second, the existing system lacks efficient processing capabilities for high-dimensional teaching scene data, resulting in insufficient real-time performance; finally, the evaluation results are often limited to surface data, making it difficult to deeply analyze the essential characteristics of teaching effectiveness. Summary of the invention

[0006] In view of this, the embodiments of the present disclosure provide an intelligent evaluation method for real-time learning and teaching effect feedback, which can solve the problems existing in the prior art such as reliance on manual observation and subjective evaluation, and one-sided and inaccurate evaluation results.

[0007] In a first aspect, the present disclosure provides an intelligent evaluation method for real-time learning and teaching effect feedback, including: Collecting multimodal teaching scene data of a target classroom, and preprocessing the multimodal teaching scene data to obtain a multimodal data stream; the multimodal data stream includes aligned target video data, target audio data and a timestamp; Pre-training a visual basic model based on the ViT architecture using an educational scene dataset to obtain a first model; and aligning and fusing the first model with a preset variational autoencoder structure through an alignment loss function to obtain a preset target model; Extracting the spatiotemporal features of the target video data by using the first model; Extracting a feature vector of the target audio data; Performing dimensionality reduction processing on the spatiotemporal features and the feature vector based on a preset target model to obtain target features; Extracting teaching behavior features, student behavior features, and teaching interaction behavior features from the target features; Based on the teaching behavior characteristics, the student behavior characteristics, and the teaching interaction behavior characteristics, the classroom atmosphere index, the student participation index, and the teaching interaction effect index are determined, and based on the classroom atmosphere index, the student participation index, and the teaching interaction effect index, a real-time teaching quality evaluation report is generated.

[0008] In a second aspect, the embodiments of the present disclosure further provide an intelligent evaluation system for real-time learning and teaching effect feedback, including: A collection unit, used for collecting multimodal teaching scene data of a target classroom, and preprocessing the multimodal teaching scene data to obtain a multimodal data stream; the multimodal data stream includes aligned target video data, target audio data, and a timestamp; A model acquisition unit, used to pre-train a visual basic model based on the ViT architecture using a number of educational scene data sets to obtain a first model, and to align and fuse the first model with a preset variational autoencoder structure through an alignment loss function to obtain a preset target model; A target feature acquisition unit, configured to extract the spatiotemporal features of the target video data through the first model; An extraction unit, used to extract a feature vector of the target audio data; A target feature acquisition unit, used for performing dimensionality reduction processing on the spatiotemporal features and the feature vector based on a preset target model to obtain target features; An identification unit, used to identify teaching behavior characteristics, student behavior characteristics, and teaching interaction behavior characteristics from the target characteristics; A report synchronization generation unit is used to determine the classroom atmosphere index, the student participation index and the teaching interaction effect index based on the teaching behavior characteristics, the student behavior characteristics and the teaching interaction behavior characteristics, and to generate a real-time teaching quality evaluation report based on the classroom atmosphere index, the student participation index and the teaching interaction effect index.

[0009] The intelligent evaluation method for real-time learning and teaching effect feedback disclosed in this application can fully capture classroom dynamics through the collection and analysis of multimodal data; effectively improve the accuracy of feature extraction and behavior recognition by using the ViT architecture and variational autoencoder; effectively improve the real-time performance of the system through dimensionality reduction processing; provide objective data support based on the classroom atmosphere, student participation and teaching interaction effect indicators of feature recognition, and can generate teaching quality evaluation reports instantly and synchronously, which is helpful for timely adjustment of teaching strategies; the method disclosed in this application can generate evaluation reports in real time through data-driven, which can provide scientific basis for the improvement of subsequent teaching and management decisions during the teaching process; through continuous evaluation and feedback, it can effectively promote the continuous improvement of teaching quality and learning effects. Through this intelligent evaluation method, educational institutions and teachers can more effectively monitor and optimize the teaching process in real time and improve students' learning experience and results.

[0010] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0012] Figure 1 A flowchart of an intelligent evaluation method for real-time learning and teaching effect feedback provided in an embodiment of the present disclosure.

[0013] Figure 2 A flowchart of a method for acquiring a multimodal data stream provided in an embodiment of the present disclosure.

[0014] Figure 3 A schematic flow chart of a method for extracting spatiotemporal features of target video data provided in an embodiment of the present disclosure.

[0015] Figure 4 A schematic flow chart of a method for extracting a feature vector of target audio data provided in an embodiment of the present disclosure.

[0016] Figure 5 A flowchart of a method for identifying teaching behavior characteristics provided in an embodiment of the present disclosure.

[0017] Figure 6A flowchart of a method for identifying student behavior characteristics provided in an embodiment of the present disclosure.

[0018] Figure 7 A flowchart of a method for identifying teaching interaction behavior characteristics provided in an embodiment of the present disclosure.

[0019] Figure 8 A flowchart of a method for generating a real-time teaching quality assessment report provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0021] Reference Figure 1 , the present application discloses an intelligent evaluation method for real-time learning and teaching effect feedback, including: S100, collecting multimodal teaching scene data of the target classroom, and preprocessing the multimodal teaching scene data to obtain a multimodal data stream; the multimodal data stream includes aligned target video data, target audio data, and timestamp.

[0022] Specifically, the multimodal teaching scene data of the target classroom can be collected through a multi-angle audio and video acquisition system.

[0023] Through multimodal data, visual and auditory information in the classroom can be captured to more comprehensively reflect the classroom status.

[0024] S200 uses several educational scene data sets to pre-train the visual basic model based on the ViT (Vision Transformer) architecture to obtain a first model, and aligns and fuses the first model with a preset variational autoencoder structure through an alignment loss function to obtain a preset target model.

[0025] Specifically, it includes: 1) using several educational scene data sets to pre-train the visual base model based on the ViT architecture to obtain the first model, that is, the first model has learned many visual features related to educational scenes; wherein, the visual base model includes 12 transformer encoder layers, each transformer encoder layer uses 8 attention heads, and the hidden layer dimension of each transformer encoder layer is 768. Such a structural design enables the model to capture complex features and patterns in the image, that is, to extract high-level feature data from the image, with a shape of [B,768], where B is Batch size; 2) Construct a preset variational autoencoder (VAE) structure, in which the encoder and decoder both adopt a multi-layer perceptron (MLP) architecture, and the output dimension of the encoder is 256, which means that the encoder can map the input data to a 256-dimensional latent space representation. The latent space representation generated by the encoder is a low-dimensional, highly generalized feature representation used to capture the main features of the input data; 3) Align and fuse the first model with the variational autoencoder structure through the alignment loss function to obtain the preset target model, that is, in the preset target model, the first model and the variational autoencoder structure can work together.

[0026] Among them, the alignment loss consists of two parts: reconstruction loss and KL divergence. The reconstruction loss is calculated using mean square error (MSE), and the reconstruction loss measures the difference between the output regenerated from the latent space by the model and the original input. The other part of the alignment loss, KL divergence, is used to measure the difference between the latent space distribution and the standard normal distribution. The weight coefficient β of the KL divergence term is set to 0.01 to balance the reconstruction quality and the normalization of the feature distribution. The VA-VAE model (i.e., the preset target model) after initialization has the ability to process multimodal teaching scene data.

[0027] Regarding “aligning and fusing the first model with the variational autoencoder structure through an alignment loss function”, the specific fusion method includes: 1) inputting the image data x of the educational scene into the first model; 2) the first model extracts a high-level feature representation h with a shape of [B, 768]; 3) inputting the output feature h of the first model into the encoder of VAE; 4) the encoder generates the mean μ and variance σ of the latent space, both with a shape of [B, 256]; 5) sampling from the latent space distribution N(μ, σ) to obtain the latent variable z with a shape of [B, 256]; 6) inputting the latent variable z into the decoder of VAE; 7) the decoder generates the reconstructed image data , the shape is the same as the original input data; 8) Calculate the original input data x and the reconstructed output The mean square error (MSE) between them; formula: ; 9) Calculate the latent space distribution N(μ,σ) and the standard normal distribution The KL divergence between (0,I) includes the following formula: ; 10) Combine the reconstruction loss and KL divergence to get the alignment loss. The specific calculation formula includes: Alignment Loss=Reconstruction Loss+0.01⋅KL Divergence When the alignment loss is the smallest, it means that the alignment fusion of the first model and the variational autoencoder structure is completed, and the aligned fusion of the first model and the variational autoencoder structure is recorded as the preset target model.

[0028] In this step, the pre-trained visual basic model (i.e., the first model) is loaded to ensure that the model has learned rich visual features; the VAE structure (i.e., the variational autoencoder structure) is constructed, that is, the encoder and decoder of the multi-layer perceptron architecture are used to generate the latent space representation; the alignment loss function is designed, specifically through the reconstruction loss and KL divergence terms, to ensure the alignment of the first model with the VAE, and obtain the initialized VA-VAE model to achieve a balance between the reconstruction quality and the normalization of the feature distribution.

[0029] S300, extracting spatiotemporal features of target video data through a first model; Extracting feature vectors of target audio data; Based on the preset target model, the spatiotemporal features and feature vectors are reduced in dimension to obtain the target features.

[0030] By reducing the dimension,of features and improving the efficiency of subsequent analysis, the target features retain key,teaching and learning behavior information for further analysis.,S400, identify teaching behavior features, student behavior features,,teaching interaction behavior features from the target features.

[0031] The identification of specific behavioral characteristics can provide insight into the dynamics of teaching and learning.

[0032] S500 determines the classroom atmosphere index, student participation index and teaching interaction effect index based on teaching behavior characteristics, student behavior characteristics and teaching interaction behavior characteristics, and generates a real-time teaching quality evaluation report based on the classroom atmosphere index, student participation index and teaching interaction effect index.

[0033] Through this step, instant feedback can be achieved, that is, teachers can immediately understand the classroom effect and adjust teaching strategies. Evaluation reports based on objective data can help teachers and managers make more scientific decisions. In addition, continuous evaluation and feedback can promote the continuous improvement of teaching quality and learning effects.

[0034] The intelligent evaluation method for real-time learning and teaching effect feedback disclosed in this application can fully capture classroom dynamics through the collection and analysis of multimodal data; effectively improve the accuracy of feature extraction and behavior recognition by using the ViT architecture and variational autoencoder; effectively improve the real-time performance of the system through dimensionality reduction processing; provide objective data support based on the classroom atmosphere, student participation and teaching interaction effect indicators of feature recognition, and can generate teaching quality evaluation reports instantly and synchronously, which is helpful for timely adjustment of teaching strategies; the method disclosed in this application can generate evaluation reports in real time through data-driven, which can provide scientific basis for the improvement of subsequent teaching and management decisions during the teaching process; through continuous evaluation and feedback, it can effectively promote the continuous improvement of teaching quality and learning effects. Through this intelligent evaluation method, educational institutions and teachers can more effectively monitor and optimize the teaching process in real time and improve students' learning experience and results.

[0035] The intelligent evaluation method for real-time learning and teaching effect feedback disclosed in this application proposes a VA-VAE (Visually Based Model Aligned Variational Autoencoder) model. By aligning the pre-trained visual model with the variational autoencoder, the ability to extract teaching scene features is significantly improved; an improved latent diffusion strategy is designed to optimize the convergence speed of the DiT model in the high-dimensional latent space, and the real-time performance of teaching behavior recognition is improved; multimodal data acquisition and synchronous processing solutions are innovatively integrated to achieve all-round perception of teaching scenes; a hierarchical teaching effect evaluation system is constructed to support multi-dimensional and multi-level teaching quality evaluation.

[0036] Reference Figure 2 , the method for acquiring the multimodal data stream in S100 specifically includes: S110, deploys a multi-angle audio and video acquisition system covering the entire scene of the target classroom.

[0037] Among them, the multi-angle audio and video acquisition system includes a multi-angle audio acquisition system and a multi-angle video acquisition system. The multi-angle audio acquisition system is used to collect audio information in the entire scene of the target classroom, and the multi-angle video acquisition system is used to collect video information in the entire scene of the target classroom.

[0038] Specifically, the multi-angle audio acquisition system includes an omnidirectional microphone array deployed in the target classroom. Furthermore, an omnidirectional microphone array can be installed on the top of the classroom. Six omnidirectional microphones can be evenly distributed in a circular shape to ensure 360-degree audio acquisition without blind spots.

[0039] The multi-angle video acquisition system includes a 4K high-definition camera array that covers the entire scene of the target classroom. The array includes at least 3 cameras, which can be installed according to the layout plan in the three directions of "front, back, and side", as long as the full scene coverage of the classroom can be achieved.

[0040] Furthermore, a 4K high-definition camera can be installed in the front of the classroom, mainly used to capture students' positive expressions and behavioral characteristics; a 4K high-definition camera can be installed at the back of the classroom to capture the teacher's teaching behavior and blackboard content; a 4K high-definition camera can be installed on the side of the classroom to capture the overall teaching atmosphere and classroom interaction scenes; among them, the field of view of each camera is not less than 100 degrees to ensure maximum scene coverage.

[0041] Through the multi-angle audio acquisition system and multi-angle video acquisition system, you can get audio streams with a sampling rate of 48kHz and 4K high-definition video streams at 30fps.

[0042] S120, collecting original video data and original audio data of the entire scene based on the multi-angle audio and video collection system.

[0043] Specifically, the original audio data in the entire scene of the target classroom is collected through a multi-angle audio acquisition system, and the original video data in the entire scene of the target classroom is collected through a multi-angle video acquisition system.

[0044] S130, performing noise reduction processing, image enhancement processing, and video stabilization processing on the original video data in sequence to obtain target video data.

[0045] In this step, the original video data adopts a three-stage processing strategy; the first is noise reduction processing. Specifically, a Gaussian filter can be used to suppress noise in each frame of the original video data. The Gaussian kernel size is set to 5×5, and the standard deviation σ is dynamically adjusted to 0.5-1.5 to balance the noise reduction effect and detail retention. The Gaussian filter is a spatial filter that performs a weighted average on the neighborhood of each pixel in the image. The weight is determined by the Gaussian function. In this process, the video is decomposed into a series of frames, and each frame is treated as an image for processing. Therefore, although the input is a video stream, the processing is performed frame by frame, and each frame is processed as an independent image.

[0046] The second stage is image enhancement processing, which can be done by using adaptive histogram equalization technology (CLAHE). Each frame of the image (i.e., single-frame image level) is divided into 8×8 grids, and histogram equalization is performed on each grid. The contrast threshold is set to 3.0 to improve the contrast and clarity of the image.

[0047] The third stage is video stabilization processing. The motion vector between adjacent frames is calculated based on the Lucas-Kanade optical flow algorithm, and then the motion trajectory is smoothed using a Kalman filter. Finally, frame alignment is performed through affine transformation, and the smoothing window size is set to 15 frames.

[0048] In this step, noise reduction and image enhancement are performed on each frame. The third stage of video stabilization is performed on the continuous frames of the original video data after the second stage of processing. The motion vectors between adjacent frames are calculated, and the Kalman filter is used to smooth these motions. Then, an affine transformation is applied to align the frames to achieve video stabilization. This process needs to consider the relationship between frames, so it is processed for the entire video sequence.

[0049] S140, performing noise elimination processing, sound source localization processing, and speech enhancement processing on the original audio data in sequence to obtain target audio data.

[0050] In this step, the processing of the original audio data includes three key links, which is a multi-step audio quality improvement process, the purpose of which is to ensure that clear speech signals can be accurately identified in subsequent analysis; first, environmental noise is eliminated, and an adaptive noise cancellation algorithm (ANC) can be used. A 32-order FIR filter is used to build a noise reference model, and the adaptive step size μ is set to 0.01 to effectively suppress steady-state environmental noise.

[0051] Next, sound source localization can be performed by using the time delay characteristics (TDOA) of the microphone array combined with the generalized cross-correlation algorithm (GCC-PHAT) to determine the sound source position. The time window length is set to 256ms, which is helpful for subsequent speech enhancement processing. That is, setting a time window can help ensure data consistency when calculating cross-correlation, thereby locating the sound source more accurately.

[0052] The last part is speech enhancement. For this purpose, a speech enhancement algorithm based on Wiener filtering can be used. The frame length is set to 20ms and the frame shift is 10ms. That is, the audio data is divided into small segments of 20ms in length for processing, and the shift is 10ms each time. This can effectively process continuous audio streams. At the same time, the voice activity detection (VAD) technology is combined to distinguish between the speech part and the non-speech part, so that the processing of the speech signal is more centralized and efficient, thereby improving the clarity and intelligibility of the speech signal.

[0053] Through the above processing scheme, a clearer and more understandable speech signal can be obtained, which helps to improve the accuracy of subsequent speech recognition and analysis tasks.

[0054] S150, aligning the target video data and the target audio data based on the timestamp with millisecond-level time accuracy to obtain a standardized multimodal data stream.

[0055] Specifically, a unified timestamp is assigned to each video frame in the target video data and each audio frame in the target audio data to accurately mark the time point of each frame of data, which is convenient for time synchronization and comparison; the time accuracy reaches the millisecond level, which can effectively ensure high-precision synchronization of data; a sliding window mechanism is adopted, and the window size is set to 100ms, and the target video data and target audio data are cached within this window.

[0056] Then, the dynamic time warping (DTW) algorithm is used to calculate the optimal alignment scheme between different modal data streams. The maximum allowable time deviation is ±20ms, that is, a certain small deviation in the time synchronization between the target video data and the target audio data is allowed, but these deviations cannot exceed 20ms.

[0057] Finally, the aligned multimodal data are encapsulated into a unified data packet, each of which contains the video frame, audio frame and timestamp information of the corresponding time point.

[0058] Furthermore, for possible data loss, linear interpolation can be used to supplement it to ensure the continuity of the data flow and provide a high-quality data foundation for subsequent steps.

[0059] Reference Figure 3 , the method for extracting the spatiotemporal features of the target video data in S300 specifically includes: A100, segment the target video data into continuous frames to obtain a plurality of sub-video information.

[0060] The number of frames of each sub-video information is 16, that is, the target video data is segmented into units of 16 frames; segmentation can reduce the amount of data processed at a single time, making it easier for the model to process and analyze each video, and can focus more on the local features in each sub-video segment, avoiding missing important details due to processing large amounts of data, while being able to capture local features and dynamic changes in the video.

[0061] A200, input each sub-video information into the first model to obtain the spatiotemporal features of each video.

[0062] Specifically, the target video data is segmented into 16 frames, and then the first model is used to output the spatiotemporal features of each video segment as a feature matrix with a dimension of 768×16, that is, 768-dimensional features are extracted for each frame, and 16 768-dimensional features are generated for 16 frames. In this step, the model processes each sub-video segment and outputs the spatiotemporal features of each sub-video segment. These features include spatial information (such as the position, shape, color, etc. of the object) and temporal information (such as the movement and change of the object, etc.) in the video.

[0063] Reference Figure 4, the method for extracting the feature vector of the target audio data in S300 specifically includes: B100, performing short-time Fourier transform on the audio data in the target audio data corresponding to each segment of sub-video information to obtain a corresponding sub-spectrum diagram.

[0064] Specifically, the target video data is segmented into 16-frame units, and the audio data corresponding to each video segment is converted into a spectrogram (i.e., a sub-spectrogram) through short-time Fourier transform (STFT).

[0065] B200, uses a 2D convolutional network to extract the features of each sub-spectrum graph and obtain the feature vector of each sub-spectrum graph.

[0066] The feature vector of each sub-spectrum graph is a feature vector with a dimension of 256.

[0067] Specifically, the 2D convolutional network gradually extracts high-order features from the sub-spectrogram through multiple convolutional layers and pooling layers, and finally generates a feature vector that can capture the time-frequency characteristics of the audio signal, such as spectral energy distribution, time changes, etc.

[0068] By performing short-time Fourier transform and 2D convolutional network feature extraction on the target audio data, the accuracy, robustness and efficiency of feature extraction can be significantly improved, providing strong support for subsequent multimodal data fusion and advanced analysis.

[0069] In this embodiment, the integration method of the preset denoising diffusion probability model (DenoisingDiffusionProbabilisticModels, DDPM) into the preset target model specifically includes: 1) Use clean multimodal data to train the DDPM model and learn the generative distribution of data; 2) Sample and generate new synthetic data from the trained DDPM, mix the generated synthetic data with real data, and use it to train the preset target model to increase data diversity; 3) Connect the trained DDPM and the preset target model to form an end-to-end model, train the DDPM and the preset target model at the same time, and update the parameters of the two models through back propagation, so that the denoising features generated by the DDPM can better optimize the performance of the target model. After both models are trained, the integration operation is completed.

[0070] The method for acquiring the target feature in S300 specifically includes: The spatiotemporal features and feature vectors are input into the encoder of the preset target model, and the dimension is reduced through two layers of MLP to obtain the reduced dimension feature vector (i.e., the target feature).

[0071] Specifically, the spatiotemporal features and feature vectors are sent to the encoder of VA-VAE, and dimensionality reduction is performed through two layers of MLP (with dimensions of 512 and 256 respectively), and finally a high-dimensional feature vector of 256 dimensions is obtained in the latent space. That is, the first layer of MLP reduces the features from a higher dimension to 512, and the second layer of MLP further reduces the features to 256, and finally a high-dimensional feature vector of 256 dimensions is obtained in the latent space.

[0072] Furthermore, in order to maintain the temporal relationship of the features, a position encoding mechanism is introduced in the encoding process, and a sinusoidal position encoding method is used to ensure that the features can retain the temporal information.

[0073] Furthermore, feature consistency loss can be introduced to ensure that the optimized features are consistent with the original features in the semantic space. The weight λ of the consistency loss is set to 0.1. The weight λ is used to balance the relative importance between feature consistency loss and other model losses (such as reconstruction loss, etc.), so that the model can take into account the semantic consistency with the original features while optimizing the features.

[0074] Through the synergy of the above steps, the VA-VAE model successfully achieved efficient feature extraction and optimization of teaching scene data. This solution not only significantly improved the expressiveness of features, but also improved the convergence speed of the model through an improved potential diffusion strategy. The feature optimization process fully considers the particularity of teaching scenes, can effectively capture key information in teaching activities, and provides high-quality feature representation for subsequent teaching behavior recognition and state analysis.

[0075] Reference Figure 5 , the method for identifying the teaching behavior characteristics in S400 specifically includes: C110, build an enhanced DiT model. The enhanced DiT model adopts a hierarchical Transformer structure. The enhanced DiT model contains 8 encoder layers, and each encoder layer is configured with 16 attention heads.

[0076] The constructed enhanced DiT model can effectively process high-dimensional feature data and capture complex temporal and spatial relationships.

[0077] Specifically, the construction method of the enhanced DiT model includes: 1) Determine the model architecture; Specifically, the model is composed of multiple encoder layers, each of which contains a multi-head self-attention mechanism and a feedforward neural network. The model as a whole adopts a hierarchical Transformer structure. Encoder layer configuration: 8 encoder layers in total, each with 16 attention heads.

[0078] 2) Set model parameters Input dimension: Assume that the dimension of the input sequence is input dim ; Hidden layer dimension: hidden dim , which is usually a multiple of the input dimension; Number of attention heads: num heads = 16; Feedforward neural network dimension: ffn dim , usually 4 times the hidden layer dimension; Layer Normalization: Layer normalization is applied before each sub-layer; Residual Connection: Add residual connections to the output of each sub-layer.

[0079] 3) Build the encoder layer Each encoder layer contains the following components: 3.1 Multi-Head Self-Attention Mechanism, whose input is a sequence tensor with a shape of [batch size , seq length , input dim ]; Operations: linear transformation to generate query, key, and value tensors; use multi-head attention mechanism to calculate attention weights; concatenate multi-head outputs and perform linear transformation to generate the final output; the output shape is [batch size , seq length , hidden dim ].

[0080] 3.2 Layer Normalization and Residual Connections; Specifically, the input is layer normalized and the output of the attention mechanism is added to the input.

[0081] 3.3 Feed-Forward Network, whose input is the output of residual connection, with shape [batch size , seq length , hidden dim The specific operation includes two linear layers with a ReLU activation function in between. The first layer expands the hidden layer dimension to ffn dim , the second layer shrinks the dimension back to hidden dim The output shape is [batch size ,seq length , hidden dim ].

[0082] 3.4 Layer Normalization and Residual Connections; Specifically, the input is layer normalized and residual connections are used to add the output of the feedforward network to the input.

[0083] 4) Stacking encoder layers; specifically, 8 encoder layers are stacked in sequence to form a complete encoder.

[0084] 5) Output layer: used to add the output layer to generate the final prediction results according to task requirements.

[0085] 6) Training and optimization; specifically including: 6.1) Loss function: select the appropriate loss function according to the task; 6.2) Optimizer: select optimizer such as Adam; 6.3) Learning rate scheduling: use the learning rate decay strategy.

[0086] 7) Adjustment and optimization; Specifically, adjust model parameters according to actual conditions, such as hidden layer dimensions, feedforward neural network dimensions, etc.; regularization techniques (such as dropout) can be introduced to prevent overfitting; pre-trained models can be used for fine-tuning to improve model performance.

[0087] Through the above steps, we can build an enhanced DiT model with a hierarchical Transformer structure, which contains 8 encoder layers and 16 attention heads in each layer. Such a model can effectively process sequence data and perform well in a variety of tasks.

[0088] C120, the target features are input into the enhanced DiT model, and the enhanced DiT model uses a sliding window to analyze continuous frames to obtain the initial behavior characteristics exhibited by the teacher at different time scales.

[0089] Among them, the length of the sliding window is greater than the window step size, and the correlation between frames is calculated through the self-attention mechanism.

[0090] In this embodiment, the length of the sliding window is preferably 32 frames, and the window step is preferably 8 frames. That is, in the enhanced DiT model, the continuous frame features are analyzed using a sliding window of length 32 through the temporal attention module, sliding once every 8 frames, and the features in the window are sent to the model for processing. Through the self-attention mechanism, the model can capture the initial behavioral characteristics of the teacher at different time scales.

[0091] C130, determines the attention weights of different spatial regions through the spatial attention module in the enhanced DiT model.

[0092] By calculating the attention weights of different regions, the spatial attention module is able to accurately track the teacher's position and movement trajectory, which helps capture the teacher's activities in different regions and enhance the accuracy of behavior recognition.

[0093] C140, identifying teaching behavior characteristics based on initial behavior characteristics and attention weights; Teaching behavior characteristics include one or more of explanation behavior, blackboard writing behavior, and interactive actions.

[0094] Specifically, by combining the initial behavioral characteristics and attention weights, it can accurately identify 15 basic teaching behaviors, including explanation postures (such as standing, walking, instructing, etc.), blackboard writing behaviors (writing, erasing, instructing, etc.), and interactive actions (asking questions, demonstrating, calling names, etc.), and determine the specific teaching behavior characteristics.

[0095] The enhanced DiT model can be used to identify teacher behaviors, which can not only accurately identify 15 types of basic teaching behaviors, but also improve the accuracy and robustness of recognition through methods such as temporal and spatial attention mechanisms and temporal smoothing algorithms. It is suitable for behavior analysis and monitoring in real-time teaching scenarios.

[0096] Furthermore, when the model identifies specific teaching behavior characteristics, it can generate corresponding confidence scores for each recognition result to measure the accuracy of recognition. At the same time, in order to eliminate the jitter phenomenon that may appear in the recognition results, a time series smoothing algorithm can be used to set the smoothing window size to 5 frames. This operation makes the recognition results more stable and reliable, effectively improves the quality of teacher behavior recognition, and provides solid data support for subsequent teaching evaluation and analysis.

[0097] Reference Figure 6 , the method for identifying the student behavior characteristics in S400 includes: C210, determine the target facial feature extraction network, the target facial feature extraction network includes the MobileNetV3 architecture, and the last three layers of the MobileNetV3 architecture are configured with an attention mechanism.

[0098] By introducing the attention mechanism, the network can focus on subtle changes in facial expressions and improve the ability to recognize subtle changes in facial expressions.

[0099] C220, inputs the target features into the target facial feature extraction network to obtain the students’ facial expressions.

[0100] In this step, the students' facial expressions can be used to analyze their seven basic emotional states, including concentration, confusion, boredom, interest, etc.

[0101] C230, inputs the target features into a three-layer fully connected network to obtain the students’ head pitch, head yaw, and head roll angles.

[0102] Specifically, a 6D rotation representation method is used to represent the head posture, and the pitch, yaw and roll angles of the head are predicted through a three-layer fully connected network; the three sets of angles (pitch, yaw, roll) of the head posture are output to determine the direction of students' attention.

[0103] C240, inputs the target features into a preset multi-resolution network, obtains several key point information of body posture, and determines the student's sitting status and behavior status based on several key point information.

[0104] Specifically, the key point information includes at least a nose key point, a left eye key point, a right eye key point, a left ear key point, a right ear key point, a left shoulder key point, a right shoulder key point, a left elbow key point, a right elbow key point, a left wrist key point, a right wrist key point, a left hip key point, and a right hip key point. In this embodiment, it preferably refers to the corresponding center position.

[0105] C250, builds a temporal state transfer model based on Markov chain, and uses students' classroom behavior history record data to train the temporal state transfer model to obtain the target transfer model.

[0106] Specifically, the dynamic change process of the student status is characterized by the Markov chain, and the transition probability matrix can be obtained by offline learning of historical data.

[0107] C260, inputs facial expression, head pitch, head yaw, head roll angle, sitting state, and behavioral state into the goal transfer model to obtain student behavior characteristics.

[0108] In this embodiment, each key point is generally represented by two coordinate values: 1) an X coordinate, which represents the horizontal position of the key point in the image; and 2) a Y coordinate, which represents the vertical position of the key point in the image.

[0109] Therefore, the coordinates of the 13 keypoints can be represented as a 2D array (or matrix) with a shape of 13 x 2, where each row corresponds to a keypoint and each column corresponds to the X and Y coordinates.

[0110] Based on the coordinate information of these key points, the students' sitting posture and behavior status can be further analyzed. The specific methods are as follows: 1) Sitting posture analysis specifically includes: Relative position of shoulders and hips: By analyzing the coordinates of the left shoulder, right shoulder, left hip and right hip, determine whether the student is sitting upright, leaning forward or leaning back; if the Y coordinates of the shoulders and hips are close, the student is sitting upright; if the Y coordinate of the shoulders is lower than the hips, the student may be leaning forward; if the Y coordinate of the shoulders is higher than the hips, the student may be leaning back; Position of knees and ankles: By analyzing the coordinates of the left knee, right knee, left ankle and right ankle, determine the student's leg posture; if the Y coordinates of the knees and ankles are close, the student may have his legs together; if the Y coordinate of the knees is higher than the ankles, the student may have his legs apart.

[0111] 2) Behavioral state analysis specifically includes: Arm and wrist position: By analyzing the coordinates of the left elbow, right elbow, left wrist and right wrist, it can be determined whether the student is writing with his hand, raising his hand or making other movements.

[0112] If the Y coordinate of the wrist is lower than the elbow, the student may be raising his hand; if the Y coordinate of the wrist is close to the height of the table, the student may be writing.

[0113] Head posture: By analyzing the coordinates of the nose, left ear, and right ear, the student's head posture is determined. If the X coordinates of the nose and ear are close, the student's head may be upright; if the X coordinate of the nose deviates from the ear, the student's head may be turned to one side.

[0114] 3) More complex behavioral analysis includes: Shoulder and hip rotation angles: By calculating the shoulder and hip rotation angles, it is possible to determine whether the student's overall posture is stable.

[0115] Body symmetry: By comparing key points of left-right symmetry (such as left and right shoulders, left and right hips, etc.), determine whether the student's body is symmetrical.

[0116] Reference Figure 7 , the method for identifying the teaching interaction behavior characteristics in S400 includes: C310, determine the interactive relationship between teachers and students or between students based on teaching behavior characteristics and student behavior characteristics.

[0117] C320, builds a dynamic interaction graph with teachers and each student as nodes. The edges of the dynamic interaction graph are interaction relationships. The weight of each edge is set to match the interaction intensity corresponding to each interaction relationship. The node characteristics of each node include one or more of teaching behavior characteristics and student behavior characteristics.

[0118] C330, inputs the dynamic interaction graph into the spatiotemporal graph convolutional network, extracts the spatiotemporal characteristics of the teacher-student interaction relationship through the spatiotemporal graph convolutional network, and obtains the interaction feature vector of each node.

[0119] Among them, the spatiotemporal graph convolutional network contains 4 layers of graph convolutional layers, each layer is followed by batch normalization and ReLU activation function.

[0120] C340,analyzes the interactive feature vector of each node through the LSTM-based temporal memory module to determine the characteristics of teaching interaction behavior.

[0121] Among them, the characteristics of teaching interaction behavior include one or more of teacher questions, student answers, group discussions, teacher patrols, students raising their hands, student discussions, teacher explanations, student independent learning, student distractions, and student interactions.

[0122] Specifically, an interaction history buffer with a capacity of 64 is maintained through an LSTM-based temporal memory module to capture long-term interaction patterns and historical information, and then identify the characteristics of teaching interaction behavior.

[0123] Furthermore, the corresponding behavioral characteristics can be confirmed through interaction duration, interaction participation, and interaction effect. The interaction duration is the duration of a single interaction scenario, the interaction participation can be measured by counting the proportion of students participating in the interaction, and the interaction effect can be evaluated by analyzing students' emotional changes and attention changes.

[0124] Furthermore, based on the duration, participation, and interaction effect, an interactive heat map can be created to vividly and intuitively display the active interactive areas and interactive modes in the classroom. The color depth in the interactive heat map should be matched with the interaction intensity. The interactive heat map can show which areas have frequent interactions and which interactive modes are dominant, helping teachers and managers understand the distribution and characteristics of classroom interactions.

[0125] The method for obtaining the classroom atmosphere index in S500 specifically includes: D110, based on teaching behavior characteristics, student behavior characteristics, and teaching interaction behavior characteristics, determine the distribution of students’ emotional states and classroom noise levels; D120, based on the preset emotional state weight, the preset noise level weight, the distribution of students' emotional states, and the classroom noise level, determine the classroom atmosphere index CAI (Classroom Activity Index).

[0126] Specifically, the emotional state score can be determined based on the distribution of students' emotional states; the noise score can be determined based on the classroom noise level; and finally, the classroom atmosphere index CAI can be determined based on the preset emotional state weight, the preset noise level weight, the emotional state score, and the noise score.

[0127] Specifically, the calculation is done by comprehensively considering the distribution of students' emotional states and the classroom noise level, using a weighted average method. For example, the emotional state weight is set to 0.7 and the noise level weight is set to 0.3. This means that when measuring classroom atmosphere, the influence of students' emotional state is relatively greater, and classroom noise level is also an important reference factor.

[0128] In order to make the weighted average calculation of the emotional state distribution and the noise level possible, the emotional state distribution needs to be normalized to the range of [0, 1]. The specific method is as follows: The proportion of each emotional state is mapped to a score (e.g., between 0 and 1), and the score is usually assigned based on the positivity of the emotional state; for positive emotions (e.g., focus, interest), the score is higher, which can be close to 1, and for negative emotions (e.g., boredom, confusion), the score is lower, which can be close to 0).

[0129] Calculate the weighted sum of emotional state scores: ,in, For the The proportion of emotional states, is the score of the corresponding emotional state, is the number of emotional states.

[0130] Then normalize the noise level: The noise level is usually a decibel value (dB), which needs to be converted to the [0, 1] range. The specific method is as follows: define a maximum and minimum value for the noise level (for example, the maximum value of classroom noise is 70 dB and the minimum value is 30 dB).

[0131] Normalization formula: , where the noise level range is [30 dB, 70 dB], after normalization The range is [0, 1].

[0132] CAI uses a weighted average approach to combine the emotional state distribution and the noise level: .

[0133] The specific formula is: .

[0134] For example, the distribution of students' emotional states includes: concentration: 50% (score 1.0), interest: 30% (score 0.8), confusion: 10% (score 0.3), other emotions: 10% (score 0.1); calculate the weighted sum of emotional state scores: ; The noise level is 50 dB, normalized noise level: ,but .

[0135] The CAI value range is [0, 1]. The higher the value, the more active the class is and the better the student participation is. By calculating CAI in real time, teachers can understand the activity of the class and the participation of students; CAI can be used as an indicator for teaching quality evaluation to help teachers adjust teaching strategies and improve classroom effectiveness.

[0136] In this embodiment, the calculation of CAI comprehensively considers the distribution of students' emotional states and the classroom noise level, and combines the two through a weighted average method to form a comprehensive classroom activity index. The specific calculation steps include normalization of the emotional state proportion and normalization of the noise level, and finally the CAI value is obtained through a weighted formula.

[0137] The method for obtaining the student engagement indicator in S500 specifically includes: D210, based on teaching behavior characteristics, student behavior characteristics, and teaching interaction behavior characteristics, determines the length of time students can maintain their attention, the enthusiasm of students in answering questions, and the students' classroom participation behavior.

[0138] D220 uses the cumulative calculation method of time-series decay to analyze the duration of students' attention, students' enthusiasm for answering questions, and students' classroom participation behavior to determine the student participation index SPI.

[0139] Specifically, SPI is an indicator that comprehensively evaluates students' classroom participation. It is calculated by analyzing students' attention span, their enthusiasm for answering questions, and their classroom participation behaviors. In order to reflect changes over time, the cumulative calculation method of time series decay is adopted, and the decay factor is set to 0.95. This method takes into account the impact of time factors on student participation. Over time, the contribution of early behaviors to the overall participation index will gradually decay.

[0140] Specifically, the input data includes the length of time students maintain their attention, the enthusiasm of students in answering questions, and the students' classroom participation behavior. The length of time students maintain their attention refers to the length of time students maintain a focused state for a certain period of time, usually measured in seconds or minutes; the enthusiasm of students in answering questions refers to the frequency and quality of students' answers in class, which can be measured by the number of answers, the correctness of the answers, or the teacher's feedback; the students' classroom participation behavior refers to the students' active participation in class, such as raising their hands to ask questions, participating in group discussions, etc.

[0141] The formula for the timing attenuation mechanism used is: SPI t =SPI t-1 × attenuation factor + current behavior contribution value; SPI t is the SPI value at the current moment, and the attenuation factor is preferably 0.95, which means that after each time unit, the impact of past behavior on the current SPI will decay to 95% of the original value. The closer the attenuation factor is to 1, the slower the impact of historical behavior decays; the closer it is to 0, the faster the impact of historical behavior decays.

[0142] Current behavior contribution value = (attention contribution value × weight of student’s attention maintenance time) + (answer question contribution value × weight of student’s enthusiasm for answering questions) + (participation behavior contribution value × weight of student’s classroom participation behavior).

[0143] Among them, attention contribution value = attention maintenance time / maximum attention maintenance time; specifically, the student's attention maintenance time is mapped to the range of [0, 1]. Assuming that the maximum attention maintenance time is 30 minutes (1800 seconds) and the minimum is 0 seconds, the normalized formula is: attention contribution value = attention maintenance time / 1800. If the attention maintenance time exceeds 1800 seconds, the value is 1.

[0144] Among them, the contribution value of answering questions = number of answers / maximum number of questions answered in each class; specifically, the students' enthusiasm for answering questions is mapped to the range of [0, 1]; assuming that a maximum of 10 questions are answered in each class, the normalized formula is: the contribution value of answering questions = number of answers / 10, if the number of answers exceeds 10 times, then the value is 1.

[0145] Among them, the contribution value of participation behavior = number of participations / maximum number of active participation behaviors in each class; specifically, the class participation behavior is mapped to the range of [0, 1]. Assuming that each class participates in a maximum of 5 active behaviors (such as raising hands to ask questions, group discussions), the normalized formula is: contribution value of participation behavior = number of participations / 5. If the number of participations exceeds 5 times, the value is 1.

[0146] Comprehensive calculation of the contribution value of current behavior: The contribution value of each behavior can be weighted and integrated. Assume that the weights are as follows: the weight of attention maintenance time is 0.5, the weight of enthusiasm for answering questions is 0.3, and the weight of classroom participation behavior is 0.2.

[0147] The calculation of SPI is a dynamic accumulation process, combining the time decay mechanism and the current behavior contribution value. At the initial moment (t=0), SPI0=0. After each time unit (such as every minute), SPI is updated according to the current behavior contribution value.

[0148] Sample calculation, assuming that SPI0=0 at the initial moment; In the first minute, the attention span is 60 seconds (contribution value = 60 / 1800 = 0.033), the number of questions answered is 1 (contribution value = 1 / 10 = 0.1), and the number of class participation is 0 (contribution value = 0 / 5 = 0). The current behavior contribution value is: (0.033×0.5)+(0.1×0.3)+(0×0.2)=0.0165+0.03+0=0.0465. Update SPI: SPI t =0×0.95+0.0465=0.0465.

[0149] In the second minute, the attention span is 120 seconds (contribution value = 120 / 1800 = 0.067), the number of questions answered is 2 times (contribution value = 2 / 10 = 0.2), and the number of class participation is 1 time (contribution value = 1 / 5 = 0.2). The current behavior contribution value is: (0.067×0.5)+(0.2×0.3)+(0.2×0.2)=0.0335+0.06+0.04=0.1335.

[0150] Update SPI: SPI t =0.0465×0.95+0.1335=0.177675; and so on, the SPI is updated every minute.

[0151] The SPI value range is [0, 1], and the higher the value, the higher the student engagement. By calculating SPI in real time, teachers can understand the classroom participation of each student. SPI can be used for personalized teaching, helping teachers adjust teaching strategies and improve student engagement.

[0152] SPI dynamically calculates students' classroom participation through the cumulative calculation method of time-series decay, combining students' attention span, enthusiasm for answering questions, and classroom participation behavior. The specific calculation steps include normalization and synthesis of behavioral contribution values, as well as the introduction of a time-series decay mechanism. Ultimately, the SPI value reflects students' classroom participation and provides data support for teaching optimization.

[0153] The method for obtaining the teaching interaction effect indicator in S500 specifically includes: D310, based on the teaching behavior characteristics, student behavior characteristics, and teaching interaction behavior characteristics, determine the teacher's questioning frequency, student response rate, and teacher-student interaction duration; D320, analyze the frequency of teachers’ questions, students’ response rate, and duration of teacher-student interaction to determine the teaching interaction effectiveness index TEI.

[0154] TEI (Teacher Engagement Index) is an indicator for evaluating the effectiveness of teachers' interaction with students in class. It is calculated based on the frequency of teachers' questions, students' response rate and interaction duration, and uses an adaptive threshold method to determine effective interaction. In addition, the system introduces a time series smoother, using the exponential moving average (EMA) method, with a smoothing window set to 5 minutes to ensure the stability of the indicator.

[0155] Specifically, the input data includes the teacher's question frequency, student response rate, and teacher-student interaction duration. The teacher's question frequency refers to the number of questions asked by the teacher within a certain period of time. It is usually measured in questions per minute. The student response rate refers to the proportion of students' responses to the teacher's questions, such as the proportion of students who answered questions to the total number of students. The interaction duration refers to the duration of each interaction, that is, the time period from the teacher's question to the student's answer.

[0156] In order to ensure the quality of interaction, adaptive thresholds can be set to determine which interactions are effective. Specifically, thresholds for question frequency, student response rate, and interaction duration can be set. The question frequency threshold can be set based on historical data. For example, a range of ±10% of the average question frequency is considered normal. For the student response rate threshold, for example, a response rate above 20% can be considered effective interaction. For the interaction duration threshold, a duration between 10 seconds and 5 minutes can be considered effective interaction.

[0157] It should be noted that the threshold can be dynamically adjusted based on real-time data and historical data to adapt to different classroom environments and teaching styles.

[0158] The number of times the teacher asks questions exceeding the question frequency threshold within a period of time is taken as the effective number of questions; the ratio of student response rate above the set threshold is taken as the effective response rate; and the total time of interaction within the set range is taken as the effective interaction duration.

[0159] TEI takes into account the number of effective questions, effective response rate and effective interaction duration, and is calculated by weighted average. The specific formula is as follows: TEI = (effective question weight × effective question score) + (response rate weight × response rate score) + (interaction duration weight × interaction duration score).

[0160] For weight distribution, you can set the effective question weight to 0.4, the response rate weight to 0.3, and the interaction duration weight to 0.3.

[0161] Valid question score = valid question number / maximum valid question number; in this embodiment, the valid question number can be normalized to the range of [0,1]. When the maximum valid question number is set to 10 times and the minimum is 0 times, the valid question score = valid question number / 10.

[0162] Response rate score = effective response rate / maximum response rate; for the response rate score, the effective response rate can be normalized to the range of [0,1]. When the maximum response rate is set to 100% and the minimum is 0%, the response rate score = effective response rate / 100%.

[0163] Interaction duration score = effective interaction duration / maximum interaction duration; for the interaction duration score, the effective interaction duration can be normalized to the [0,1] range. When the maximum interaction duration is set to 5 minutes and the minimum is 0 minutes, the interaction duration score = effective interaction duration / 5 minutes.

[0164] Furthermore, in order to reduce the volatility of the indicator, the exponential moving average (EMA) method is introduced for smoothing, and the smoothing window is set to 5 minutes.

[0165] EMA t =(current TEI value × smoothing coefficient) + (previous EMA value × (1 - smoothing coefficient)); smoothing coefficient α = 2 / (window size + 1). For a 5-minute window, α = 2 / (5+1) = 0.333. For the initial value, EMA0 can be set to the first TEI value.

[0166] Example calculation, assuming that the number of valid questions is 8, the effective response rate is 80%, and the effective interaction duration is 3 minutes; the effective question score = 8 / 10=0.8, the response rate score = 80% / 100%=0.8, and the interaction duration score = 3 / 5=0.6. TEI = (0.4×0.8) + (0.3×0.8) + (0.3×0.6) = 0.32 + 0.24 + 0.18 = 0.74.

[0167] Assuming the previous EMA value is 0.7, the current TEI value is 0.74, and the smoothing coefficient α is 0.333, then EMA t =(0.74×0.333) + (0.7×(1 - 0.333)) = 0.246 + 0.467 = 0.713. The TEI value range is [0,1]. The higher the value, the better the interaction between teachers and students. By calculating TEI in real time, teachers can understand their teaching interaction and adjust their teaching strategies to improve the quality of classroom interaction. TEI can be used as part of teacher performance evaluation to help teachers improve their teaching methods. TEI combines the frequency of teachers' questions, students' response rate and interaction duration, and uses an adaptive threshold to determine effective interaction. Then, the exponential moving average (EMA) method is used for time series smoothing to obtain a stable teacher interaction index. This index helps teachers understand and optimize the interactive effect of classroom teaching in real time.

[0168] Reference Figure 8 , the method for generating a real-time teaching quality evaluation report in S500 specifically includes: E100, establish an evaluation report template library, which includes a classroom performance overview section, a key event analysis section, and a teaching suggestion section.

[0169] The classroom performance overview section is used to provide a summary of the overall performance of the entire class, including the comprehensive scores and overall evaluation of key indicators; the key event analysis section is used to focus on important events or key moments in the class, such as teachers' questions, students' active participation, etc., and the impact of these events on teaching effectiveness; the teaching suggestions section is used to provide teachers with specific suggestions for improving teaching methods and improving the quality of classroom interaction based on data analysis. These templates provide a framework for the structure and content of the report to ensure the completeness and consistency of the report.

[0170] E200 uses natural language generation technology to convert classroom atmosphere indicators, student participation indicators, and teaching interaction effect indicators into corresponding descriptive texts.

[0171] Specifically, in order to make the data more understandable and intuitive, natural language generation technology can be used to convert classroom atmosphere indicators, student engagement indicators, and teaching interaction effect indicators into descriptive text. This means that instead of simply listing numbers and charts, the meaning and impact of these indicators are explained through text, allowing teachers and administrators to more easily understand the information behind the data.

[0172] For example, instead of displaying “Student engagement index is 85%”, the report can output: “Students were highly engaged in class, actively participated in discussions and answered questions, and showed interest and understanding of the course content.” E300, based on classroom atmosphere index, student participation index, and teaching interaction effect index, obtains the comprehensive score of first-level indicators, teacher teaching behavior analysis, student learning status analysis, classroom interaction effect analysis, key teaching moments, and abnormal situations.

[0173] In this step, the comprehensive score of the first-level indicators can be calculated based on previously collected indicators such as classroom atmosphere, student participation, and teaching interaction effect.

[0174] At the same time, the following analyses will also be conducted: by evaluating teachers' teaching methods, classroom management capabilities, and ways of interacting with students, we can obtain the results of teacher teaching behavior analysis; by analyzing students' concentration, level of understanding, and enthusiasm for participation in class, we can obtain the results of student learning status analysis; by evaluating the quality of interaction between teachers and students, including the frequency of questions and answers, and cooperative learning among students, we can obtain the results of classroom interaction effect analysis.

[0175] In addition, key teaching moments can be identified, that is, specific time points or events that have a significant impact on teaching effectiveness, as well as any abnormal situations, such as students' sudden loss of interest or problems with classroom order.

[0176] E400 generates a multi-level real-time teaching quality evaluation report based on the comprehensive score of the first-level indicators, analysis of teacher teaching behavior, analysis of student learning status, analysis of classroom interaction effects, key teaching moments, and abnormal situations.

[0177] Among them, the comprehensive score of the first-level indicators is used as the top-level structure in the real-time teaching quality evaluation report; The analysis of teacher teaching behavior, student learning status and classroom interaction effect is used as the multi-dimensional middle-level structure in the real-time teaching quality evaluation report; Use key teaching moments and abnormal situations as the underlying structure in real-time teaching quality evaluation reports.

[0178] In this embodiment, the top-level structure in the real-time teaching quality evaluation report represents the overall evaluation of the class, which provides teachers with an intuitive understanding of the overall situation of the class based on the comprehensive scores of the three first-level indicators: classroom atmosphere index (CAI), student participation index (SPI) and teaching interaction effect index (TEI). The middle-level structure in the real-time teaching quality evaluation report represents a detailed analysis of different dimensions, including an in-depth analysis of teachers' teaching behaviors, students' learning status and classroom interaction effects, helping teachers understand the classroom situation from different angles. The bottom-level structure in the real-time teaching quality evaluation report represents a record of specific events, covering key teaching moments, abnormal situations, etc., providing teachers with specific event references to facilitate their review and analysis of classroom details.

[0179] The real-time teaching quality evaluation report generation method aims to provide teachers with comprehensive and easy-to-understand teaching feedback through structured templates, natural language generation technology, and multi-dimensional data analysis. This method not only helps teachers understand classroom dynamics in real time, but also provides a scientific basis for continuous improvement of teaching methods.

[0180] Furthermore, it also includes an integrated intelligent recommendation engine to provide teachers with targeted improvement suggestions based on historical teaching data and current classroom performance.

[0181] Furthermore, it also includes displaying the three core indicators of CAI, SPI and TEI in a real-time monitoring area in the form of a dashboard. Specifically, a circular progress bar design can be used, and the color is dynamically adjusted as the indicators change.

[0182] Furthermore, it also includes the use of a variety of charts to display the statistical analysis area, such as using a line chart to show the changing trend of indicators over time, a bar chart to show the comparative analysis of various indicators, and a radar chart to show the balance of various dimensions of teaching.

[0183] Furthermore, it also supports interactive data exploration, where teachers can view historical data through the timeline and deeply analyze the teaching situation in a specific period through zooming and panning operations.

[0184] Furthermore, it also includes the use of a multi-level early warning mechanism and an intelligent recommendation system. Specifically, the early warning mechanism is set at three levels: prompt level (yellow warning), warning level (orange warning) and intervention level (red warning); real-time monitoring of the changing trends of various indicators, triggering an early warning when the indicator is lower than the preset threshold or an abnormal fluctuation occurs. The early warning rules use fuzzy logic reasoning, comprehensively considering the absolute value, change rate and duration of the indicator.

[0185] For example, a yellow warning is triggered when the student participation index is below 0.6 for 10 minutes, and an orange warning is triggered when it is below 0.4 for 20 minutes. The intelligent recommendation system can learn the optimal teaching strategy by analyzing historical teaching data based on the deep reinforcement learning model.

[0186] Furthermore, it also includes: according to the current classroom status, the most suitable improvement suggestions are selected from the teaching strategy library, including specific measures such as adjusting the teaching rhythm, optimizing the interactive method, and awakening attention. The suggestion generation adopts a personalized recommendation algorithm, taking into account the teacher's teaching style and the characteristics of the students to ensure the feasibility and pertinence of the suggestions.

[0187] The intelligent evaluation method for real-time learning and teaching effect feedback disclosed in the application uses chain logic design, and the output of each step is used as the input of the next step, ensuring the continuity of the data flow of the entire system and the integrity of the processing logic. At the same time, through clear data flow and conversion relationships, the various modules of the system can work closely and efficiently together.

[0188] Next, a specific classroom teaching scenario is used to illustrate the workflow of the entire system: Suppose in a high school physics class, Mr. Wang is explaining the knowledge point of "composition and decomposition of forces". The working process of the system is as follows: Data collection and preprocessing stage: The 4K camera at the front of the classroom captures the front view of the students, clearly recording the changes in each student's facial expressions and behavioral reactions. The camera at the back of the classroom records the whole process of Teacher Wang drawing a schematic diagram of the parallelogram law of forces on the blackboard, while the side camera captures the entire classroom scene. At the same time, the top microphone array captures the voice of Teacher Wang's explanation: "Let's look at this example, an object is acted upon by two forces..." The system preprocesses these raw data, removes ambient noise in the classroom (such as footsteps in the corridor), optimizes the clarity of the video, and accurately aligns the video and audio data to the millisecond level.

[0189] Feature extraction stage: The preprocessed data is input into the VA-VAE model. The system recognizes the content of Mr. Wang’s blackboard writing, including the schematic diagrams of various forces and related formulas he drew. At the same time, the model captures the teacher’s teaching behavior characteristics such as the rhythm of explanation and the layout of the blackboard. These features are optimized through the improved potential diffusion strategy to form a high-quality scene feature representation.

[0190] Behavior recognition and state analysis stage: Based on the extracted features, the system identified that Mr. Wang adopted the teaching mode of "demonstration-explanation-interaction". For example, after Mr. Wang finished drawing the force decomposition diagram, he turned around and asked: "How do you calculate the magnitude of these two component forces?" The system captured that Xiao Zhang, a student in the front row, raised his hand to answer the question, and at the same time found that Xiao Li, a student in the back row, was distracted for a short time. The system also identified the reactions of other students: some were taking notes seriously, some showed a thoughtful expression, and some were discussing in a low voice.

[0191] Teaching effect evaluation and feedback stage: The system calculates various teaching indicators in real time. When it detects that the attention of students in the back row is declining, the system intuitively displays the distribution of students' attention through a heat map on the teacher's interface and triggers a yellow warning. The system suggests to Teacher Wang: "It is recommended to increase classroom interaction and let students try to draw force decomposition diagrams at different angles on the blackboard." Teacher Wang adopted the suggestion and invited students to demonstrate on stage. The classroom atmosphere became noticeably more active, and the student participation index rose from 0.75 to 0.89.

[0192] The system generates a teaching report after class, pointing out that the highlights of this class are the teacher's standard blackboard writing, clear examples, and reasonable interactive design; the areas that need improvement are that the course rhythm can be more compact, and it is recommended to add real-life examples when explaining basic concepts. The report also includes complete classroom data statistics and visual charts to help Teacher Wang better understand his teaching results.

[0193] Through this example, we can see how the system achieves all-round intelligent monitoring and evaluation of classroom teaching through multimodal data collection, intelligent analysis and processing, real-time evaluation feedback, etc., and provides teachers with effective teaching support and improvement suggestions. The various modules of the system work closely together to form a complete closed loop for improving teaching quality.

[0194] In the prior art, a single video analysis model can only perform simple face detection and behavior recognition, and cannot effectively handle complex dynamic scenes such as teacher-student interactions, and lacks the ability to integrate multimodal data. The solution of this application collects multimodal teaching scene data (video, audio, timestamp) of the target classroom and performs preprocessing to form a multimodal data stream. This method not only covers visual information, but also introduces audio information, which can more comprehensively capture dynamic scenes in the classroom; through the combination of the visual basic model based on the ViT architecture and the variational autoencoder, it is possible to extract spatiotemporal features and audio feature vectors from video and audio, and perform deep fusion. This fusion of multimodal features can better reflect complex teaching behaviors and student interactions, and improve the ability to capture complex scenes.

[0195] In the prior art, the ability to process high-dimensional data is limited, resulting in deficiencies in the real-time performance of the system, making it difficult to cope with the analysis needs of large-scale teaching scene data. The present application scheme uses a variational autoencoder (VAE) to reduce the dimensionality of the extracted spatiotemporal features and audio feature vectors, compressing high-dimensional data into low-dimensional target features. This method not only reduces the complexity of data processing, but also retains the essential characteristics of the data, thereby improving the real-time performance of the system. The visual base model based on the ViT architecture has high computational efficiency and feature extraction capabilities, can quickly process and analyze large-scale video data, and improve the real-time response speed of the system.

[0196] In the prior art, the evaluation results can often only reflect the surface phenomena of the classroom (such as students' expressions or simple behaviors), and it is difficult to deeply analyze the deep-level characteristics of the teaching effect (such as the nature of teaching interaction, student participation, etc.). This application scheme can conduct an in-depth analysis of the classroom from multiple dimensions by extracting teaching behavior characteristics, student behavior characteristics, and teaching interaction behavior characteristics. For example, teaching behavior characteristics can reflect the teacher's teaching methods, student behavior characteristics can reflect the student's participation, and teaching interaction behavior characteristics can reflect the interaction effect between teachers and students. Based on the above characteristics, specific evaluation indicators can be generated, such as classroom atmosphere indicators, student participation indicators, and teaching interaction effect indicators. These indicators can not only reflect the surface phenomena of the classroom, but also reveal the deep-level characteristics of the teaching effect, providing a scientific basis for teaching improvement.

[0197] The solution of this application effectively solves the problems existing in the prior art through multimodal data collection and processing, application of advanced deep learning models, high-dimensional data dimensionality reduction processing, and multi-dimensional feature recognition and evaluation. Specifically: multimodal data fusion solves the problem that a single video analysis model is difficult to capture complex scenes; high-dimensional data dimensionality reduction processing improves the real-time performance of the system; multi-dimensional feature recognition and evaluation realizes the analysis of the deep-level characteristics of teaching effects, thereby providing a more comprehensive and scientific teaching quality evaluation. These innovations enable the solution of this application to accurately evaluate classroom teaching effects on a large-scale, objective and real-time basis, and promote the development of the field of intelligent education.

[0198] In a second aspect, the embodiments of the present disclosure further provide an intelligent evaluation system for real-time learning and teaching effect feedback, including: A collection unit is used to collect multimodal teaching scene data of a target classroom, and preprocess the multimodal teaching scene data to obtain a multimodal data stream; the multimodal data stream includes aligned target video data, target audio data, and a timestamp; A model acquisition unit, used to pre-train a visual basic model based on the ViT architecture using a number of educational scene data sets to obtain a first model, and to align and fuse the first model with a preset variational autoencoder structure through an alignment loss function to obtain a preset target model; A target feature acquisition unit, used for extracting the spatiotemporal features of the target video data through a first model; An extraction unit, used for extracting a feature vector of target audio data; A target feature acquisition unit is used to perform dimensionality reduction processing on the spatiotemporal features and feature vectors based on a preset target model to obtain target features; An identification unit, used to identify teaching behavior characteristics, student behavior characteristics, and teaching interaction behavior characteristics from target characteristics; The report synchronization generation unit is used to determine the classroom atmosphere index, student participation index and teaching interaction effect index based on the teaching behavior characteristics, student behavior characteristics and teaching interaction behavior characteristics, and generate a real-time teaching quality evaluation report based on the classroom atmosphere index, student participation index and teaching interaction effect index.

[0199] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

Claims

1. An intelligent evaluation method for real-time learning and teaching effect feedback, characterized in that: include: Collecting multimodal teaching scene data of a target classroom, and preprocessing the multimodal teaching scene data to obtain a multimodal data stream; The multimodal data stream includes aligned target video data, target audio data and a timestamp; The first model is obtained by pre-training the visual base model based on the ViT architecture using the educational scene dataset; And aligning and fusing the first model with a preset variational autoencoder structure through an alignment loss function to obtain a preset target model; Extracting the spatiotemporal features of the target video data by using the first model; Extracting a feature vector of the target audio data; Performing dimensionality reduction processing on the spatiotemporal features and the feature vector based on a preset target model to obtain target features; Extracting teaching behavior features, student behavior features, and teaching interaction behavior features from the target features; Based on the teaching behavior characteristics, the student behavior characteristics, and the teaching interaction behavior characteristics, the classroom atmosphere index, the student participation index, and the teaching interaction effect index are determined, and based on the classroom atmosphere index, the student participation index, and the teaching interaction effect index, a real-time teaching quality evaluation report is generated.

2. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 1 is characterized in that: The collecting of multimodal teaching scene data of the target classroom and preprocessing of the multimodal teaching scene data to obtain a multimodal data stream includes: Deploy a multi-angle audio and video acquisition system that covers the entire scene of the target classroom; Collecting original video data and original audio data of the whole scene based on the multi-angle audio and video acquisition system; The original video data is subjected to noise reduction processing, image enhancement processing, and video stabilization processing in sequence to obtain target video data; The original audio data is processed in sequence by noise elimination, sound source localization, and speech enhancement to obtain target audio data; The target video data and the target audio data are aligned based on a timestamp with millisecond-level time accuracy to obtain a standardized multimodal data stream.

3. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 2 is characterized in that: The step of extracting the spatiotemporal features of the target video data by using the first model includes: Segmenting the target video data into consecutive frames to obtain a plurality of sub-video information segments; The information of each sub-video is input into the first model to obtain the spatiotemporal features of each video.

4. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 3 is characterized in that: The step of extracting the feature vector of the target audio data comprises: Performing short-time Fourier transform on the audio data in the target audio data corresponding to each segment of the sub-video information to obtain a corresponding sub-spectrum graph; A 2D convolutional network is used to extract features of each sub-spectrum graph to obtain a feature vector of each sub-spectrum graph.

5. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 4 is characterized in that: The process of performing dimensionality reduction processing on the spatiotemporal features and the feature vectors based on a preset target model to obtain target features includes: inputting the spatiotemporal features and the feature vectors into an encoder of the preset target model, performing dimensionality reduction through a two-layer MLP, and obtaining target features.

6. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 1 is characterized in that: The method for identifying the teaching behavior characteristics comprises: The target features are input into the constructed enhanced DiT model, which uses a sliding window to analyze continuous frames to obtain the initial behavior features exhibited by the teacher at different time scales; the length of the sliding window is greater than the window step length, and the correlation between frames is calculated through a self-attention mechanism; Determining the attention weights of different spatial regions through the spatial attention module in the enhanced DiT model; Based on the initial behavior characteristics and the attention weights, teaching behavior characteristics are identified; the teaching behavior characteristics include one or more of explanation behavior, blackboard writing behavior, and interactive actions.

7. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 6 is characterized in that: The method for identifying the student's behavioral characteristics includes: Determine a target facial feature extraction network, wherein the target facial feature extraction network includes a MobileNetV3 architecture, and the last three layers of the MobileNetV3 architecture are configured with an attention mechanism; Inputting the target features into the target facial feature extraction network to obtain the student's facial expression; Inputting the target features into a three-layer fully connected network to obtain the student's head pitch, head yaw and head roll angles; Input the target features into a preset multi-resolution network to obtain a number of key point information of body posture, and determine the sitting state and behavior state of the student based on the key point information; Building a time series state transfer model based on Markov chain, and using students' classroom behavior history record data to train the time series state transfer model to obtain a target transfer model; The facial expression, the head pitch, the head yaw, the head roll angle, the sitting state, and the behavioral state are input into the target transfer model to obtain student behavioral characteristics.

8. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 7 is characterized in that: The method for identifying the teaching interaction behavior characteristics comprises: Determining the interactive relationship between teachers and students or between students according to the teaching behavior characteristics and the student behavior characteristics; A dynamic interaction graph is constructed with the teacher and each student as nodes, wherein the edges of the dynamic interaction graph are the interaction relationships, the weight of each edge is set to match the interaction intensity corresponding to each interaction relationship, and the node characteristics of each node include one or more of the teaching behavior characteristics and the student behavior characteristics; Inputting the dynamic interaction graph into a spatiotemporal graph convolutional network, extracting the spatiotemporal features of the teacher-student interaction relationship through the spatiotemporal graph convolutional network, and obtaining the interaction feature vector of each node; The interactive feature vector of each node is analyzed through the LSTM-based temporal memory module to determine the characteristics of teaching interactive behavior; The teaching interaction behavior characteristics include one or more of teacher questions, student answers, group discussions, teacher patrols, students raising hands, student discussions, teacher explanations, student independent learning, student distractions, and student interactions.

9. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 4 is characterized in that: The determining of the classroom atmosphere index, the student participation index and the teaching interaction effect index based on the teaching behavior characteristics, the student behavior characteristics and the teaching interaction behavior characteristics includes: Based on the teaching behavior characteristics, the student behavior characteristics, and the teaching interaction behavior characteristics, determine the distribution of students' emotional states, classroom noise level, students' attention span, students' enthusiasm for answering questions, students' classroom participation behavior, teacher's questioning frequency, student response rate, and duration of teacher-student interaction; Determining the classroom atmosphere index based on a preset emotional state weight, a preset noise level weight, the student emotional state distribution, and the classroom noise level; The cumulative calculation method of time series decay is used to analyze the duration of the student's attention, the student's enthusiasm for answering questions, and the student's classroom participation behavior to determine the student participation index; The teacher's question frequency, the student response rate, and the duration of the teacher-student interaction are analyzed to determine the teaching interaction effect index.

10. The intelligent evaluation method for real-time learning and teaching effect feedback according to claim 9, characterized in that: The generating of a real-time teaching quality evaluation report based on the classroom atmosphere index, the student participation index, and the teaching interaction effect index includes: Establishing an evaluation report template library, wherein the evaluation report template library includes a classroom performance overview section, a key event analysis section, and a teaching suggestion section; The classroom atmosphere index, the student participation index, and the teaching interaction effect index are converted into corresponding descriptive texts by using natural language generation technology; Based on the classroom atmosphere index, the student participation index, and the teaching interaction effect index, obtain the comprehensive score of the first-level index, the teacher's teaching behavior analysis, the student's learning status analysis, the classroom interaction effect analysis, the key teaching moments, and the abnormal situations; The comprehensive score of the first-level indicators is used as the top-level structure; The teacher's teaching behavior analysis, the student's learning status analysis, and the classroom interaction effect analysis are used as a multi-dimensional middle-level structure; Taking the key teaching moments and the abnormal situations as the underlying structure; Based on the top-level structure, the multi-dimensional middle-level structure and the bottom-level structure, a real-time teaching quality evaluation report is generated.

Citation Information

Patent Citations

  • Personalized active news recommending service system and method for mobile phone user

    CN102611785A

  • User portrait updating method, apparatus and system

    CN105005587A

  • Method and apparatus for obtaining Web browsing interest of user

    CN105930507A

  • Unmanned aerial vehicle group sensing data security sharing method based on federal learning

    CN113268920A

  • Image-based auto-encoder training method and device, equipment and storage medium

    CN116664444A

Cited By

  • Classroom behavior analysis method and system, electronic equipment and storage medium

    CN120180388A

  • Education robot voice signal processing method

    CN120472924A

  • Education data processing method based on coexistence of standard reference and norm reference

    CN120495026A

  • Education data processing method based on coexistence of standard reference and norm reference

    CN120495026B

  • Adaptive learning evaluation method and system based on large model driving

    CN120508791A