Resonant mode-based high school mathematics classroom learning state feedback method and device
By employing a resonant modal learning state feedback method, this approach utilizes a neural network model to extract features from video, audio, and behavioral event streams. Combined with adaptive weighted fusion, it overcomes the limitations of learning state analysis in high school mathematics classroom teaching, achieving high-precision learning state identification and real-time intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WENHUA UNIV
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-05
AI Technical Summary
In current high school mathematics classroom teaching, the methods for analyzing students' cognitive states have limitations. It is difficult to distinguish specific bottlenecks, the modal weights are not adjusted in real time, and typical errors in derivatives cannot be identified, resulting in insufficient teaching support.
We design a learning state feedback method based on resonant modalities. This method extracts multi-dimensional features from video, audio, and behavioral event streams using a neural network model, combines them with adaptive weighted fusion, identifies fine-grained learning cognitive states, and provides real-time intervention suggestions.
It achieves high-precision learning status recognition, provides reliable data support, improves the accuracy and timeliness of teaching, and can identify students' bottlenecks in the process of solving derivatives and make precise interventions.
Smart Images

Figure CN122155906A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of educational technology, specifically to a method and device for feedback on learning status in high school mathematics classrooms based on resonant modes. Background Technology
[0002] In classroom teaching of high school mathematics (for example, the knowledge point of "derivatives and solving for the maximum and minimum values of functions"), the inventors of this application have found that the existing technology has limitations in the methods for analyzing students' learning and cognitive states.
[0003] First, current technologies based on single gestures or voices are insufficient to distinguish the specific points where students get stuck in the reasoning chain of "differentiation-stationary point-monotonicity-maximum-minimum comparison", and it is difficult to associate the learner's calculation sketches, whispered derivations, etc., resulting in a disconnect between state judgment and the actual problem-solving process. Secondly, most existing methods adopt a fixed weight fusion strategy, which fails to adjust modal weights in real time based on classroom teaching links such as teacher drawing and student practice, as well as data quality, resulting in a decrease in the reliability of state recognition during function image analysis. Third, most systems stop at state classification and fail to associate with typical derivative error types such as "discussion with parameters" and "endpoint ignoring". They cannot identify when students are blocked in specific steps and provide step-by-step, hierarchical teaching support. Therefore, it is difficult to pay attention to the learning situation of each learner and implement precise intervention in the context of teaching.
[0004] Therefore, there is an urgent need for a technical solution that can deeply align with the cognitive path of derivative solving, integrate classroom information through neural coordination, and implement step-by-step, graded intervention to improve the accuracy and timeliness of high school mathematics teaching. Summary of the Invention
[0005] This application provides a method and device for feedback on learning status in high school mathematics classrooms based on resonant modes. By designing a non-intrusive, high-precision processing architecture that can identify fine-grained learning and cognitive states, it extracts features from multiple dimensions based on video streams, audio streams, and behavioral event streams collected in smart classrooms. Then, it combines the resonant mode feature vector obtained by adaptive weighted fusion to achieve accurate identification of various cognitive states. This can provide reliable data support for real-time intervention in high school mathematics classroom teaching and in-depth analysis of post-class learning.
[0006] Firstly, this application provides a high school mathematics classroom learning status feedback method based on resonant modes, the method including: The video stream, audio stream, and behavioral event stream of the current high school mathematics class are obtained by processing the video stream, audio stream, and behavioral event stream through the camera, microphone, and user interaction event logs on the teaching software, respectively. The first neural network model extracts features from the video stream to obtain the head pose angle sequence; the second neural network model extracts features from the audio stream to obtain speech feature vectors and key semantic labels; and the third neural network model extracts features from the behavioral event stream to obtain structured behavioral pattern vectors. The obtained multi-dimensional feature vectors are adaptively weighted and fused to obtain the resonant mode feature vector; The resonant modal feature vectors are used to learn cognitive state recognition through a classifier, and the severity score is calculated based on the learned cognitive state recognition results. Based on the severity score, the optimal intervention recommendation is determined and fed back to different subjects.
[0007] The second unit is a feedback device for high school mathematics classroom learning status based on resonant modes. The device includes: The data acquisition unit is used to acquire the video stream, audio stream, and behavioral event stream of the current high school mathematics class. The video stream, audio stream, and behavioral event stream are obtained by processing the user interaction event logs on the camera, microphone, and teaching software, respectively. The feature extraction unit is used to extract features from the video stream using a first neural network model to obtain a head pose angle sequence, to extract features from the audio stream using a second neural network model to obtain speech feature vectors and key semantic labels, and to extract features from the behavioral event stream using a third neural network model to obtain a structured behavioral pattern vector. The feature fusion unit is used to adaptively weight and fuse the obtained multi-dimensional feature vectors to obtain the resonant mode feature vectors. The recognition and calculation unit is used to learn the cognitive state recognition of the resonant mode feature vector through a classifier, and calculate the severity score based on the learned cognitive state recognition result. A feedback unit is defined to determine the optimal intervention recommendation based on the severity score and to provide feedback to different subjects.
[0008] Thirdly, this application provides a processing device, including a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the method provided in the first aspect of this application when it invokes the computer program in the memory.
[0009] Fourthly, this application provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the method provided in the first aspect of this application.
[0010] From the above, it can be concluded that this application has the following beneficial effects: To address the goal of providing high-quality feedback on learning status in high school mathematics classrooms, this application designs a non-intrusive, high-precision processing architecture capable of identifying fine-grained learning and cognitive states. Based on video streams, audio streams, and behavioral event streams collected in smart classrooms, it extracts features from multiple dimensions and then combines them with resonant modal feature vectors obtained through adaptive weighted fusion to achieve accurate identification of various cognitive states. This provides reliable data support for real-time intervention in high school mathematics classroom teaching and in-depth analysis of post-class learning. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a high school mathematics classroom learning status feedback method based on resonant modes, as described in this application. Figure 2 This is a schematic diagram of a classroom sensing device according to this application. Figure 3 This is a schematic diagram of one of the architectures for extracting speech feature vectors in this application; Figure 4 This is a schematic diagram of the processing architecture of the solution in this application; Figure 5 This is a schematic diagram of another architecture for the processing architecture of this application. Figure 6 This is a schematic diagram of a high school mathematics classroom learning status feedback device based on resonant modes, as described in this application. Figure 7 This is a schematic diagram of one type of processing equipment used in this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved.
[0015] The module division described in this application is a logical division. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual coupling, direct coupling, or communication connections may be through interfaces, and the indirect coupling or communication connections between modules may be electrical or other similar forms, none of which are limited in this application. Moreover, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed across multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this application.
[0016] Before introducing the high school mathematics classroom learning status feedback method based on resonant modes provided in this application, we will first introduce the background content involved in this application.
[0017] The method, device, and computer-readable storage medium for feedback on learning status in high school mathematics classrooms based on resonant modes provided in this application can be applied to processing devices. Targeting the goal of providing high-quality feedback on learning status in high school mathematics classrooms, this application designs a non-intrusive, high-precision processing architecture capable of identifying fine-grained learning and cognitive states. It extracts features from multiple dimensions based on video streams, audio streams, and behavioral event streams collected in smart classrooms, and then combines these with resonant mode feature vectors obtained through adaptive weighted fusion to achieve accurate identification of various cognitive states. This provides reliable data support for real-time intervention in high school mathematics classroom teaching and in-depth post-class learning analysis.
[0018] The resonant mode-based high school mathematics classroom learning status feedback method mentioned in this application can be implemented by a resonant mode-based high school mathematics classroom learning status feedback device, or by different types of processing devices such as servers, physical hosts, or user equipment (UE) that integrate the resonant mode-based high school mathematics classroom learning status feedback device. The resonant mode-based high school mathematics classroom learning status feedback device can be implemented in hardware or software. The UE can specifically be a terminal device such as a smartphone, tablet, laptop, desktop computer, or personal digital assistant (PDA). The processing devices can also be configured in a device cluster.
[0019] It is understandable that the proposed solution is usually based on existing data or data that has already been collected. Therefore, the processing equipment that implements the high school mathematics classroom learning status feedback method based on the resonant mode of this application or that is equipped with the corresponding application service of the high school mathematics classroom learning status feedback method based on the resonant mode of this application usually only needs to meet the required data processing capabilities. The specific equipment type and equipment deployment form are quite flexible.
[0020] If the direct acquisition of existing data mentioned above is also involved, then further hardware and software adaptation configurations are needed for the processing equipment to enable it to acquire data. For example, devices that acquire video streams, audio streams, and behavioral event streams can be included in the processing equipment cluster, or the processing equipment itself can be the control part of these devices. Alternatively, these external devices can be triggered to perform real-time data acquisition operations through third-party calls.
[0021] In addition, if there is a need to display the processing progress (including the processing results), the processing device itself can be configured with the required display screen (including touch screen) to display the specific content. Of course, the processing device can also display the specific content through an external display device or other devices with a display screen.
[0022] At the same time, this application also involves the use of teaching software on learning terminals in the application scenario of high school mathematics classroom. The teaching software itself can be independent of the system on the processing device, or it may be a nested system or an integrated system. In this case, it is quite flexible.
[0023] The following section introduces the high school mathematics classroom learning status feedback method based on resonant modes provided in this application.
[0024] First, refer to Figure 1 , Figure 1 This paper illustrates a flowchart of a high school mathematics classroom learning state feedback method based on resonant modes, as described in this application. The high school mathematics classroom learning state feedback method based on resonant modes provided in this application may specifically include the following steps S101 to S105: Step S101: Obtain the video stream, audio stream, and behavioral event stream of the current high school mathematics class. The video stream, audio stream, and behavioral event stream are obtained by processing the user interaction event logs on the camera, microphone, and teaching software, respectively. Understandably, the implementation of this application involves acquiring raw data streams from three aspects: video stream, audio stream, and behavioral event stream.
[0025] Video and audio streams are easy to understand; they are captured by cameras and microphones deployed in the classroom. User interaction event logs, on the other hand, are obtained by the teaching software installed on the student's terminal. This teaching software corresponds to the now-familiar smart classroom and digital learning. During class, students can learn by interacting with the teaching software and accessing the learning resources provided by the software.
[0026] In practice, the student terminal can be configured as a laptop, tablet, desk with a built-in computer, or desktop computer. Furthermore, this application solution can directly utilize existing student terminals without requiring further adaptation to those already configured in the classroom; that is, this application does not require the deployment of additional learning equipment. Similarly, the camera and microphone can be directly adopted from the sensing devices typically deployed in modern smart classrooms.
[0027] As a practical implementation method, refer to Figure 2 The schematic diagram shown here illustrates one scenario of the classroom sensing device of this application, including: 1.1) For cameras, specifically, it can include two high-definition cameras, namely Cam_Front_L and Cam_Front_R, deployed on both sides of the blackboard and covering the front half of the classroom. They can use wide-angle lenses to cover learners in the front half of the classroom. In addition, it can also include two high-definition variable-focus cameras, namely Cam_Rear_L and Cam_Rear_R, deployed in the middle of both sides of the classroom and covering the back half of the classroom. They can cover learners in the back half of the classroom and eliminate blind spots caused by front-row obstruction. In addition, it can also include the built-in camera on the student terminal, so as to collect close-up facial images as a supplement.
[0028] In practice, it can be done at a frame rate of 30fps. and 1920 1080 resolution To capture video streams .
[0029] And video stream In subsequent signal processing, illumination normalization and automatic white balance can be performed first to reduce the impact of changes in ambient lighting. Then, a lightweight face detector can be used to detect face regions in each frame of the image and output face bounding boxes. , Then, the face bounding box Crop the image by expanding the bounding box by 20% and then uniformly scale it to a fixed size of 128×128 pixels. To generate standardized facial image sequences. This serves as the specific input for the subsequent head pose estimation network.
[0030] Furthermore, the model training for this lightweight face detector uses bounding box regression loss, and this loss metric... It can be represented as: , in, To smooth out the conventional L1 loss, where i is the coordinate dimension, These represent the x-coordinate of the bounding box center point, the y-coordinate of the bounding box center point, the width of the bounding box, and the height of the bounding box, respectively. For the true value of the i-th coordinate dimension, This is the predicted value for the i-th coordinate dimension.
[0031] Furthermore, when the system is deployed in a multi-camera environment or needs to be evaluated across datasets, there may be inherent coordinate system biases between different data sources. In this case, the system can also use a rotation matrix alignment method to eliminate the corresponding system errors. The system error elimination process based on the rotation matrix alignment method can include the following processing: 1.1.1) Assume the set of rotation matrices predicted by the test set. There are corresponding real labels. Then there exists a fixed alignment rotation matrix Δ such that In this case, by solving To optimize the problem estimation Δ; 1.1.2) Solve effectively using iterative algorithms such as Karcher's mean.
[0032] After alignment, use Conduct performance evaluations to ensure that the results are fair and comparable.
[0033] All data is timestamped and transmitted to the central server via the classroom LAN.
[0034] 1.2) For the microphone, it can be deployed on the student terminal; Understandably, by learning the microphone array built into the terminal, specifically at a sampling rate of 16kHz... and 16-bit depth To synchronously capture audio streams .
[0035] audio stream In subsequent signal processing, the signal can first pass through a pre-emphasis filter (usually a first-order high-pass filter with a transfer function of...). Among them, the preset constant =0.97, z is a variable) to improve signal quality, followed by the Hamming window function ( Windowing is performed on each frame of the signal after windowing (where n is the sampling point index and N is the window length) to reduce spectral leakage. Voice Activity Detection (VAD) can then be performed on each frame, using short-time energy and zero-crossing rate to distinguish between speech and non-speech segments, retaining only the valid speech frames. This information is used for subsequent detailed analysis, such as providing a reference for the effective content location in the feature extraction work of the second and third neural network models.
[0036] 1.3) As described above, the user interaction event log specifically records the interaction events of students using the teaching software through the student terminal.
[0037] In practice, the event streams directly related to math lessons on finding the maximum and minimum values of derivatives can be captured in real time by monitoring the user interaction event logs of the teaching software on the learner's terminal. The events involved include, but are not limited to: 1.3.1) Recording of Extreme Values for Differentiation: Record the submission time, time taken, and correctness of the answer for the in-class problem of finding the extreme values for differentiation; 1.3.2) Content browsing events: Record the number of times the derivative chapter of "Selective Compulsory Course II" or "Step by Step - Mathematics" is searched, suggest video playback, suggest video pause and playback speed adjustment; 1.3.3) Annotation events: Record the content and location of text highlighting, function graphing and stationary point finding, adding notes during calculation, and adding annotations during calculation; 1.3.4) Mouse trajectory: Records the mouse's movement speed, trajectory complexity, and dwell points within a specific area, including the question area and answer area.
[0038] All events are timestamped and can be scheduled according to preset event windows. (For example, 10 seconds) aggregation is performed to achieve preliminary encoding and obtain a sequence of discrete-continuous hybrid behavior vectors. .
[0039] Step S102: The first neural network model is used to extract features from the video stream to obtain the head pose angle sequence; the second neural network model is used to extract features from the audio stream to obtain speech feature vectors and key semantic labels; and the third neural network model is used to extract features from the behavioral event stream to obtain structured behavioral pattern vectors. Understandably, this application has specifically built three neural network models based on machine learning methods to perform corresponding feature extraction processing on the preceding video stream, audio stream, and behavioral event stream, in order to obtain multi-dimensional feature vectors including head pose angle sequences, speech feature vectors, key semantic labels, and structured behavioral pattern vectors, which serve as inputs for subsequent feature fusion.
[0040] As an example of model types, the first neural network model can be the HPNet model, the second neural network model can be the ASNet model, and the third neural network model can be the BGNet model.
[0041] Furthermore, the specific tasks of the three models can be summarized as follows: 2.1) For the first neural network model, the following configuration may be included: Apply a lightweight face detector to the video stream The cropped face image sequence obtained by performing face bounding box localization A lightweight head pose estimation network with GhostNeXtNet as the backbone feature extractor is used to estimate head pose angles, resulting in a head pose angle sequence. , The lightweight face detector model is trained using bounding box regression loss, while the lightweight head pose estimation network employs the h-swish activation function to balance accuracy and computational efficiency. The lightweight head pose estimation network's terminals are connected to an Euler angle regression head and a 6D continuous representation regression head. The input 3D vector corresponding to the Euler angle regression head is represented as follows: The model training for the Euler angle regression head uses a combination of Euler regression and smoothing. The loss, 6D continuous representation, is used in the model training corresponding to the regression head, which employs geodesic loss. Specifically, the six-dimensional output vector corresponding to the 6D continuous representation regression head is mapped back to a complete, orthogonal vector through a Gram-Schmidt orthogonalization process. Rotation matrix .
[0042] Furthermore, for combining Eulerian regression and smoothing... Loss of loss , can be represented as: , , Where N is the total number of samples, and i is the sample index. To smooth out L1 loss, Let be the predicted value for the i-th sample. Let x be the true value of the i-th sample in the Euler complex space, and let x be the prediction error (i.e., the error between the predicted value and the true value).
[0043] Geodesic loss is directly in the rotation matrix. Metric errors on manifolds can effectively avoid the discontinuities and metric distortions in Euler angle representation under extreme conditions, and can be expressed as follows: , in, For geodesic loss, To estimate the rotation matrix, For the true rotation matrix, For the rotation trajectory, for and The relative rotation calculation is performed, where T is the matrix transpose.
[0044] 6D continuous representation represents the six-dimensional output vector corresponding to the regression head. ( The first two columns of the rotation matrix are mapped back to a complete, orthogonal 3×3 rotation matrix through the Gram-Schmidt orthogonalization process. , can be represented as: , , , , , in, For anti-flattening functions, To output a six-dimensional vector The first and second 3D vectors obtained by deflattening. Let T be three orthogonal unit vectors, and T be the matrix transpose.
[0045] Meanwhile, to meet the needs of real-time classroom analysis, this application can also adopt a dynamic frame skipping strategy for video streams, which involves the following processing: 1) If continuous The predicted head pose angle change within a frame (e.g., 10 frames) is less than a threshold. If the posture is stable, then the subsequent steps can be skipped. Frames (e.g., 30 frames) are used to skip subsequent frames once the attitude is stable; 2) If continuous If a frame (e.g., 5 frames) fails to detect a valid face or the confidence level of key points is too low, subsequent frames can be skipped. Frames are used to skip subsequent frames when key points are missing.
[0046] Thus, this application's system can dynamically select or fuse the two representations of the output based on the current classroom's real-time and accuracy requirements (e.g., when solving derivative maximum / minimum problems, learners need to concentrate highly during the first listening to the new lesson to hear and absorb the minor, error-prone points, and complete a comprehensive derivative-related question worth 12 points in class). Ultimately, the head feature vector... It is constructed as a vectorized representation of a normalized sequence of attitude angles or a rotation matrix.
[0047] 2.2) For the second neural network model, the overall process involves speaker separation, speech recognition, and sentiment analysis of the audio stream, and may include the following configuration: 2.2.1) Global sentiment feature vector
[0048] In the speech feature extraction stage, acoustic features are extracted from each frame of speech containing mathematical terms related to self-discussion derivative evaluation for the entire speech segment. These acoustic features are then input into a BiLSTM-based acoustic encoder to capture the temporal dynamics of the speech and output frame-level emotion embeddings. The frame-level sentiment embeddings of the entire speech segment are aggregated through a temporal neural concerto pooling layer to generate a global sentiment feature vector for the entire speech segment. ; The feature extraction performed can extract the acoustic features of each frame of speech from mathematical terms related to derivative calculations, such as "On an open interval, first determine monotonicity... If there is no maximum or minimum value due to unidirectionality; if there is a unique extreme point, then the extreme point is the maximum or minimum value."
[0049] The extracted acoustic features may include Mel-Frequency Cepstral Coefficients (MFCC, which takes 13 fundamental coefficients, first-order difference and second-order difference), fundamental frequency, short-time energy and spectral centroid, etc., to form the underlying acoustic feature vector.
[0050] BiLSTM stands for Bidirectional Long Short-Term Memory.
[0051] 2.2.2) Key semantic tags Speech feature vectors
[0052] In the semantic feature extraction stage, the entire speech segment is input into the Automatic Speech Recognition (ASR) engine to obtain transcribed text containing both mathematical terminology and the learner's speech-semantic emotion when solving derivative problems. ; For transcribed text Semantic representation vectors are extracted using an LLM model and a fine-tuned BERT model. ; For transcribed text Key semantic tags highly correlated with learning and cognitive states are extracted using the keyword extraction module. ; The global sentiment feature vector Semantic representation vector and key semantic tags Deep concerto fusion is performed using the cross-neural concerto module to obtain speech feature vectors. Key semantic tags and speech feature vectors As the model output (i.e., sentiment semantic features).
[0053] LLM stands for Large Language Model. An LLM model can process transcribed text... Targeted analysis is used to capture complex cognitive states that cannot be reflected by a single modality, such as: {You are a high school math classroom intelligent analysis system. Please analyze the following student speech content and output the structured analysis results.} { Student Intent: "Asking Questions (Methodological Assistance)" Related Knowledge Points: ["General Steps for Finding Maximum and Minimum Values of Derivatives"], "Affective Tendency": "Confused (Negative)" Key phrases: ["How to find", "maximum value"] Semantic Interpretation: "Learners may not have a clear understanding of the overall process framework for finding maximum and minimum values, and may not have systematically mastered the general problem-solving path of 'differentiation → finding stationary points / non-differentiable points → listing and analyzing monotonicity → comparing function values of all candidate points'." } For example: { "system_role": "You are a high school math classroom intelligent analysis system. Please analyze the following student voice content:", "student_utterance": "How exactly do I find the extreme value of this derivative?" "structured_analysis": { "student_intent": "Asking a question (clarifying method requirements)", "related_knowledge_points": ["Derivatives and Solving for Maximum and Minimum Values of Functions", "Steps for Solving for Maximum and Minimum Values on Closed Intervals", "Relationship between Extreme Values and Maximum / Minimum Values"] "emotional_tendency": "Confusion (negative)", "key_phrases": ["how to find", "maximum value"], Interpretation: "Students showed uncertainty about the overall method for finding extrema, possibly due to a lack of clear understanding of the general process of 'differentiation → finding stationary and non-differentiable points → analyzing monotonicity → comparing function values at candidate points.' Their emotional confusion suggests they may be experiencing a cognitive bottleneck." } } Among them, the semantic representation vector output by LLM Specifically, it can contain structured tags. With sentiment score .
[0054] The fine-tuned BERT model is a BERT model pre-trained in a general domain and fine-tuned using educational domain corpora. BERT stands for Bidirectional Encoder Representations from Transformers. During the fine-tuning process, the final hidden state of the [CLS]-tagged layer can be extracted as the semantic representation vector of the text. .
[0055] Global sentiment feature vector Semantic representation vector and key semantic tags Deep concerto fusion is performed using the cross-neural concerto module to obtain speech feature vectors. ,refer to Figure 3 The schematic diagram shown here illustrates one possible architecture for extracting speech feature vectors in this application, which may include: , , , in, For learnable parameter matrix, Scaling factor It is a feedforward network.
[0056] Under these conditions, the corresponding conditions also include: , in, This is the cross-attention function.
[0057] 2.3) For the third neural network model, the following configuration can be included: This refers to the behavior event stream, which includes derivative maximum / minimum value recording, content browsing events, annotation events, and mouse trajectory events. According to the preset event window The resulting sequence of behavior vectors The behavioral pattern features are obtained by encoding using the T-GCN model and serving as the model output. The T-GCN model captures the spatial dependencies between behavioral nodes at each time step through the GCN layer, and then captures the dynamic evolution pattern of behavior along the time dimension through GRU to obtain the behavioral pattern feature processing results. For the GCN layer, coherent behavior is used as a node, and the node attributes are behavioral quantification features including frequency, duration and result. The edges between nodes are constructed by the temporal relationship of behavior occurrence.
[0058] T-GCN stands for Temporal Graph Convolutional Network. This application involves specific configuration work for the T-GCN model. The T-GCN model uses learners' actions such as taking notes on key, easily mistaken knowledge points during lectures (e.g., "first define the domain, then find the extreme points, and finally compare endpoints / extreme values"), reviewing common errors in the derivative chapter of the People's Education Press's "Selective Compulsory Course 2" textbook, and manually drawing function graphs in text to model nodes. These actions are represented as nodes, with node attributes being the quantitative characteristics of the behavior, such as frequency, duration, and result. Edges between nodes are constructed based on the temporal relationships of the behaviors, such as sequence and co-occurrence frequency. Thus, the T-GCN model captures the preceding audio stream at each time step through Graph Convolution Neural (GCN) layers. Spatial dependencies between behavioral nodes related to sentiment semantic features; Then, the dynamic evolution pattern of behavior is captured along the time dimension through the gated recurrent unit (GRU); Ultimately, T-GCN can output a low-dimensional, dense feature vector of behavioral patterns. This vector comprehensively represents the learner's learning engagement, pace, and strategies within the current time window.
[0059] Step S103: Adaptively weighted fuse the obtained multi-dimensional feature vectors to obtain the resonant mode feature vector; Understandably, the previous feature extraction process yielded multi-dimensional feature vectors, namely the head pose angle sequence. Key semantic tags Speech feature vectors and behavioral pattern characteristics Then, the adaptive weighted fusion can be performed here to achieve resonant mode fusion and obtain the corresponding resonant mode feature vector.
[0060] Specifically, this application also provides a practical implementation method. This adaptive weighted fusion (the corresponding processing model can be referred to as the ResGFNet model) can include the following processing at the detailed operation level: 3.1) Sequence of head pose angles Speech feature vectors and behavioral pattern characteristics The corresponding features are obtained by mapping to a common feature space through independent learnable linear projection layers. ,feature and characteristics , ; 3.2) Employing a gated neural concerto mechanism, different modalities, i.e., features, are represented. ,feature and characteristics Calculate a dynamic weight coefficient that depends on the current modality content. The corresponding representation is: , , in, , , , , and Learnable parameters that can be configured; This gated neural concerto mechanism can adaptively evaluate the reliability and information content of each modality at the current moment, and dynamically adjust the fusion strategy according to actual conditions such as audio quality degradation, thereby improving stability and reliability in real noisy classroom environments. For example, it can automatically reduce the weight of speech modalities when the environment is noisy.
[0061] 3.3) Using weighting coefficients For the original mode, i.e., the head pose angle sequence Key semantic tags Speech feature vectors and behavioral pattern characteristics A weighted summation is performed to fuse the features and obtain a unified feature representation as the feature vector of the resonant modes. The corresponding representation is: , in, , and It is a nonlinear transformation.
[0062] Step S104: The resonant modality feature vector is used to learn cognitive state recognition through a classifier, and the severity score is calculated based on the learned cognitive state recognition result. Thus, a unified feature representation of the resonant mode feature vector was obtained. In cases where it is easy to understand, it can be represented by a unified feature. As input, the classifier specifically identifies seven fine-grained learning cognitive states: deep focus, shallow focus, active thinking, confusion, acceptance, mild distraction, and severe distraction, determining which pre-defined fine-grained cognitive state the learner is currently in. probability distribution And in real time, select the state category with the highest probability. As output.
[0063] The learning state classifier comprises seven stages. As depth increases, the spatial size of the feature map gradually decreases, while the number of channels gradually increases. It can simultaneously capture the spatial path and global contextual information of the learner's notes—plotting function graphs, flipping through math textbooks / study materials, and self-discussion—including mathematical terminology and derivative-related statements. This can be represented as follows: .
[0064] At the same time, this application also designed a severity score to quantify the intensity or severity of cognitive states. , It is a comprehensive function of state probability, duration, and feature intensity; correspondingly, the severity score can be expressed as: , in, For severity score, The sigmoid function is used to normalize scores to the interval [0,1]. , and These are learnable weight parameters that can be optimized using historical feedback data. For state category The probability, For state category The duration has been continuous. This represents the characteristic intensity of the current resonant mode.
[0065] Step S105: Based on the severity score, determine the optimal intervention recommendation and provide feedback to different subjects.
[0066] Understandably, this application designs different intervention recommendations corresponding to different severity scores. Thus, through adaptation processing, the optimal intervention recommendation can be obtained for the current severity score, providing data support for more precise intervention recommendations while accurately identifying fine-grained learning and cognitive states.
[0067] As a practical implementation method, this application can divide the severity score into three intervention levels according to a dual threshold. It can also be based on the structured intervention strategy knowledge base maintained by the system. Search for optimal intervention recommendations The corresponding representation is: , in, For the similarity function, restructure the intervention strategy knowledge base. In the middle, each strategy Represented as a tuple, .
[0068] For using the double threshold The three intervention levels are divided into It is easy to understand, and includes: , , .
[0069] As an example, 0.4 can be taken. 0.7 can be taken.
[0070] Regarding the optimal intervention recommendation As can be seen, the state is quantified through a similarity function. The intervention recommendation is selected based on its matching degree with the strategies in the knowledge base, and the one with the highest similarity is taken as the optimal intervention recommendation. of.
[0071] Once the optimal intervention recommendation is determined, it can be output according to the corresponding output strategy, such as local storage, off-site storage, result display, result push, or further data analysis.
[0072] Taking the result push as an example, this application can push the current solution processing results to different target groups such as teachers, learners and parents, and through differentiated feedback content and methods, different target groups can be informed of the specific situation in a friendly manner.
[0073] Understandably, for different types of objects, as well as objects of the same type but with different circumstances, the specific design of differentiated feedback content and differentiated feedback methods can be further configured.
[0074] As one example, we can have: 5.1) On the classroom teacher's dashboard, a head alert will be displayed in an orange pop-up window, along with a concise text suggestion such as "Learner #3 seems to be confused about 'function monotonicity,' so we suggest paying attention." 5.2) Present learners with non-intrusive personalized prompts, encouraging words, or learning resources such as "'Points where the derivative does not exist but the function is defined' are a bit difficult? No problem! Click here to see the diagram again." 5.3) Push learning status information such as "good" and "good grasp of derivative evaluation knowledge points" to parents' mobile phone APP.
[0075] Regarding the various intervention suggestions that may be involved, this application also provides an example of a tuning table for the learning status profiles of learners in high school mathematics derivative classes at different times: Table 1 - Profiles and Tuning Tables of Learners' Learning Status in High School Mathematics Derivatives Classrooms at Different Times
[0076] Furthermore, it can be further combined Figure 4 and Figure 5 The diagrams shown below illustrate different architectures of the solution proposed in this application, providing a more intuitive understanding.
[0077] Finally, regarding the above solutions, overall, for the goal of providing high-quality feedback on learning status in high school mathematics classrooms, this application designs a non-intrusive, high-precision processing architecture capable of identifying fine-grained learning cognitive states. Based on video streams, audio streams, and behavioral event streams collected in smart classrooms, it extracts features from multiple dimensions and then combines them with resonant modal feature vectors obtained through adaptive weighted fusion to achieve accurate identification of various cognitive states. This can provide reliable data support for real-time intervention in high school mathematics classroom teaching and in-depth analysis of post-class learning.
[0078] In terms of details, there are: 1. A complete resonant language model is proposed, which is based on the constructed adaptive weighted fusion processing model (ResGFNet model) and is tuned to realize fine-grained perception of the learning status of high school mathematics students listening to lectures and doing classroom exercises. 2. By tuning three strongly related and complementary overall processing models (which can be referred to as the MAGNet model) – head pose tuning (first neural network model, HPNet), speech emotion and semantics tuning (second neural network model, ASNet), and log interaction behavior tuning (third neural network model, BGNet) – the system can capture complex cognitive states that cannot be reflected by a single modality. Through the gated neural concerto mechanism, the system can dynamically adjust the fusion strategy according to actual conditions such as audio quality degradation, thereby improving the stability and reliability in real noisy classroom environments. 3. It can fully utilize learners' existing terminal devices without deploying additional dedicated equipment, enabling natural perception of the cognitive state of listening to and solving problems related to finding the maximum and minimum values of derivatives. At the same time, the lightweight design ensures real-time processing capabilities, enabling timely warnings and recording of fine-grained state sequences. This helps teachers comprehensively grasp the class's learning situation and knowledge absorption during instruction, and provides multi-dimensional data support for post-class analysis and teaching evaluation, significantly improving the timeliness and effectiveness of intervention, reducing the burden on teachers, and providing feedback to parents on learners' learning status and specific problems, indirectly promoting the improvement of learners' mathematical learning abilities.
[0079] The above is an introduction to the high school mathematics classroom learning status feedback method based on resonant modes provided in this application. In order to facilitate the better implementation of the high school mathematics classroom learning status feedback method based on resonant modes provided in this application, this application also provides a high school mathematics classroom learning status feedback device based on resonant modes from the perspective of functional modules.
[0080] See Figure 6 , Figure 6 This is a schematic diagram of a high school mathematics classroom learning status feedback device based on resonant modes according to this application. In this application, the high school mathematics classroom learning status feedback device 600 based on resonant modes may specifically include the following structure: The data acquisition unit 601 is used to acquire the video stream, audio stream and behavioral event stream of the current high school mathematics class. The video stream, audio stream and behavioral event stream are obtained by processing the user interaction event logs on the camera, microphone and teaching software, respectively. The feature extraction unit 602 is used to extract features from the video stream through a first neural network model to obtain a head pose angle sequence, and to extract features from the audio stream through a second neural network model to obtain speech feature vectors and key semantic labels, and to extract features from the behavioral event stream through a third neural network model to obtain a structured behavioral pattern vector. The feature fusion unit 603 is used to adaptively weight and fuse the obtained multi-dimensional feature vectors to obtain the resonant mode feature vectors. The recognition and calculation unit 604 is used to learn the cognitive state recognition by passing the resonant mode feature vector through a classifier, and calculate the severity score based on the learned cognitive state recognition result. The feedback unit 605 is used to determine the optimal intervention recommendation based on the severity score and provide feedback to different subjects.
[0081] In one exemplary embodiment, the cameras specifically include two high-definition cameras deployed on both sides of the blackboard and covering the front half of the classroom, two high-definition zoom cameras deployed in the middle of both sides of the classroom and covering the rear half of the classroom, and a built-in camera on the student terminal. The microphones are specifically deployed on student terminals; The user interaction event log specifically records the interaction events that students engage with while using the teaching software through their student terminals.
[0082] In yet another exemplary embodiment, the first neural network model includes the following configuration: Apply a lightweight face detector to the video stream The cropped face image sequence obtained by performing face bounding box localization A lightweight head pose estimation network with GhostNeXtNet as the backbone feature extractor is used to estimate head pose angles, resulting in a head pose angle sequence. , The lightweight face detector model is trained using bounding box regression loss, while the lightweight head pose estimation network employs the h-swish activation function to balance accuracy and computational efficiency. The lightweight head pose estimation network's terminals are connected to an Euler angle regression head and a 6D continuous representation regression head. The input 3D vector corresponding to the Euler angle regression head is represented as follows: The model training for the Euler angle regression head uses a combination of Euler regression and smoothing. The loss, 6D continuous representation, is used in the model training corresponding to the regression head, which employs geodesic loss. Specifically, the six-dimensional output vector corresponding to the 6D continuous representation regression head is mapped back to a complete, orthogonal vector through a Gram-Schmidt orthogonalization process. Rotation matrix .
[0083] In yet another exemplary embodiment, the second neural network model includes the following configuration: In the speech feature extraction stage, acoustic features are extracted from each frame of speech containing mathematical terms related to self-discussion derivative evaluation for the entire speech segment. These acoustic features are then input into a BiLSTM-based acoustic encoder to capture the temporal dynamics of the speech and output frame-level emotion embeddings. The frame-level sentiment embeddings of the entire speech segment are aggregated through a temporal neural concerto pooling layer to generate a global sentiment feature vector for the entire speech segment. ; In the semantic feature extraction stage, the input speech segment will be integrated into the automatic speech recognition engine to obtain transcribed text containing both mathematical terminology and the learner's speech-semantic emotion when solving derivative problems. ; For transcribed text Semantic representation vectors are extracted using an LLM model and a fine-tuned BERT model. ; For transcribed text Key semantic tags highly correlated with learning and cognitive states are extracted using the keyword extraction module. ; The global sentiment feature vector Semantic representation vector and key semantic tags Deep concerto fusion is performed using the cross-neural concerto module to obtain speech feature vectors. Key semantic tags and speech feature vectors As model output.
[0084] In yet another exemplary embodiment, the third neural network model includes the following configuration: This refers to the behavior event stream, which includes derivative maximum / minimum value recording, content browsing events, annotation events, and mouse trajectory events. According to the preset event window The resulting sequence of behavior vectors The behavioral pattern features are obtained by encoding using the T-GCN model and serving as the model output. The T-GCN model captures the spatial dependencies between behavioral nodes at each time step through the GCN layer, and then captures the dynamic evolution pattern of behavior along the time dimension through GRU to obtain the behavioral pattern feature processing results. For the GCN layer, the relevant behavior is treated as a node, and the node attributes are behavioral quantitative features including frequency, duration and result. The edges between nodes are constructed by the temporal relationship of behavior occurrence.
[0085] In yet another exemplary embodiment, adaptive weighted fusion includes the following processing: Head pose angle sequence Speech feature vectors and behavioral pattern characteristics The corresponding features are obtained by mapping to a common feature space through independent learnable linear projection layers. ,feature and characteristics ; Employing a gated neural concerto mechanism, for characteristic ,feature and characteristics Calculate a dynamic weight coefficient that depends on the current modality content. The corresponding representation is: , , in, , , , , and These are learnable parameters; Using weighting coefficients For head pose angle sequence Key semantic tags Speech feature vectors and behavioral pattern characteristics A weighted summation is performed to fuse the features and obtain a unified feature representation as the feature vector of the resonant modes. The corresponding representation is: , in, , and It is a nonlinear transformation.
[0086] In yet another exemplary embodiment, it is represented by a uniform feature. As input, the classifier identifies seven fine-grained cognitive states: deep focus, shallow focus, active thinking, confusion, acceptance, mild distraction, and severe distraction, and selects the state category with the highest probability in real time. As output; The severity score is specifically expressed as follows: , in, For severity score, For the Sigmoid function, , and For learnable weight parameters, For state category The probability, For state category The duration has been continuous. The characteristic intensity of the current resonant mode; For the severity score, three intervention levels are determined based on a dual threshold. And based on the structured intervention strategy knowledge base maintained by the system. Search for optimal intervention recommendations The corresponding representation is: , in, For the similarity function, restructure the intervention strategy knowledge base. In the middle, each strategy Represented as a tuple, .
[0087] This application also provides a processing device from a hardware architecture perspective. As mentioned earlier, in practice, a processing device may exist as a device cluster. In this case, each device in the device cluster can also be referred to as a processing device. See [reference needed]. Figure 7 , Figure 7 This diagram illustrates a structural schematic of the processing device of this application. Specifically, the processing device may include a processor 701, a memory 702, and an input / output device 703. The processor 701 executes the computer program stored in the memory 702 to implement, for example... Figure 1 The corresponding steps of the high school mathematics classroom learning state feedback method based on resonant modes in the embodiment; or, when the processor 701 executes the computer program stored in the memory 702, it implements as follows: Figure 6 Corresponding to the functions of each unit in the embodiment, the memory 702 is used to store the functions executed by the processor 701 as described above. Figure 1 The corresponding embodiment includes the computer program required for the high school mathematics classroom learning status feedback method based on resonant modes.
[0088] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 702 and executed by processor 701 to complete this application. One or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a computer device.
[0089] The processing device may include, but is not limited to, processor 701, memory 702, and input / output device 703. Those skilled in the art will understand that the illustrations are merely examples of the processing device and do not constitute a limitation on the processing device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the processing device may also include network access devices, buses, etc., and processor 701, memory 702, input / output device 703, etc., are connected via a bus.
[0090] The processor 701 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the processing device, connecting various parts of the device through various interfaces and lines.
[0091] The memory 702 can be used to store computer programs and / or modules. The processor 701 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 702 and by calling data stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the processing device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0092] When processor 701 executes a computer program stored in memory 702, it can specifically perform the following functions: The video stream, audio stream, and behavioral event stream of the current high school mathematics class are obtained by processing the video stream, audio stream, and behavioral event stream through the camera, microphone, and user interaction event logs on the teaching software, respectively. The first neural network model extracts features from the video stream to obtain the head pose angle sequence; the second neural network model extracts features from the audio stream to obtain speech feature vectors and key semantic labels; and the third neural network model extracts features from the behavioral event stream to obtain structured behavioral pattern vectors. The obtained multi-dimensional feature vectors are adaptively weighted and fused to obtain the resonant mode feature vector; The resonant modal feature vectors are used to learn cognitive state recognition through a classifier, and the severity score is calculated based on the learned cognitive state recognition results. Based on the severity score, the optimal intervention recommendation is determined and fed back to different subjects.
[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described high school mathematics classroom learning state feedback device, processing equipment, and its corresponding units based on resonant modes can be found in the following reference: Figure 1 The description of the high school mathematics classroom learning status feedback method based on resonant modes in the corresponding embodiment will not be repeated here.
[0094] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0095] Therefore, this application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the present application. Figure 1 The steps of the high school mathematics classroom learning status feedback method based on resonant modes in the corresponding embodiment can be referred to as follows for specific operations. Figure 1 The description of the high school mathematics classroom learning status feedback method based on resonant modes in the corresponding embodiment will not be repeated here.
[0096] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0097] Because of the instructions stored in the computer-readable storage medium, the present application can be executed as described above. Figure 1 The steps of the high school mathematics classroom learning state feedback method based on resonant modes in the corresponding embodiment can therefore achieve the results of this application. Figure 1 The beneficial effects that the high school mathematics classroom learning status feedback method based on resonant modes can achieve in the corresponding embodiment are detailed in the preceding description and will not be repeated here.
[0098] The foregoing has provided a detailed description of the high school mathematics classroom learning status feedback method, device, processing equipment, and computer-readable storage medium based on resonant modes provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the core ideas of this application; furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A high school mathematics classroom learning status feedback method based on resonant modes, characterized in that, The method includes: The video stream, audio stream, and behavioral event stream of the current high school mathematics class are obtained, wherein the video stream, audio stream, and behavioral event stream are obtained by processing the user interaction event logs on the camera, microphone, and teaching software, respectively. The video stream is subjected to feature extraction by a first neural network model to obtain a head pose angle sequence; the audio stream is subjected to feature extraction by a second neural network model to obtain a speech feature vector and key semantic labels; and the behavioral event stream is subjected to feature extraction by a third neural network model to obtain a structured behavioral pattern vector. The obtained multi-dimensional feature vectors are adaptively weighted and fused to obtain the resonant mode feature vector; The resonant modal feature vectors are used to learn cognitive state recognition through a classifier, and a severity score is calculated based on the learned cognitive state recognition results. Based on the severity score, the optimal intervention recommendation is determined and fed back to different subjects.
2. The method according to claim 1, characterized in that, Specifically, the cameras include two high-definition cameras deployed on both sides of the blackboard and covering the front half of the classroom, two high-definition variable-focus cameras deployed in the middle of both sides of the classroom and covering the rear half of the classroom, and a built-in camera on the student terminal. The microphone is specifically deployed in the student terminal; The user interaction event log specifically records the interaction events of students using the teaching software through the student terminal.
3. The method according to claim 1, characterized in that, For the first neural network model, the following configuration is included: Apply a lightweight face detector to the video stream The cropped face image sequence obtained by performing face bounding box localization A lightweight head pose estimation network with GhostNeXtNet as the backbone feature extractor is used to estimate head pose angles, resulting in a head pose angle sequence. , The lightweight face detector model is trained using bounding box regression loss, and the lightweight head pose estimation network uses the h-swish activation function to balance accuracy and computational efficiency. The lightweight head pose estimation network is connected at its terminals to an Euler angle regression head and a 6D continuous representation regression head. The input 3D vector corresponding to the Euler angle regression head is represented as follows: The model training corresponding to the Euler angle regression head uses a combination of Euler regression and smoothing. The loss mechanism, where the 6D continuous representation of the regression head is trained using geodesic loss, specifically maps the output six-dimensional vector of the regression head back to a complete, orthogonal vector through a Gram-Schmidt orthogonalization process. Rotation matrix .
4. The method according to claim 3, characterized in that, The second neural network model includes the following configuration: In the speech feature extraction stage, acoustic features of each frame of speech in the entire speech segment are extracted, including mathematical terms related to the self-discussion of derivative evaluation. The acoustic features are input into a BiLSTM-based acoustic encoder to capture the temporal dynamics of speech and output frame-level emotion embeddings. The frame-level sentiment embeddings of the entire speech segment are aggregated through a temporal neural concerto pooling layer to generate a global sentiment feature vector for the entire speech segment. ; In the semantic feature extraction stage, the integrated speech segment is input into the automatic speech recognition engine to obtain transcribed text containing both mathematical terminology and the semantic and emotional aspects of the learner's speech when solving derivative problems. ; for the transcribed text Semantic representation vectors are extracted using an LLM model and a fine-tuned BERT model. ; for the transcribed text Key semantic tags highly correlated with learning and cognitive states are extracted using the keyword extraction module. The global sentiment feature vector The semantic representation vector and the key semantic tags Deep concerto fusion is performed using the cross-neural concerto module to obtain speech feature vectors. The key semantic tags and the speech feature vector As model output.
5. The method according to claim 4, characterized in that, The third neural network model includes the following configuration: This refers to the behavior event stream, which includes derivative maximum / minimum value recording, content browsing events, annotation events, and mouse trajectory events. According to the preset event window The resulting sequence of behavior vectors The behavioral pattern features are obtained by encoding using the T-GCN model and serving as the model output. The T-GCN model captures the spatial dependencies between behavioral nodes at each time step through the GCN layer, and then captures the dynamic evolution pattern of behavior along the time dimension through GRU to obtain the behavioral pattern feature processing result. For the GCN layer, the relevant behavior is used as a node, and the node attributes are behavioral quantification features including frequency, duration and result. The edges between nodes are constructed by the temporal relationship of behavior occurrence.
6. The method according to claim 5, characterized in that, The adaptive weighted fusion includes the following processing steps: The head posture angle sequence The speech feature vector and the behavioral pattern features The corresponding features are obtained by mapping to a common feature space through independent learnable linear projection layers. ,feature and characteristics ; The gated neural concerto mechanism is used to provide the features The features and the features Calculate a dynamic weight coefficient that depends on the current modality content. The corresponding representation is: , , in, , , , , and These are learnable parameters; Using the weighting coefficients For the head pose angle sequence The key semantic tags The speech feature vector and the behavioral pattern features A weighted summation is performed to fuse the features into a unified feature representation that serves as the feature vector of the resonant modes. The corresponding representation is: , in, , and It is a nonlinear transformation.
7. The method according to claim 6, characterized in that, Represented by the aforementioned unified features As input, the classifier identifies seven fine-grained learning and cognitive states: deep focus, shallow focus, active thinking, confusion, acceptance, mild distraction, and severe distraction, and selects the state category with the highest probability in real time. As output; The severity score is specifically represented as follows: , in, The severity score, For the Sigmoid function, , and For learnable weight parameters, For the state category The probability, For the state category The duration has been continuous. The characteristic intensity of the current resonant mode; Based on the severity score, three intervention levels are determined using a dual threshold method. And based on the structured intervention strategy knowledge base maintained by the system. Search for optimal intervention recommendations The corresponding representation is: , in, The similarity function is used in the structured intervention strategy knowledge base. In the middle, each strategy Represented as a tuple, .
8. A feedback device for high school mathematics classroom learning status based on resonant modes, characterized in that, The device includes: The data acquisition unit is used to acquire the video stream, audio stream, and behavioral event stream of the current high school mathematics class, wherein the video stream, audio stream, and behavioral event stream are obtained by processing the user interaction event logs on the camera, microphone, and teaching software, respectively. The feature extraction unit is used to extract features from the video stream using a first neural network model to obtain a head pose angle sequence, and to extract features from the audio stream using a second neural network model to obtain speech feature vectors and key semantic labels, and to extract features from the behavioral event stream using a third neural network model to obtain a structured behavioral pattern vector. The feature fusion unit is used to adaptively weight and fuse the obtained multi-dimensional feature vectors to obtain the resonant mode feature vectors. The recognition calculation unit is used to learn the cognitive state recognition of the resonant mode feature vector through a classifier, and calculate the severity score based on the obtained learning cognitive state recognition result; A feedback unit is defined to determine the optimal intervention recommendation based on the severity score and to provide feedback to different subjects.
9. A processing device, characterized in that, The method includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the method as described in any one of claims 1 to 7 when it invokes the computer program in the memory.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the method of any one of claims 1 to 7.