Method and system for evaluating classroom teaching effectiveness of teacher based on AI assistance
By using AI-assisted methods and leveraging the 3D-LCBQN neural network to process multimodal data on teachers' classroom teaching behavior and generate quantitative evaluation results, we have solved the problems of traditional assessments being time-consuming, labor-intensive, and highly subjective, and achieved real-time and objective evaluation of teaching effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DIGITAL TIMES BIG DATA TECH CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional methods for evaluating the effectiveness of classroom teaching rely on manual observation and scoring, which is time-consuming and makes it difficult to reflect the dynamic characteristics of teaching in real time, leading to problems with the consistency and objectivity of evaluation results.
Using an AI-assisted approach, a 3D-LCBQN neural network is used to generate quantitative sets of teacher and student behaviors by acquiring a denoised clean speech dataset, a clear image dataset optimized by lighting, and a filtered effective audio clip. These quantitative sets include effective teacher behaviors, effective student behaviors, and teacher-student classroom teaching task progress sets.
It enables real-time, multi-dimensional quantitative evaluation of teachers' classroom teaching activities and students' learning responses, reducing the time-consuming and subjective biases of manual operations and improving the objectivity and consistency of evaluation.
Smart Images

Figure CN121996972A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to an AI-assisted method and system for evaluating the effectiveness of teachers' classroom teaching. Background Technology
[0002] The field of instructional analytics technology primarily involves the research and application of systematic analysis of teachers' teaching activities and students' learning outcomes using educational data such as teaching behaviors and learning processes. Core aspects of this technology include instructional process data collection, instructional behavior modeling, instructional quality assessment, learning effectiveness analysis, and intelligent feedback mechanisms. The aim is to achieve a quantitative description of teaching patterns and data monitoring and evaluation of instructional quality through data-driven methods. Traditional methods for evaluating the effectiveness of teachers' classroom teaching refer to evaluations based on methods such as instructional observation scales. These methods typically involve qualitative or semi-quantitative analysis of teachers' classroom performance through manual observation and scoring. The evaluation process includes manual recording of classroom behavior, subjective judgment of teaching segments, and data processing of teaching results. It relies on manual operation, has a single data source, and the evaluation process is time-consuming and fails to reflect the dynamic characteristics of teachers' teaching in real time. Summary of the Invention
[0003] The purpose of this invention is to provide an AI-assisted method for evaluating the effectiveness of teachers' classroom teaching, which at least solves one of the aforementioned technical problems.
[0004] One aspect of the present invention provides a method for evaluating the effectiveness of teacher classroom teaching based on AI assistance, the method comprising:
[0005] Obtain the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments;
[0006] Obtain the trained 3D-LCBQN neural network;
[0007] The denoised clean speech dataset, the light-optimized clear image dataset, and the selected effective audio segments are input into the 3D-LCBQN neural network to generate a quantitative set of effective teacher behaviors, a quantitative set of effective student behaviors, and a quantitative set of effective behaviors in promoting classroom teaching tasks between teachers and students.
[0008] Optionally, obtaining the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments includes:
[0009] Acquire classroom visual data and classroom audio data;
[0010] The YOLOv8 object detection algorithm is used to identify classroom video data, thereby obtaining a frame-by-frame object detection result set;
[0011] Facial key points were extracted from classroom video data using Dlib facial key point extraction, thereby obtaining a frame-by-frame facial key point dataset.
[0012] A classroom original image dataset with frame numbers is generated based on the frame-by-frame target detection result set and the frame-by-frame facial key point dataset.
[0013] Preprocess the classroom audio data to obtain a time-stamped raw classroom speech dataset;
[0014] Multimodal data timestamp alignment is performed on the original classroom image dataset numbered by frame and the original classroom audio dataset with timestamps to obtain the timestamp-aligned original classroom image dataset and original classroom audio dataset.
[0015] A real-time denoising algorithm is performed on the original classroom audio dataset to obtain a clean audio dataset after denoising.
[0016] The image data is corrected for lighting using a lighting compensation algorithm, thereby obtaining a clear image dataset with optimized lighting.
[0017] By using speech semantic analysis algorithms to filter out invalid speech in the original classroom speech dataset, valid audio segments are obtained after filtering.
[0018] Optionally, the 3D-LCBQN neural network structure includes:
[0019] A multimodal basic feature extraction module is used to generate basic features based on the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered effective audio segments. The basic features include visual features, audio features, text features, and timestamp feature sets.
[0020] The temporal-semantic-knowledge point anchoring coding layer is used to generate visual anchoring coding features, audio anchoring coding features, text anchoring coding features, and a three-dimensional label set based on visual features, audio features, text features, and timestamp feature sets.
[0021] The teaching behavior-cognitive state linkage extraction layer is used to generate teacher guidance-behavior linkage features and student cognition-behavior linkage features based on the visual anchoring coding features, audio anchoring coding features, and text anchoring coding features.
[0022] The knowledge point progression-assessment dimension dynamic fusion layer is used to generate teacher-specific fusion features, student-specific fusion features, and teacher-student task progression dimension-specific fusion features based on visual anchoring coding features, audio anchoring coding features, and text anchoring coding features.
[0023] The indicator association-rule embedding quantification layer is used to generate the teacher's final dynamic quantification features, the student's final dynamic quantification features, and the teacher-student classroom teaching task progress quantification score information based on the teacher-dimensional exclusive fusion features, the student-dimensional exclusive fusion features, the teacher-student task progress exclusive fusion features, the teacher guidance-behavior linkage features, and the student cognition-behavior linkage features.
[0024] The three-dimensional linkage output layer is used to calibrate the teacher's final dynamic quantitative characteristics, the student's final dynamic quantitative characteristics, and the quantitative score information of the progress of teacher and student classroom teaching tasks, and output the quantitative set of effective teacher behavior with traceability labels, the quantitative set of effective student behavior with traceability labels, and the quantitative set of effective teacher and student classroom teaching task progress with traceability labels.
[0025] Optionally, the visual features include: movement trajectory features, head-up rate, frowning duration, writing action, front row seating rate, blackboard area features, teaching aids, and projection features;
[0026] Audio features include: speech rate features, frequency of spoken words features, frequency of questions features, student speaking rate features, and interactive response delay features;
[0027] Textual features include: teaching keyword features, teaching segment keyword features, and key and difficult point keyword features.
[0028] Optionally, the temporal-semantic-knowledge point anchoring coding layer includes:
[0029] A visual channel, which is used to generate visually encoded features based on a timestamp feature set;
[0030] A temporal anchoring unit is used to generate a temporal anchoring coding vector based on a light-optimized clear image dataset.
[0031] A semantic anchoring unit is used to generate a semantic anchoring encoding vector based on the filtered valid audio segments;
[0032] A knowledge point anchoring unit, which is used to generate a knowledge point anchoring encoding vector based on a semantic anchoring encoding vector;
[0033] A cross-axis fusion gating unit is used to fuse temporal anchoring encoding vectors, semantic anchoring encoding vectors, knowledge point anchoring encoding vectors, and visual encoding features to generate visual anchoring encoding features, audio anchoring encoding features, text anchoring encoding features, and a three-dimensional tag set.
[0034] Optionally, the teaching behavior-cognitive state linkage extraction layer includes:
[0035] The teacher-side collaborative extraction unit is used to generate teacher guidance features based on visual features, audio features, and text features.
[0036] The student-side linkage extraction unit is used to generate student cognitive engagement features based on visual features, audio features, and text features.
[0037] A bidirectional association unit is used to generate teacher guidance-behavior linkage features and student cognitive-behavior linkage features based on teacher guidance characteristics and student cognitive engagement characteristics.
[0038] Optionally, the knowledge point progression-evaluation dimension dynamic fusion layer includes:
[0039] A modal adaptation and fusion unit is used to generate modal fusion features based on visual anchoring coding features, audio anchoring coding features, and text anchoring coding features;
[0040] A dimension binding fusion unit is used to generate initial fusion features for the teacher dimension, initial fusion features for the student dimension, and initial fusion features for the teacher-student task progression dimension based on the modal fusion features.
[0041] The knowledge point progressive adjustment unit is used to generate teacher-specific fusion features, student-specific fusion features, and teacher-student task progress-specific fusion features based on the teacher-dimensional initial fusion features, student-dimensional initial fusion features, and teacher-student task progress-specific fusion features.
[0042] Optionally, the indicator association-rule embedding quantization layer includes:
[0043] A rule preset unit, which is used to store a set of indicator evaluation rules;
[0044] The rule embedding unit is used to generate an indicator rule embedding weight set based on the stored indicator evaluation rule set, teacher guidance-behavior linkage features, student cognition-behavior linkage features, teacher-dimensional exclusive fusion features, student-dimensional exclusive fusion features, and teacher-student task advancement dimension exclusive fusion features.
[0045] The indicator association matrix unit is used to generate intra-dimensional association matrices and cross-dimensional association matrices based on teacher guidance-behavior linkage characteristics and student cognition-behavior linkage characteristics.
[0046] The dynamic quantification unit is used to generate the final dynamic quantification features of teachers, the final dynamic quantification features of students, and the quantitative score information of the progress of classroom teaching tasks of teachers and students based on the teacher-specific fusion features, student-specific fusion features, teacher-student task progress-specific fusion features, indicator rule embedded weight set, teacher guidance-behavior linkage features, and student cognition-behavior linkage features.
[0047] Optionally, the three-dimensional linkage output layer includes:
[0048] A three-dimensional calibration unit is used to generate a calibrated three-dimensional unified feature set and a calibrated multi-dimensional feature set based on the teacher's final dynamic quantitative characteristics, the student's final dynamic quantitative characteristics, and the quantitative score information of the teacher and student's classroom teaching task progress.
[0049] Based on the calibrated three-dimensional unified feature set and the calibrated multi-dimensional feature set, generate a quantitative set of effective teacher behavior, a quantitative set of effective student behavior, and a quantitative set of effective behavior for promoting classroom teaching tasks between teachers and students.
[0050] The traceability marking unit is used to generate traceability information for each set of effective teacher behaviors, set of effective student behaviors, and set of effective behaviors in promoting classroom teaching tasks.
[0051] This application also provides an AI-assisted evaluation system for assessing the effectiveness of teachers' classroom teaching, the AI-assisted evaluation system for assessing the effectiveness of teachers' classroom teaching includes:
[0052] The information acquisition module is used to acquire the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments.
[0053] A neural network acquisition module, which is used to acquire a trained 3D-LCBQN neural network;
[0054] The effective behavior quantity acquisition module is used to input the denoised clean speech dataset, the light-optimized clear image dataset, and the selected effective audio segments into the 3D-LCBQN neural network, thereby generating the teacher effective behavior quantification set, the student effective behavior quantification set, and the teacher-student classroom teaching task promotion effective behavior quantification set.
[0055] The AI-assisted evaluation method for assessing the effectiveness of classroom teaching proposed in this application simultaneously observes and evaluates teacher classroom activities and behavioral responses, student classroom activities and behavioral responses, and the progress of classroom teaching tasks by introducing feature fusion, image recognition, and speech recognition. This solves the problem that traditional classroom teaching evaluation is often based on the evaluator's manual classroom records and personal subjective judgment. Due to differences in the evaluator's focus, evaluation criteria, subjective experience, and other factors, it is not only time-consuming and labor-intensive, but also often causes disputes due to issues of consistency and objectivity in the evaluation results. Attached Figure Description
[0056] Figure 1 This is a flowchart illustrating an embodiment of an AI-assisted method for evaluating the effectiveness of classroom teaching by teachers.
[0057] Figure 2 This is a schematic diagram of the evaluation criteria for the effectiveness of teachers' classroom teaching according to an embodiment of this application.
[0058] Figure 3 This is a schematic diagram of speech rate variation feature acquisition according to an embodiment of this application.
[0059] Figure 4 This is a schematic diagram of the movement trajectory features according to an embodiment of this application.
[0060] Figure 5 This is a schematic diagram of head-up ratio according to an embodiment of this application.
[0061] Figure 6 This is a schematic diagram illustrating the student speaking rate characteristics according to an embodiment of this application. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0063] like Figure 1 The methods for evaluating the effectiveness of AI-assisted classroom teaching, as shown, include:
[0064] Obtain the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments;
[0065] Obtain the trained 3D-LCBQN neural network;
[0066] The denoised clean speech dataset, the light-optimized clear image dataset, and the selected effective audio segments are input into the 3D-LCBQN neural network to generate a quantitative set of effective teacher behaviors, a quantitative set of effective student behaviors, and a quantitative set of effective behaviors in promoting classroom teaching tasks between teachers and students.
[0067] In this embodiment, obtaining the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments includes:
[0068] Acquire classroom visual data (including teacher and student actions, postures, seating arrangement, blackboard writing dynamics, classroom lighting environment, etc.) and classroom audio data;
[0069] The YOLOv8 object detection algorithm was used to identify objects in classroom video data, thereby obtaining a frame-by-frame object detection result set. Specifically, the YOLOv8 model input resolution was set to 640×640, the confidence threshold to 0.5, and the IOU threshold to 0.45. The object detection output format (object coordinates, category label, confidence score) was defined to match the data collection requirements and detection accuracy requirements. Based on the classroom labeled dataset, the backbone network of the model was frozen, the head detection layer was fine-tuned, and iterative training was conducted for 100 epochs. The cross-entropy loss function was used for optimization to ensure that the detection accuracy for classroom-specific objects was ≥92%. GPU parallel computing was enabled, parameters and weights were loaded, and the camera video stream was read in real time (or analyzed through post-recording recording). Object detection was performed frame by frame at a frame rate of 30fps, and the object coordinates, category, and confidence score of each frame were output, thereby obtaining a frame-by-frame object detection result set (including frame number, object information, and frame generation timestamp).
[0070] Facial key points in classroom video data are extracted using Dlib facial key point extraction, thereby obtaining a frame-by-frame facial key point dataset. Based on the coordinate frames of individual teachers and students in each frame, the Dlib 68-point facial key point detector is called to extract the coordinates of key points such as eyes and facial contours for each person and record the dynamic changes of key points.
[0071] A classroom original image dataset with frame numbers is generated based on the frame-by-frame target detection result set and the frame-by-frame facial key point dataset. Specifically, the frame number is associated with the target ID, the target detection information of the same teacher and student in the same frame is bound with the facial key point features, the frame timestamp format is unified, and structured image data is generated.
[0072] The classroom audio data is preprocessed to obtain a timestamped raw classroom speech dataset. Specifically, the Whisper Medium model is selected, with a sampling rate of 16kHz and an audio input format of mono PCM. The output includes raw audio data and basic speech features (energy value, frequency distribution). A spectral subtraction algorithm is used to extract the spectral features of environmental noise, and the difference between this and the spectrum of the raw audio data stream is calculated to filter out noise components while preserving the speech signal spectrum. This ensures that the signal-to-noise ratio (SNR) after denoising is ≥35dB, resulting in a clean, denoised audio data stream. The audio data stream is segmented into segments with a fixed 100ms time window. Each segment is labeled with a start and end index to ensure no overlap or omissions, thus obtaining a segmented audio dataset (containing each PCM segment and start / end indices). Based on the system clock, a unique timestamp is assigned to each audio segment, recording the actual start and end times of the segment. The timestamp format is uniformly "hour:minute:second:millisecond," ultimately yielding a timestamped raw classroom speech dataset (containing segmented audio, timestamp labels, and audio energy values).
[0073] Multimodal data timestamp alignment is performed on the original classroom image dataset numbered by frame and the original classroom audio dataset with timestamps to obtain the timestamp-aligned original classroom image dataset and original classroom audio dataset.
[0074] In this embodiment, multimodal data timestamp alignment is performed on the frame-numbered original classroom image dataset and the timestamped original classroom audio dataset to obtain the timestamp-aligned original classroom image dataset and original classroom audio dataset, including:
[0075] Using an NTP (Network Time Protocol) client, synchronize the system clocks of the camera and microphone respectively, compare the calibrated clocks with the NTP standard time, and ensure that the time error between the two is ≤20ms, thereby obtaining the calibrated camera clock and the calibrated microphone clock.
[0076] Set the image frame timestamp as the alignment reference, define the alignment error threshold (≤100ms), and define the matching success rule (the difference between the audio segment timestamp and the image frame timestamp ≤100ms).
[0077] Centered on the timestamp of each frame, construct a ±100ms sliding window, traverse the timestamps of all audio segments, determine whether they meet the matching rules, and record the index of the successfully matched audio segment.
[0078] Based on the matching index table, each frame of image data is bound one-to-one with the corresponding audio segment data to form a pairing data structure between a single frame of image and the corresponding audio segment.
[0079] The original timestamps of image frames and audio segments in the paired data are uniformly converted into Unix millisecond-level timestamps to ensure that the timestamp indexes of the same paired data are consistent, thereby obtaining a timestamp-aligned bimodal dataset (containing frame-level image data, corresponding audio segment data, and a unified Unix timestamp index).
[0080] A real-time denoising algorithm is executed on the original classroom audio dataset to filter out interference signals such as moving desks and chairs and external environmental noise, thereby obtaining a clean audio dataset after denoising.
[0081] The image data is processed using a light compensation algorithm to correct the imaging deviations in strong light and low light environments, thereby obtaining a clear image dataset after light optimization.
[0082] By using speech semantic analysis algorithms to filter out invalid speech in the original classroom speech dataset (such as filtering out invalid speech like teachers debugging equipment or chatting), the effective audio segments are obtained after filtering.
[0083] In this embodiment, the 3D-LCBQN neural network structure includes:
[0084] A multimodal basic feature extraction module is used to generate basic features based on the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered effective audio segments. The basic features include visual features, audio features, text features, and timestamp feature sets.
[0085] The temporal-semantic-knowledge point anchoring coding layer is used to generate visual anchoring coding features, audio anchoring coding features, text anchoring coding features, and a three-dimensional label set based on visual features, audio features, text features, and timestamp feature sets.
[0086] The teaching behavior-cognitive state linkage extraction layer is used to generate teacher guidance-behavior linkage features and student cognition-behavior linkage features based on the visual anchoring coding features, audio anchoring coding features, and text anchoring coding features.
[0087] The knowledge point progression-assessment dimension dynamic fusion layer is used to generate teacher-specific fusion features, student-specific fusion features, and teacher-student task progression dimension-specific fusion features based on visual anchoring coding features, audio anchoring coding features, and text anchoring coding features.
[0088] The indicator association-rule embedding quantification layer is used to generate the teacher's final dynamic quantification features, the student's final dynamic quantification features, and the teacher-student classroom teaching task progress quantification score information based on the teacher-dimensional exclusive fusion features, the student-dimensional exclusive fusion features, the teacher-student task progress exclusive fusion features, the teacher guidance-behavior linkage features, and the student cognition-behavior linkage features.
[0089] The three-dimensional linkage output layer is used to calibrate the teacher's final dynamic quantitative characteristics, the student's final dynamic quantitative characteristics, and the quantitative score information of the progress of teacher and student classroom teaching tasks, and output the quantitative set of effective teacher behavior with traceability labels, the quantitative set of effective student behavior with traceability labels, and the quantitative set of effective teacher and student classroom teaching task progress with traceability labels.
[0090] In this embodiment, the visual features include:
[0091] In this embodiment, the movement trajectory features are obtained by using YOLOv8 to track the target coordinates of teachers / students in real time, recording the position at a sampling interval of 0.5 seconds, and calculating the proportion of dwell time in different areas of the classroom (front row / back row / left / right) and the path complexity (variance of coordinate changes).
[0092] In this embodiment, the head-up rate is determined by Dlib detection of key points on the student's face (36-47 points around the eyes). A head-up rate is defined as a line of sight with the blackboard / projection screen that is ≤30° and lasts for more than 2 seconds. The percentage is counted every 5 minutes.
[0093] In this embodiment, the duration of frowning is calculated by extracting key points (points 17-21 and 22-26) of the eyebrows from 68 facial points using Dlib, and then calculating the duration of the eyebrow contraction state.
[0094] The writing action was recorded by using YOLOv8 to detect the student's hand posture (holding pen) and changes in the texture of the notebook page, to determine the writing behavior and count the participation rate.
[0095] Front row seating rate: YOLOv8 is used to divide the front row area of the classroom (the first 1 / 3 of the seats) and count the percentage of students in that area.
[0096] The whiteboard area features are identified by using YOLOv8 to locate the whiteboard area and combining it with OCR to recognize text clarity and layout structure, extracting the core knowledge points and their presentation features.
[0097] The teaching aids and projection features are identified using YOLOv8 target classification to recognize teaching aids (such as experimental devices) and projection content (such as videos / PPTs) and generate visual semantic tags.
[0098] Audio features include speech rate features, which are generated by using Whisper to count the number of characters transcribed per minute and producing real-time speech rate (words per minute).
[0099] The frequency of spoken words is determined by matching the text transcribed by Whisper with a pre-set database of spoken words (such as "um," "then," etc.) and counting the number of times they appear in a single class.
[0100] Question frequency characteristics: Whisper is used to identify question keywords (what, why, how, etc.) in the text, and the number of times teachers ask questions and the interval between them are counted.
[0101] Student speaking rate characteristics were determined by using Whisper to detect speech energy (excluding weak noise ≤-30dB) and combining it with image recognition to identify students who spoke, and then calculating the percentage and the number of times per student.
[0102] Interactive response delay characteristics are determined by using Whisper to mark the time of teacher questioning and student response, and then calculating the time interval.
[0103] Text-based features include: teaching keyword features, which are extracted by segmenting effective text fragments into words related to the teaching domain, matching them with the teaching syllabus keyword database, and extracting vocabulary related to knowledge points;
[0104] Key words and features of teaching segments are identified by recognizing key words for segments such as introduction, summary, and group discussion, which are used to divide teaching segments.
[0105] Key words and phrases that highlight key points and difficulties are identified by matching them with a pre-set vocabulary database of key and difficult points, and the number of times teachers emphasize these points (the frequency of repeated occurrences in the text) is counted.
[0106] In this embodiment, the temporal-semantic-knowledge point anchoring coding layer includes:
[0107] A visual channel, used to generate visually encoded features based on the denoised clean speech dataset;
[0108] In this embodiment, the specific processing procedure of the visual channel is as follows:
[0109] Conv1 (first convolutional layer): uses 64 3×3 convolutional kernels, stride 1, padding=1, to perform preliminary feature extraction on a clear 640×640×3 image and output a 640×640×64 feature map.
[0110] BN layer + ReLU activation: After Conv1 output, batch normalization (BN) is performed to eliminate dimensional differences, and then the nonlinear expression is enhanced by the ReLU activation function to generate... ;
[0111] Conv2 (second convolutional layer): uses 128 3×3 convolutional kernels, stride 2, padding=1. Perform feature compression and enhancement to output a 320×320×128 feature map;
[0112] Repeated BN layer + ReLU activation generation ;
[0113] Conv3 (third convolutional layer): uses 256 3×3 convolutional kernels, stride 2, padding=1. Further compression yields a 160×160×256 feature map;
[0114] Finally, activation and normalization are performed to obtain the visual encoding features (160×160×256).
[0115] A temporal anchoring unit is used to generate a temporal anchoring coding vector based on a light-optimized clear image dataset.
[0116] In this embodiment, the dataset of clear images after light optimization is divided into continuous, non-overlapping time segments at a granularity of 0.5 minutes per segment, and each segment is labeled with a unique index (e.g., segment 1: =0-30 seconds, second segment: 31-60 seconds);
[0117] Each time segment is bound to the corresponding interval of multimodal data: clear image frames (at a frame rate of 30fps, each segment contains 1500 frames), clean audio segments, and valid text segments, generating a segment-data association table.
[0118] Extract the relative timestamp range of each time segment and compare it with the reasonable range in the teaching node time planning table (this table is a known preset, such as the reasonable time range of each link preset by the school (such as 5-8 minutes for introduction and 20-25 minutes for new knowledge teaching), which meets the quantitative requirements for classroom pace control), and calculate the time overlap (such as the overlap between the relative time of the segment of 10-10.5 minutes and the new knowledge teaching of 20-25 minutes = 0%, and the overlap between the segment and the introduction of 5-8 minutes = 0%, which is judged as a transition segment).
[0119] By combining keywords related to teaching segments in effective text fragments (such as "Let's review first" corresponding to the introduction, and "Let's discuss in groups" corresponding to interactive exercises) and teaching behaviors in clear images (such as "projected courseware" corresponding to the teaching of new knowledge, and "blackboard summary" corresponding to the summary), the teaching nodes with the highest overlap are matched for each fragment, thereby obtaining a fragment-teaching node overlap table (including fragment index, corresponding node, and overlap value), initial node labels for fragments (marking the teaching nodes with the highest overlap in each fragment), and a list of non-overlapping fragments (fragments with overlap = 0 are temporarily marked as transitional fragments).
[0120] Specifically, for each valid text segment, the keyword similarity KeyMatch(t) is calculated by matching it with a keyword library for the teaching process (this library is a pre-set library in which each keyword is stored as a vector). The similarity is 1 if the target node keyword is included, otherwise it is 0.
[0121] For each segment, identify visual features of teaching behaviors (such as screen-projected courseware corresponding to the teaching of new knowledge, and group discussion postures corresponding to interactive exercises) and generate visual behavior labels.
[0122] If the overlap of the initial node labels of the segment is ≥0.5, combine the keyword matching degree and visual behavior labels to confirm the node labels (e.g., overlap 0.6 + group discussion keywords = confirmed interactive practice node).
[0123] For non-overlapping segments (transitional segments), if they contain explicit keywords or visual behavior tags, they should be corrected to the corresponding node tags (e.g., if they contain the keyword "summarize at the end", they should be corrected to the summary node); otherwise, the transitional segment tags should be retained.
[0124] Finally, obtain the final teaching node tags for each segment (including 6 types of nodes or transition segments such as introduction and new knowledge teaching) and the node matching basis table (record the overlap, keyword matching results, and visual behavior tags of each segment).
[0125] For segments with explicit node labels, calculate the initial fit value using the formula: (Overlap(t) represents temporal overlap, and KeyMatch(t) represents keyword similarity).
[0126] For transition segments, the initial adaptation value is set to 0.3 (between the core and invalid phases).
[0127] Ensure the initial adaptation value is within the range of 0-1 (truncate to 0 or 1 if it exceeds this range) to obtain the initial adaptation value for each segment. (0-1 interval);
[0128] Each segment is labeled with a stage type (core / transitional / invalid) according to the grading criteria (preset as: (core teaching stage: S(t)≥0.8, transition stage: 0.3≤S(t)<0.8, ineffective stage: S(t)<0.3)).
[0129] For segments within the preset invalid time period, force S(t) = 0.1 and mark the stage type as invalid stage;
[0130] For adjacent segments at node switching points (such as import → new knowledge teaching), linear interpolation is used to adjust the adaptation value to avoid abrupt changes (such as adaptation value of segment i is 0.7, adaptation value of segment i+1 is 0.9, and intermediate transition value is 0.8), thereby obtaining the final stage adaptation value S(t) and stage type label of each segment.
[0131] Extract the core information of each segment: segment index, relative time range start and end values, S(t), Overlap(t), KeyMatch(t), and stage type encoding (core=1, transition=2, invalid=3);
[0132] After normalizing the above information to the 0-1 range, it is encoded into a 64-dimensional vector through a fully connected layer. L2 normalization is then performed on the encoded vector to ensure that the vector magnitude is 1, ultimately yielding the temporal anchoring encoded vector. (64-dimensional, each segment corresponds to one vector);
[0133] A semantic anchoring unit is used to generate a semantic anchoring encoding vector based on the filtered valid audio segments;
[0134] In this embodiment, generating semantic anchoring encoding vectors based on the filtered valid audio segments includes:
[0135] Using classroom time segments as units, the original datasets for each semantic feature were traversed, and the eight basic semantic features (speech rate, frequency of spoken words, frequency of questions, student speaking rate, interaction response delay, matching degree of teaching keywords, matching degree of keywords in teaching segments, and frequency of emphasis on key and difficult points) were statistically quantified. The specific quantification method is as follows:
[0136] For speech rate characteristics, frequency of spoken words characteristics, frequency of questioning characteristics, student speaking rate characteristics, and interactive response delay characteristics: the weighted average of each characteristic value is calculated according to the timestamp sub-interval, with the weight being the proportion of the duration of the corresponding timestamp sub-interval to the total duration of the class time segment, to obtain the mean of the corresponding characteristic under the class time segment;
[0137] For the keyword matching degree features, teaching segment keyword matching degree features, and key point and difficult point keyword emphasis frequency features: The matching frequency / occurrence frequency of corresponding keywords in the effective text segments under this class time segment is statistically analyzed. This is then combined with the total length of the effective text segments for preliminary normalization to obtain the mean value of the corresponding features under this class time segment. The mean value of the above 8 features is the initial value of the basic semantic features for this class time segment. For semantically empty segments, their preset initial values of basic semantic features are directly called, without the above statistical calculations.
[0138] The min-max normalization method was used to uniformly calibrate the correction values of the eight basic semantic features of all classroom time segments to ensure that all feature values fall within the 0-1 range;
[0139] Special case handling: If the maximum and minimum values of the correction for a certain basic semantic feature are equal across all class time segments, then the normalized value of that feature is set to 0.5 across all class time segments.
[0140] After calibration, an 8-dimensional normalized basic semantic feature vector is generated for each classroom time segment. The vector dimension corresponds one-to-one with the 8 basic semantic features, and the vector elements are the normalized values of the corresponding features. The vectors are sorted in the order of [F1, F2, ..., F8], where F1 is the normalized value of speech rate feature, F2 is the normalized value of spoken word frequency feature, F3 is the normalized value of question frequency feature, F4 is the normalized value of student speaking rate feature, F5 is the normalized value of interactive response delay feature, F6 is the normalized value of teaching keyword matching degree feature, F7 is the normalized value of teaching segment keyword matching degree feature, and F8 is the normalized value of emphasis frequency feature of key and difficult keywords. The order of the vector dimensions remains fixed.
[0141] Calculation of the semantic relevance matrix specific to time segments:
[0142] Using a classroom time segment as a unit, the 8-dimensional normalized basic semantic feature vector of the segment is multiplied element-wise by a pre-set 8×8 semantic relevance matrix for the teaching domain to generate an 8×8 specific semantic relevance matrix for that classroom time segment. The value of the element in the i-th row and j-th column of the matrix is calculated using the following formula:
[0143] ;
[0144] For the element in the i-th row and j-th column of the semantic relevance matrix, The value of the i-th element in the 8-dimensional normalized basic semantic feature vector of this class time segment. Let j be the value of the element in the j-th dimension of the vector. The value of the element in the i-th row and j-th column of the semantic relevance matrix in the teaching domain. (The pre-defined semantic relevance matrix for the teaching domain is an 8×8 real symmetric matrix, calibrated by experts in the teaching domain, with matrix elements...) The matrix represents the inherent correlation between the i-th basic semantic feature and the j-th basic semantic feature, with values ranging from 0 to 1. The diagonal elements of the matrix... =1 (the feature's self-association degree is 1), the matrix is stored in tensor format and can be directly called by the neural network); after calculation, the specific semantic relevance matrix is still an 8×8 real symmetric matrix, the element values are retained to 6 decimal places and stored in tensor format.
[0145] Extraction of semantic relevance coefficient:
[0146] Perform real symmetric matrix eigenvalue decomposition on the 8×8 specific semantic relevance matrix to extract the maximum eigenvalue of the matrix. and the corresponding 8-dimensional feature vector ;in The overall semantic relevance of this class time segment is characterized. The feature vector represents the contribution of eight basic semantic features to the overall semantics of the segment, and the element values of the feature vector correspond one-to-one with the basic semantic features.
[0147] The maximum eigenvalue of all class time segments The minimum-maximum normalization method is used to uniformly process the data to the 0-1 interval, resulting in a semantic relevance coefficient α (0≤α≤1). This coefficient is a scalar and is bound to the corresponding classroom time segment.
[0148] 8-dimensional feature vector Multiplying the transpose of the 8-dimensional normalized basic semantic feature vector of this class time segment element-wise generates a 16-dimensional fused semantic feature vector, as shown in the formula:
[0149] ;
[0150] Where ⊙ represents the element-wise multiplication operation of vectors, [F1,F2,…,F8]T is the column vector form of the 8-dimensional normalized basic semantic feature vector, and U is the 16-dimensional fused semantic feature vector in tensor format, where the vector elements combine the basic semantic feature values and feature contributions.
[0151] Extraction of 32-dimensional local semantic feature vectors:
[0152] A 1D convolutional layer is constructed to extract local semantic features from a 16-dimensional fused semantic feature vector U. The specific parameters and calculation logic of the convolutional layer are as follows: 32 1×3 convolutional kernels are used, with a stride of 1, padding of 1, and ReLU activation function. The convolutional kernels are initialized using He normal initialization, and the bias term is initialized to 0. The 16-dimensional fused semantic feature vector U is input into this 1D convolutional layer. After convolution operation and activation function processing, a 32-dimensional local semantic feature vector U32 is output in tensor format with a dimension of 32×1.
[0153] Generation of 64-dimensional intermediate feature vectors: Two fully connected layers are constructed to achieve global fusion and dimensionality enhancement of local semantic features. The specific parameters of the fully connected layers are as follows:
[0154] The first fully connected layer has 64 neurons, the activation function is ReLU, the weight matrix is initialized using He normal initialization, and the bias term is initialized to 0. The 32-dimensional local semantic feature vector U32 is input into this layer, and after linear transformation and activation function processing, the 64-dimensional feature vector is output.
[0155] The second fully connected layer has 64 neurons, no activation function, and the weight matrix is initialized using Xavier normal initialization, with the bias term initialized to 0. The output vector from the first fully connected layer is input into this layer, and after linear transformation, it outputs a 64-dimensional intermediate feature vector. Tensor format, with a dimension of 64×1, preserves the global distribution features of the features.
[0156] Calculation of the 64-dimensional weighted eigenvector:
[0157] The 64-dimensional intermediate feature vector The semantic relevance coefficient α of the class time segment is multiplied element-wise to achieve weighted control of the semantic relevance on the encoding vector. The formula is as follows:
[0158] ;
[0159] The eigenvectors are 64-dimensional weighted features in tensor format, with dimensions 64×1; α is the semantic relevance coefficient (scalar); ⊙ represents the element-wise multiplication operation between α and α. Multiply each element separately; It is a 64-dimensional intermediate feature vector.
[0160] Feature calibration of semantically null fragments:
[0161] Iterate through the list of semantically nullable fragments, and for each semantically nullable fragment, calculate its 64-dimensional weighted feature vector. All elements are uniformly set to 0.1, and the vector dimension remains unchanged at 64×1 after calibration. This ensures that a distinguishable feature distribution is formed with the encoding vector of the effective semantic segment, while avoiding subsequent fusion anomalies caused by null values.
[0162] After the above processing, the 64-dimensional weighted feature vectors of each class time segment are obtained. That is, a 64-dimensional unnormalized semantic anchoring encoding vector. This vector is bound to the corresponding class time segment index ID, stored in tensor format, and retains the original calculated values of all elements.
[0163] Normalization:
[0164] 64-dimensional unnormalized semantic anchoring encoding vectors for each class time segment L2 normalization is performed to generate This ensures that the normalized encoded vector has a magnitude of 1; the normalized vector element values retain 6 decimal places, with no positive or negative overflow.
[0165] For each 64-dimensional semantic anchor encoding vector Generate structured traceability information tags in JSON format, containing the following core fields. All field information is retrieved from the input structured table to ensure authenticity and unambiguity:
[0166] Basic identifiers: Class time segment index ID, timestamp start and end range (Unix millisecond level), teaching node label, and stage type label;
[0167] Semantic feature information: semantic relevance coefficient α, normalized values of 8 basic semantic features, and top 3 contributions of basic semantic features (from V). max Extract the three features with the highest contribution, and label the feature type and contribution value.
[0168] Data traceability information: number of valid audio segments under this segment, average audio energy value, and whether it is a semantically null segment (0 / 1, 0 for no, 1 for yes);
[0169] Calculate the source information: the maximum eigenvalue λmax of the specific semantic relevance matrix. Then, map the source information labels to their corresponding 64-dimensional semantic anchoring encoding vectors. A unique binding is performed, and the vector retains its tensor format after binding. The traceability information label is stored as an additional attribute of the vector.
[0170] Generation of the classroom semantic anchoring encoding vector set: Sort the 64-dimensional semantic anchoring encoding vectors bound to the traceability information tags in ascending order by classroom time segment index ID. The vectors are integrated into a semantic anchoring encoding vector set for the classroom. This vector set is in tensor array format with a dimension of n×64 (n is the total number of classroom time segments). The order of each vector in the vector set is consistent with the chronological order of the classroom time segments, with no disorder or omissions.
[0171] The 64-dimensional semantic anchoring encoding vector generated by this semantic anchoring unit The dimensions of the 64-dimensional time-anchored encoding vector generated by the time-anchoring unit are completely matched, and both are bound to the same classroom time segment index ID and timestamp information. This provides a standardized, structured, and traceable feature foundation for the subsequent cross-axis fusion of time-semantic-knowledge point multi-feature cross-axis fusion of the cross-axis fusion gating unit, ensuring that the fused feature vector can simultaneously represent the time dimension features and semantic dimension features of classroom teaching.
[0172] A knowledge point anchoring unit, which is used to generate a knowledge point anchoring encoding vector based on a semantic anchoring encoding vector;
[0173] In this embodiment, generating knowledge point anchoring encoding vectors based on semantic anchoring encoding vectors includes:
[0174] Keyword extraction and word vector generation from effective text corpora:
[0175] For the effective text corpus of each class time segment, the TF-IDF algorithm is used to extract core feature keywords (the number of extracted keywords is 1%~5% of the effective character length of the corpus, with a minimum of 3 keywords), and function words and modal particles without pedagogical semantics are removed. The Word2Vec model, which is from the same source as the knowledge point database, is used to convert the extracted core feature keywords into 128-dimensional word vectors, generating a keyword vector set (tensor format, dimension k×128, where k is the number of core feature keywords) for the corpus of that class time segment. The keyword vector set is then subjected to mean pooling to generate the overall semantic vector S (128-dimensional, tensor format) for the corpus of that class time segment, representing the overall semantic features of the effective text corpus.
[0176] Quantification of similarity matching between knowledge points and effective text corpora:
[0177] Using a classroom time segment as a unit, the semantic similarity Sim is calculated between each knowledge point in the lightweight knowledge point subset (a pre-set teaching syllabus knowledge point database: developed by subject teaching experts in conjunction with the teaching syllabus of the corresponding grade level and subject, a structured hierarchical knowledge base, divided into three levels: chapter-core knowledge point-sub-knowledge point, with fields including the unique identifier ID of the knowledge point, knowledge point level, knowledge point name, set of core keywords of the knowledge point, knowledge point weight γ (representing the importance of the knowledge point in the teaching syllabus, with values from 0 to 1, core knowledge point γ≥0.7, sub-knowledge point γ<0.7), and the teaching link associated with the knowledge point (introduction / new knowledge teaching / interactive practice / summary). The knowledge base also stores the word vectors of the core keywords of each knowledge point (trained by the Word2Vec model, with a dimension of 128, tensor format) and the effective text corpus of the segment. Sim is obtained by mean pooling the set of word vectors of the core keywords of the knowledge point to obtain the semantic vector K of the knowledge point (128). The similarity between the semantic vector K of a knowledge point and the overall semantic vector S of the corpus is calculated using cosine similarity, where Sim∈[0,1]. A higher value indicates a higher semantic match between the knowledge point and the text corpus.
[0178] Set similarity threshold If Sim ≥ 0.5, the knowledge point is considered to have successfully matched the valid text corpus of the current class time segment; if Sim < 0.5, the match is considered to have failed, and the knowledge point is removed. A segment-matched knowledge point table is generated based on the matching results, with fields including class time segment index ID, successfully matched knowledge point ID, semantic similarity Sim, knowledge point weight γ, and the number of core keyword matches.
[0179] Extraction and quantification of anchor features for basic knowledge points:
[0180] For each class time segment, based on the successfully matched knowledge point information, four basic knowledge point anchoring features are extracted and quantified to characterize the knowledge point anchoring characteristics of that segment. The four features are: number of knowledge point matches, mean knowledge point match degree, proportion of core knowledge point matches, and knowledge point anchoring relevance. The specific quantification method is as follows:
[0181] Knowledge point matching count: The total number of knowledge points that were successfully matched in the current segment, normalized to the range of 0-1;
[0182] Mean of knowledge point matching: The semantic similarity Sim of successfully matched knowledge points is weighted and averaged according to the knowledge point weight γ. ;
[0183] Core knowledge point matching percentage: Number of successfully matched core knowledge points / Total number of core knowledge points in the knowledge point subset. This value is 0 when there are no core knowledge points.
[0184] Knowledge point anchoring relevance: After dimensional adaptation of the mean semantic vector of the successfully matched knowledge points and the 16-dimensional fused semantic feature vector U of the semantic anchoring unit, the cosine similarity is calculated to represent the degree of relevance between knowledge point anchoring and semantic features.
[0185] The quantified values of the above four features are the initial values of the anchor features of the basic knowledge points for this class time segment. For class time segments without matching knowledge points (such as transition / ineffective phases), the initial values of the four anchor features of the basic knowledge points are all set to 0.
[0186] Calculation of anchor feature correction value (basic knowledge point):
[0187] The initial values of the anchor features for basic knowledge points are adjusted by incorporating the knowledge point weight γ. The adjusted values for each feature are calculated using a specific formula: ;
[0188] The correction value for anchoring features for a single basic knowledge point. Set an initial value for the anchor feature of a basic knowledge point for this class time segment. The average weight of the knowledge points that were successfully matched to this segment ( , where n is the number of successfully matched knowledge points, and when there are no matched knowledge points... =0);
[0189] The calculation results are rounded to 6 decimal places. The minimum-maximum normalization method is used to uniformly calibrate the anchor feature correction values of the four basic knowledge points for all classroom time segments to ensure that all feature values fall within the 0-1 range.
[0190] After calibration, a 4-dimensional normalized anchor feature vector P for basic knowledge points is generated for each class time segment. The vector dimension corresponds one-to-one with the four anchor features of basic knowledge points, and the vector elements are the normalized values of the corresponding features. The vectors are sorted in [P1, P2, P3, P4], where P1 is the normalized value of the number of knowledge point matches, P2 is the normalized value of the mean of knowledge point matching degree, P3 is the normalized value of the proportion of core knowledge point matching, and P4 is the normalized value of the knowledge point anchor relevance. The order of the vector dimensions remains unchanged, and the tensor format is 4×1.
[0191] Specific knowledge points for class time segments - Calculation of semantic relevance matrix:
[0192] Using a segment of class time as a unit, the 4-dimensional normalized basic knowledge point anchoring feature vector P and the 16-dimensional fused semantic feature vector U of that segment are multiplied element-wise with a pre-set 4×16-dimensional knowledge point-semantic relevance matrix to generate a unique knowledge point-semantic relevance matrix C (4×16-dimensional) for that class time segment. The value of the element in the p-th row and q-th column of the matrix is calculated using the following formula:
[0193] ;
[0194] For the element value in the p-th row and q-th column of the semantic relevance matrix, which represents a specific knowledge point. To anchor the p-th dimension element value in the feature vector of 4-dimensional basic knowledge points, The value of the q-th element in the 16-dimensional fused semantic feature vector. This refers to the value of the element in the p-th row and q-th column of the pre-defined knowledge point-semantic relevance matrix (the pre-defined knowledge point-semantic relevance matrix is a 4×16 dimensional matrix, calibrated by experts in the field of education in conjunction with the characteristics of the subject). This represents the inherent correlation between the anchor feature of the p-th knowledge point and the fused semantic feature of the q-th knowledge point, with a value range of 0-1, p∈[1,4], q∈[1,16]. The matrix is stored in tensor format and can be directly accessed by the neural network. The average weight of the knowledge points that were successfully matched to this segment;
[0195] After calculation, the exclusive knowledge point-semantic relevance matrix is still 4×16 dimensions, with element values retained to 6 decimal places and stored in tensor format. For classroom time segments without matching knowledge points, all element values in the matrix are 0.
[0196] Extraction and quantification of knowledge point anchoring coefficients:
[0197] Singular Value Decomposition (SVD) is performed on the 4×16 dimensional semantic relevance matrix of specific knowledge points to extract the maximum singular value of the matrix. This value represents the overall correlation between the knowledge point anchoring features and semantic features of the current class time segment; the maximum singular value of all class time segments. The min-max normalization method is used to uniformly process the data to the 0-1 interval to obtain the knowledge point anchoring coefficient β (0≤β≤1). This coefficient is a scalar and is bound to the corresponding classroom time segment. The higher the β value, the higher the degree of anchoring between the teaching content of the segment and the knowledge points of the teaching syllabus.
[0198] 16-Dimensional Knowledge Point - Generation of Semantic Fusion Anchored Feature Vectors:
[0199] The dimensionality of the knowledge point-semantic relevance matrix is increased and features are fused. Through matrix transposition and fully connected layer mapping (the fully connected layer has 16 neurons, no activation function, and the weight matrix is initialized using Xavier normal initialization), the 4×16 dimensional matrix is mapped to a 16-dimensional feature vector. This 16-dimensional feature vector is then element-wise multiplied with the 16-dimensional fused semantic feature vector U of the semantic anchoring unit to generate a 16-dimensional knowledge point-semantic fusion anchoring feature vector W, achieving deep fusion of knowledge point features and semantic features. The formula is as follows:
[0200] ;
[0201] ⊙ represents the element-wise multiplication operation for vectors. U is the 16-dimensional feature vector mapped from the knowledge point-semantic relevance matrix, and W is the 16-dimensional knowledge point-semantic fusion anchor feature vector. The tensor format is 16×1 with no dimensional redundancy.
[0202] Extraction of 32-dimensional local knowledge point anchored feature vectors:
[0203] Constructing 1D convolutional and fully connected layers with parameters identical to the semantic anchoring unit to perform local feature extraction on the 16-dimensional knowledge point-semantic fusion anchoring feature vector W (the specific structure is not described in detail here), and outputting a 64-dimensional intermediate knowledge point anchoring feature vector. The tensor format has a dimension of 64×1 and preserves the global distribution features of knowledge point anchoring features.
[0204] Anchoring 64-dimensional intermediate knowledge points to feature vectors The anchoring coefficient β of the knowledge points in this class time segment is multiplied element-wise to achieve weighted control of the anchoring degree of the knowledge points on the encoding vector, calculated according to the formula:
[0205] ;
[0206] This is a 64-dimensional weighted anchor feature vector for knowledge points, in tensor format, with dimensions 64×1; β is the anchor coefficient (scalar) for knowledge points; ⊙ represents the element-wise multiplication operation of the vectors, i.e., β and... Multiply each element separately; Anchor feature vectors for 64-dimensional intermediate knowledge points.
[0207] Feature calibration for fragments without matching knowledge points:
[0208] Iterate through the list of semantically null value fragments and the list of class time fragments without matching knowledge points. For all such fragments, anchor the knowledge points to feature vectors based on their 64-dimensional weighted summation. All elements are uniformly set to 0.1. After calibration, the vector dimension remains unchanged at 64×1 to ensure that the encoding vector of the effective knowledge point anchor segment forms a distinguishable feature distribution, while avoiding subsequent cross-axis fusion anomalies caused by null values.
[0209] Determination of Unnormalized Knowledge Point Anchoring Coding Vectors: After the above processing, the 64-dimensional weighted knowledge point anchoring feature vectors for each class time segment are obtained. This is the 64-dimensional unnormalized knowledge point anchoring encoding vector. This vector is bound to the corresponding class time segment index ID, stored in tensor format, and retains the original calculated values of all elements.
[0210] right Normalization and binding of structured traceability information tags are performed to ultimately obtain a 64-dimensional knowledge point anchoring encoding vector. The 64-dimensional temporal anchoring encoding vector generated by the temporal anchoring unit and the 64-dimensional semantic anchoring encoding vector generated by the semantic anchoring unit are completely consistent in dimension, and are all bound to the same classroom time segment index ID, Unix millisecond-level timestamp, and structured traceability information label. This provides a standardized, structured, and traceable knowledge point feature foundation for the subsequent cross-axis fusion of the temporal-semantic-knowledge point three features of the cross-axis fusion gating unit, ensuring that the fused feature vector can simultaneously represent the temporal dimension features, semantic dimension features, and knowledge point anchoring dimension features of classroom teaching.
[0211] A cross-axis fusion gating unit is used to fuse temporal anchoring encoding vectors, semantic anchoring encoding vectors, knowledge point anchoring encoding vectors, and visual encoding features to generate visual anchoring encoding features, audio anchoring encoding features, text anchoring encoding features, and a three-dimensional tag set.
[0212] In this embodiment, the temporal anchoring encoding vector, semantic anchoring encoding vector, knowledge point anchoring encoding vector, and visual encoding features are fused to generate visual anchoring encoding features, audio anchoring encoding features, text anchoring encoding features, and a 3D tag set, including:
[0213] Using the fragment index ID as the core identifier, , , Perform structured integration to generate a fragment-trimodal vector integration table. This table is a structured two-dimensional table with fields including fragment index ID, time-series vector storage address, semantic vector storage address, knowledge point vector storage address, vector magnitude, verification result, and whether it is a completed fragment.
[0214] Simultaneously, the three 64×1 tensor vectors are concatenated according to the channel dimension to generate the trimodal anchored coding tensor M for this class time segment. The tensor format is 64×3, where the first channel is... The second channel is The third channel is The channel sequence remains fixed to ensure dimensional correspondence in subsequent gating fusion.
[0215] Generation of classroom-level trimodal anchored coding tensor sets:
[0216] Arranged in ascending order by classroom time segment index ID, all classroom time segment trimodal anchored coding tensors M are integrated into a classroom-level trimodal anchored coding tensor set M. The tensor format is n×64×3, where n is the total number of classroom time segments. This tensor set maintains the same time order as the previous vector sets.
[0217] Calculation of basic weights for single-feature dimension gating:
[0218] Using classroom time segments as units, and combining the temporal stage adaptation value S(t), semantic relevance coefficient α, and knowledge point anchoring coefficient β, the basic weight of single-feature cross-axis gating is calculated. These correspond to temporal, semantic, and knowledge point anchoring features, respectively, representing the core contribution of these three types of features in the current segment, calculated using a specific formula:
[0219] ;
[0220] As the basic weights for time-series feature gating, As the basic weights for semantic feature gating, The basic weights for knowledge point feature gating; S(t) is the time-series adaptation value, α is the semantic relevance coefficient, and β is the knowledge point anchoring coefficient; if Then let The calculation result should be rounded to 6 decimal places and satisfy the following conditions: This achieves a normalized allocation of basic weights.
[0221] By introducing a pre-defined cross-axis feature inherent correlation matrix Ω, the basic weights of single-feature gating are corrected, and the cross-axis dynamic gating weights are calculated. This enables adaptive adjustment of the gating weights based on the correlation between features, calculated using the following formula:
[0222] ;
[0223] in, These are cross-axis dynamic gating weights for temporal, semantic, and knowledge point features, respectively. These are the element values of the cross-axis feature inherent correlation matrix; for the calculated... Perform min-max normalization to ensure that the corrected values still meet the requirements. and , keep 6 decimal places.
[0224] In this embodiment, Ω is a preset cross-axis feature inherent correlation matrix Ω, which is a 3×3 real symmetric matrix. It is calibrated by experts in the teaching field in conjunction with the feature correlation characteristics of classroom teaching effectiveness assessments. The matrix elements... This represents the inherent correlation between the x-th anchor feature and the y-th anchor feature, with a value ranging from 0 to 1. (1 is the temporal feature, 2 is the semantic feature, and 3 is the knowledge point feature), the diagonal element Ωxx=1 (the correlation degree of the feature itself is 1), the semantic-knowledge point correlation degree Ω23 of the core teaching stage is ≥0.8, stored in tensor format, and can be directly called by the neural network.
[0225] Based on cross-axis dynamic gating weights A 64-dimensional cross-axis gated fusion matrix G is constructed to match the dimension of the three-modal anchored coding tensor M. The tensor format is 64×3, which realizes the gating weight allocation of each dimension. The element value of the first column of the k-th row of the matrix is W1, the element value of the second column of the k-th row is W2, and the element value of the third column of the k-th row is W3 (k∈[1,64]). That is, the weight distribution of each row of the matrix is consistent with the cross-axis dynamic gating weight, ensuring the uniformity of the gating fusion rules of the 64 feature dimensions. The gating fusion matrix G is stored in 32-bit floating-point type and bound to the corresponding classroom time segment index ID, which provides the basis for subsequent element-wise weighted fusion.
[0226] Element-wise cross-axis gated weighted fusion of three-modal features:
[0227] Using classroom time segments as units, the trimodal anchored coding tensor M is multiplied element-wise with the cross-axis gated fusion matrix G to obtain the weighted trimodal anchored coding tensor. The tensor format remains 64×3, and the element-by-element values are calculated using the formula:
[0228] ;
[0229] M(k,x) is the element value of the k-th row and x-th column of the weighted tensor, M(k,x) is the element value of the k-th row and x-th column of the original trimodal anchor coding tensor, and G(k,x) is the element value of the k-th row and x-th column of the gated fusion matrix, k∈[1,64], x∈[1,3]; the calculation result is retained to 6 decimal places and is kept as a 32-bit floating-point number.
[0230] Generation of 64-dimensional fused feature vectors:
[0231] Weighted trimodal anchored coding tensor Perform a summation operation along the channel dimension to compress the 64×3 tensor into a 64×1 tensor, generating a 64-dimensional cross-axis fusion intermediate feature vector. To achieve integrated fusion of three modal features, the calculation is performed according to the formula:
[0232] ;
[0233] To fuse the element values of the k-th dimension of the intermediate feature vector, The value of the element in the k-th row and x-th column of the weighted tensor is used. The calculation result is rounded to 6 decimal places. The tensor format is 64×1 and it is bound to the corresponding classroom time segment index ID. This vector initially integrates the core information of three types of features: time sequence, semantics, and knowledge points.
[0234] To enhance the local correlation representation ability of the fused features, a pre-defined 1×3 feature fusion enhancement convolution kernel is used to fuse the 64-dimensional cross-axis intermediate feature vector. Perform local one-dimensional convolution enhancement; the specific parameters for the convolution operation are as follows:
[0235] The convolution kernel stride is set to 1, padding is set to 1, and there is no activation function; only linear convolution is performed. The 64×1 vector is expanded into a 1×64 one-dimensional feature map and then input into the convolution kernel. After convolution, it is restored to a 64×1 tensor format, generating a 64-dimensional enhanced cross-axis fusion intermediate feature vector. During convolution operations, the feature dimensions remain unchanged; only the feature values of adjacent dimensions are weighted and enhanced to strengthen the local continuity of the fused features. Retain 6 decimal places and bind the fragment index ID.
[0236] Generation of classroom-level enhanced fusion feature vector sets:
[0237] Sort by class time segment index ID in ascending order, and retrieve the 64-dimensional enhanced cross-axis fusion intermediate feature vectors of all class time segments. Integrate into a classroom-level enhanced fusion feature vector set The tensor format is n×64, where n is the total number of class time segments. This vector set is the basis for batch processing, with no disorder and no omissions.
[0238] Fusion feature calibration of special stage segments:
[0239] By combining the temporal stage type labels and the list of null / unmatched segment markers, feature value calibration is performed on transition / invalid stage segments, semantic null segment segments, and knowledge point unmatched segment segments to ensure that the fusion features of special stages form a distinguishable feature distribution with the core teaching stages. The specific calibration rules are as follows:
[0240] Invalid phase fragment: All element values are multiplied by 0.1 to keep the dimension unchanged and reduce the impact of invalid stage features on subsequent layers.
[0241] Transition phase segment: All element values are multiplied by 0.5 to maintain the same dimension and represent the weak feature contribution during the transition phase.
[0242] Semantic null values / no matching fragments for knowledge points: Multiply the fusion value of the corresponding feature channel (semantic is channel 2, knowledge point is channel 3) by 0.3, while keeping the other channels unchanged, to achieve targeted feature weakening.
[0243] After calibration, the tensor format remains 64×1, with element values retained to 6 decimal places and no positive or negative overflow. It is labeled as a 64-dimensional unnormalized cross-axis anchored fusion feature vector. This is the basic output vector after the integration of this unit, bound to the corresponding classroom time segment index ID.
[0244] right By performing normalization and source information fusion and binding, a 64-dimensional cross-axis anchored fusion feature vector corresponding to each classroom time segment is finally obtained. (L2 normalized, modulus of 1, tensor format 64×1, bound to cross-dimensional structured traceability information tags); Classroom-level cross-axis anchored fusion feature vector set F normalized set (tensor format n×64, sorted in ascending order by segment index ID, the final output of the anchoring encoding layer); Cross-dimensional structured traceability information tags for each classroom time segment (JSON format, uniquely bound to the fusion feature vector); Classroom-level fusion feature traceability information set (JSON array format, sorted in ascending order by segment index ID).
[0245] In this embodiment, the teaching behavior-cognitive state linkage extraction layer includes:
[0246] The teacher-side collaborative extraction unit is used to generate teacher guidance features based on visual features, audio features, and text features.
[0247] In this embodiment, generating teacher guidance features based on visual features, audio features, and text features includes:
[0248] Obtain a teacher-specific basic feature set: including visual features (teacher movement trajectory features, blackboard writing area features, teaching aids / projection features), audio features (teacher speech rate features, frequency of spoken words features, frequency of questions features), and text features (teaching keyword features, frequency of emphasis of key and difficult keywords features). All features are quantized values in the range of 0-1.
[0249] Using the classroom time segment index ID as an identifier, nine core features were selected from the teacher-specific basic feature set, sorted by visual-audio-text dimensions, to generate a nine-dimensional teacher-specific basic feature vector. The vector elements are 0-1 quantized values of each feature, and the tensor format is 9×1.
[0250] right Perform dimensionality scaling by mapping 9 dimensions to 64 dimensions using a 1×1 convolution kernel, resulting in the same... 64-dimensional teacher-specific adaptive feature vectors based on dimension matching Preserve the semantic features, tensor format 64×1;
[0251] Using class time segments as units, and The teacher-side feature tensor is concatenated into a 64×2 structure based on the channels. This completes the structured integration of fusion features and teacher-specific features.
[0252] Obtain the preset teacher behavior-guidance correlation matrix P: a 64×64 real symmetric matrix, calibrated by experts in the field of education, with elements... This represents the guidance-behavior correlation between the m-th fusion feature and the n-th teacher-specific feature, with values ranging from 0 to 1, and diagonal elements being 1. It is stored in tensor format.
[0253] From the teacher's feature tensor Extract cross-axis anchored fusion feature vectors from the middle Teacher-specific adaptive feature vector Calculate the teacher-side association feature matrix based on the preset association matrix. The formula represents the bidirectional correlation between fusion features and teacher-specific features:
[0254] ;
[0255] A transpose of the feature vector specifically adapted for teachers. The teacher-side association feature matrix is 64×64, in tensor format, with element values retained to 6 decimal places.
[0256] right Perform trace extraction and eigenvalue aggregation to extract the trace of the matrix. With the largest eigenvalue The teacher guidance coefficient η is generated by weighted summation, which characterizes the strength of the teacher's teaching guidance in the current segment. The formula is as follows:
[0257] ;
[0258] η∈[0,1], the higher the value, the stronger the matching degree and correlation between the teacher's guidance and teaching behavior; the trace and the feature value are both normalized to 0-1 before being used in the calculation.
[0259] Teacher-side correlation feature matrix Perform global pooling to compress the 64×64 matrix into a 64-dimensional teacher behavior-related feature vector. The tensor format is 64×1, which preserves the core correlation information between teacher behavior and characteristics;
[0260] Introduce the teacher guidance coefficient η (which can be set as needed) to... Weighted regulation is applied, and cross-axis anchoring is used to fuse feature vectors. Generate the initial linkage feature vector for the teacher side. The formula is:
[0261] ;
[0262] ⊙ represents the element-wise multiplication operation for vectors. The initial linkage feature vector is 64-dimensional, which integrates the core correlation between teacher guidance and teaching behavior;
[0263] Using lightweight network layers Feature enhancement: A 1D convolutional layer (32 1×3 kernels, stride 1, padding=1, ReLU activation) is constructed to extract local correlation features. These features are then passed through a fully connected layer (64 neurons, no activation) to maintain dimensionality, outputting a 64-dimensional teacher-guided initial behavior linkage feature vector. , tensor format 64×1.
[0264] right Normalization and source binding were performed to obtain 64-dimensional teacher-guided behavior linkage feature vectors for each classroom time segment. (L2 normalized, modulus length is 1, tensor format 64×1, bound to simple traceability labels); Classroom-level teacher guidance-behavior linkage feature vector set (tensor array format n×64, n is the total number of classroom time segments, sorted in ascending order by index ID).
[0265] The student-side linkage extraction unit is used to generate student cognitive engagement features based on visual features, audio features, and text features.
[0266] In this embodiment, generating student cognitive engagement features based on visual features, audio features, and text features includes:
[0267] Obtain a student-specific basic feature set, including visual features (student posture features, attention focus features, classroom interaction gesture features), audio features (student response frequency features, response duration ratio features, classroom feedback sound features), and text features (answer keyword matching features, note core word features). All features are quantized values in the range of 0-1.
[0268] Using the classroom time segment index ID as an identifier, eight core features are selected from the student-specific basic feature set, sorted by visual-audio-text dimensions, to generate an 8-dimensional student-specific basic feature vector. The tensor format is 8×1, and the elements are feature quantization values;
[0269] Dimensionality upscaling is performed using a 1×1 convolution kernel, mapping an 8-dimensional vector to 64 dimensions, resulting in... 64-dimensional student-specific adaptive feature vectors based on dimension matching Preserves the core semantics of student characteristics, tensor format 64×1;
[0270] Assemble by channel and Generate a 64×2 student-side feature tensor This completes the structured integration of the two types of features and binds them to the corresponding fragment index IDs.
[0271] Obtain the preset student behavior-cognition correlation matrix Q: a 64×64 real symmetric matrix, labeled by experts in the field of education based on students' cognitive patterns, with elements... This represents the behavioral-cognitive correlation between the m-th fused feature and the n-th student-specific feature, with values ranging from 0 to 1, and diagonal elements being 1. It is stored in tensor format.
[0272] Split student-side feature tensor ,get and Calculate the student-side association feature matrix based on the preset matrix. The formula represents the bidirectional association between fused features and student-specific features:
[0273] ;
[0274] This is the transpose of the cross-axis anchored fusion feature vector. The matrix is 64×64, in tensor format, with element values retained to 6 decimal places.
[0275] extract traces With the largest eigenvalue After 0-1 normalization, a weighted sum is used to generate a student cognitive state coefficient θ, which represents the comprehensive level of the student's cognitive state (focus, comprehension) in the current segment. The formula is:
[0276] ;
[0277] θ∈[0,1], the higher the value, the stronger the match and coherence between the student's learning behavior and cognitive state.
[0278] right Perform global pooling to compress the 64×64 matrix into a 64-dimensional student behavior-related feature vector. The tensor format is 64×1, which preserves the core information related to behavior and cognition;
[0279] Introducing θ pairs Weighted regulation, combined with Generate initial linkage feature vectors for students This achieves a deep integration of behavioral and cognitive characteristics, as shown in the formula:
[0280] ;
[0281] ⊙ represents the element-wise multiplication operation for vectors. It is a 64-dimensional vector that integrates student behavior, cognition, and cross-axis fusion features;
[0282] The system employs a lightweight network layer consistent with the teacher's approach to enhance features: a 1D convolutional layer (32 1×3 convolutional kernels, stride 1, padding=1, ReLU activation) extracts local associations, and a 1 fully connected layer (64 neurons, no activation) preserves dimensionality, outputting a 64-dimensional student behavior-cognitive initial linkage feature vector. .
[0283] right By performing normalization and source tracing, a 64-dimensional student behavior-cognition linkage feature vector for each classroom time segment can be obtained. (L2 normalization, modulus of 1, bound to source tag); Classroom-level student behavior-cognition linkage feature vector set (tensor array format n×64, where n is the total number of classroom time segments, sorted in ascending order by index ID).
[0284] A bidirectional association unit is used to generate teacher guidance-behavior linkage features and student cognitive-behavior linkage features based on teacher guidance characteristics and student cognitive engagement characteristics.
[0285] In this embodiment, the generation of teacher guidance-behavior linkage features and student cognitive-behavior linkage features based on teacher guidance characteristics and student cognitive engagement characteristics includes:
[0286] Using the class time segment index ID as a unique identifier, the same segment is grouped together. and By concatenating the channels, a 64×2 bidirectional feature tensor between teachers and students is generated. The first channel of the tensor represents the teacher's features, and the second channel represents the student's features. The channel order is fixed and is bound to the corresponding segment index ID.
[0287] Integrating η and θ from the same segment, the basic coefficient μ for teacher-student feature adaptation is calculated, representing the initial adaptation degree of the core states at both ends of the teacher-student relationship. The formula is:
[0288] ;
[0289] The higher the value, the stronger the initial compatibility, and the more it binds to... As an additional attribute.
[0290] Split get and Calculate the bidirectional correlation feature matrix between teachers and students based on a preset matrix. The formula for mining the mutual information correlation between the features at both ends is:
[0291] ;
[0292] This is the transpose of the teacher-side vector. It is a 64×64 matrix in tensor format, with element values retained to 6 decimal places, representing the strength of the association between teacher and student characteristics in each dimension;
[0293] extract The maximum singular value σmax (normalized to 0-1) is used to generate the bidirectional teacher-student correlation coefficient ν, which is then expressed by the formula:
[0294] The higher the value, the closer the two-way connection between teachers and students;
[0295] Based on ν, generate cross-platform dynamic fusion weights To achieve relevance-driven weight allocation, the formula is:
[0296] ;
[0297] After normalization, it satisfies These correspond to the fusion weights of features from the teacher's and student's ends, respectively.
[0298] Perform a weighted fusion operation to generate an initial 64-dimensional bidirectional correlation feature vector. The formula, which integrates core features from both teacher and student perspectives with bidirectional correlation information, is as follows:
[0299] ;
[0300] ⊙ represents element-wise multiplication of vectors. It is a 64-dimensional vector, with a tensor format of 64×1;
[0301] right Perform global pooling to obtain a 64-dimensional association enhancement vector. to Adding elements together strengthens the expression of related features;
[0302] Lightweight network enhancement: A 1D convolutional layer (32 1×3 kernels, stride 1, padding=1, ReLU activation) extracts local correlations, and a 1 fully connected layer (64 neurons, no activation) preserves dimensionality, outputting a 64-dimensional enhanced bidirectional correlation feature vector. .
[0303] right Normalization and source binding were performed to obtain 64-dimensional teacher-student bidirectional correlation feature vectors for each class time segment. (L2 normalization, module length 1, bound to integrated traceability tag); Classroom-level teacher-student bidirectional correlation feature vector set (tensor array format n×64, sorted in ascending order by segment index ID, which is the final output of the linkage extraction layer).
[0304] In this embodiment, the knowledge point progression-evaluation dimension dynamic fusion layer includes:
[0305] A modal adaptation and fusion unit is used to generate modal fusion features based on visual anchoring coding features, audio anchoring coding features, and text anchoring coding features;
[0306] In this embodiment, generating modality fusion features based on visual anchoring coding features, audio anchoring coding features, and text anchoring coding features includes:
[0307] Using the class time segment index ID as a unique identifier, the same segment is grouped together. (Visual anchoring encoding features, 64-dimensional tensor (64×1), L2 normalized modulus length is 1, bound to source tags (including core dimensions of visual features and fragment index ID)). (Audio anchor coding features, 64-dimensional tensor (64×1), L2 normalized modulus length is 1, bound to source tags (including audio feature core dimension, segment index ID)). (Text anchoring encoding features, 64-dimensional tensor (64×1), L2 normalized modulus length of 1, bound to source tags (including core dimensions of text features and fragment index ID)) concatenated by channel to generate a 64×3 trimodal anchoring feature tensor. The first channel of the tensor represents visual features, the second channel represents audio features, and the third channel represents text features. The channel attributes are bound to the features one by one, with no overlap or confusion.
[0308] Calculation of global correlation degree in a single mode:
[0309] Based on the pre-defined modal inherent correlation matrix (Preset modal intrinsic correlation matrix) A 3×3 real symmetric matrix, labeled by experts in the field of education based on the modal characteristics of classroom teaching, with elements... This represents the intrinsic correlation between the i-th modality and the j-th modality (i / j=1 for visual, 2 for audio, 3 for text), taking values from 0 to 1, with diagonal elements Mii=1. For example, the visual-audio correlation M12≥0.7 and the audio-text correlation M23≥0.8 (stored in tensor format for direct access). It calculates the global correlation coefficient for each modality. , , The overall correlation strength between a single mode and the other two modes is represented by the formula:
[0310] , , The coefficient ranges from 0 to 1, and is rounded to 6 decimal places. A higher value indicates a stronger correlation between the mode and other modes, and a higher characterization value.
[0311] Modal adaptation dynamic weight allocation:
[0312] Based on the single-modal global correlation coefficient, the modality adaptation dynamic weights of the three types of anchored coding features are calculated. , , This achieves feature weight allocation driven by relevance, ensuring that modalities with high relevance and strong representational value receive higher fusion weights. The formula is as follows:
[0313] ;
[0314] ;
[0315] All weights have been normalized to meet the requirements. Retain 6 decimal places and bind to the three-modal anchored feature tensor. As an additional attribute.
[0316] Weighted adaptation and fusion of three-modal anchoring features:
[0317] Dimension-wise weighted fusion: anchoring the three-modal feature tensor The modality-adaptive dynamic weights are subjected to element-wise weighted operations to generate a 64-dimensional initial modality fusion feature vector. To achieve modal adaptation and fusion of three types of anchored coding features, the formula is:
[0318] ;
[0319] k∈[1,64] is the feature dimension index. Let the value of the k-th dimension of the initial fusion feature be... , , These are the k-th element values of the three types of anchored coding features. The calculation results are retained to 6 decimal places and the tensor format is 64×1.
[0320] Modality fusion local enhancement: To strengthen the local correlation representation after the fusion of three modal features, a 1×3 lightweight fusion convolution kernel (weights preset to [0.3, 0.4, 0.3], stride 1, padding=1, no activation function) is used to enhance the local correlation representation after the fusion of three modal features. Local convolutional enhancement is performed, which only strengthens the feature associations of adjacent dimensions without changing the feature dimensions and core representations, and outputs a 64-dimensional enhanced modality fusion feature vector. , tensor format 64×1.
[0321] Standardization of fusion features and binding of traceability tags:
[0322] Enhanced modality fusion feature vector Perform L2 normalization to ensure that the normalized feature vector magnitude is 1, thus obtaining the core output 64-dimensional modality fusion feature of this unit. Tensor format 64×1, element value range [−1,1], no positive or negative overflow;
[0323] Modal fusion traceability tag binding: for Generate structured traceability tags, retaining at least the core fields related to trimodal anchoring fusion, with no redundant information. Fields include: fragment index ID, timestamp start and end range, visual / audio / text modality adaptation weights, unimodal global correlation coefficient, and core representation dimensions of trimodal features. Tags and... Unique binding ensures end-to-end traceability; sorted in ascending order by classroom time segment index ID, all bound traceability tags are fused with 64-dimensional modal features. It is integrated into a classroom-level modal fusion feature vector set, in tensor array format n×64.
[0324] A dimension binding fusion unit is used to generate initial fusion features for the teacher dimension, initial fusion features for the student dimension, and initial fusion features for the teacher-student task progression dimension based on the modal fusion features.
[0325] In this embodiment, generating initial fusion features for the teacher dimension, initial fusion features for the student dimension, and initial fusion features for the teacher-student task progression dimension based on modal fusion features includes:
[0326] Based on the dimension-modal feature mapping rule table, and using the classroom time segment index ID as the identifier, from , , Sub-features matching the three dimensions of teacher, student, and teacher-student task progress are extracted respectively. Sub-features without corresponding relationships are removed to obtain the modal sub-feature sets of each dimension (visual / audio / text sub-feature sets for teacher dimension, similarly for student dimension, and similarly for teacher-student task progress dimension).
[0327] Based on dimension-specific feature weight matrix The corresponding extraction weights are applied to the visual, audio, and text sub-feature sets of each dimension, and the elements are summed in a weighted manner to generate the initial basic feature vectors (64-dimensional, 64×1 tensor) for each dimension:
[0328] Initial basic feature vector for teacher dimension:
[0329] ;
[0330] Initial basic feature vector for student dimension:
[0331] ;
[0332] Initial basic feature vector for the teacher-student task advancement dimension:
[0333] ;
[0334] Where ⊙ represents the element-wise multiplication of vectors, and all initial basic feature vectors are 64-dimensional, which perfectly matches the feature dimension of modality fusion.
[0335] In this embodiment, the knowledge point progressive adjustment unit is used to generate teacher-specific fusion features, student-specific fusion features, and teacher-student task advancement-specific fusion features based on the teacher-specific initial fusion features, student-specific initial fusion features, and teacher-student task advancement-specific initial fusion features.
[0336] In this embodiment, the generation of teacher-specific fusion features, student-specific fusion features, and teacher-student task progression-specific fusion features based on the initial fusion features of the teacher dimension, the initial fusion features of the student dimension, and the initial fusion features of the teacher-student task progression dimension includes:
[0337] right Perform 0-1 interval normalization calibration to eliminate dimensional bias caused by differences in sub-feature extraction weights, and obtain three-dimensional exclusive basic feature vectors (teacher). ,student Teacher and student task advancement (Each is a 64×1 tensor).
[0338] In this embodiment, the anchored coding sub-features corresponding to the teacher dimension include:
[0339] Visual anchoring coding sub-features: teacher movement trajectory features, blackboard writing area features, teaching aids and projection features;
[0340] Audio anchoring coding sub-features: teacher's speech rate features, verbal frequency features, and questioning frequency features;
[0341] Text anchoring encoding sub-features: teaching keyword features, emphasis frequency features of key and difficult keywords;
[0342] The anchored coding sub-features corresponding to the student dimension include:
[0343] Visual anchoring coding features: student head-up rate, frowning duration, writing action, and front-row seating rate;
[0344] Audio anchoring coding sub-features: student speaking rate feature, interactive response delay feature;
[0345] Text anchoring encoding sub-features: No directly corresponding text sub-features (relying solely on the text association representation of teacher-student interaction, without extracting text sub-features separately).
[0346] The anchored coding sub-features corresponding to the teacher-student task advancement dimension include:
[0347] Visual anchoring coding sub-features: features of the teacher's blackboard writing area, features of teaching aids and projection screen, and student writing actions;
[0348] Audio anchoring coding sub-features: teacher questioning frequency, student speaking rate, and interactive response delay;
[0349] Text anchoring encoding sub-features: teaching keyword features, teaching process keyword features, and emphasis frequency features of key and difficult points keywords.
[0350] To achieve targeted adaptation of modal fusion features to the specific basic features of each dimension, and considering the core differences in representation across various dimensions of classroom teaching, modal fusion feature adaptation weights are calculated for teachers, students, and the teacher-student task progression dimension. At the same time, the specific basic feature weights for each dimension are obtained. The formula is:
[0351] ;
[0352] ;
[0353] ;
[0354] =0.3 is the baseline value for modality fusion feature weights; S(t) is the temporal stage adaptation value of the temporal anchoring unit, α is the semantic relevance coefficient of the semantic anchoring unit, and β is the knowledge point anchoring coefficient of the knowledge point anchoring unit; all weights satisfy 0.3≤ ≤0.4, 0.6≤ ≤0.7, ensuring that dimension-specific basic features are the core contributors, and modal fusion features are auxiliary enhancements.
[0355] Using dimension-specific basic features as the core and modality fusion features as an auxiliary, targeted binding fusion is performed on the teacher, student, and teacher-student task advancement dimensions respectively. Then, feature enhancement is performed through a lightweight network, maintaining the 64 dimensions throughout the process. The operation is as follows:
[0356] Targeted weighted fusion: The specific basic feature vectors of each dimension are summed element-wise with the modal fusion feature F according to their corresponding weights to generate the initial specific fusion feature vectors of each dimension. The formula is as follows:
[0357] ;
[0358] ;
[0359] ;
[0360] Where ⊙ represents the element-wise multiplication of vectors. All are 64×1 tensors, preserving the core representations of each dimension and the general features of modality fusion;
[0361] Employing a lightweight network structure fully compatible with the parameters of the preceding units, local correlation enhancement is performed on the initial dedicated fusion feature vectors for each dimension (avoiding cross-dimensional feature interference). The network parameters are: 1D convolutional layer (32 1×3 convolutional kernels, stride 1, padding=1, ReLU activation) + 1 fully connected layer (64 neurons, no activation, weights initialized with a Xavier normal distribution), outputting enhanced dedicated fusion feature vectors for each dimension. (All are 64×1 tensors).
[0362] right Perform L2 normalization on each vector to ensure that the magnitude of each vector is 1 after normalization, and obtain the core output of this unit:
[0363] 64-dimensional teacher-specific integrated features ;
[0364] 64-dimensional student-specific fusion features ;
[0365] 64-dimensional teacher-student task advancement dimension exclusive integration features ;
[0366] All output features are 64×1 tensors with element values in the range [−1,1], no positive or negative overflow, and perfectly match the input requirements of the subsequent index-rule embedding quantization layer.
[0367] Source tracing tag binding: for Generate source tracing tags separately, with the tag fields uniformly containing four main categories:
[0368] Basic identifiers: segment index ID, Unix millisecond-level timestamp start and end range, teaching node label, and stage type label;
[0369] Feature source: This dimension comes from The extracted sub-feature names and weights, and the mean / variance of the dimension-specific basic feature vectors;
[0370] Fusion parameters: Modality fusion feature adaptation weights, specific basic feature weights, and preceding core coefficients (S(t) / α / β) for this dimension;
[0371] Enhancement and standardization: core parameters of convolutional / fully connected layers, modulus of feature values before and after L2 normalization, and normalization coefficients; each label and its corresponding dimension's exclusive fusion feature are uniquely bound through fragment index IDs, and the label is stored as an additional attribute of the feature to ensure traceability throughout the entire chain.
[0372] Sort by class time segment index ID in ascending order, and bind all those with complete traceability tags. The features are integrated into three sets: a teacher-specific feature vector set, a student-specific feature vector set, and a teacher-student task progression feature vector set. Each set is an n×64 tensor array. A dimension-binding fusion process record table is output synchronously and provided to subsequent network layers along with the three feature sets.
[0373] In this embodiment, the indicator association-rule embedding quantization layer includes:
[0374] A rule preset unit, which is used to store a set of indicator evaluation rules;
[0375] See Figure 2 In this embodiment, the indicator evaluation rule set includes the following: 3, 4, 5, 6.
[0376] A set of 23 evaluation rules (including specific judgment criteria, which can be set as needed):
[0377] Teacher Dimensions (9 items): Punctuality (arriving early ≤ -3 minutes / arriving on time -3 to 0 minutes / being late > 0 minutes), Lesson Preparation Adequacy (knowledge point coverage ≥ 80% is excellent), Roll Call Efficiency (time ≤ 3 minutes is efficient), Pace Reasonableness (time percentage of each segment conforms to the preset range), Speech Speed Achievement Rate (200-250 words / minute), Frequency of Useful Words (≤ 15 times / lesson), Fairness of Interaction (balanced time spent in different areas ≥ 80%), Diversity of Teaching Methods (≥ 5 types is excellent), Blackboard Writing Standardization (coverage of core knowledge points ≥ 90%)
[0378] Student dimensions (8 items): attendance rate (percentage of students actually attending class), front row seating rate (percentage of students sitting in the front row), head-up rate (average percentage of students looking up), speaking rate (percentage of students speaking and average number of times speaking per student), textbook browsing rate (percentage of students actively browsing textbooks), note-taking rate (participation in note-taking), presentation participation rate (percentage of students presenting on stage), and concentration loss rate (percentage of time spent daydreaming).
[0379] Teacher-student task progress dimensions (6 items): Achievement of teaching objectives (keyword matching degree + quiz accuracy rate), Completion of teaching content (progress deviation rate ≤10%), Resolution of key and difficult points in teaching (emphasis ≥3 times + practice accuracy rate ≥70%), Richness of teaching methods (less than 3 types / 3-5 types / more than 5 types), Completeness of teaching nodes (no missing 6 standard nodes), Activity level of teaching process (weighted sum score ≥7 points);
[0380] The rule embedding unit is used to generate an indicator rule embedding weight set based on the stored indicator evaluation rule set, teacher guidance-behavior linkage features, student cognition-behavior linkage features, teacher-dimensional exclusive fusion features, student-dimensional exclusive fusion features, and teacher-student task advancement dimension exclusive fusion features.
[0381] In this embodiment, the indicator rule embedding weight set generated based on the stored indicator evaluation rule set, teacher guidance-behavior linkage features, student cognition-behavior linkage features, teacher-specific fusion features, student-specific fusion features, and teacher-student task advancement-specific fusion features includes:
[0382] The natural language judgment criteria of 23 indicators are transformed into quantitative logic that can be compared with features, as shown in the following example:
[0383] Numerical threshold categories (such as lesson preparation adequacy and speech rate compliance rate): These directly use rule-based thresholds, comparing feature component values with the thresholds to output excellent / qualified / unqualified levels and corresponding scores (excellent 1 point, qualified 0.6 points, unqualified 0.3 points). For example, using two indicators—speech rate compliance rate and frequency of spoken words—from the teacher's perspective, fixed numerical threshold ranges and corresponding scores are set. Speech rate compliance rate corresponds to the teacher-specific fusion feature set components, with quantification standards of 1 point for 200-250 words / minute, 0.6 points for 180-199 words / minute or 251-280 words / minute, and 0.3 points for less than 180 words / minute or more than 280 words / minute. Speech frequency corresponds to the audio category of the teacher-specific fusion feature set, with quantification standards of 1 point for ≤5 times / lesson, 0.7 points for 6-10 times / lesson, 0.4 points for 11-15 times / lesson, and 0.2 points for >15 times / lesson.
[0384] Range-based categories (such as punctuality and rhythmic rationality): Feature components are mapped to rule ranges, corresponding to output grade scores. Taking punctuality, roll call efficiency, and rhythmic rationality as examples, time / proportion ranges and gradient scores are set. The punctuality quantification standard is 1 point for arriving ≥3 minutes early, 0.8 points for arriving 1-2 minutes early or on time (-3 to 0 minutes), 0.5 points for arriving 1-3 minutes late, and 0.3 points for arriving >3 minutes late. The roll call efficiency quantification standard is 1 point for taking ≤1 minute, 0.8 points for taking 1-2 minutes, 0.5 points for taking 2-3 minutes, and 0.3 points for taking >3 minutes. The rhythmic rationality quantification standard is 1 point for the deviation of the teaching segment's duration from the preset standard ≤5%, 0.7 points for a deviation of 6%-10%, 0.4 points for a deviation of 11%-15%, and 0.2 points for a deviation >15%.
[0385] Statistical calculations (such as attendance rate and interactive fairness): Based on basic classroom data, statistical values are calculated according to rules, then adapted to feature components, and quantitative results are output. Taking interactive fairness and blackboard writing standardization as examples, scores are calculated using direct statistical formulas, with the scoring formulas being standardized as follows: (S represents the indicator score,) This refers to the actual statistical value of the indicator. (This is the total statistical count for the indicators), and the score is rounded to two decimal places; each indicator in this category corresponds to a fixed feature component index.
[0386] Composite judgment categories (such as the degree of resolution of key and difficult points): quantified by two conditions: the number of times emphasis is emphasized and the accuracy rate of practice, and the final score is obtained by weighted summation (each weight is 50%).
[0387] Understandably, the specific conversion method can be set according to needs, as long as it can be converted into an executable vector in the end.
[0388] For the three categories of dimensional features—teacher-specific integrated features, student-specific integrated features, and teacher-student task progression-specific integrated features—indicator matching calculations were performed item by item with the corresponding indicators for each dimension.
[0389] The corresponding components of a single-dimensional feature are extracted and compared with the transformed quantification rules to output the score and level of each indicator. Using the quantification scores of each indicator as the basis, normalization is applied to convert them into indicator rule embedding weights, ensuring that the sum of the weights of all indicators within the same dimension is 1. After the above transformation, the weights of each dimension's indicators are arranged in order according to their corresponding feature component indices, forming an indicator rule embedding weight set.
[0390] The indicator association matrix unit is used to generate intra-dimensional association matrices and cross-dimensional association matrices based on teacher guidance-behavior linkage characteristics and student cognition-behavior linkage characteristics.
[0391] In this embodiment, the generation of intra-dimensional correlation matrices and cross-dimensional correlation matrices based on teacher guidance-behavior linkage features and student cognition-behavior linkage features includes:
[0392] Construction of the indicator correlation matrix:
[0393] Matrix dimension definition: Construct a 23×23 square matrix, with each row and column corresponding to 23 indicators (arranged in the order of 9 for teachers → 8 for students → 6 for task progress), clearly labeling the indicator name and the dimension to which each row / column belongs;
[0394] Quantitative calculation of inter-indicator correlation: For the three dimensions, the correlation degree within each dimension is calculated first, followed by the correlation degree across dimensions. The Pearson correlation coefficient formula is used to calculate the correlation between indicators, and the absolute value is taken as the initial correlation degree (range [0,1]). The cross-dimensional correlation degree is weighted and corrected by incorporating the rule embedding weights of the corresponding indicators. The correction formula is as follows: ( The correlation after cross-dimensional correction. Embed weights into the rules for two cross-dimensional metrics. (The absolute value of the initial Pearson correlation coefficient).
[0395] Association Matrix Construction: Construct a 23×23 global indicator association matrix. The rows and columns of the matrix correspond to the 23 indicators in the order of "9 items in the teacher dimension → 8 items in the student dimension → 6 items in the teacher-student task progress dimension". The elements on the diagonal of the matrix represent the correlation degree of the indicators themselves, which is uniformly set to 1. In the off-diagonal elements, the positions of indicators in the same dimension are filled with the intra-dimensional correlation degree, and the positions of cross-dimensional indicators are filled with the weighted and corrected cross-dimensional correlation degree. Elements between indicators with no actual correlation are uniformly set to 0. Extract the corresponding areas in the matrix to obtain the intra-dimensional association matrices of teachers, students, and teacher-student task progress, as well as the cross-dimensional association matrices of teachers-students, teachers-teacher-student task progress, and students-teacher-student task progress.
[0396] • Matrix standardization: Normalization is performed on the global index correlation matrix and the sub-matrices within and across dimensions to ensure that the value range of all matrix elements is uniformly [0,1], which meets the numerical adaptation requirements of subsequent quantification calculations; at the same time, matching traceability labels are generated for all correlation matrices to record the correlation calculation parameters and the correspondence between matrix indicators to ensure traceability.
[0397] In this embodiment, the dynamic quantification unit is used to generate the final dynamic quantification features of teachers, the final dynamic quantification features of students, and the quantitative score information of the progress of classroom teaching tasks of teachers and students based on the teacher-specific fusion features, student-specific fusion features, teacher-student task progress-specific fusion features, indicator rule embedded weight set, teacher guidance-behavior linkage features, and student cognition-behavior linkage features.
[0398] In this embodiment, based on teacher-specific fusion features, student-specific fusion features, teacher-student task progress-specific fusion features, indicator rule embedded weight set, teacher guidance-behavior linkage features, and student cognition-behavior linkage features, the following are generated: Teacher's final dynamic quantitative features, student's final dynamic quantitative features, and teacher-student classroom teaching task progress quantitative score information:
[0399] Alignment of basic quantitative information: The three types of exclusive fusion features, two types of linkage features and indicator rules are embedded into the weight set and the 23×23 global indicator association matrix for index alignment, ensuring that the feature components, indicator weights and the degree of correlation between indicators correspond one-to-one. The alignment standard follows the feature-indicator mapping rule mentioned above, with no mismatches.
[0400] Dynamic adjustment of indicator weights: Based on the correlation degree of the global indicator correlation matrix, the embedding weights of indicator rules are dynamically adjusted. The calculation formula is as follows:
[0401] ; ( The dynamic adjustment weight of the i-th indicator, Weights are embedded into the original rules of this indicator, where m is the total number of indicators in the corresponding dimension. (For the correlation between the i-th indicator and the j-th indicator in the same dimension), ensure that the weights are dynamically adapted to the correlation between the indicators. After correction, the weights of each dimension are normalized to ensure that the sum of the weights in the same dimension is 1, and finally obtain the dynamic corrected weights of each of the three dimensions.
[0402] The calculation is performed separately for teachers, students, and teacher-student task progression dimensions, with the dynamically adjusted weights for each dimension weighted and weighted by the corresponding dedicated fusion feature components. Calculation formula:
[0403] ; ( Let be the initial dynamic quantization feature vector (64 dimensions) for the d-th time segment. The weights of the i-th indicator in this dimension are dynamically adjusted. Let be the feature component value corresponding to the i-th indicator in the t-th time segment. After time segment-by-time calculation, the initial dynamic quantization feature vector sets (n×64) of each of the three dimensions are obtained.
[0404] Standardization and integration of quantization results: L2 normalization was performed on the three initial dynamic quantization feature vector sets to ensure that the vector magnitude was 1 and the element value range was [0,1]. The results were then sorted and integrated according to the classroom time segment index to obtain the final dynamic quantization feature sets of the three dimensions of teacher, student, and teacher-student task progress, maintaining the n×64 tensor format and consistent with the storage format of the feature sets mentioned above.
[0405] Quantitative traceability record: Generate traceability tags for dynamic quantization feature sets of each dimension, record dynamic adjustment parameters of weights, basis for use of correlation, quantization calculation formula and standardization results, and bind them one by one to the feature set to ensure that the quantization process is traceable.
[0406] In this embodiment, the three-dimensional linkage output layer includes:
[0407] Fully connected layer feature fusion (FC1 hidden layer):
[0408] The n×192 fused feature vector is input into the first-level fully connected hidden layer (FC1), which is the core implementation of nonlinear fusion and correlation mining of three-dimensional features:
[0409] Neuron configuration: 128 neurons, with ReLU activation function to meet the requirements of non-linear feature mapping;
[0410] Operational logic: Based on the weight matrix fixed by model training, linear transformation and nonlinear activation are performed on the input features to explore the potential correlation patterns of effective behaviors in three dimensions, while preserving the independence of each dimension's features;
[0411] Finally, a 128-dimensional fusion feature vector is obtained (this is only an intermediate feature of the network, not output externally, and is only used for subsequent dimensionality numerical mapping).
[0412] Effective behavior numerical mapping (FC2 output layer):
[0413] The 128-dimensional fused feature vector from the FC1 layer is input into the second fully connected output layer (FC2), which is a linear fully connected layer (without an activation function). This completes the dimensionality mapping of the fused features to the quantized values of the three types of effective behaviors, directly producing the core numerical values of the multi-dimensional effective behaviors.
[0414] Neuron configuration: 3 output neurons, each corresponding one-to-one with one of the three dimensions of effective teacher behavior, effective student behavior, and effective teacher-student task progress behavior, ensuring that the mapping results accurately correspond to the target quantification set;
[0415] Numerical mapping: The fused features are linearly mapped to specific numerical values through the training weight matrix, then normalized to a 0-10 scale (keeping one decimal place). Simultaneously, combined with the index of preceding feature components, detailed quantified values (proportion, key behavior item values) are generated for each type of effective behavior. Specific outputs include:
[0416] Core Quantitative Values of Effective Teacher Behaviors ( ): Represents the overall level of teachers' effective classroom behaviors;
[0417] Core Quantitative Values of Effective Student Behaviors ( ): Represents the overall level of effective classroom participation among students;
[0418] Core Quantitative Values of Effective Behaviors in Teacher-Student Task Promotion ( ): Represents the overall level of effective behavior of teachers and students in collaborating to complete teaching tasks.
[0419] The specific training process of the 3D-LCBQN neural network in this application is as follows:
[0420] Training dataset construction:
[0421] (1) Data acquisition and preprocessing
[0422] Following the multimodal data processing workflow of the technical solution, collect no less than 1000 classroom video / audio data sessions from different subjects and grade levels, and perform the following preprocessing:
[0423] Image data: The original image dataset, numbered by frame, was generated by YOLOv8 object detection (input resolution 640×640, confidence threshold 0.5) and Dlib facial key point extraction. The dataset was then optimized into a clear image dataset by a light compensation algorithm.
[0424] Speech data: Transcribed using the Whisper Medium model (sampling rate 16kHz), denoised by spectral subtraction (signal-to-noise ratio ≥35dB), and invalid speech was filtered by speech semantic analysis to generate a clean speech dataset with timestamps and valid audio segments;
[0425] Timestamp alignment: The device clock is synchronized using the NTP protocol (error ≤ 20ms), and image frames and audio segments are bound in a ±100ms sliding window to generate a timestamp-aligned bimodal dataset.
[0426] (2) Labeling
[0427] Three teaching experts with at least eight years of teaching experience used the 23 indicators in the technical solution to assign three categories of quantifiable labels to each 0.5-minute segment of class time:
[0428] Core tag: Overall score of effective behaviors in promoting tasks by teachers / students / teacher-student relationships (0-10 points, rounded to one decimal place);
[0429] Sub-labels: numerical values for key behavioral items (such as teacher's speaking speed compliance rate score, student head-up rate percentage, etc., 0-10 points or percentage);
[0430] Label calibration: The weighted average method was used to integrate the expert annotation results (with equal weight for each expert) to ensure label consistency (Cohen's Kappa ≥ 0.85).
[0431] (3) Dataset partitioning:
[0432] Divide the dataset in a 7:2:1 ratio to ensure a consistent data distribution.
[0433] Training set: 700 class sessions of data, used for model parameter updates;
[0434] Early Stop Validation Set: Data from 200 class sessions, used to monitor overfitting and trigger early stop;
[0435] Hyperparameter validation set: 100 class sessions of data, used to select the optimal hyperparameters (such as learning rate and batch size).
[0436] 2. Model Initialization
[0437] Based on the network structure of the technical solution, initialize the parameters of each module:
[0438] Convolutional layers: Conv1 / Conv2 / Conv3 use 64 / 128 / 256 3×3 convolutional kernels, stride / padding is set according to the technical solution, and weights are initialized using He normal distribution;
[0439] Fully connected layers: The weights of all fully connected layers (including the encoding layer, the linked extraction layer, the fusion layer, and the output layer) are initialized using the Xavier normal distribution, and the bias terms are initialized to 0.
[0440] Gating Unit: The inherent correlation matrix Ω and modal correlation matrix M of the cross-axis fusion gating unit are initialized using values calibrated by teaching experts (stored in tensor format).
[0441] Activation function: The ReLU activation function is used in all hidden layers, while the output layer has no activation function (linear mapping).
[0442] 3. Hyperparameter settings
[0443] Optimizer: Adam optimizer, initial learning rate 1e-4, weight decay coefficient 1e-5;
[0444] Batch size: 32 (each batch contains time segment features of 32 classes, with 20 segments taken from each class, for a total dimension of 32×20×64×3).
[0445] Number of iteration rounds: Maximum 150 rounds, early stopping strategy (stop if the MAE of the early stopping validation set does not decrease for 5 consecutive rounds);
[0446] Loss function: Weighted MSE loss is adopted, and the formula is: teaching task, where teaching task is the MSE loss of the teacher / student / teacher-student task progress quantification value and label respectively.
[0447] Layered iterative training process
[0448] 1. First stage: Training of the basic feature extraction module (rounds 1-30)
[0449] Input: a clear image dataset, a clean speech dataset, and valid audio clips from the training set;
[0450] Training objective: To enable the module to accurately output visual features (movement trajectory, head-up rate, etc.), audio features (speech rate, frequency of spoken words, etc.), text features (teaching keywords, etc.), and timestamp feature sets;
[0451] Monitoring metric: Feature extraction accuracy (accuracy ≥ 90% compared with manually labeled feature values);
[0452] Iteration logic: The learning rate of the convolutional layer is adjusted every 5 rounds (decreasing by 0.95 times), and the basic feature extraction module converges in the 30th round.
[0453] 2. Second stage: Training of the temporal-semantic-knowledge point anchoring encoding layer (31-60 rounds)
[0454] Input: The basic feature set output from the first stage;
[0455] Training objective: The visual / audio / text anchoring coding features (64 dimensions) generated by the temporal-semantic-knowledge point anchoring coding layer have a correlation ≥ 0.7 with the temporal stage adaptation value S(t), semantic correlation coefficient α, and knowledge point anchoring coefficient β of the label;
[0456] Key operations:
[0457] The parameters of the basic feature extraction module are fixed, and only the parameters of the encoding layer (temporal / semantic / knowledge point anchoring unit, cross-axis fusion gating unit) are updated;
[0458] The weight allocation is constrained according to the gating weight formula ω1 / ω2 / ω3 in the technical solution to ensure that ω1+ω2+ω3=1;
[0459] Monitoring indicator: Normalization error of anchored coding features (deviation of L2 modulus from 1 ≤ 0.01).
[0460] 3. Third stage: Training of the linkage extraction layer + dynamic fusion layer (61-90 rounds)
[0461] (1) Teaching behavior-cognitive state linkage retrieval layer
[0462] Input: Anchored coding features output from the second stage;
[0463] Training objective: The generated teacher-guided behavior linkage features and student cognitive-behavioral linkage features (both 64-dimensional, L2 normalized) have a Pearson correlation coefficient ≥ 0.75 with the expert-annotated guidance force coefficient η and cognitive state coefficient θ.
[0464] Constraints: The bidirectional association coefficient ν∈[0,1] of the teacher-student bidirectional association unit satisfies the teaching requirements.
[0465] (2) Knowledge point progression - dynamic integration layer of assessment dimensions
[0466] Input: Anchored coding features, linked features;
[0467] Training objective: The weighted fusion error between the generated teacher / student / teacher-student task progression dimension-specific fusion features (all 64-dimensional) and the corresponding dimension-based basic features should be ≤0.05.
[0468] Key operation: According to the weight formula of the technical solution, the model education and basic education constraints are integrated to ensure the model education and basic education.
[0469] 4. Fourth stage: Training of the indicator association-rule embedding quantization layer (91-120 rounds)
[0470] Input: The exclusive fusion features and linkage features output from the third stage, as well as the preset set of 23 indicator evaluation rules;
[0471] Training objectives:
[0472] The indicator rule embedding weight set generated by the rule embedding unit has a weight sum of 1 within the same dimension, and the MAE of the weights is ≤0.02 with that of the expert-calibrated weights.
[0473] The 23×23 global correlation matrix generated by the indicator correlation matrix unit has elements with values ∈ [0,1] and the correlation error with the sample statistics is ≤0.03.
[0474] The final dynamic quantization feature of the teacher / student / teacher-student task progress output by the dynamic quantization unit, after L2 normalization, has a modulus of 1 and an MSE of ≤0.09 with the label.
[0475] 5. Fifth stage: Training of the three-dimensional linked output layer (121-150 rounds)
[0476] Input: The final dynamic quantization features output from the fourth stage;
[0477] Training objectives:
[0478] The 128-dimensional fused feature vector output by the FC1 layer can effectively represent the correlation patterns of the three dimensions.
[0479] The core quantitative values (0-10 points) of the three types of effective behaviors output by the FC2 layer have an MAE of ≤0.3 for expert labels and an MAE of ≤0.04 for subdivided quantitative items (such as the percentage of effective teaching time for teachers).
[0480] Key operations:
[0481] Unfreeze all module parameters and perform full-network fine-tuning;
[0482] In each round, the training set loss and the early stopping validation set loss are calculated. If the validation set MAE does not decrease for 5 consecutive rounds, early stopping is triggered, and the current optimal model is saved.
[0483] IV. Training Process Monitoring and Calibration
[0484] Feature distribution monitoring: Every 10 rounds, the mean and variance of the output features of each module are calculated to ensure that the feature values fall within the [0,1] interval (except for intermediate layer features) and there is no overflow;
[0485] Outlier sample handling: For outlier segments in the training set with feature values < 0.05, they are calibrated to 0.1 according to the technical solution to avoid affecting model convergence;
[0486] Source tag consistency verification: During training, source tags of each module are bound synchronously to ensure that the tag fields correspond one-to-one with the features without misalignment;
[0487] Hyperparameter optimization: The optimal learning rate (1e-4 / 5e-5 / 1e-5) and batch size (16 / 32 / 64) are selected through the validation set of hyperparameters, and finally the combination with the smallest MAE on the validation set is chosen.
[0488] V. Model Saving and Output
[0489] After training, save the complete 3D-LCBQN neural network model file, which includes:
[0490] Weight parameters for each module (convolutional layer, fully connected layer, gating unit, etc.);
[0491] Hyperparameter configuration (learning rate, batch size, loss function weights, etc.);
[0492] Training logs (training / validation loss per round, number of rounds triggered by early stopping, optimal metrics);
[0493] When deploying the model, the saved parameters are loaded and the preprocessed multimodal data is input. The model can then directly output a set of three types of valid behaviors with traceability labels, which is fully compatible with the reasoning process of the technical solution.
[0494] This application has the following advantages:
[0495] Multimodal data processing closed loop: Integrating multiple algorithms such as YOLOv8 object detection, Dlib key point extraction, and Whisper speech transcription, it realizes standardized processing of visual, audio, and text data across the entire chain of "acquisition-alignment-cleaning-feature extraction", solving the pain points of traditional evaluation data being single and having many interferences, and significantly improving data purity (signal-to-noise ratio ≥35dB after denoising) and time consistency (timestamp error ≤20ms).
[0496] The network architecture is designed in a scenario-based manner: the 3D-LCBQN neural network is designed in layers according to the teaching and assessment logic of "data → features → association → quantification → output". The input and output dimensions of each layer are precisely matched (all are 64-dimensional standardized features). This not only conforms to the feature deepening law of neural networks, but also closely follows the core classroom relationship of "teacher-student-teacher-student task advancement", breaking through the limitation of traditional models that "emphasize data and neglect scenarios".
[0497] Rigorous quantitative assessment logic: Construct a complete logical chain of 23 indicator rule set - correlation matrix - dynamic weight - quantitative output. Through mathematical methods such as Pearson correlation coefficient and normalized weight allocation, objective quantification of subjective rules is achieved. At the same time, traceability tags are bound to the entire process to ensure that the assessment results are traceable and verifiable, solving the defects of traditional manual assessment that are subjective and lack evidence.
[0498] Practical application of assessment results: Directly outputs a structured set of three types of effective behavior quantifications, including core scores on a 0-10 scale, as well as key behavioral sub-items (such as teacher speaking speed compliance rate and student head-up rate), while also marking the dimensions of shortcomings, realizing a closed loop of "assessment-feedback-optimization", and adapting to the needs of multiple scenarios such as teaching management and teacher research.
[0499] Strong cross-scenario adaptability: All core modules have reserved adjustable parameter interfaces (such as indicator rule sets and modal fusion weights), which can adapt to classroom assessments of different grades and subjects without reconstructing the model. At the same time, it is compatible with real-time video streams and offline recording analysis, and has high feasibility for engineering implementation.
[0500] The multimodal basic feature extraction module systematically extracts visual (7 types, such as movement trajectory and head-up rate), audio (5 types, such as speech rate and question frequency), text (3 types, such as teaching keywords), and timestamp features, comprehensively covering the core behaviors of teaching and learning. Each feature has a clear extraction algorithm (such as head-up rate through Dlib68 key point detection), resulting in high feature representation accuracy.
[0501] Standardized operations such as spectral subtraction denoising, light compensation, and NTP timestamp alignment ensure that the input data is pure and free of interference, providing a high-quality foundation for subsequent feature encoding and avoiding evaluation bias caused by defects in the original data.
[0502] The temporal-semantic-knowledge point anchoring coding layer has created a unique three-dimensional anchoring mechanism that integrates temporal, semantic, and knowledge point dimensions. By dynamically allocating weights through cross-axis fusion gating units (based on S(t), α, and β coefficients), it achieves for the first time an integrated representation of classroom data in terms of "time dimension (teaching stage), semantic dimension (voice and text), and knowledge point dimension (teaching syllabus matching)," which is more in line with the essence of teaching than single-dimensional coding.
[0503] The teaching behavior-cognitive state linkage extraction layer extracts the characteristics of teacher guidance and student cognitive engagement respectively. Through bidirectional correlation units, a teacher-student characteristic correlation matrix is constructed to quantify the dynamic adaptation relationship of "teacher guidance-student response" and break through the limitations of traditional "one-way assessment".
[0504] Strong feature enhancement adaptability: By employing operations such as 1×1 convolution dimensionality increase and lightweight network layer enhancement, the output linkage features (64 dimensions) are ensured to match the dimensions of the subsequent fusion layer, while retaining the core correlation information between teacher and student behavior and cognition, making the feature expression more targeted.
[0505] The knowledge point progression-assessment dimension dynamic fusion layer integrates the weight γ of the knowledge points in the teaching syllabus into the feature fusion process. Through the knowledge point progression adjustment unit, the feature fusion closely follows the teaching progression logic (introduction → new knowledge teaching → summary), avoiding the blindness of traditional fusion.
[0506] Modal sub-features are bound to the dimensions of teacher, student, and teacher-student task advancement, and dynamic weights are assigned with "dimensional-specific features as the core (weight 0.6-0.7) and modal fusion features as auxiliary (weight 0.3-0.4)" to ensure that the feature representation of each dimension is accurate and free from cross-interference.
[0507] The indicator association-rule embedding quantification layer categorizes and quantifies 23 natural language indicator rules into "numerical threshold, range interval, statistical calculation, and composite judgment" categories. Each rule has a clear scoring standard and corresponding feature component relationship, realizing the objective implementation of subjective rules.
[0508] Dynamic weighting through correlation matrix: Construct a 23×23 global indicator correlation matrix, quantify intra-dimensional / cross-dimensional correlations through Pearson correlation coefficient, dynamically adjust indicator weights based on correlation degree, ensure that weights adapt to the inherent laws between indicators, and improve the rationality of quantification results.
[0509] The three-dimensional linkage output layer achieves nonlinear fusion of three-dimensional dynamic quantization features through two fully connected layers (FC1+FC2). The FC1 layer (128 neurons) mines potential correlations of features, and the FC2 layer accurately maps to three types of effective behavioral quantization values, balancing fusion and mapping efficiency.
[0510] This application also provides an AI-assisted evaluation system for assessing the effectiveness of teachers' classroom teaching, the AI-assisted evaluation system for assessing the effectiveness of teachers' classroom teaching includes:
[0511] The information acquisition module is used to acquire the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments.
[0512] A neural network acquisition module, which is used to acquire a trained 3D-LCBQN neural network;
[0513] The effective behavior quantity acquisition module is used to input the denoised clean speech dataset, the light-optimized clear image dataset, and the selected effective audio segments into the 3D-LCBQN neural network, thereby generating the teacher effective behavior quantification set, the student effective behavior quantification set, and the teacher-student classroom teaching task promotion effective behavior quantification set.
[0514] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for evaluating the effectiveness of AI-assisted classroom teaching, characterized in that, The methods for evaluating the effectiveness of AI-assisted classroom teaching include: Obtain the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments; Obtain the trained 3D-LCBQN neural network; The denoised clean speech dataset, the light-optimized clear image dataset, and the selected effective audio segments are input into the 3D-LCBQN neural network to generate a quantitative set of effective teacher behaviors, a quantitative set of effective student behaviors, and a quantitative set of effective behaviors in promoting classroom teaching tasks between teachers and students.
2. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 1, characterized in that, The acquisition of the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments includes: Acquire classroom visual data and classroom audio data; The YOLOv8 object detection algorithm is used to identify classroom video data, thereby obtaining a frame-by-frame object detection result set; Facial key points were extracted from classroom video data using Dlib facial key point extraction, thereby obtaining a frame-by-frame facial key point dataset. A classroom original image dataset with frame numbers is generated based on the frame-by-frame target detection result set and the frame-by-frame facial key point dataset. Preprocess the classroom audio data to obtain a time-stamped raw classroom speech dataset; Multimodal data timestamp alignment is performed on the original classroom image dataset numbered by frame and the original classroom audio dataset with timestamps to obtain the timestamp-aligned original classroom image dataset and original classroom audio dataset. A real-time denoising algorithm is performed on the original classroom audio dataset to obtain a clean audio dataset after denoising. The image data is corrected for lighting using a lighting compensation algorithm, thereby obtaining a clear image dataset with optimized lighting. By using speech semantic analysis algorithms to filter out invalid speech in the original classroom speech dataset, valid audio segments are obtained after filtering.
3. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 2, characterized in that, The 3D-LCBQN neural network structure includes: A multimodal basic feature extraction module is used to generate basic features based on the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered effective audio segments. The basic features include visual features, audio features, text features, and timestamp feature sets. The temporal-semantic-knowledge point anchoring coding layer is used to generate visual anchoring coding features, audio anchoring coding features, text anchoring coding features, and a three-dimensional label set based on visual features, audio features, text features, and timestamp feature sets. The teaching behavior-cognitive state linkage extraction layer is used to generate teacher guidance-behavior linkage features and student cognition-behavior linkage features based on the visual anchoring coding features, audio anchoring coding features, and text anchoring coding features. The knowledge point progression-assessment dimension dynamic fusion layer is used to generate teacher-specific fusion features, student-specific fusion features, and teacher-student task progression dimension-specific fusion features based on visual anchoring coding features, audio anchoring coding features, and text anchoring coding features. The indicator association-rule embedding quantification layer is used to generate the teacher's final dynamic quantification features, the student's final dynamic quantification features, and the teacher-student classroom teaching task progress quantification score information based on the teacher-dimensional exclusive fusion features, the student-dimensional exclusive fusion features, the teacher-student task progress exclusive fusion features, the teacher guidance-behavior linkage features, and the student cognition-behavior linkage features. The three-dimensional linkage output layer is used to calibrate the teacher's final dynamic quantitative characteristics, the student's final dynamic quantitative characteristics, and the quantitative score information of the progress of teacher and student classroom teaching tasks, and output the quantitative set of effective teacher behavior with traceability labels, the quantitative set of effective student behavior with traceability labels, and the quantitative set of effective teacher and student classroom teaching task progress with traceability labels.
4. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 3, characterized in that, The visual features include: movement trajectory features, head-up rate, frowning duration, writing action, front row seating rate, blackboard area features, teaching aids, and projection features. Audio features include: speech rate features, frequency of spoken words features, frequency of questions features, student speaking rate features, and interactive response delay features; Textual features include: teaching keyword features, teaching segment keyword features, and key and difficult point keyword features.
5. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 4, characterized in that, The time-series-semantic-knowledge point anchoring encoding layer includes: A visual channel, used to generate visually encoded features based on the denoised clean speech dataset; A temporal anchoring unit is used to generate a temporal anchoring coding vector based on a light-optimized clear image dataset. A semantic anchoring unit is used to generate a semantic anchoring encoding vector based on the filtered valid audio segments; A knowledge point anchoring unit, which is used to generate a knowledge point anchoring encoding vector based on a semantic anchoring encoding vector; A cross-axis fusion gating unit is used to fuse temporal anchoring encoding vectors, semantic anchoring encoding vectors, knowledge point anchoring encoding vectors, and visual encoding features to generate visual anchoring encoding features, audio anchoring encoding features, text anchoring encoding features, and a three-dimensional tag set.
6. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 5, characterized in that, The teaching behavior-cognitive state linkage extraction layer includes: The teacher-side collaborative extraction unit is used to generate teacher guidance features based on visual features, audio features, and text features. The student-side linkage extraction unit is used to generate student cognitive engagement features based on visual features, audio features, and text features. A bidirectional association unit is used to generate teacher guidance-behavior linkage features and student cognitive-behavior linkage features based on teacher guidance characteristics and student cognitive engagement characteristics.
7. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 6, characterized in that, The knowledge point progression-assessment dimension dynamic fusion layer includes: A modal adaptation and fusion unit is used to generate modal fusion features based on visual anchoring coding features, audio anchoring coding features, and text anchoring coding features; A dimension binding fusion unit is used to generate initial fusion features for the teacher dimension, initial fusion features for the student dimension, and initial fusion features for the teacher-student task progression dimension based on the modal fusion features. The knowledge point progressive adjustment unit is used to generate teacher-specific fusion features, student-specific fusion features, and teacher-student task progress-specific fusion features based on the teacher-dimensional initial fusion features, student-dimensional initial fusion features, and teacher-student task progress-specific fusion features.
8. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 7, characterized in that, The indicator association-rule embedding quantization layer includes: A rule preset unit, which is used to store a set of indicator evaluation rules; The rule embedding unit is used to generate an indicator rule embedding weight set based on the stored indicator evaluation rule set, teacher guidance-behavior linkage features, student cognition-behavior linkage features, teacher-dimensional exclusive fusion features, student-dimensional exclusive fusion features, and teacher-student task advancement dimension exclusive fusion features. The indicator association matrix unit is used to generate intra-dimensional association matrices and cross-dimensional association matrices based on teacher guidance-behavior linkage characteristics and student cognition-behavior linkage characteristics. The dynamic quantification unit is used to generate the final dynamic quantification features of teachers, the final dynamic quantification features of students, and the quantitative score information of the progress of classroom teaching tasks of teachers and students based on the teacher-specific fusion features, student-specific fusion features, teacher-student task progress-specific fusion features, indicator rule embedded weight set, teacher guidance-behavior linkage features, and student cognition-behavior linkage features.
9. The method for evaluating the effectiveness of AI-assisted classroom teaching as described in claim 8, characterized in that, The three-dimensional linkage output layer includes: The first-level fully connected hidden layer is used to generate a 128-dimensional fusion feature vector based on the teacher's final dynamic quantitative features, the student's final dynamic quantitative features, and the quantitative score information of the teacher and student's classroom teaching tasks. The second-level fully connected output layer is used to generate core quantitative values of effective teacher behavior, core quantitative values of effective student behavior, and core quantitative values of effective teacher-student task promotion behavior based on the 128-dimensional fused feature vector.
10. A system for evaluating the effectiveness of teacher classroom teaching based on AI assistance, characterized in that, The AI-assisted teacher classroom teaching effectiveness evaluation system includes: The information acquisition module is used to acquire the denoised clean speech dataset, the light-optimized clear image dataset, and the filtered valid audio segments. A neural network acquisition module, which is used to acquire a trained 3D-LCBQN neural network; The effective behavior quantity acquisition module is used to input the denoised clean speech dataset, the light-optimized clear image dataset, and the selected effective audio segments into the 3D-LCBQN neural network, thereby generating the teacher effective behavior quantification set, the student effective behavior quantification set, and the teacher-student classroom teaching task promotion effective behavior quantification set.