A teacher classroom behavior recognition and teaching content analysis method and system
Patent Information
- Application Number
- CN202611150442.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-31
AI Technical Summary
[0006]本发明的一个目的在于提出一种教师课堂行为识别与教学内容分析方法及系统,针对现有技术中课堂多模态数据受遮挡、噪声、视角变化和时间不同步影响,且行为识别结果难以与具体教学内容及教学环节对应的问题,提出了基于教学事件锚点对齐、模态可靠度调节、行为内容耦合图构建和多模态行为内容耦合图Transformer推理的技术方案,本发明具有提高复杂课堂场景下行为片段识别稳定性、内容对应准确性和课堂过程分析连续性的技术效果
[0051]1、通过依据课件翻页、板书新增、教师指向、语音关键词、提问句式、停顿和教师朝向变化生成教学事件锚点,并结合锚点相似度、局部时间窗口和模态可靠度向量进行动态时间对齐,使异步采集的教师视频、课堂音频、课件画面和板书图像能够形成统一的多模态教学事件单元,降低时间不同步对行为片段识别的影响。
Smart Images

Figure CN122657941B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational informatization and multimodal classroom intelligent analysis, and in particular to a method and system for identifying teacher classroom behavior and analyzing teaching content. Background Technology
[0002] With the development of recording studios, smart classrooms, and online teaching platforms, classroom process analysis is gradually shifting from manual observation and evaluation to automated analysis based on video, audio, courseware images, and blackboard writing. Existing solutions typically utilize single technologies such as posture recognition, speech recognition, or text recognition to detect teacher behaviors such as lecturing, blackboard writing, questioning, and movement, or to extract knowledge points from courseware and blackboard writing, thereby providing a data foundation for classroom quality evaluation and teaching process review.
[0003] However, in real classroom recordings or smart classroom scenarios, teachers' postures and gestures are easily obstructed by students, the podium, and equipment. Audio is also affected by environmental noise and multiple people talking. Furthermore, the presentation slides and blackboard images may produce recognition errors due to changes in viewing angle, glare, and unclear writing increments. Additionally, the timestamps of different acquisition devices are not entirely consistent, and there are often short delays between teacher pointing, voice keywords, slide page turning, and new blackboard writing, making it difficult to stably align multi-source data.
[0004] Furthermore, existing classroom analytics often separates behavior recognition from the understanding of teaching content, resulting in only discrete behavior labels or content texts. It is difficult to determine which specific knowledge point, courseware area, blackboard area, or teaching segment a particular lecturing, questioning, blackboard writing, or pointing behavior corresponds to. This leads to fragmented classroom process analysis results that cannot accurately reflect the actual teaching status of teachers' lecturing, interaction, blackboard writing, and content organization.
[0005] Therefore, there is a need for a method and system for identifying teacher classroom behavior and analyzing teaching content that can overcome the shortcomings of existing technologies. Summary of the Invention
[0006] One objective of this invention is to propose a method and system for teacher classroom behavior recognition and teaching content analysis. Addressing the problems in existing technologies where multimodal classroom data is affected by occlusion, noise, perspective changes, and time asynchrony, and where behavior recognition results are difficult to correlate with specific teaching content and steps, this invention proposes a technical solution based on teaching event anchor point alignment, modal reliability adjustment, behavior-content coupling graph construction, and Transformer inference for the multimodal behavior-content coupling graph. This invention effectively improves the stability of behavior segment recognition, the accuracy of content correspondence, and the continuity of classroom process analysis in complex classroom scenarios.
[0007] This invention provides a method for identifying teacher classroom behavior and analyzing teaching content, including:
[0008] S1. Acquire teacher videos, classroom audio, courseware images and blackboard images during the same class period, extract teacher posture, movement, gestures, voice text, speech rhythm, courseware text, blackboard increments and knowledge point phrases to form a multimodal feature sequence, and generate a modal reliability vector from the acquisition quality parameters of each modality;
[0009] S2. Based on the changes in courseware page turning, blackboard writing, teacher pointing, voice keywords, question sentence structure, pauses and teacher orientation, generate a teaching event anchor sequence in the multimodal feature sequence. Each teaching event anchor includes anchor time, source modality, feature representation and reliability value.
[0010] S3. Based on the anchor point feature representation similarity, local time window and modal reliability vector, perform dynamic time alignment on the asynchronous multimodal feature sequence to obtain multimodal teaching event units;
[0011] S4. Generate behavior nodes, content nodes, and teaching link nodes based on the multimodal teaching event units, and determine the edge weights between nodes according to time overlap, spatial orientation, semantic similarity, content progression order, and the modal reliability vector to construct a behavior-content coupling graph.
[0012] S5. Input the behavior content coupling graph into the trained multimodal behavior content coupling graph Transformer, and output the teacher behavior segment, behavior category, corresponding knowledge point, corresponding courseware area or blackboard area, teaching link and confidence level.
[0013] Optionally, S1 includes:
[0014] Human key point detection and target tracking are performed on teacher videos to obtain posture integrity rate, tracking continuity, movement trajectory and gesture pointing features;
[0015] Speech recognition, sentence structure recognition, and prosodic analysis were performed on classroom audio to obtain speech-text features, question sentence structure tags, pause duration, and speech prosodic features.
[0016] The courseware images and blackboard images are subjected to text recognition, page stability detection and image difference to obtain courseware text features, courseware page numbers, incremental areas of blackboard writing and incremental text of blackboard writing;
[0017] The modal reliability vector is formed by normalizing the posture integrity rate, tracking continuity, hand clarity, speech recognition word confidence, text recognition character confidence, courseware page stability, and blackboard incremental clarity to the interval of 0 to 1.
[0018] Optionally, S2 includes:
[0019] When the courseware page number changes, the area of the blackboard writing increment is greater than the preset area threshold, the gesture pointing ray intersects with the courseware area or the blackboard writing area, the voice text contains knowledge point phrases, the voice text conforms to the preset interrogative sentence pattern, the pause duration is greater than the preset pause threshold, or the change in the teacher's facing angle is greater than the preset angle threshold, a corresponding teaching event anchor point is generated.
[0020] The timestamp of the triggering event is used as the anchor time, the modality to which the triggering event belongs is used as the source modality, the features within a preset time before and after the triggering event are concatenated and encoded into the feature representation, and the corresponding reliability value is read from the modality reliability vector.
[0021] Optionally, S3 includes:
[0022] For any two source modalities of teaching event anchor points, calculate the cosine similarity of the feature representations;
[0023] The anchor point time difference is not more than the preset window length as the candidate alignment condition;
[0024] The anchor point matching score is obtained by multiplying the cosine similarity by the corresponding reliability value.
[0025] Dynamic programming is used to select the anchor matching path that maximizes the sum of matching scores;
[0026] The modal feature sequences are resampled over time according to the anchor point matching path to obtain the multimodal teaching event unit.
[0027] Optionally, S4 includes:
[0028] Encode the teacher's posture, movement, gestures, and speech behavior as behavior nodes;
[0029] Encode the text blocks in the courseware, the incremental text on the blackboard, and the knowledge point phrases into content nodes;
[0030] Encode one of the following identification results—import, lecture, questioning, blackboard writing, interaction, and summary—as a teaching segment node;
[0031] For any two nodes, calculate the temporal overlap ratio, spatial orientation hit value, semantic similarity, and content-advancement adjacency value respectively, and then sum the above four types of values according to the modal reliability vector to obtain the edge weight between the nodes;
[0032] Furthermore, before constructing the behavior content coupling graph, the method further includes: writing knowledge points, blackboard phrases, courseware titles and voice keywords that have been confirmed in the historical teaching event unit and have a confidence level greater than a preset confirmation threshold into the knowledge point memory queue;
[0033] The memory weights are obtained by exponentially decaying the entries in the knowledge point memory queue according to the difference between the current time and the entry time.
[0034] The cosine similarity between the current content features and the semantic vector of the content node is determined as the content node matching score;
[0035] When the matching score of the content node is less than the preset matching threshold, the entries with a memory weight greater than the preset memory threshold are added as candidate content nodes to the behavior content coupling graph.
[0036] Furthermore, the calculation of the spatial pointing hit value includes: fitting a hand gesture pointing ray based on the key points of the teacher's wrist and index finger;
[0037] Project the gesture pointing ray onto the coordinate system of the courseware screen or the coordinate system of the blackboard image;
[0038] When the projection ray intersects the outer frame of the text block, set the space pointer hit value to 1;
[0039] When the projection ray does not intersect the outer frame of the text block and the shortest distance between them is not greater than a preset distance threshold, a spatial pointing hit value between 0 and 1 is determined according to the ratio of the shortest distance to the preset distance threshold.
[0040] When the shortest distance is greater than the preset distance threshold, the spatial pointing hit value is set to 0.
[0041] Optionally, S5 includes:
[0042] The node input representation is obtained by adding the node features, node type encoding, and time position encoding of each node in the behavior content coupling graph;
[0043] In the attention layer of the multimodal behavior content coupling graph Transformer, the edge weights between nodes are used as attention bias terms, and the reliability value of the source modality is used as a multiplicative adjustment term for the attention score.
[0044] Based on the output node representation, decode the start and end time of the behavior, the behavior category, the knowledge point identifier, the coordinates of the courseware area, the coordinates of the blackboard area, the teaching link label, and the confidence level respectively;
[0045] Furthermore, the training of the multimodal behavior content coupling graph Transformer includes: randomly occluding at least one modality among teacher video, classroom audio, courseware screen or blackboard image in the training samples, inputting the teaching event anchor point representation of the unoccluded modality into the compensation encoder, and reconstructing the anchor point representation of the occluded modality;
[0046] The model parameters are updated by combining the distance loss between the reconstructed anchor representation and the real anchor representation, the behavior classification loss, the knowledge point matching loss, and the teaching link classification loss.
[0047] During the inference phase, when the reliability value of a certain mode is less than the preset reliability threshold, the compensation encoder is used to generate the compensation anchor point representation of that mode, and the compensation anchor point representation is input into steps S3 and S4.
[0048] On the other hand, the present invention also provides a teacher classroom behavior recognition and teaching content analysis system, comprising:
[0049] The module for feature extraction and reliability generation acquires teacher videos, classroom audio, courseware images, and blackboard images to generate multimodal feature sequences and modal reliability vectors. The module for anchor point generation generates teaching event anchor point sequences based on courseware page turning, new blackboard entries, teacher pointing, voice keywords, question formats, pauses, and changes in teacher orientation. The module for dynamic time alignment generates multimodal teaching event units based on anchor point feature representation similarity, local time windows, and modal reliability vectors. The module for coupled graph construction generates behavior nodes, content nodes, and teaching segment nodes and determines the edge weights between nodes. The module for graph Transformer inference outputs teacher behavior fragments, behavior categories, corresponding knowledge points, corresponding courseware or blackboard areas, teaching segments, and confidence levels based on the behavior-content coupled graph.
[0050] The beneficial effects of this invention are:
[0051] 1. By generating teaching event anchor points based on changes in courseware page turning, blackboard additions, teacher pointing, voice keywords, question sentence structure, pauses, and teacher orientation, and combining anchor point similarity, local time windows, and modal reliability vectors for dynamic time alignment, asynchronously acquired teacher videos, classroom audio, courseware images, and blackboard images can form a unified multimodal teaching event unit, reducing the impact of time asynchrony on behavioral segment recognition.
[0052] 2. By constructing a behavior-content coupling graph based on the recognition results of posture, movement, gesture, speech behavior, courseware text blocks, blackboard incremental text, knowledge point phrases, and teaching links, and determining the edge weights based on temporal overlap, spatial orientation, semantic similarity, and content progression order, the teacher's behavior, knowledge points, courseware areas, blackboard areas, and teaching links can be jointly reasoned within the same graph structure, reducing the problem of the separation between behavior labels and content understanding results.
[0053] 3. By introducing modal reliability vectors, occlusion reconstruction training, low-reliability modal compensation, and knowledge point memory queues, reliable modalities and recently confirmed content can still be used to supplement the reasoning basis when there is occlusion, noise, decreased recognition confidence, or insufficient content matching in teacher videos, classroom audio, courseware images, or blackboard images. This improves the stability of analysis results and the accuracy of behavioral content correspondence in complex classroom environments. Attached Figure Description
[0054] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0055] Figure 1 This is a flowchart of a method for identifying teacher classroom behavior and analyzing teaching content.
[0056] Figure 2 This is a flowchart for constructing the behavior-content coupling graph in step S4 of the present invention. Detailed Implementation
[0057] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0058] refer to Figures 1-2 A method for identifying teacher classroom behavior and analyzing teaching content, comprising:
[0059] S1. Acquire teacher videos, classroom audio, courseware images and blackboard images during the same class period, extract teacher posture, movement, gestures, voice text, speech rhythm, courseware text, blackboard increments and knowledge point phrases to form a multimodal feature sequence, and generate a modal reliability vector from the acquisition quality parameters of each modality;
[0060] S2. Based on the changes in courseware page turning, blackboard writing, teacher pointing, voice keywords, question sentence structure, pauses and teacher orientation, generate a teaching event anchor sequence in the multimodal feature sequence. Each teaching event anchor includes anchor time, source modality, feature representation and reliability value.
[0061] S3. Based on the anchor point feature representation similarity, local time window and modal reliability vector, perform dynamic time alignment on the asynchronous multimodal feature sequence to obtain multimodal teaching event units;
[0062] S4. Generate behavior nodes, content nodes, and teaching link nodes based on the multimodal teaching event units, and determine the edge weights between nodes according to time overlap, spatial orientation, semantic similarity, content progression order, and the modal reliability vector to construct a behavior-content coupling graph.
[0063] S5. Input the behavior content coupling graph into the trained multimodal behavior content coupling graph Transformer, and output the teacher behavior segment, behavior category, corresponding knowledge point, corresponding courseware area or blackboard area, teaching link and confidence level.
[0064] In this specific embodiment, S1 includes:
[0065] Four raw sequences were generated by using teacher video, classroom audio, courseware images, and blackboard images from the same class period as input data and assigning each a timestamp. The teacher video was captured at a specific resolution. Furthermore, the RGB frame sequence has a frame rate of 25fps, the classroom audio uses a mono waveform with a sampling rate of 16kHz and a quantization bit width of 16bit, the courseware image is captured from the same screen as the teacher's computer and outputs a frame sequence at 5fps while simultaneously parsing to obtain the courseware page number, and the blackboard image uses a fixed-position camera to output a frame sequence at 1fps and pre-calibrates to obtain the quadrilateral ROI of the blackboard in the image for subsequent difference analysis.
[0066] Human detection is performed frame by frame on the teacher video, and only the teacher's human bounding boxes with a confidence threshold of 0.50 are retained. The human detection network adopts a single-stage convolutional detector structure and outputs bounding boxes of the category "person" and their corresponding confidence scores.
[0067] Human keypoint detection is performed within the bounding box to obtain the teacher's pose features. The keypoint detection network adopts a top-down heatmap regression structure with ResNet-50 as the backbone network. The input is a cropped and scaled image. The teacher's human body bounding box image is generated and 17 key point heatmaps are output. Soft-argmax is used to obtain the two-dimensional coordinates of the key points from each heatmap and the confidence of the key points is obtained simultaneously. The number of key points with a confidence of not less than 0.30 is counted from the 17 key point confidence scores and divided by 17 to obtain the pose integrity rate.
[0068] Cross-frame correlation is performed on the teacher target to obtain the movement trajectory and tracking continuity. Target tracking adopts a multi-target tracking structure of "Kalman filter state prediction + appearance feature matching + Hungarian matching", where the Kalman filter state vector is taken as... And the observation is the detection frame The appearance feature extraction network adopts a lightweight convolutional network and outputs a 128-dimensional L2 normalized feature vector. The appearance cosine distance threshold is set to 0.20 and the detection box IoU threshold is set to 0.30 to complete the association. The tracking continuity is defined as the ratio of the number of frames in which the teacher's trajectory is effectively associated within the effective class period to the total number of frames. The movement trajectory is defined as the two-dimensional position sequence of the center point of the teacher's human body frame in the image coordinate system in each frame.
[0069] When extracting gesture pointing features, the wrist key points and the index finger tip key points are read from the key points and the line connecting the two is used as the image plane representation of the gesture pointing ray. At the same time, the mean confidence of the wrist and index finger tip key points is used as the hand sharpness.
[0070] The classroom audio is first subjected to speech activity detection to remove silent segments and the pause duration is calculated accordingly. The speech activity detection uses short-time energy and zero-crossing rate with a frame length of 20ms and a frame shift of 10ms, and the silence threshold is fixed at 0.35 times the average energy. The pause duration is defined as the duration of the silence interval between two adjacent speech activities.
[0071] Speech recognition is performed on the speech activity segment to obtain speech text features and word confidence. The speech recognition adopts an end-to-end Conformer structure with a 12-layer encoder, 4 attention heads, and a model dimension of 256. CTC and attention are used for joint decoding. The posterior probability of each word is used as the word confidence, and the word sequence concatenated in time order is used as the speech text feature.
[0072] Sentence recognition is performed on the speech text to obtain question sentence labels. The sentence recognition adopts a classifier structure of "word embedding after word segmentation + bidirectional LSTM + fully connected Softmax". The word embedding dimension is 128, the bidirectional LSTM hidden state dimension is 128, and the output category set is {declarative sentence, interrogative sentence}. The category with the highest probability in Softmax is used as the question sentence label.
[0073] When extracting prosodic features, the fundamental frequency is calculated for each speech activity segment. mean Variance, short-time energy mean, and energy variance are combined with the corresponding speech text segments to form speech prosodic features;
[0074] Page stability detection is performed on the courseware screen to obtain the page stability and text features within the page. Page stability detection uses ORB feature point matching between adjacent frames and estimates the homography matrix using RANSAC. The page stability is defined as the ratio of the number of points in RANSAC to the total number of matching points, and the page number is directly output by the screen acquisition interface. The text features of the courseware are obtained through text recognition, which uses a CRNN structure consisting of a 7-layer convolutional feature extractor and a 2-layer bidirectional LSTM sequence modeler. The bidirectional LSTM hidden state dimension is 256, and CTC decoding is used to output the character sequence. At the same time, the mean posterior probability of characters on the CTC path is used as the confidence of the character recognition, and the character sequence is organized into the text features of the courseware according to the bounding box of the text block.
[0075] Registration and image differencing are performed on the blackboard image within the blackboard ROI to obtain the incremental regions and incremental text of the blackboard. Registration is performed by ORB matching of adjacent frames and alignment is completed by estimating the homography matrix using RANSAC. The difference image is binarized with a fixed threshold, which is the mean of the difference grayscale plus twice the standard deviation. Connectivity analysis is then performed. Connectivity regions with an area greater than 300 pixels are defined as the incremental regions of the blackboard, and CRNN text recognition is performed on them to obtain the incremental text of the blackboard. At the same time, the normalized value of the Laplacian variance within the incremental region is used as the incremental sharpness of the blackboard.
[0076] The voice text, courseware text features, and incremental blackboard text are all fed into the knowledge point phrase extraction module to obtain knowledge point phrase features. The knowledge point phrase extraction module uses maximum matching word segmentation based on the course knowledge point vocabulary and filters out noun phrases and proper noun phrases that match the vocabulary in the segmentation results. The top 5 phrases with the highest TF-IDF scores in each time window are taken as knowledge point phrases and a 128-dimensional semantic vector is generated for each phrase. The semantic vector is obtained by taking the average of the Skip-gram word vectors trained on the textbook corpus and the word vectors within the phrase.
[0077] Teacher posture, movement, gestures, audio text, speech rhythm, courseware text, blackboard increments, and knowledge point phrases are organized into multimodal feature sequences according to their respective timestamps, and modal reliability vectors are generated. The modal reliability vectors are constructed as follows:
[0078] ;
[0079] in This represents the modal reliability vector, and each dimension takes values ranging from 1 to 2. The pose integrity rate is represented by linear normalization and clipped to [value]. The value after that, The tracking continuity is linearly normalized and clipped to... The value after that, This indicates that the hand sharpness has been linearly normalized and cropped to... The value after that, This represents the mean of the speech recognition word confidence score within the current time window, linearly normalized and cropped to... The value after that, This represents the mean of the character recognition confidence score within the current time window, linearly normalized and cropped to [value missing]. The value after that, This indicates that the stability of the courseware pages has been linearly normalized and cropped to... The value after that, This indicates that the incremental clarity of the blackboard writing has been linearly normalized and cropped to [a certain value]. The value after that.
[0080] In this specific embodiment, S2 includes:
[0081] Based on multimodal feature sequences and modal reliability vectors Generate a sequence of teaching event anchor points and arrange them in ascending order of anchor point time. The anchor point triggering rules are executed on the same time axis and the timestamps are all converted to the relative time of the class start time.
[0082] When the page number changes between two adjacent courseware frames, the timestamp of the courseware frame where the page number change occurs is set as the anchor time, and a "courseware page turn" anchor is generated. Simultaneously, the source modality is set to the courseware frame, and the reliability value is set to [value missing]. ;
[0083] When the proportion of the incremental area of the whiteboard writing region within the whiteboard writing ROI to the total area of the whiteboard writing ROI is greater than a threshold Furthermore, when this condition is met for the first time, the timestamp of the whiteboard image frame is set as the anchor time and a "new whiteboard" anchor point is generated. Simultaneously, the source modality is set to the whiteboard image and the reliability value is set to [value missing]. The incremental area of the whiteboard and the whiteboard ROI both come from S1, and the area ratio is obtained by dividing the number of pixels in the binary differential connected component by the number of pixels in the whiteboard ROI.
[0084] When the teacher's gesture points to the ray in continuous When a frame of the teacher's video intersects with the bounding box of the current courseware text block or the bounding box of the current whiteboard increment area, the timestamp of the first frame of the teacher's video that meets the intersection condition is set as the anchor time and a "teacher pointing" anchor point is generated. Simultaneously, the source modality is set to teacher video and the reliability value is set to [value missing]. The gesture pointing to the ray is from The wrist key points and the index fingertip key points are fitted together, and the bounding box is the bounding box of the courseware text block or the bounding box of the blackboard incremental area. The intersection determination adopts the intersection point determination of the two-dimensional ray and the axis-aligned rectangle, and the existence of an intersection point is taken as the intersection is established.
[0085] When the speech text output by speech recognition contains a knowledge point phrase within a sentence time period, the start timestamp of that knowledge point phrase in the speech recognition time stamp is set as the anchor time, and a "speech keyword" anchor is generated. Simultaneously, the source modality is set to classroom audio, and the reliability value is set to [value missing]. ;
[0086] When the sentence structure recognition output is a question tag, the timestamp at the end of the sentence is set as the anchor time, and a "question sentence" anchor is generated. Simultaneously, the source modality is set to classroom audio, and the reliability value is set to [value missing]. ;
[0087] When the duration of the silence interval between adjacent speech segments detected by speech activity exceeds the threshold At that time, the end timestamp of the silence interval is set as the anchor time and a "pause" anchor is generated. Simultaneously, the source modality is set to classroom audio and the reliability value is set to [value missing]. ;
[0088] When the teacher's facing angle is continuous The change in a teacher's video frame relative to the previous frame is greater than a threshold. At that time, the timestamp of the first teacher video frame that meets the threshold condition is set as the anchor time and an "orientation change" anchor is generated. Simultaneously, the source modality is set to the teacher video and the reliability value is set to [value missing]. The teacher's facing angle is obtained from the yaw angle output by the head pose estimation network, which uses a ResNet-18 backbone and regresses 3D Euler angles in the teacher's face region. , with yaw angle As the orientation angle and calculated for adjacent frames. As a variable;
[0089] For each trigger anchor point, a preset duration will be set before and after the trigger event. The multimodal features within the range are concatenated and encoded to form anchor feature representations. Specifically, the anchor time is used as the center within the interval. Internal sampling frequency Construction length For each time index, the feature vector with the closest timestamp and a time difference of no more than 0.1s is retrieved from the teacher video feature sequence, classroom audio feature sequence, courseware screen feature sequence and blackboard image feature sequence, and is then concatenated in a fixed order to form a time feature. If a feature vector that meets the conditions is not retrieved for a certain modality at that time, it is filled with a zero vector of the same dimension.
[0090] The result The features at each time step are input into the anchor encoder in chronological order to represent the output anchor features. The anchor encoder adopts a two-layer TransformerEncoder structure with 4 self-attention heads in each layer, a model dimension of 256, a feedforward network hidden dimension of 512, and the activation function is GELU. Dropout with a dropout rate of 0.1 is applied after each layer. The aggregated label vector at the beginning of the sequence is used as the output anchor feature representation.
[0091] Each anchor point is uniformly represented as ,in Indicates the first An anchor point for a teaching event, This indicates the anchor time and takes the timestamp of the triggering event. This indicates the source modality and takes one of the following: teacher video, classroom audio, courseware images, or blackboard images. This represents the anchor point feature representation and is taken as the 256-dimensional vector output by the anchor point encoder described above. Represents the reliability value and is derived from the modal reliability vector. The values are read from the source modality mapping and their range is: ;
[0092] When in time interval When multiple anchor points exist, only the reliability value of one of them is retained. The largest anchor point is selected and the rest are discarded to ensure the temporal sparsity and stability of the anchor point sequence.
[0093] In this specific embodiment, S3 includes:
[0094] The teaching event anchors are grouped into four anchor sequence groups according to their source modality and arranged in ascending order of anchor time. The teacher video anchor sequence is denoted as the anchor set whose source modality is teacher video, the classroom audio anchor sequence is denoted as the anchor set whose source modality is classroom audio, the courseware screen anchor sequence is denoted as the anchor set whose source modality is courseware screen, and the blackboard image anchor sequence is denoted as the anchor set whose source modality is blackboard image.
[0095] Using the classroom audio anchor sequence as a reference sequence and its anchor time as the alignment time axis, dynamic time alignment is performed on the teacher video anchor sequence, the courseware screen anchor sequence, and the blackboard image anchor sequence with the classroom audio anchor sequence in turn;
[0096] For any set of two source modes to be aligned and The anchor point sequence, take the first... Anchor points are And take the first Anchor points are ,in and These represent the anchor timestamps, and These represent the source modality identifiers, and These represent the 256-dimensional anchor point feature representations output by S2. and These represent the modal reliability vectors respectively. The reliability value read and its range is: ;
[0097] With local time window length Constrain candidate alignment pairs only if Anchor point matching scores are calculated only when the conditions are met, and other combinations are considered unmatchable; anchor point matching scores are calculated for anchor point pairs that meet the candidate criteria. And used for subsequent dynamic programming, the anchor point matching score is defined by the following formula:
[0098] ;
[0099] in Indicates source mode The Anchor points and source modes The Matching score for each anchor point This represents the dot product of the vectors representing the features of two anchor points. and Let L2 and L3 respectively represent the feature representations of the corresponding anchor points. Represents cosine similarity. and These represent the reliability values of the two anchor points and are used to adjust the reliability of the cosine similarity.
[0100] Based on the above matching scores, a size of [size] is constructed. The score matrix is used to search for the matching path with the largest sum of cumulative matching scores under the constraints that the anchor time is monotonically non-decreasing and the anchor index is monotonically increasing. The dynamic programming allows three transition operations, including "matching", which means advancing simultaneously. and And accumulate ,"jump over "That is, only advance" And accumulate fixed skip penalties ,"jump over "That is, only advance" And accumulate fixed skip penalties And record the optimal predecessor of each state in the dynamic programming table to complete backtracking;
[0101] The output of the matching path obtained by backtracking through dynamic programming consists of a set of ordered matching pairs. The anchor point matching relationship is constructed and used as a time alignment map between two modes, where the order of the matching pairs satisfies and And it is used to ensure the consistency of the order of events across modalities;
[0102] After three alignments with the classroom audio as the reference sequence, each classroom audio anchor point and its matching anchor points in the teacher video, courseware screen, and blackboard image are merged into the same alignment group, and the time of the classroom audio anchor point is used as the alignment group time. At the same time, empty anchor points are filled in the alignment group for unmatched modalities, and the reliability value of the empty anchor point is set to 0 to explicitly indicate the absence.
[0103] Time resampling is performed on each modal feature sequence according to the alignment group time to generate multimodal teaching event units. Specifically, for each alignment group time... In the interval Internally with a fixed sampling frequency Generate time sampling points and retrieve the feature vector in each modal feature sequence that is closest to the timestamp of each sampling point and has a time difference of no more than 0.1s as the modal feature of that sampling point. If a sampling point does not have a feature vector that satisfies the time difference constraint in a certain mode, it is filled with a zero vector of the same dimension and the modal reliability value corresponding to that sampling point is set to 0 at the same time.
[0104] The sampling point features of each modality within the same alignment group time window are stacked in chronological order and together with the reliability value of the modality in the alignment group to form the multimodal teaching event unit of the alignment group. The multimodal teaching event units of all alignment groups are output in chronological order for S4 to construct the behavior content coupling graph.
[0105] In this specific embodiment, S4 includes:
[0106] Using a sequence of multimodal teaching event units arranged in chronological order as input and the most recent The current graph construction window is composed of multiple multimodal teaching event units. For each multimodal teaching event unit in the window, behavior nodes, content nodes and teaching link nodes are generated and a behavior-content coupling graph is uniformly constructed in the window.
[0107] Behavioral nodes are obtained by encoding the teacher's posture, movement, gesture pointing, and speech behavior. The teacher's posture is calculated from the sequence of key points in the teacher's video and is based on the head yaw angle. Angle with torso Together, we determine the three-value attitude labels {facing students, facing the blackboard, turning to the side}, and the head yaw angle. The network output is used in conjunction with the head pose estimation method of S2 and is based on... Determined to be student-oriented The orientation is determined to be facing the blackboard, and otherwise turned to the side. This orientation label is then one-hot encoded into an orientation state vector. The movement state is calculated from the movement trajectory obtained by teacher target tracking, and the displacement velocity of adjacent frames is divided into thresholds of 0.02 and 0.08 image widths per second. Still, moving slowly, moving quickly The ternary movement label is one-hot encoded into a movement state vector. The speech behavior state is jointly determined by speech activity detection and sentence pattern recognition in the classroom audio, and is set to [value] when in a speech activity segment and the question sentence pattern label is an interrogative sentence. Question When in a speech activity segment and the question type is labeled as a declarative sentence, set it to {lecture}; when in a silent segment and the pause duration is greater than [missing information], set it to {lecture}. The pause threshold is set to {pause} and encoded one-hot into a speech behavior state vector. The gesture pointing state is based on the spatial pointing hit value calculation result and simultaneously records the pointing target type as {courseware area, blackboard area, no pointing}.
[0108] The content nodes are obtained by encoding courseware text blocks, whiteboard incremental text, and knowledge point phrases. The courseware text blocks are taken from the text recognition output of S1 on the courseware screen, and each text block is composed of bounding box coordinates and text strings. The whiteboard incremental text is taken from the text recognition output of S1 on the differential connected region of the whiteboard ROI, and each incremental region is also composed of bounding box coordinates and text strings. The knowledge point phrases are taken from the knowledge point phrase extraction output of S1.
[0109] To ensure consistency in semantic measurement, the courseware text blocks, the incremental text on the blackboard, and the knowledge point phrases all use Skip-gram word vectors trained by S1 to generate 128-dimensional semantic vectors, and the mean of word vectors within the phrase is used as the semantic vector of the content node. At the same time, the coordinates of the center point of the bounding box of the courseware text block and the coordinates of the center point of the bounding box of the incremental region on the blackboard are concatenated into the features of the corresponding content node to preserve spatial information.
[0110] The teaching segment nodes are encoded by the six-class classification results output by the teaching segment identifier, and the category set is fixed as {introduction, lecture, questioning, blackboard writing, interaction, summary}. The teaching segment identifier adopts a two-layer bidirectional LSTM structure and encodes the time window feature sequence of the multimodal teaching event unit with a hidden state dimension of 128 in each layer. The time window feature sequence is obtained by concatenating the teacher's posture state vector, walking state vector, voice behavior state vector, courseware page number change marker, blackboard writing increment area ratio and knowledge point phrase semantic vector in the unit in chronological order. The forward hidden state and backward hidden state of the last time step of the bidirectional LSTM are concatenated and input into the fully connected layer and output as six probabilities through Softmax. The category corresponding to the highest probability is taken as the teaching segment label and encoded as the teaching segment node feature with one-hot encoding.
[0111] After constructing the nodes, the edge weights between the nodes are calculated to form a weighted edge set of the behavior-content coupling graph, where for any two nodes... and Each node's time interval is defined as the time window of its corresponding multimodal teaching event unit. Based on this, the time overlap ratio is calculated. The time overlap ratio is the ratio of the intersection duration to the union duration of the two intervals and is used to characterize the co-occurrence relationship.
[0112] The spatial pointing hit value is calculated only between the behavior node corresponding to the gesture pointing state and the content node containing the bounding box, and is determined according to the following process: First, based on the key points of the teacher's wrist and the key points of the index finger tip, a gesture pointing ray is fitted in the teacher's video coordinate system, with the wrist point as the ray origin and the index finger tip direction as the ray direction. Second, this ray is projected onto the courseware screen coordinate system or the blackboard image coordinate system through a homography matrix. The courseware screen projection homography matrix is calculated at the beginning of the class by detecting the four corners of the projection screen in the teacher's video and establishing corresponding points with the four corners of the courseware screen, and remains unchanged throughout the class. The blackboard image projection homography matrix is calculated at the beginning of the class by detecting the four corners of the blackboard in the teacher's video and establishing corresponding points with the four corners of the blackboard ROI in the blackboard image, and remains unchanged throughout the class. Then, in the target coordinate system, it is determined whether the projected ray intersects with the bounding box of the content node. If they intersect, the spatial pointing hit value is set to 1; if they do not intersect, the shortest Euclidean distance between the projected ray and the bounding box is calculated and compared with a preset distance threshold. Pixel comparison, when the shortest distance is not greater than The space pointer is then set to the hit value. And among them For the shortest distance, when the shortest distance is greater than... The space pointer is then set to 0 for the hit value.
[0113] Semantic similarity is calculated between any two content nodes and the cosine similarity of their 128-dimensional semantic vectors is taken. When calculating between behavior nodes and content nodes, the cosine similarity of the semantic vector of the speech text segment corresponding to the behavior node and the semantic vector of the content node is taken, and the semantic vector of the speech text segment is obtained by the average of all word vectors in the segment.
[0114] The content advancement adjacency value is calculated between content node pairs and is used to represent the content order relationship. When two courseware text blocks come from the same courseware page and their outer boxes are adjacent in the reading order from top to bottom and then from left to right, the content advancement adjacency value is set to 1. When two courseware text blocks come from adjacent courseware pages and the page numbers differ by 1, the content advancement adjacency value is set to 0.5. When two blackboard incremental texts come from two adjacent multimodal teaching event units in time, the content advancement adjacency value is set to 0.5. In all other cases, it is set to 0.
[0115] To incorporate modal acquisition quality into edge weight calculation, node reliability is first defined for each node. And from the modal reliability vector of S1:
[0116] ;
[0117] Where the attitude state node is taken Position status node Gesture pointing to state node Voice behavior state node The text block node in the courseware is retrieved. Incremental text node retrieval on the blackboard Knowledge point phrase nodes Teaching process nodes ;
[0118] Based on this, for any node pair and Calculate edge weights And use it as a reference in the figure. arrive The weighted edge weights are calculated using a single weighting formula:
[0119] ;
[0120] in Represents a node With nodes The edge weights between them and the range of values are: The base weights represent the proportion of time overlap and are taken as follows: The base weights of the space pointing to the hit value are taken as follows: The base weights represent semantic similarity and are taken as follows: The base weight representing the content advancement adjacency value is set to 0.10. Indicates the proportion of time overlap between node time intervals. This indicates the hit value of the spatial pointer and is set to 0 when the node pair does not satisfy the relationship of "gesture pointer state node - content node including bounding box". Indicates semantic similarity. Indicates the content-advancement value. Indicates the reliability gating of the time term and takes Indicates the reliability gating of the space pointer item and when The gesture points to the state node and When retrieving text block nodes in courseware And when The gesture points to the state node and When adding incremental text nodes to the whiteboard And in other cases, take Represents the reliability gating of semantic items and takes Indicates the reliability gating of the content advancement item and when and When all are text block nodes in the courseware, take And when and When all are incremental text nodes on the whiteboard, take And in all other cases, the value is 0;
[0121] To implement the knowledge point memory queue mechanism, before each processing of the current graph construction window, knowledge points, blackboard phrases, courseware titles and voice keywords that have been output by S5 in the historical multimodal teaching event unit and have a confidence level greater than the confirmation threshold of 0.80 are read and written into the knowledge point memory queue. Each entry in the knowledge point memory queue contains the entry text, a corresponding 128-dimensional semantic vector, an entry writing timestamp and an entry type, and the queue capacity is fixed at 200. When the capacity is exceeded, the entry is eliminated according to the earliest writing time.
[0122] For each entry in the knowledge point memory queue, the memory weight is calculated by exponential decay based on the time difference between the current time and the entry, with a decay coefficient of 0.03 per second. The obtained memory weight and entry type are then retained together.
[0123] When generating content nodes for the current graph, the cosine similarity between the semantic vector of each content node in the current multimodal teaching event unit and the semantic vector of the existing content nodes in the window is calculated, and the maximum similarity is taken as the content node matching score. When the matching score of the content node is less than the matching threshold of 0.55, entries with memory weights greater than the memory threshold of 0.30 are selected from the knowledge point memory queue, and the top 10 entries are selected from high to low according to the "cosine similarity between the semantic vector of the entry and the semantic vector of the content node" as candidate content nodes and added to the behavioral content coupling graph. At the same time, each candidate content node is assigned the node time interval as the time window of the current multimodal teaching event unit, and its memory weight is multiplied by its node reliability as the initial node confidence of the candidate content node for subsequent use by S5.
[0124] After generating the node set and edge set, the output includes node characteristics, node type identifier, node time interval, node reliability, and weighted edge weight. The behavioral content coupling graph is used for S5's multimodal behavioral content coupling graph Transformer inference.
[0125] In this specific embodiment, S5 includes:
[0126] The behavior-content coupling graph is used as the inference input and denoted as the graph structure. ,in This represents a set of nodes, where each node corresponds to an action node, content node, or teaching segment node, and carries a node feature vector, node type, and node time interval. This represents a set of edges, with each edge carrying its weight calculated using S4. And the node reliability obtained by node reliability mapping and ;
[0127] For each node The node input representation is constructed and used as the input sequence of the multimodal behavior-content coupled graph Transformer. The node input representation is obtained by adding three parts, all of which have the same 256 dimensions. The first part is the node feature vector obtained from S4, which is zero-padded when the dimension is less than 256 and projected to 256 dimensions using a linear layer when the dimension exceeds 256. The second part is the node type encoding, which is generated by a trainable embedding table of size 3 and integrated with... The behavior nodes, content nodes, and teaching link nodes are in one-to-one correspondence. The third part is the time position encoding, which maps the center time of the node time interval to a discrete time index and uses fixed sine and cosine position encoding to generate a 256-dimensional vector to represent the time order.
[0128] The above-mentioned node input represents the input of the multimodal behavior content coupled graph Transformer to complete joint inference. The multimodal behavior content coupled graph Transformer employs a 6-layer stacked graph attention encoding structure, with each layer containing a multi-head self-attention sublayer and a feedforward network sublayer. The number of heads in the self-attention sublayer is set to 8, and each head has a dimension of 32. In the attention calculation, edge weights are used as attention bias terms, and node reliability is used as a multiplicative adjustment term for the attention score. Specifically, for any target node... Its adjacent nodes Calculate attention weights satisfy:
[0129] ;
[0130] in This indicates that within the same attention head at the same layer, nodes... Pointing to node Attention weights Indicates a node All adjacent nodes The weight distribution is obtained by exponentially normalizing the values within the parentheses. Indicates that by node The node input represents the query vector obtained through linear mapping and has a dimension of . Indicates that by node The node input represents the key vector obtained through linear mapping and has a dimension of . This represents the transpose of a vector. Represents the dot product of two vectors. Let the vector dimension of the single-head attention be 32. This represents the trainable scalar parameter used to scale the edge weight bias term and is initialized to 1.0. The nodes calculated by S4 represent With nodes The edge weights between them and the range of values are: , Represents a node The node reliability is determined by the node reliability mapping rule of S4 from the modal reliability vector. Determine and take the range of values as follows Represents a node The node reliability and the value range are also . ;
[0131] The feedforward network sublayers adopt a two-layer fully connected structure with a hidden layer dimension of 1024 and use the GELU activation function. Each sublayer in the 6 layers uses residual connections and LayerNorm, and dropout is set to 0.1 to suppress overfitting.
[0132] After the 6-layer Transformer outputs the node representations of each node, multi-task decoding is performed on these node representations to output teacher behavior fragments, behavior categories, corresponding knowledge points, corresponding courseware or blackboard areas, teaching segments, and confidence levels. Behavior fragments and behavior categories are obtained by decoding the output node representations of the behavior nodes. Two layers of MLP are used to output the normalized start and end times of the behavior relative to the time interval of the behavior node, which are then mapped to absolute timestamps. The output behavior category set (lecturing, asking questions, writing on the blackboard, pointing, walking, interacting) is also output with a Softmax probability distribution. Knowledge points and content areas are jointly determined by the output node representations of the behavior nodes and content nodes. For each behavior node, the similarity between its representation and that of all content nodes is calculated, and the content node with the highest similarity is selected as the corresponding knowledge point node. Based on whether the content node carries a courseware bounding box or a blackboard bounding box, the corresponding courseware or blackboard area coordinates are output, and these coordinates are normalized relative to the image width and height. The output is in the form of the teaching segment, which is decoded and output by the output node of the teaching segment node. Introduction, lecture, questioning, blackboard writing, interaction, summary The Softmax probability distribution is used to select the label with the highest probability. The confidence level is obtained by regressing the output node representation of the action node using an independent MLP, and the Sigmoid output is used as the confidence level. Confidence level of the interval;
[0133] To enable the aforementioned multimodal behavior-content coupled graph Transformer to possess robust inference capabilities under conditions of occlusion, noise, and modality loss, a compensation encoder is trained simultaneously during the training phase, and occlusion reconstruction training is performed. Specifically, for each training sample, at least one modality from teacher video, classroom audio, courseware images, and blackboard images is randomly selected with a probability of 1.0 as the occluded modality. The teaching event anchor point feature representation of this modality is set to a zero vector, and its reliability value is set to 0 to form the occluded input. Simultaneously, the teaching event anchor point feature representation of the unoccluded modality is input into the compensation encoder to reconstruct the teaching event anchor point feature representation of the occluded modality. The compensation encoder adopts a 2-layer TransformerEncoder structure with a model dimension of 256, an attention head of 4, a feedforward network hidden dimension of 512, and a dropout of 0.1. The mean squared error between the reconstructed anchor point representation and the true anchor point representation is used as the distance loss. At the same time, cross-entropy losses are calculated for the behavior classification, knowledge point matching, and teaching segment classification outputs of the graph Transformer, and the distance loss, behavior classification loss, knowledge point matching loss, and teaching segment classification loss are weighted accordingly. The weighted summation update of all parameters of the compensated encoder and the graph Transformer is performed, with the optimizer using AdamW and the learning rate set. The weight decay is 0.01, the batch size is 16, and the number of training rounds is 30.
[0134] During the inference phase, when the reliability value of a certain mode is less than a preset reliability threshold... At that time, the compensation encoder is used to generate the compensation anchor point feature representation of the modality and the compensation anchor point feature representation is backfilled into S3 and S4 to complete the alignment and graph construction. Then, the graph Transformer inference of this S5 is executed to output the teacher behavior segment, teaching link and confidence level corresponding to the specific knowledge point and courseware or blackboard area.
[0135] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0136] This invention utilizes teaching event anchors to organize teacher posture, movement, gestures, voice text, speech rhythm, courseware text, blackboard increments, and knowledge point phrases into the same temporal analysis framework. It also describes the temporal, spatial, and semantic relationships between behavior nodes, content nodes, and teaching link nodes through a behavior-content coupling diagram. This enables classroom behavior recognition results to be directly associated with specific knowledge points, courseware areas, or blackboard areas, thus solving the problem of fragmented classroom process analysis from a technical structural perspective.
[0137] Modal reliability assessment, missing modality reconstruction, and knowledge point memory queue change the way low-quality data is processed in multimodal alignment and graph reasoning. This allows anchor point alignment attention and graph edge weights to be dynamically adjusted according to the quality of data collection. When the current content matching is insufficient, recently confirmed knowledge points, blackboard phrases, courseware titles, and voice keywords are introduced as candidate content nodes, thereby better achieving stable recognition and accurate correspondence under conditions of occlusion, noise, missing modalities, and time asynchrony.
Claims
1. A method for identifying teacher classroom behavior and analyzing teaching content, characterized in that, include: S1. Acquire teacher videos, classroom audio, courseware images and blackboard images during the same class period, extract teacher posture, movement, gestures, voice text, speech rhythm, courseware text, blackboard increments and knowledge point phrases to form a multimodal feature sequence, and generate a modal reliability vector from the acquisition quality parameters of each modality; S2. Based on the changes in courseware page turning, blackboard writing, teacher pointing, voice keywords, question sentence structure, pauses and teacher orientation, generate a teaching event anchor sequence in the multimodal feature sequence. Each teaching event anchor includes anchor time, source modality, feature representation and reliability value. S3. Based on the anchor point feature representation similarity, local time window and modal reliability vector, perform dynamic time alignment on the asynchronous multimodal feature sequence to obtain multimodal teaching event units; S4. Generate behavior nodes, content nodes, and teaching link nodes based on the multimodal teaching event units, and determine the edge weights between nodes according to time overlap, spatial orientation, semantic similarity, content progression order, and the modal reliability vector to construct a behavior-content coupling graph. S5. Input the behavior content coupling graph into the trained multimodal behavior content coupling graph Transformer, and output the teacher behavior segment, behavior category, corresponding knowledge point, corresponding courseware area or blackboard area, teaching link and confidence level. S1 includes: Human body key point detection and target tracking were performed on teacher videos to obtain posture integrity rate, tracking continuity, movement trajectory, gesture pointing features and hand clarity; Speech recognition, sentence structure recognition, and prosodic analysis were performed on classroom audio to obtain speech-text features, question sentence structure labels, pause duration, speech prosodic features, and speech recognition word confidence. The courseware screen is subjected to text recognition and page stability testing to obtain the text features, page number, text recognition confidence score, and page stability. Image difference and text recognition are performed on the blackboard image to obtain the incremental area of the blackboard, the incremental text of the blackboard, and the incremental clarity of the blackboard. The modal reliability vector is formed by normalizing the posture integrity rate, tracking continuity, hand clarity, speech recognition word confidence, text recognition character confidence, courseware page stability, and blackboard incremental clarity to the interval between 0 and 1. S3 includes: For any two source modalities of teaching event anchor points, calculate the cosine similarity of the feature representations; The anchor point time difference is not more than the preset window length as the candidate alignment condition; The anchor point matching score is obtained by multiplying the cosine similarity by the corresponding reliability value. Dynamic programming is used to select the anchor matching path that maximizes the sum of matching scores; The modal feature sequences are resampled over time according to the anchor point matching path to obtain the multimodal teaching event unit; S5 includes: The node input representation is obtained by adding the node features, node type encoding, and time position encoding of each node in the behavior content coupling graph; In the attention layer of the multimodal behavior content coupling graph Transformer, the edge weights between nodes are used as attention bias terms, and the reliability value of the source modality is used as a multiplicative adjustment term for the attention score. Based on the output node representation, decode the start and end time of the behavior, the behavior category, the knowledge point identifier, the coordinates of the courseware area, the coordinates of the blackboard area, the teaching link label, and the confidence level respectively; The training of the multimodal behavior-content coupling graph Transformer includes: In the training samples, at least one modality among teacher video, classroom audio, courseware screen or blackboard image is randomly occluded. The teaching event anchor representation of the unoccluded modality is input into the compensation encoder to reconstruct the anchor representation of the occluded modality. The model parameters are updated by combining the distance loss between the reconstructed anchor representation and the real anchor representation, the behavior classification loss, the knowledge point matching loss, and the teaching link classification loss. During the inference phase, when the reliability value of a certain mode is less than the preset reliability threshold, the compensation encoder is used to generate the compensation anchor point representation of that mode, and the compensation anchor point representation is input into steps S3 and S4.
2. The method according to claim 1, characterized in that, S2 include: When the courseware page number changes, the area of the blackboard writing increment is greater than the preset area threshold, the gesture pointing ray intersects with the courseware area or the blackboard writing area, the voice text contains knowledge point phrases, the voice text conforms to the preset interrogative sentence pattern, the pause duration is greater than the preset pause threshold, or the change in the teacher's facing angle is greater than the preset angle threshold, a corresponding teaching event anchor point is generated. The timestamp of the triggering event is used as the anchor time, the modality to which the triggering event belongs is used as the source modality, the features within a preset time before and after the triggering event are concatenated and encoded into the feature representation, and the corresponding reliability value is read from the modality reliability vector.
3. The method according to claim 1, characterized in that, S4 include: Encode the teacher's posture, movement, gestures, and speech behavior as behavior nodes; Encode the text blocks in the courseware, the incremental text on the blackboard, and the knowledge point phrases into content nodes; Encode one of the following identification results—import, lecture, questioning, blackboard writing, interaction, and summary—as a teaching segment node; For any two nodes, calculate the temporal overlap ratio, spatial orientation hit value, semantic similarity, and content-advancing adjacency value respectively, and then sum the above four types of values according to the modal reliability vector to obtain the edge weight between the nodes.
4. The method according to claim 3, characterized in that, Before constructing the behavior-content coupling graph, the following is also included: Knowledge points, blackboard phrases, courseware titles, and audio keywords that have been confirmed in the historical teaching event unit and have a confidence level greater than the preset confirmation threshold are written into the knowledge point memory queue. The memory weights are obtained by exponentially decaying the entries in the knowledge point memory queue according to the difference between the current time and the entry time. The cosine similarity between the current content features and the semantic vector of the content node is determined as the content node matching score; When the matching score of the content node is less than the preset matching threshold, the entries with a memory weight greater than the preset memory threshold are added as candidate content nodes to the behavior content coupling graph.
5. The method according to claim 3, characterized in that, The calculation of the spatial pointing hit value includes: The hand gesture pointing ray was fitted based on the key points of the teacher's wrist and index finger; Project the gesture pointing ray onto the coordinate system of the courseware screen or the coordinate system of the blackboard image; When the projection ray intersects the outer frame of the text block, set the space pointer hit value to 1; When the projection ray does not intersect the outer frame of the text block and the shortest distance between them is not greater than a preset distance threshold, a spatial pointing hit value between 0 and 1 is determined according to the ratio of the shortest distance to the preset distance threshold. When the shortest distance is greater than the preset distance threshold, the spatial pointing hit value is set to 0.
6. A teacher classroom behavior recognition and teaching content analysis system, used to execute the method described in any one of claims 1 to 5, characterized in that, include: The feature extraction and reliability generation module is used to acquire teacher videos, classroom audio, courseware images, and blackboard images to generate multimodal feature sequences and modal reliability vectors. The anchor point generation module is used to generate a sequence of teaching event anchor points based on changes in courseware page turning, new blackboard writing, teacher pointing, voice keywords, question sentence structure, pauses, and changes in teacher orientation. The dynamic time alignment module is used to generate multimodal teaching event units based on anchor point feature representation similarity, local time windows, and modal reliability vectors. The coupling graph construction module is used to generate behavior nodes, content nodes, and teaching link nodes, and to determine the edge weights between nodes. The Graph Transformer reasoning module is used to output teacher behavior fragments, behavior categories, corresponding knowledge points, corresponding courseware areas or blackboard areas, teaching links, and confidence levels based on the behavior-content coupling graph.
Citation Information
Patent Citations
STEM teacher intelligent research and repair method and system fusing knowledge graph and graph neural network
CN121213314A
Method and system for automatically constructing course knowledge graph based on multi-modal classroom teaching resources
CN122366596A