A classroom behavior analysis method and system based on an artificial intelligence model
Patent Information
- Application Number
- CN202610983033.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]目前,现有课堂行为分析技术仍存在多方面的不足:其一,多数方案仅针对学生端的专注度、课堂纪律进行单向监测,未关联教师的教学行为与授课节奏,仅能输出学生状态的统计结果,无法解释学生状态变化的教学端成因,分析结果难以直接指导教学改进;其二,现有方案普遍采用固定统一的行为评价规则与判定阈值,无法适配课堂讲授、互动提问、小组讨论、实验操作等不同教学环节的目标差异,容易在互动性较强的教学环节产生大量误判,评价结果的场景适配性较差;其三,现有分析多以整节课为单位输出整体统计结论,分析颗粒度较粗,无法精准定位课堂内的问题时段与高效时段,给出的优化建议较为空泛,可落地性不足;其四,部分方案依赖单一模态数据进行教学环节识别,受课堂环境噪声、人员走动遮挡等因素干扰较大,识别准确率与运行稳定性不足,难以支撑后续的精细化分析流程
1、相对于现有技术采用仅针对学生端的单维度课堂行为监测方案,仅能独立输出学生专注度与课堂行为的统计结果,无法建立学生状态变化与教师教学行为之间的内在关联,分析结果始终停留在状态呈现的表层,无法解释学生状态波动的教学端成因,难以直接支撑教学方法与授课节奏的优化调整,分析价值局限于事后记录与粗略评估。本发明采用师生双端同步识别与时序维度耦合关联的分析方案,同步完成教师教学行为与学生课堂状态的双端并行识别,基于统一时序基准对两类数据进行关联度量化计算,建立不同教学行为模式与学生专注度参与度之间的对应关系,实现从单纯状态描述向成因归因诊断的升级,让分析结果能够直接指向教学端的可优化方向,为教学改进提供明确且可落地的数据依据。
Smart Images

Figure CN122817962A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of behavior analysis technology, and in particular to a classroom behavior analysis method and system based on an artificial intelligence model. Background Technology
[0002] With the rapid development of the smart education industry, AI-based classroom behavior analysis technology has become an important support for teaching quality assessment and teaching method optimization. Various classroom behavior monitoring and analysis systems are gradually being implemented in primary and secondary schools and universities.
[0003] Currently, existing classroom behavior analysis technologies still have several shortcomings: First, most solutions only monitor students' focus and classroom discipline in a one-way manner, without linking them to teachers' teaching behaviors and pacing. They can only output statistical results of student status, but cannot explain the teaching-side causes of changes in student status, and the analysis results are difficult to directly guide teaching improvement. Second, existing solutions generally use fixed and uniform behavior evaluation rules and judgment thresholds, which cannot adapt to the different objectives of different teaching links such as classroom lectures, interactive questioning, group discussions, and experimental operations. This can easily lead to a large number of misjudgments in highly interactive teaching links, and the evaluation results have poor scenario adaptability. Third, existing analyses mostly output overall statistical conclusions on a whole-lesson basis, with coarse analysis granularity. They cannot accurately locate problem periods and high-efficiency periods in the classroom, and the optimization suggestions given are rather vague and lack practicality. Fourth, some solutions rely on single-modal data for teaching link identification, which is greatly affected by factors such as classroom environmental noise and occlusion caused by people moving around. The identification accuracy and operational stability are insufficient, making it difficult to support subsequent refined analysis processes.
[0004] Overall, existing classroom behavior analysis technologies have limited analytical dimensions, rigid evaluation standards, and insufficient analytical granularity, failing to meet the needs of refined and adaptive classroom teaching diagnosis and optimization in the current smart education scenario. Summary of the Invention
[0005] To overcome the problems mentioned in the background art, the present invention proposes a classroom behavior analysis method and system based on an artificial intelligence model.
[0006] The technical solution of the present invention is as follows: In one aspect, a classroom behavior analysis method based on an artificial intelligence model includes the following steps: S1: Collect multimodal perception data from the entire classroom, perform unified clock synchronization and timing alignment processing on all data, and generate a structured multimodal timing feature sequence; S2: Based on a pre-trained artificial intelligence model, the time-series feature sequences are recognized in parallel at both ends, and the time-series sequences of teacher teaching behavior and student classroom status with timestamps are output respectively. S3: Integrates three types of information—teacher's speech and semantic features, teacher's action features, and class group behavior features—to jointly identify the current teaching segment in the classroom. S4: Based on the identified teaching segments, dynamically load the corresponding segment's behavior evaluation rules, indicator weighting system, and anomaly judgment threshold, and complete the quantitative evaluation of teachers' and students' classroom behavior based on the matched evaluation system; S5: Based on a unified temporal dimension, a coupled correlation analysis is conducted on the temporal sequence of teachers' teaching behaviors and the temporal sequence of students' classroom status to quantify the correspondence between different teaching behaviors, teaching links and students' focus and participation. S6: Based on the judgment threshold of the corresponding teaching segment, locate the high-incidence period of inattentiveness and the period of high-efficiency interaction in the classroom, and match the teaching behavior characteristics of the corresponding period to complete the state attribution. S7: Based on the full-process analysis results, generate teaching optimization suggestions that include overall classroom evaluation, segment-by-segment optimization guidance, and precise time-slot adjustment plans.
[0007] Preferably, step S1 specifically includes: S11: Deploy video and audio equipment covering the teacher and student areas in the classroom to collect classroom video data and full-area audio data respectively; S12: Configure a unified clock synchronization unit for all acquisition devices, and add a millisecond-level timestamp to each frame of video and each segment of audio to achieve a unified timing reference for all modal data; S13: After extracting frames from the video data at a preset frequency, extract key points of the human skeleton, facial features and limb posture features; after denoising and separating the human voice from the audio data, extract acoustic features and convert them to obtain semantic text features. S14: Using a preset duration as a basic analysis unit, all features are spliced and aligned according to their corresponding timestamps to form a structured multimodal time series feature sequence.
[0008] Preferably, the process of outputting the time sequence of teacher teaching behavior in step S2 is as follows: Based on the spatiotemporal motion recognition model, the teacher's body movement types are identified, including standing lecturing, writing on the blackboard, walking around and patrolling, giving gestures, operating equipment and demonstrating experiments; Based on the speech semantic recognition model, the teaching semantic types of teachers are classified, including knowledge delivery, question initiation, Q&A guidance, classroom instructions and summary comments. At the same time, the teaching rhythm features of speech rate, pause frequency and intonation changes are extracted. By integrating action recognition results and semantic recognition results, a time sequence of teacher teaching behaviors with corresponding timestamps is generated.
[0009] Preferably, the process of outputting the student classroom status time sequence in step S2 is as follows: S21: Perform individual behavior recognition for each student, determine the type of body movement through the posture recognition model, and output the facial expression state classification through the facial expression recognition model; S22: Aggregate the individual identification results of the whole class and calculate the statistical characteristics of group behavior such as class head-raising rate, hand-raising rate, interaction rate, and percentage of abnormal behavior; S23: Integrating individual action and facial expression characteristics, the system quantifies and outputs each student's real-time focus and participation scores according to a preset score range, while simultaneously generating time-series curves for the class's overall average focus and participation.
[0010] Preferably, step S3 specifically includes: S31: Pre-set multiple standard teaching segments, including classroom lectures, interactive questioning, group discussions, independent practice, experimental operations, and classroom summaries; S32: Input three types of data: teacher speech and semantic features, teacher action features, and class group behavior features. Use a late-stage fusion classification model to make joint decisions and output the teaching link to which the current basic analysis unit belongs and the recognition confidence. S33: Based on the continuity constraint of the classroom teaching process, the instantaneous recognition results are subjected to temporal smoothing filtering to eliminate meaningless high-frequency link switching and obtain the final teaching link temporal segmentation result.
[0011] Preferably, step S4 specifically includes: S41: Pre-configure an independent evaluation rule package for each type of teaching segment. The rule package includes the core evaluation indicators for the corresponding segment, the weight ratio of each indicator, and the threshold for judging abnormal behavior. S42: When a teaching segment switch is detected, the evaluation rule package for the corresponding segment is automatically loaded, and the weight parameters and judgment thresholds for behavior quantification are adjusted in real time. S43: Based on the completed evaluation system, output quantitative evaluation results of teacher and student behavior that match the current teaching scenario.
[0012] Preferably, step S5 specifically includes: S51: Using basic analysis units as the granularity, the time sequence of teachers' teaching behaviors, the sequence of teaching links labels, and the time sequence of students' attention and participation are mapped one by one to construct a full-scale time-series dataset of the classroom. S52: Using the sliding window analysis method, the correlation coefficients between different teachers' teaching behaviors and students' average concentration and participation are calculated window by window with a preset window size and step size, so as to quantify the degree of influence of various teaching behaviors on students' classroom status. S53: Statistically analyze the average and fluctuation range of student focus and participation for each teaching segment to obtain the overall classroom performance differences for different teaching segments.
[0013] Preferably, step S6 specifically includes: S61: Based on the anomaly judgment threshold corresponding to the current teaching segment, mark the time period when the average concentration of the class in multiple consecutive basic analysis units is lower than the threshold as the high-incidence period of inattention; S62: Based on the excellent judgment threshold corresponding to the current teaching link, the time period when the average class participation rate of multiple consecutive basic analysis units is higher than the threshold is marked as the high-efficiency interaction period; S63: For each period of high incidence of student distraction, reverse match the teacher's teaching characteristics and teaching segment attributes within that period to form corresponding attribution results of teaching behavior characteristics and changes in student state.
[0014] Preferably, step S7 specifically includes: S71: Generate overall classroom evaluation content, including the overall score of the whole lesson, the time ratio of each teaching segment, and the average overall focus and participation. S72: For the evaluation results of each type of teaching segment, output optimization suggestions for the corresponding teaching segment; S73: Based on the identified periods of high inattention and high-efficiency interaction, provide precise time-based suggestions for adjusting the teaching rhythm and summaries of reusable teaching experiences.
[0015] On the other hand, a classroom behavior analysis system based on an artificial intelligence model is characterized by comprising: The multimodal data acquisition module is used to collect video and audio data from the entire classroom area, complete clock synchronization and timing alignment processing, and output structured timing feature sequences. The teacher-student dual-end behavior recognition module is used to identify teachers' teaching behavior and students' classroom status based on artificial intelligence models, and output the corresponding time-series sequences with timestamps. The teaching segment joint identification module is used to integrate three types of features: teacher's voice and semantics, teacher's actions, and class group behavior, to jointly determine the current teaching segment in the classroom. The scenario-adaptive evaluation module is used to dynamically load matching evaluation rules, weighting systems, and judgment thresholds based on the identified teaching segments, thereby completing the quantitative evaluation of teacher and student behavior. The temporal coupling correlation analysis module is used to perform correlation analysis on the behavioral data of teachers and students based on a unified temporal dimension, and to quantify the correspondence between teaching behavior and students' focus and participation. The key period location and attribution module is used to combine the judgment thresholds of the corresponding links to locate the high-incidence period of inattention and the high-efficiency interaction period, and to complete the attribution of teaching behavior of students' status. The optimization suggestion generation module is used to generate tiered teaching optimization suggestions and classroom analysis reports based on the full-process analysis results.
[0016] The beneficial effects of this invention are: 1. Compared to existing technologies that employ single-dimensional classroom behavior monitoring solutions targeting only the student end, which can only independently output statistical results of student focus and classroom behavior, failing to establish an intrinsic correlation between changes in student state and teacher teaching behavior, the analysis results remain at the surface level of state presentation, unable to explain the teaching-side causes of student state fluctuations, and cannot directly support the optimization and adjustment of teaching methods and pacing. The analytical value is limited to post-event recording and rough evaluation. This invention adopts an analysis scheme that simultaneously identifies both teachers and students and couples them along a temporal dimension. It simultaneously completes parallel identification of teacher teaching behavior and student classroom state, performs correlation quantification calculations on the two types of data based on a unified temporal benchmark, establishes a correspondence between different teaching behavior patterns and student focus and participation, and upgrades from simple state description to causal attribution diagnosis. This allows the analysis results to directly point to areas for optimization at the teaching end, providing clear and actionable data support for teaching improvement.
[0017] 2. Compared to existing technologies that employ uniform and fixed behavior evaluation rules and threshold schemes, using the same standard to judge behavior in all classroom scenarios fails to adapt to the differences in teaching objectives across different teaching stages. This leads to numerous misjudgments in highly interactive stages such as discussions and experiments, resulting in a significant deviation between the evaluation results and the actual teaching objectives, greatly reducing the practical reference value of the evaluation results. This invention adopts a scenario-adaptive scheme that automatically identifies teaching stages and dynamically loads evaluation rules. By jointly determining the current teaching stage of the classroom through multiple features, it automatically switches the matching evaluation indicator weight system and anomaly judgment threshold based on the stage attributes. This ensures that the behavior judgment standards always remain consistent with the target requirements of the current teaching scenario, eliminating cross-scenario misjudgment problems at the root and significantly improving the accuracy and scenario adaptability of classroom behavior evaluation.
[0018] 3. Compared to existing technologies that use single-modal features for teaching segment identification, most rely solely on speech and semantics or single visual features for segment determination. The identification results are significantly affected by environmental noise, visual occlusion, and personnel movement, resulting in insufficient accuracy and operational stability, making it difficult to support subsequent dynamic evaluation and correlation analysis across the entire process. This invention employs a joint decision-making scheme based on three types of features: speech, semantics, teacher actions, and class group behavior. It integrates multi-dimensional modal features for late-stage fusion classification and introduces teaching process continuity constraints to temporally smooth the identification results. This effectively reduces the interference from single modalities, improves the accuracy and robustness of segment identification in complex classroom environments, and provides a reliable foundation for subsequent scene-adaptive evaluation and deep correlation analysis.
[0019] 4. Compared to existing technologies that use coarse-grained analysis methods that statistically analyze the entire lesson, this approach only outputs an overall evaluation and generalized optimization suggestions for the whole lesson. It fails to accurately pinpoint specific problem and strength periods during the lesson, and the optimization suggestions are vague and general, making them difficult for teachers to directly apply and hindering the refined refinement and adjustment of the teaching rhythm. This invention employs a key period precise positioning and targeted attribution scheme based on temporal units. Combined with the judgment criteria for corresponding teaching segments, it automatically identifies high-incidence periods of inattentiveness and high-efficiency interaction periods in the classroom. Simultaneously, it reverse-matches the teaching behavior characteristics of the corresponding periods to complete state attribution, ultimately outputting adjustment suggestions and reusable teaching experiences precise to specific time nodes. This makes the optimization suggestions more operable and can directly guide teachers to make refined adjustments to the classroom rhythm and teaching methods. Attached Figure Description
[0020] Figure 1 The flowchart shown is a classroom behavior analysis method based on an artificial intelligence model according to the present invention. Figure 2 The diagram shown is an architecture diagram of the classroom behavior analysis system based on an artificial intelligence model according to the present invention. Figure 3 The diagram shown illustrates the temporal coupling data relationship between teachers and students in this invention. Detailed Implementation
[0021] The present invention will be further described in detail below with reference to specific embodiments. This embodiment takes a standard 45-minute routine general classroom as an application scenario to fully and clearly illustrate the technical solution of the present invention. The described embodiments are only for explaining the present invention and are not intended to limit the scope of protection of the present invention. Parameter adjustments and scenario adaptations that can be completed by those skilled in the art based on the core concept of the present invention without creative effort are all within the scope of protection of the present invention.
[0022] Example 1 This embodiment provides a classroom behavior analysis method based on an artificial intelligence model, such as... Figure 1 and Figure 3 As shown, the details are as follows: S1: Collect multimodal sensing data from the entire classroom area, perform unified clock synchronization and timing alignment processing on all data, and generate a structured multimodal timing feature sequence. This step is implemented through the following sub-steps: S11: Deploy camera and audio equipment covering the teacher and student areas in the classroom to collect classroom video data and full-area audio data respectively.
[0023] Specifically, a panoramic network camera is deployed at the center of the top rear of the classroom, with a focal length covering the entire student seating area and a resolution of no less than 1080P, to capture the students' body movements and facial images. A close-up camera for the teacher is deployed above the blackboard at the front of the classroom, facing the podium area, to capture the teacher's body movements, blackboard writing, and teaching posture. A directional microphone is deployed at the podium, pointing towards the teacher's position. Simultaneously, an omnidirectional microphone array is deployed at the center of the top of the classroom to capture the entire classroom's audio. All acquisition devices are connected to the classroom's local edge computing terminal, and data transmission uses a wired LAN to ensure transmission stability and timing accuracy.
[0024] S12: Configure a unified clock synchronization unit for all acquisition devices, and add a millisecond-level timestamp to each frame of video and each segment of audio to achieve a unified timing reference for all modal data.
[0025] Specifically, the edge computing terminal has a built-in NTP network time synchronization module that periodically synchronizes its clock with a public time server. Simultaneously, it sends synchronization clock signals to all camera and audio pickup devices. Each frame of video data is embedded with a millisecond-level timestamp at the acquisition end, and each audio segment is also labeled with its corresponding start and end timestamps, ensuring that all modal data share the same time base, with timing deviations controlled within 10ms. In this embodiment, the raw audio and video data is processed only locally on the edge computing terminal. After feature extraction is completed, the raw audio and video files are immediately destroyed, and only structured feature data is stored, balancing analytical capabilities and privacy security.
[0026] S13: After extracting frames from the video data at a preset frequency, extract key points of the human skeleton, facial features, and limb posture features; after denoising and separating the human voice from the audio data, extract acoustic features and convert them into semantic text features.
[0027] Specifically, both video streams are processed by frame extraction at a frequency of 3 frames per second. The extracted frames are then preprocessed by denoising and scale normalization. For the student area, 17 human skeleton key points are extracted for each student using a human key point detection algorithm. At the same time, the face area is detected and facial expression feature vectors are extracted. For the teacher area, human skeleton key points and action posture features of the teacher are extracted.
[0028] For the audio data, environmental noise reduction is first performed using spectral subtraction, followed by a voice separation algorithm to distinguish between the teacher's voice and the students' group voice. Twelve acoustic features, including Mel-frequency cepstral coefficients (MFCC), volume, speech rate, and pause percentage, are extracted from the separated speech. Simultaneously, a pre-trained speech recognition model converts the teacher's speech into text data, and semantic features are obtained through keyword extraction and sentence structure classification. In this embodiment, the human keypoint detection uses the publicly available OpenPose pre-trained model, and the speech recognition uses the publicly available Whisper-base pre-trained model. Those skilled in the art can directly call upon the publicly available pre-trained weights to complete the deployment.
[0029] S14: Using a preset duration as a basic analysis unit, all features are spliced and aligned according to their corresponding timestamps to form a structured multimodal time series feature sequence.
[0030] Specifically, in this embodiment, 10 seconds is set as a basic analysis unit. The average value of the video features extracted from 30 frames within 10 seconds is taken as the video feature vector of the unit. The audio features and semantic features within 10 seconds are aggregated and statistically analyzed to obtain the audio feature vector and semantic feature vector of the unit. Finally, the video features, audio features, and semantic features within the same basic unit are concatenated one by one according to the timestamp to form a multimodal temporal feature sequence with unified dimensions and continuous temporal sequence, which serves as the unified input for all subsequent analysis steps.
[0031] S2: Based on a pre-trained artificial intelligence model, the time-series feature sequences are recognized in parallel from both ends, outputting time-series sequences of teacher teaching behavior and student classroom status with timestamps respectively. This step is divided into parallel processing on the teacher end and the student end. The specific process of outputting the time-series sequence of teacher teaching behavior is as follows: The spatiotemporal action recognition model identifies the types of teachers' body movements, including standing lectures, blackboard writing, walking around, hand gestures, equipment operation, and experimental demonstrations. The speech and semantic recognition model classifies the semantic types of teachers' teaching, including knowledge delivery, question initiation, Q&A guidance, classroom instructions, and summary comments. At the same time, it extracts the teaching rhythm features of speech rate, pause frequency, and intonation changes. The action recognition results and semantic recognition results are integrated to generate a time sequence of teachers' teaching behaviors with corresponding timestamps.
[0032] Specifically, this embodiment uses the SlowFast pre-trained spatiotemporal action recognition model, inputs the continuous frame posture features of the teacher region, outputs the classification probabilities of 6 types of limb actions, and takes the item with the highest probability as the teacher action recognition result of the current unit; at the same time, based on the BERT pre-trained text classification model, the text transcribed from the teacher's speech is semantically classified to obtain 5 semantic type results; finally, the action recognition results, semantic classification results and teaching rhythm features are combined together to form teacher teaching behavior unit data, which are arranged in chronological order to form a complete time sequence of teacher teaching behavior, with each data unit corresponding to a basic analysis duration of 10 seconds.
[0033] The process of outputting the time sequence of students' classroom status is implemented through the following sub-steps: S21: Perform individual behavior recognition for each student, determine the type of body movement through the posture recognition model, and output the facial expression state classification through the facial expression recognition model.
[0034] Specifically, based on the skeletal keypoint data of each student, a posture classification model is used to identify five types of body movements: raising hands, looking down while writing, leaning on a table, turning around, and whispering. For the detected facial region, a facial expression recognition model is used to output four types of facial expression states: focused, confused, tired, and happy, with a confidence level of 0-1 for each state. In this embodiment, facial expression recognition uses a publicly available pre-trained expression classification model based on ResNet-50.
[0035] S22: Aggregate the individual identification results of the whole class and calculate the statistical characteristics of group behavior, such as the class head-raising rate, hand-raising rate, interaction rate, and the proportion of abnormal behavior.
[0036] Specifically, the percentage of students in the current basic unit who are looking up is recorded as the class head-up rate; the percentage of students who raise their hands is recorded as the hand-raising rate; the percentage of students who actively interact (raising their hands, speaking) is recorded as the interaction rate; and the percentage of students who exhibit abnormal behaviors such as lying on their desks or whispering is recorded as the abnormal behavior rate. These four statistical characteristics of class group behavior are obtained.
[0037] S23: Integrating individual action and facial expression characteristics, the system quantifies and outputs each student's real-time focus and participation scores according to a preset score range, while simultaneously generating time-series curves for the class's overall average focus and participation.
[0038] Specifically, the focus score is calculated on a 100-point scale, as follows: Facial focus confidence × 60 + Positive posture confidence × 40, where the positive posture confidence refers to the recognition confidence of focused postures such as looking up and writing. The participation score is calculated as: Interactive action confidence × 70 + Pleasant facial expression confidence × 30. The arithmetic mean of the scores of each student in the class is taken to obtain the class average focus and participation for the current basic unit. The averages of all units are arranged in chronological order to generate the overall class focus and participation time-series curves.
[0039] S3: Integrating three types of information—teacher's speech and semantic features, teacher's action features, and class group behavior features—to jointly identify the current teaching segment in the classroom. This step is implemented through the following sub-steps: S31: Pre-set multiple standard teaching segments, including classroom lectures, interactive questioning, group discussions, independent practice, experimental operations, and classroom summaries.
[0040] This embodiment pre-sets six standard teaching steps, each corresponding to a typical teaching behavior pattern: classroom lectures correspond to continuous explanation by the teacher and quiet listening by the students; interactive questioning corresponds to the teacher raising questions and students raising their hands to answer; group discussions correspond to the teacher issuing discussion instructions and multiple students engaging in interactive conversations; independent practice corresponds to the teacher assigning tasks and students writing with their heads down; experimental operations correspond to the teacher demonstrating operations and students conducting experiments; and classroom summaries correspond to the teacher summarizing key points and students organizing and recording them.
[0041] S32: Input three types of data: teacher speech and semantic features, teacher action features, and class group behavior features. Use a late-stage fusion classification model to make joint decisions and output the teaching link to which the current basic analysis unit belongs and the recognition confidence.
[0042] Specifically, a late-stage fusion strategy is adopted, whereby the three types of features are input into their respective sub-classification models, each outputting a probability distribution for six types of teaching segments. The three probabilities are then weighted and fused, with the following weights: teacher speech and semantic features weighted at 0.5, teacher action features weighted at 0.2, and class group behavior features weighted at 0.3. The teaching segment with the highest fusion probability is taken as the recognition result of the current basic unit, and the corresponding confidence score is output. Those skilled in the art can adjust the fusion weights of the three types of features according to the characteristics of the classroom scenario without requiring any creative effort.
[0043] S33: Based on the continuity constraint of the classroom teaching process, the instantaneous recognition results are subjected to temporal smoothing filtering to eliminate meaningless high-frequency link switching and obtain the final teaching link temporal segmentation result.
[0044] Specifically, a sliding window majority voting method is used for smoothing. The window size is set to 3 consecutive basic analysis units (i.e., 30 seconds). If more than 2 / 3 of the units in the window are identified as the same teaching segment, then all units in the window are uniformly identified as that segment. At the same time, a minimum duration constraint for segment switching is set, and the duration of a single teaching segment must not be less than 30 seconds to avoid high-frequency meaningless jumps of "lecture-discussion-lecture". The final output is a teaching segment time sequence division result that conforms to the conventional teaching logic.
[0045] S4: Based on the identified teaching segments, dynamically load the corresponding behavioral evaluation rules, indicator weighting system, and anomaly detection thresholds. Based on the matched evaluation system, complete the quantitative evaluation of teacher and student classroom behavior. This step is implemented through the following sub-steps: S41: Pre-configure an independent evaluation rule package for each type of teaching segment. The rule package includes the core evaluation indicators for the corresponding segment, the weight ratio of each indicator, and the threshold for judging abnormal behavior.
[0046] Specifically, this embodiment configures dedicated evaluation rule packages for each of the six types of teaching segments, with the core configuration as follows: In the classroom teaching segment: the core evaluation indicators are focus (70% weight), participation (20% weight), and the percentage of abnormal behavior (10% weight); the threshold for judging abnormal behavior is that an individual's focus score is below 60 points, which is considered as the individual being distracted, and the class average focus score is below 65 points, which is considered as a period of low focus in the class. Interactive Q&A Session: The core evaluation indicators are participation (weight 60%), focus (weight 30%), and response rate (weight 10%); the threshold for anomaly detection is that a class hand-raising rate of less than 10% is considered a period of low participation. Group discussion session: The core evaluation indicators are interactive participation (70% weight), discussion activity (20% weight), and percentage of disorderly behavior (10% weight); the threshold for judging abnormal behavior is that when the percentage of students with no interactive behavior is higher than 30%, it is judged as a period of low participation. Independent practice session: The core evaluation indicators are focus (80% weight) and the percentage of disordered behavior (20% weight); the threshold for judging disorder is that whispering and frequent looking around more than twice per minute are considered as individual distraction; Experimental operation phase: The core evaluation indicators are operation participation (weight 50%), action standardization (weight 30%), and orderliness (weight 20%); the abnormal judgment threshold is that when the percentage of students without operation behavior is higher than 25%, it is judged as a low participation period; Class summary: The core evaluation indicators are focus (weight 60%), percentage of recorded behavior (weight 30%), and percentage of orderliness (weight 10%); the threshold for anomaly detection is the same as that for the classroom teaching session.
[0047] S42: When a teaching segment switch is detected, the evaluation rule package for the corresponding segment is automatically loaded, and the weight parameters and judgment thresholds for behavior quantification are adjusted in real time.
[0048] Specifically, the timing division results of the teaching segments output in step S3 are monitored in real time. When a switch in the teaching segment is detected and the recognition confidence is higher than 0.8, the rule package switching logic is triggered. The evaluation indicators, weights and threshold parameters of the corresponding segment are loaded from the local rule base to replace the evaluation parameters of the current operation, so as to achieve seamless switching of the evaluation system.
[0049] S43: Based on the completed evaluation system, output quantitative evaluation results of teacher and student behavior that match the current teaching scenario.
[0050] Specifically, according to the currently effective evaluation rule package, the teacher and student behavior data of each basic analysis unit are weighted and calculated, and the overall class performance score, individual behavior evaluation results and abnormal behavior alarms are output. All evaluation results match the target requirements of the current teaching link, thus avoiding cross-scenario misjudgment from the root.
[0051] S5: Based on a unified temporal dimension, a coupled correlation analysis is performed on the temporal sequence of teachers' teaching behaviors and the temporal sequence of students' classroom states to quantify the correspondence between different teaching behaviors, teaching links, and students' attention and participation. This step is specifically implemented through the following sub-steps: S51: Using basic analysis units as the granularity, the time sequence of teachers' teaching behaviors, the sequence of teaching links labels, and the time sequence of students' attention and participation are mapped one-to-one to construct a full-scale time-series dataset of the classroom.
[0052] Specifically, using 10-second basic units as the smallest granularity, the fields of teacher behavior type, teaching rhythm characteristics, teaching segment labels, average class focus, and average class participation under the same timestamp are concatenated to form a structured time-series data record; all records are arranged in the order of class time to construct a full time-series dataset for a single lesson.
[0053] S52: Using the sliding window analysis method, the correlation coefficients between different teacher teaching behaviors and students' average concentration and participation are calculated window by window with a preset window size and step size, so as to quantify the degree of influence of various teaching behaviors on students' classroom status.
[0054] Specifically, the sliding window size was set to 6 basic analysis units (i.e., 1 minute), and the sliding step size was 1 basic analysis unit (i.e., 10 seconds). For the data in each window, the Pearson correlation coefficient between the teacher's continuous teaching time and the class's focus, and the Pearson correlation coefficient between the teacher's questioning frequency and the class's participation were calculated to obtain the quantitative correlation between different teaching behaviors and student status. The correlation coefficient ranged from -1 to 1, with positive values indicating positive correlation and negative values indicating negative correlation. The larger the absolute value, the stronger the correlation.
[0055] S53: Statistically analyze the average and fluctuation range of student focus and participation for each teaching segment to obtain the overall classroom performance differences for different teaching segments.
[0056] Specifically, the full set of time-series data was grouped according to the teaching segment labels, and the mean and standard deviation of the average attention level and the average participation level of the class were calculated for each teaching segment. The mean was used to compare the overall performance of different segments, and the standard deviation was used to reflect the stability of the students' state under different segments.
[0057] S6: Based on the judgment thresholds of the corresponding teaching segments, identify the high-incidence periods of inattentiveness and the high-efficiency interaction periods in the classroom, and match the teaching behavior characteristics of the corresponding periods to complete the state attribution. This step is specifically implemented through the following sub-steps: S61: Based on the anomaly judgment threshold corresponding to the current teaching segment, mark the time period when the average concentration of the class in multiple consecutive basic analysis units is lower than the threshold as the high-incidence period of inattentiveness.
[0058] Specifically, the entire time series data is traversed, and the low focus threshold of the teaching segment to which the current data belongs is used as the benchmark. When the average focus of the class is detected to be lower than the corresponding threshold for three or more consecutive basic analysis units (i.e., for more than 30 seconds), the continuous period is marked as a high-incidence period of distraction. The start time, end time, teaching segment to which the period belongs, and average focus value are recorded.
[0059] S62: Based on the excellent judgment threshold corresponding to the current teaching link, mark the time period when the average class participation rate of multiple consecutive basic analysis units is higher than the threshold as the high-efficiency interaction period.
[0060] Specifically, preset excellent participation thresholds for each stage: participation in classroom lectures is above 75 points, participation in interactive questioning is above 80 points, and participation in group discussions is above 85 points. When the average class participation of three or more consecutive basic analysis units is detected to be above the corresponding excellent threshold, the time period is marked as an efficient interactive time period, and the corresponding time and teaching behavior characteristics are recorded.
[0061] S63: For each period of high incidence of student distraction, reverse match the teacher's teaching characteristics and teaching segment attributes within that period to form corresponding attribution results of teaching behavior characteristics and changes in student state.
[0062] Specifically, for each period of high concentration, the teaching characteristics of teachers during that period are extracted, including parameters such as whether it is continuous one-way teaching, average speaking speed, frequency of questions, and percentage of pauses. Combined with the results of correlation analysis, the core teaching factors that lead to decreased concentration are matched, such as "continuous one-way teaching for more than 8 minutes without interaction" and "speaking speed that is too fast and exceeds the normal range", to form a corresponding result of "problem period - cause attribution".
[0063] S7: Based on the full-process analysis results, generate teaching optimization suggestions that include overall classroom evaluation, segment-by-segment optimization guidelines, and precise time-slot adjustment plans. This step is implemented through the following sub-steps: S71: Generate overall classroom evaluation content, including the overall score of the whole lesson, the time ratio of each teaching segment, and the average overall focus and participation.
[0064] Specifically, the weighted average of the overall performance of all basic units in the whole lesson is used to obtain the overall score of the whole lesson; the total time and proportion of each type of teaching segment are calculated; and the global average of the class's average focus and average participation are calculated to form an overall picture of the classroom.
[0065] S72: For the evaluation results of each type of teaching segment, output optimization suggestions for the corresponding teaching segment.
[0066] Specifically, by comparing with general classroom benchmark data, directional optimization suggestions are generated for teaching segments that perform below the benchmark. For example, if the focus level during the lecture segment is lower than the benchmark, the suggestion is "The overall focus level during the lecture segment is low. It is recommended to optimize the pace of knowledge output and insert an interactive session every 8-10 minutes." If the participation in the question-and-answer segment is insufficient, the suggestion is "The participation in the question-and-answer segment is low. It is recommended to adopt a tiered questioning approach to accommodate the participation willingness of students at different levels."
[0067] S73: Based on the identified periods of high inattention and high-efficiency interaction, provide precise time-based suggestions for adjusting the teaching rhythm and summaries of reusable teaching experiences.
[0068] Specifically, for each period of high student distraction, specific adjustment suggestions are provided based on the attribution results. For example, "The 12th to 15th minute is a continuous lecture period, during which students' concentration drops significantly. It is recommended to insert a short classroom question or a one-minute pause for reflection at this time to bring students' attention back." For each period of highly interactive learning, corresponding teaching behavior patterns are extracted. For example, "The group discussion period from the 25th to the 28th minute has high participation. The corresponding teacher adopted a model of clear task assignment + circulating guidance, which can be reused in subsequent classes."
[0069] Example 2 like Figure 2 As shown, this embodiment provides a classroom behavior analysis system based on an artificial intelligence model, used to execute the analysis method described in Embodiment 1. The system includes the following modules: The multimodal data acquisition module is used to collect video and audio data from the entire classroom area, complete clock synchronization and timing alignment processing, and output structured timing feature sequences. The teacher-student dual-end behavior recognition module is used to identify teachers' teaching behavior and students' classroom status based on artificial intelligence models, and output the corresponding time-series sequences with timestamps. The teaching segment joint identification module is used to integrate three types of features: teacher's voice and semantics, teacher's actions, and class group behavior, to jointly determine the current teaching segment in the classroom. The scenario-adaptive evaluation module is used to dynamically load matching evaluation rules, weighting systems, and judgment thresholds based on the identified teaching segments, thereby completing the quantitative evaluation of teacher and student behavior. The temporal coupling correlation analysis module is used to perform correlation analysis on the behavioral data of teachers and students based on a unified temporal dimension, and to quantify the correspondence between teaching behavior and students' focus and participation. The key period location and attribution module is used to combine the judgment thresholds of the corresponding links to locate the high-incidence period of inattention and the high-efficiency interaction period, and to complete the attribution of teaching behavior of students' status. The optimization suggestion generation module is used to generate tiered teaching optimization suggestions and classroom analysis reports based on the full-process analysis results.
[0070] This system adopts a layered deployment architecture of edge + cloud: the multimodal data acquisition module, the teacher and student dual-terminal behavior recognition module, the teaching link joint recognition module and the scene adaptive evaluation module are deployed on the local edge computing terminal in the classroom to achieve low-latency real-time analysis; the time-series coupling correlation analysis module, the key time period location attribution module and the optimization suggestion generation module are deployed on the cloud server to complete in-depth analysis and report generation, taking into account both real-time performance and analysis depth.
[0071] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A classroom behavior analysis method based on an artificial intelligence model, characterized in that, Includes the following steps: S1: Collect multimodal perception data from the entire classroom, perform unified clock synchronization and timing alignment processing on all data, and generate a structured multimodal timing feature sequence; S2: Based on a pre-trained artificial intelligence model, the time-series feature sequences are recognized in parallel at both ends, and the time-series sequences of teacher teaching behavior and student classroom status with timestamps are output respectively. S3: Integrates three types of information—teacher's speech and semantic features, teacher's action features, and class group behavior features—to jointly identify the current teaching segment in the classroom. S4: Based on the identified teaching segments, dynamically load the corresponding segment's behavioral evaluation rules, indicator weighting system, and anomaly judgment threshold, and complete the quantitative evaluation of teachers' and students' classroom behavior based on the matched evaluation system; S5: Based on a unified temporal dimension, a coupled correlation analysis is conducted on the temporal sequence of teachers' teaching behaviors and the temporal sequence of students' classroom status to quantify the correspondence between different teaching behaviors, teaching links and students' focus and participation. S6: Based on the judgment threshold of the corresponding teaching segment, locate the high-incidence period of inattentiveness and the period of high-efficiency interaction in the classroom, and match the teaching behavior characteristics of the corresponding period to complete the state attribution. S7: Based on the full-process analysis results, generate teaching optimization suggestions that include overall classroom evaluation, segment-by-segment optimization guidance, and precise time-slot adjustment plans.
2. The classroom behavior analysis method based on an artificial intelligence model according to claim 1, characterized in that, Step S1 specifically includes: S11: Deploy video and audio equipment covering the teacher and student areas in the classroom to collect classroom video data and full-area audio data respectively; S12: Configure a unified clock synchronization unit for all acquisition devices, and add a millisecond-level timestamp to each frame of video and each segment of audio to achieve a unified timing reference for all modal data; S13: After extracting frames from the video data at a preset frequency, extract key points of the human skeleton, facial features and limb posture features; after denoising and separating the human voice from the audio data, extract acoustic features and convert them to obtain semantic text features. S14: Using a preset duration as a basic analysis unit, all features are spliced and aligned according to their corresponding timestamps to form a structured multimodal time series feature sequence.
3. The classroom behavior analysis method based on an artificial intelligence model according to claim 1, characterized in that, The process of outputting the time sequence of teacher teaching behaviors in step S2 is as follows: Based on the spatiotemporal motion recognition model, the teacher's body motion types are identified, including standing lecturing, writing on the blackboard, walking around and patrolling, giving gestures, operating equipment and demonstrating experiments; Based on the speech semantic recognition model, the teaching semantic types of teachers are classified, including knowledge delivery, question initiation, Q&A guidance, classroom instructions and summary comments. At the same time, the teaching rhythm features of speech rate, pause frequency and intonation changes are extracted. By integrating action recognition results and semantic recognition results, a time sequence of teacher teaching behaviors with corresponding timestamps is generated.
4. The classroom behavior analysis method based on an artificial intelligence model according to claim 1, characterized in that, The process of outputting the student classroom status time sequence in step S2 is as follows: S21: Perform individual behavior recognition for each student, determine the type of body movement through the posture recognition model, and output the facial expression state classification through the facial expression recognition model; S22: Aggregate the individual identification results of the whole class and calculate the statistical characteristics of group behavior such as class head-raising rate, hand-raising rate, interaction rate, and percentage of abnormal behavior; S23: Integrating individual action and facial expression characteristics, the system quantifies and outputs each student's real-time focus and participation scores according to a preset score range, while simultaneously generating time-series curves for the class's overall average focus and participation.
5. The classroom behavior analysis method based on an artificial intelligence model according to claim 1, characterized in that, Step S3 specifically includes: S31: Pre-set multiple standard teaching segments, including classroom lectures, interactive questioning, group discussions, independent practice, experimental operations, and classroom summaries; S32: Input three types of data: teacher speech and semantic features, teacher action features, and class group behavior features. Use a late-stage fusion classification model to make joint decisions and output the teaching link to which the current basic analysis unit belongs and the recognition confidence. S33: Based on the continuity constraint of the classroom teaching process, the instantaneous recognition results are subjected to temporal smoothing filtering to eliminate meaningless high-frequency link switching and obtain the final teaching link temporal segmentation result.
6. The classroom behavior analysis method based on an artificial intelligence model according to claim 1, characterized in that, Step S4 specifically includes: S41: Pre-configure an independent evaluation rule package for each type of teaching segment. The rule package includes the core evaluation indicators for the corresponding segment, the weight ratio of each indicator, and the threshold for judging abnormal behavior. S42: When a teaching segment switch is detected, the evaluation rule package for the corresponding segment is automatically loaded, and the weight parameters and judgment thresholds for behavior quantification are adjusted in real time. S43: Based on the completed evaluation system, output quantitative evaluation results of teacher and student behavior that match the current teaching scenario.
7. The classroom behavior analysis method based on an artificial intelligence model according to claim 1, characterized in that, Step S5 specifically includes: S51: Using basic analysis units as the granularity, the time sequence of teachers' teaching behavior, the sequence of teaching links labels, and the time sequence of students' attention and participation are mapped one-to-one to construct a full-scale time-series dataset of the classroom. S52: Using the sliding window analysis method, the correlation coefficients between different teachers' teaching behaviors and students' average concentration and participation are calculated window by window with a preset window size and step size, so as to quantify the degree of influence of various teaching behaviors on students' classroom status. S53: Statistically analyze the average and fluctuation range of student focus and participation for each teaching segment to obtain the overall classroom performance differences for different teaching segments.
8. The classroom behavior analysis method based on an artificial intelligence model according to claim 7, characterized in that, Step S6 specifically includes: S61: Based on the anomaly judgment threshold corresponding to the current teaching segment, mark the time period when the average concentration of the class in multiple consecutive basic analysis units is lower than the threshold as the high-incidence period of inattention; S62: Based on the excellent judgment threshold corresponding to the current teaching link, the time period when the average class participation rate of multiple consecutive basic analysis units is higher than the threshold is marked as the high-efficiency interaction period; S63: For each period of high incidence of student distraction, reverse match the teacher's teaching characteristics and teaching segment attributes within that period to form corresponding attribution results of teaching behavior characteristics and changes in student state.
9. The classroom behavior analysis method based on an artificial intelligence model according to claim 1, characterized in that, Step S7 specifically includes: S71: Generate overall classroom evaluation content, including the overall score of the whole lesson, the time ratio of each teaching segment, and the average overall focus and participation. S72: For the evaluation results of each type of teaching segment, output optimization suggestions for the corresponding teaching segment; S73: Based on the identified periods of high inattention and high-efficiency interaction, provide precise time-based suggestions for adjusting the teaching rhythm and summaries of reusable teaching experiences.
10. A classroom behavior analysis system based on an artificial intelligence model, used to implement the classroom behavior analysis method according to any one of claims 1-9, characterized in that, include: The multimodal data acquisition module is used to collect video and audio data from the entire classroom area, complete clock synchronization and timing alignment processing, and output structured timing feature sequences. The teacher-student dual-end behavior recognition module is used to identify teachers' teaching behavior and students' classroom status based on artificial intelligence models, and output the corresponding time-series sequences with timestamps. The teaching segment joint identification module is used to integrate three types of features: teacher's voice and semantics, teacher's actions, and class group behavior, to jointly determine the current teaching segment in the classroom. The scenario-adaptive evaluation module is used to dynamically load matching evaluation rules, weighting systems, and judgment thresholds based on the identified teaching segments, thereby completing the quantitative evaluation of teacher and student behavior. The temporal coupling correlation analysis module is used to perform correlation analysis on the behavioral data of teachers and students based on a unified temporal dimension, and to quantify the correspondence between teaching behavior and students' focus and participation. The key period location and attribution module is used to combine the judgment thresholds of the corresponding links to locate the high-incidence period of inattention and the high-efficiency interaction period, and to complete the attribution of teaching behavior of students' status. The optimization suggestion generation module is used to generate tiered teaching optimization suggestions and classroom analysis reports based on the full-process analysis results.