Intelligent assessment methods, devices, equipment, and storage media for classroom teacher-student interaction

By performing speech recognition and target detection on classroom teaching videos, a tensor of classroom teacher-student interaction features is generated. Using a tensor clustering model for intelligent evaluation, the problem of ignoring multimodal information in existing technologies is solved, and a more accurate and fair evaluation of teacher-student interaction is achieved.

CN121640156BActive Publication Date: 2026-06-30HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG NORMAL UNIV
Filing Date
2025-11-28
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing technologies neglect multimodal information in classroom teacher-student interaction assessments, resulting in lower accuracy and fairness in assessments, and human observation is greatly affected by subjective factors.

Method used

By performing speech recognition and target detection on classroom teaching videos, the verbal and non-verbal interaction features of teachers and students in the classroom are determined, an interaction quantification matrix is ​​generated, and an intelligent evaluation is performed using a target tensor clustering model to generate a comprehensive evaluation report.

Benefits of technology

It improves the accuracy and fairness of classroom teacher-student interaction evaluation, overcomes the shortcomings of traditional methods that are highly subjective and difficult, and provides a scientific basis for teaching quality monitoring and classroom behavior research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640156B_ABST
    Figure CN121640156B_ABST
Patent Text Reader

Abstract

This application belongs to the field of educational informatization technology, specifically disclosing an intelligent evaluation method, device, equipment, and storage medium for classroom teacher-student interaction. Through this application, verbal interaction features of teachers and students in the classroom are determined based on recognition results; non-verbal interaction features are determined based on detection results; a quantitative matrix of classroom interaction is generated based on the verbal and non-verbal interaction features, and a classroom interaction feature tensor is constructed based on the quantitative matrix; intelligent evaluation of classroom teacher-student interaction is performed based on the classroom interaction feature tensor using a target tensor clustering model. By using the above method, the verbal and non-verbal interaction features of teachers and students in the classroom are determined based on the teaching video to be evaluated, and then intelligent evaluation is performed based on a target tensor clustering model that can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure, thereby effectively improving the accuracy and fairness of evaluating teacher-student interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of educational informatization technology, and more specifically, relates to intelligent evaluation methods, devices, equipment and storage media for classroom teacher-student interaction. Background Technology

[0002] Classroom teacher-student interaction is not only the core of the teaching process but also a concrete manifestation of teacher-student communication, significantly impacting student participation and the quality of classroom teaching. In the 1960s, Flanders proposed the Flanders' Interaction Analysis System (FIAS), marking the beginning of a systematic and scientific phase in research on teaching interaction. FIAS involves periodically sampling and recording verbal interactions between teachers and students, then performing statistical analysis using matrix tables. The results can be used to assess teaching quality, identify teaching patterns, and guide teaching improvements. Since then, the FIAS system has been widely applied and adapted and expanded across various disciplines, such as the Information Technology-based Interaction Analysis System (ITIAS), the improved Flanders' Interaction Analysis System, and interactive electronic dual-board-based interaction analysis systems.

[0003] Classroom teacher-student interaction is a process of multimodal information generation, transmission, and reception. However, the aforementioned analysis systems directly ignore multimodal information, leading to incomplete and inaccurate analysis. Furthermore, recording verbal interactions involves human observation, which is obviously heavily influenced by subjective factors and is difficult to perform. Therefore, the accuracy and fairness of these methods in evaluating teacher-student interaction are low. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the purpose of this application is to provide an intelligent evaluation method, device, equipment and storage medium for classroom teacher-student interaction, which aims to solve the problems of low accuracy and fairness in evaluating teacher-student interaction due to the direct neglect of multimodal information and the large influence of subjective factors on the human observation involved, as well as the high difficulty of operation.

[0005] To achieve the above objectives, firstly, this application provides an intelligent evaluation method for classroom teacher-student interaction, including:

[0006] Speech recognition is performed on the classroom teaching videos to be evaluated, and the characteristics of teacher-student verbal interaction in the classroom are determined based on the recognition results;

[0007] The classroom teaching video to be evaluated is subjected to target detection, and the non-verbal interaction characteristics of teachers and students in the classroom are determined based on the detection results;

[0008] A classroom interaction quantification matrix is ​​generated based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and a classroom interaction feature tensor is constructed based on the classroom interaction quantification matrix.

[0009] Based on the target tensor clustering model, intelligent evaluation of classroom teacher-student interaction is performed according to the classroom interaction feature tensor, resulting in a comprehensive evaluation report of classroom teacher-student interaction.

[0010] In one embodiment, the step of performing speech recognition on the classroom teaching video to be evaluated and determining the characteristics of teacher-student verbal interaction based on the recognition results includes:

[0011] Obtain the classroom teaching videos to be evaluated, and perform data cleaning on the classroom teaching videos to be evaluated;

[0012] Extracting classroom audio data from cleaned classroom teaching videos;

[0013] The classroom teaching audio data is transcribed, and the transcribed classroom teaching audio data is subjected to speech recognition using a target speaker recognition model to obtain speech segments of the teacher and students;

[0014] The speech segments of the teacher and students are detected using a target speech detection algorithm, and semantic recognition is performed on the detected classroom text.

[0015] The semantically recognized classroom texts are identified, and the cognitive levels are classified according to the identified question texts.

[0016] The characteristics of verbal interaction between teachers and students in the classroom were determined based on the results of the cognitive level classification.

[0017] In one embodiment, the step of performing target detection on the classroom teaching video to be evaluated and determining the nonverbal interaction features between teachers and students based on the detection results includes:

[0018] Extract the teacher-facing and student-facing classroom video footage data from the classroom teaching videos to be evaluated;

[0019] Target detection is performed on the classroom video footage data for teachers, and the teacher's face region is determined based on the detection results of the first frame.

[0020] The image data of the teacher's face region is used to perform facial expression recognition, and the proportion of each type of teacher's facial expression recognized in the class time is determined.

[0021] Target detection is performed on the classroom video footage data facing students, and the student face regions are determined based on the detection results of the second screen.

[0022] Perform facial expression recognition on the image data of the student's face area and statistically analyze the distribution ratio of the recognized student facial expressions;

[0023] Using time series analysis, the duration of each student's facial expression and the trend of the distribution ratio change are determined based on the distribution ratio.

[0024] The characteristics of nonverbal interaction between teachers and students in the classroom are obtained based on the proportion of each type of teacher's facial expression in the class duration, the duration of each student's facial expression, and the trend of the distribution ratio.

[0025] In one embodiment, the step of performing target detection on the classroom teaching video to be evaluated and determining the nonverbal interaction features between teachers and students based on the detection results includes:

[0026] Extract the target classroom video frame data from the classroom teaching video to be evaluated;

[0027] Target detection is performed on the target classroom video frame data, and the teacher's body area bounding box is determined based on the third detection result;

[0028] The proportion of time a teacher spends leaving the podium and entering the student area is determined based on the spatial overlap between the teacher's body area frame and the podium area frame.

[0029] Target detection is performed on the target classroom video data, and the attitude recognition is performed on the third detection result by the head posture estimation algorithm to obtain the pitch angle and yaw angle of all students.

[0030] The pitch angle and yaw angle are used to determine the head-up ratio of all students within a unit of time, and the head-up ratio is used to determine the head-up rate of students in the classroom.

[0031] The characteristics of nonverbal interaction between teachers and students in the classroom are determined based on the proportion of time the teacher spends away from the podium and enters the student area and the rate at which students look up in the classroom.

[0032] In one embodiment, the step of intelligently evaluating classroom teacher-student interaction based on the target tensor clustering model and the classroom interaction feature tensor to obtain a comprehensive evaluation report of classroom teacher-student interaction includes:

[0033] Determine the target core feature tensor based on the classroom interaction feature tensor;

[0034] Calculate the target similarity matrix based on the target core feature tensor, the classroom interaction feature tensor, and the feature space selection vector;

[0035] Cluster analysis of the target similarity matrix is ​​performed based on the target clustering algorithm to obtain clustering results of different classrooms in terms of interaction feature dimension;

[0036] Based on the target tensor clustering model, the classroom teacher-student interaction is intelligently evaluated according to the clustering results of different classrooms in the dimension of interaction features, and a comprehensive evaluation report of classroom teacher-student interaction is obtained.

[0037] In one embodiment, the step of determining the target core feature tensor based on the classroom interaction feature tensor includes:

[0038] The classroom interaction feature tensors are fused to obtain a global correlation tensor;

[0039] The global correlation tensor is normalized to obtain the transition probability tensor;

[0040] The weight ranking vector is calculated using the target iteration algorithm based on the model parameters of the transition probability tensor and the target tensor clustering model.

[0041] The target core feature tensor is determined based on the weight ranking vector and the feature space selection vector.

[0042] Secondly, this application provides an intelligent evaluation device for classroom teacher-student interaction, comprising:

[0043] The determination module is used to perform speech recognition on the classroom teaching videos to be evaluated, and to determine the characteristics of the verbal interaction between teachers and students in the classroom based on the recognition results;

[0044] The determining module is also used to perform target detection on the classroom teaching video to be evaluated, and determine the non-verbal interaction features between teachers and students in the classroom based on the detection results;

[0045] A construction module is used to generate a classroom interaction quantification matrix based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and to construct a classroom interaction feature tensor based on the classroom interaction quantification matrix;

[0046] The evaluation module is used to intelligently evaluate classroom teacher-student interaction based on the target tensor clustering model and the classroom interaction feature tensor, and obtain a comprehensive evaluation report of classroom teacher-student interaction.

[0047] Thirdly, this application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation thereof.

[0048] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to perform the method described in the first aspect or any possible implementation thereof.

[0049] Fifthly, this application provides a computer program product that, when run on a processor, causes the processor to perform the method described in the first aspect or any possible implementation thereof.

[0050] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0051] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art:

[0052] (1) When determining multi-dimensional interaction features, this application uses an object detection strategy, which can accurately identify and locate people, objects, and positions in video footage. In classroom teaching scenarios, the object detection strategy can realize the identification, location tracking, and behavior capture of teachers and students, providing data support for subsequent pose estimation and non-verbal behavior analysis. Compared with traditional methods based on background subtraction or motion detection, object detection based on deep learning convolutional neural network models can maintain stable recognition performance under complex lighting, occlusion, and multi-object scenarios, effectively improving the accuracy of detection in real classroom environments, and thus effectively improving the accuracy of evaluating teacher-student interactions.

[0053] (2) In determining the characteristics of nonverbal interaction between teachers and students in the classroom, this application also uses a facial expression recognition (FER) strategy, which uses a deep convolutional neural network to analyze facial images or video sequences to determine the individual's emotional state. By statistically analyzing the temporal distribution and proportion of changes in teachers' and students' facial expressions, it can effectively reflect the characteristics of nonverbal interaction between teachers and students in the classroom, such as classroom atmosphere, teaching emotions, and student participation. Compared with single posture or voice signals, facial expression recognition plays an irreplaceable role in capturing emotional dimensions and willingness to interact.

[0054] (3) This application is based on intelligent evaluation using a target tensor clustering model that can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure. Tensor clustering, as an important method of multidimensional data mining, can simultaneously process multidimensional information such as time, space, individuals, and features, and realize joint modeling and clustering analysis of complex behavioral patterns. Compared with traditional two-dimensional matrix clustering, tensor clustering can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure. It is suitable for multimodal fusion analysis of classroom teacher-student interaction. By constructing a classroom interaction tensor model, multiple modal data can be aggregated and analyzed in the same space to complete the intelligent evaluation of classroom teacher-student interaction. This can effectively improve the accuracy and fairness of evaluating teacher-student interaction, overcome the defects of strong human subjectivity and high difficulty in traditional methods, and provide scientific basis and technical support for teaching quality monitoring and classroom behavior research.

[0055] In summary, this application performs speech recognition on the classroom teaching video to be evaluated and determines the verbal interaction features between teachers and students based on the recognition results; it performs target detection on the classroom teaching video to be evaluated and determines the nonverbal interaction features between teachers and students based on the detection results; it generates a classroom interaction quantification matrix based on the verbal and nonverbal interaction features, and constructs a classroom interaction feature tensor based on the quantification matrix; based on the target tensor clustering model, it performs intelligent evaluation of classroom teacher-student interaction based on the classroom interaction feature tensor, and obtains a comprehensive evaluation report of classroom teacher-student interaction. Through the above method, by determining the verbal and nonverbal interaction features between teachers and students based on the classroom teaching video to be evaluated, and then performing intelligent evaluation based on the target tensor clustering model that can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure, the accuracy and fairness of evaluating teacher-student interaction can be effectively improved. Attached Figure Description

[0056] Figure 1 This is one of the flowcharts of the intelligent evaluation method for classroom teacher-student interaction provided in the embodiments of this application;

[0057] Figure 2 This is the second flowchart of the intelligent evaluation method for classroom teacher-student interaction provided in the embodiments of this application;

[0058] Figure 3 This is a schematic diagram of the module structure of the intelligent evaluation device for classroom teacher-student interaction provided in the embodiments of this application;

[0059] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0061] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.

[0062] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first response message" and "second response message," etc., are used to distinguish different response messages, not to describe a specific order of response messages.

[0063] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0064] Based on this, embodiments of this application provide an intelligent evaluation method for classroom teacher-student interaction, referring to... Figure 1 , Figure 1 This is one of the flowcharts illustrating the intelligent evaluation method for classroom teacher-student interaction provided in this application embodiment. In this embodiment, the intelligent evaluation method for classroom teacher-student interaction includes steps S10 to S40:

[0065] Step S10: Perform speech recognition on the classroom teaching video to be evaluated, and determine the characteristics of teacher-student verbal interaction based on the recognition results.

[0066] It should be noted that the classroom teaching video to be evaluated refers to the teaching time of the teacher lecturing to the students in the classroom. This classroom teaching video can be captured using microphones and high-definition cameras in the classroom. The characteristics of teacher-student verbal interaction in the classroom refer to the characteristics of interaction between teachers and students in the verbal dimension, such as the number of questions and answers between teachers and students, the number of rounds, etc.

[0067] Further, step S10 includes: acquiring the classroom teaching video to be evaluated and performing data cleaning on the classroom teaching video; extracting classroom teaching audio data from the data-cleaned classroom teaching video; transcribing the classroom teaching audio data and performing speech recognition on the transcribed classroom teaching audio data using a target speaker recognition model to obtain speech segments of teachers and students; detecting the speech segments of teachers and students using a target speech detection algorithm and performing semantic recognition on the detected classroom corpus text; recognizing the semantically recognized classroom corpus text and classifying it according to the recognized question text; and determining the characteristics of classroom teacher-student verbal interaction based on the cognitive level classification results.

[0068] Understandably, in order to effectively improve the quality and integrity of classroom teaching videos, data cleaning is required after obtaining the videos to be evaluated. This data cleaning process includes, but is not limited to, image stabilization, audio denoising, face detection, format standardization, duration trimming, and audio-video synchronization, which can provide standardized input data for subsequent operations.

[0069] It should be understood that, in order to provide a semantic foundation for classroom speech interaction analysis, after extracting classroom teaching audio data from the cleaned classroom teaching videos, the audio data needs to be transcribed. Since voiceprints are similar to human biometric information such as fingerprints, facial features, gait, and iris, they can be used to identify a person's physiological attributes. Therefore, key teacher and student speech segments can be identified using a target speaker recognition model. The classroom corpus text includes teacher question texts and student answer texts. After semantically recognizing the classroom corpus text, semantic features and keywords can be used to classify the cognitive levels based on the identified question texts. Specifically, if the question text contains words such as "what," "who," "when," "point out," and "definite," etc., then... Factual keywords are identified as belonging to the memory level; if the question text contains explanatory keywords such as "why," "reason," "how to explain," and "what does it represent," it is identified as belonging to the comprehension level; if the question text contains application keywords such as "give an example," "how to use," "how to apply," and "use...to," it is identified as belonging to the application level; if the question text contains analytical keywords such as "compare," "difference," "discrepancy," "connection," and "cause," it is identified as belonging to the analysis level; if the question text contains evaluative keywords such as "good or bad," "reasonable or unreasonable," and "what do you think," it is identified as belonging to the evaluation level; if the question text contains creative keywords such as "design," "propose a solution," "think of a method," and "predict the result," it is identified as belonging to the creative level.

[0070] Step S20: Target detection is performed on the classroom teaching video to be evaluated, and the non-verbal interaction features between teachers and students in the classroom are determined based on the detection results.

[0071] Understandably, object detection strategies, based on deep learning convolutional neural network models, are designed for high-precision identification and localization of people, objects, and locations in video footage. In classroom teaching scenarios, object detection strategies can identify, track, and capture the behavior of teachers and students, providing data support for subsequent pose estimation and nonverbal behavior analysis. Compared to traditional methods based on background subtraction or motion detection, object detection based on deep learning convolutional neural network models maintains stable recognition performance even under complex lighting, occlusion, and multi-object scenarios, making it more suitable for real-world classroom environments. Nonverbal interaction characteristics between teachers and students in the classroom refer to the nonverbal interaction features, such as the proportion of various teacher expressions in class time, the duration of each student's expressions, and trends in their distribution.

[0072] It should be noted that Facial Expression Recognition (FER) is a strategy that uses deep convolutional neural networks to analyze facial images or video sequences to determine an individual's emotional state. Common expression categories are categorized as positive, neutral, and negative. By statistically analyzing the temporal distribution and proportion of changes in teachers' and students' facial expressions, it is possible to effectively reflect nonverbal interaction characteristics between teachers and students in the classroom, such as classroom atmosphere, teaching emotions, and student participation. Compared to single gestures or voice signals, facial expression recognition plays an irreplaceable role in capturing emotional dimensions and willingness to interact.

[0073] Further, step S20 includes: extracting classroom video frame data facing the teacher and classroom video frame data facing the students from the classroom teaching video to be evaluated; performing target detection on the classroom video frame data facing the teacher, and determining the teacher's face region based on the first frame detection result; performing expression recognition on the frame data of the teacher's face region, and determining the proportion of each type of teacher expression in the class duration; performing target detection on the classroom video frame data facing the students, and determining the student's face region based on the second frame detection result; performing expression recognition on the frame data of the student's face region, and statistically analyzing the distribution ratio of the identified student expressions; using a time series analysis strategy, determining the duration of each student's expression and the trend of the distribution ratio change based on the distribution ratio; and obtaining the nonverbal interaction features between teachers and students in the classroom based on the proportion of each type of teacher expression in the class duration, the duration of each student's expression, and the trend of the distribution ratio change.

[0074] It should be noted that, for teachers, the best indicator reflecting emotional engagement in classroom interaction can be a parameter strongly correlated with facial expressions, such as the proportion of various teacher expressions in the class duration. Therefore, after extracting classroom video footage data for teachers from the teaching videos to be evaluated, it is necessary to identify the teacher's facial region and perform facial expression recognition on the footage data of the teacher's facial region. Since the teacher's facial expressions are different throughout the teaching process, the proportion of the identified teacher expressions in the class duration can be determined to form a teacher emotion change curve, which can be used to reflect the teacher's emotional engagement in classroom interaction.

[0075] Understandably, during class time, a teacher's facial expression score can be calculated based on a grading system, specifically:

[0076]

[0077] in, The teacher's facial expressions are scored. A positive expression from the teacher. This indicates a negative expression from the teacher.

[0078] It should be emphasized that after calculating the teacher's expression score using the above teacher grading rules, the teacher's expression score also needs to be normalized to a five-level scoring range, which can be [1, 5].

[0079] It should be understood that, for students, the best indicators reflecting their overall emotional state and classroom participation are parameters strongly correlated with facial expressions, such as the duration and distribution trend of facial expressions. Therefore, after extracting student-facing classroom video footage from the teaching videos to be evaluated, it is necessary to perform face detection and facial expression recognition on each frame, and then statistically analyze the distribution of student facial expressions, such as the distribution of positive, neutral, and negative expressions. Then, time series analysis strategies should be used to determine the duration and distribution trend of each student's facial expressions to reflect the students' overall emotional state and classroom participation.

[0080] It should be understood that, during classroom time, the average percentage of positive, neutral, and negative facial expressions can be expressed as:

[0081]

[0082] in, This indicates the average percentage of students displaying positive facial expressions. Indicates class time period, The number of students displaying positive expressions. This indicates the total number of people tested. This indicates the average percentage of students with neutral facial expressions. The number of students who expressed neutral expressions. This indicates the average percentage of students with negative facial expressions. The number of students displaying negative expressions.

[0083] Understandably, after obtaining the average percentage of each of the above-mentioned student expressions, the student expression score can be calculated using the following formula:

[0084]

[0085] in, The score is based on the student's facial expression. The weighting coefficients representing students' positive facial expressions This represents the weighting coefficient of students' neutral facial expressions. The weighting coefficient representing the negative facial expressions of students.

[0086] It should be noted that the weighting coefficients for the students' positive facial expressions mentioned above... The weighting factor for neutral student expressions can be preferably set to 1.0, and the weighting factor for negative student expressions can be preferably set to 0.5. It can be preferably set to 1.0. In addition, after calculating the birth expression score, the teacher expression score also needs to be normalized to a five-level rating.

[0087] Further, step S20 includes: extracting target classroom video frame data from the classroom teaching video to be evaluated; performing target detection on the target classroom video frame data and determining the teacher's body area bounding box based on the third detection result; determining the duration ratio of the teacher leaving the podium and entering the student area based on the spatial overlap result of the teacher's body area bounding box and the podium area bounding box; performing target detection on the target classroom video frame data and performing posture recognition on the third detection result through a head posture estimation algorithm to obtain the pitch angle and yaw angle of all students; determining the head-up ratio of all students per unit time based on the pitch angle and the yaw angle, and determining the student head-up rate in the teaching classroom based on the head-up ratio; and determining the non-verbal interaction features between teachers and students in the classroom based on the duration ratio of the teacher leaving the podium and entering the student area and the student head-up rate in the teaching classroom.

[0088] Understandably, to effectively improve the accuracy of determining the proportion of time a teacher leaves the podium to enter the student area, it is necessary to perform target detection on the target classroom video footage data and determine the teacher's body area bounding box based on the third detection result. Then, it is necessary to judge in real time whether there is spatial overlap between the teacher's body area bounding box and the podium area bounding box. If so, it indicates that the teacher has not left the podium; otherwise, it indicates that the teacher has left the podium and entered the student area, which is considered close proximity. At this point, the proportion of time the teacher leaves the podium to enter the student area can be determined based on the spatial overlap result between the teacher's body area bounding box and the podium area bounding box. The higher this proportion, the closer the teacher-student interaction distance, thereby assessing the degree of spatial interaction and classroom affinity between teachers and students. For example, through... Indicates the duration percentage, in At that time, it is divided into level 1, in At that time, it is divided into two levels, in At that time, it was divided into 3 levels. At that time, it was divided into 4 levels. The time is divided into 4 levels.

[0089] It should be understood that, in order to accurately measure the level of student attention during interaction, target detection is required on the target classroom video data, and posture recognition is performed using a head pose estimation algorithm. The pitch angle describes the angle of the head nodding motion; specifically, when the pitch angle is within a preset angle range and the yaw angle is less than a preset angle threshold, it indicates that the student is looking up. Through this method, it is possible to determine whether a student is looking up, and also to count the number of frames in which each student is looking up per unit time. Combined with the total number of detected frames, the proportion of each student looking up per unit time can be calculated.

[0090]

[0091] in, This indicates the percentage of students who look up within a given unit of time. This represents the number of frames each student is in a head-up position within a unit of time. This indicates the total number of detected frames.

[0092] Understandably, after obtaining the head-up ratio, the student head-up rate in the classroom can be determined by averaging the results. Specifically:

[0093]

[0094] in, This indicates the rate at which students look up in the classroom. This indicates the percentage of students who look up within a given unit of time. This indicates the total number of students.

[0095] It should be noted that the student head-up rate in the classroom can represent the students' attention level. After obtaining the student head-up rate in the classroom, the numerical range of the student head-up rate can be normalized and scored, and the score results can be divided into five levels. Among them, the higher the student head-up rate, the higher the corresponding score.

[0096] Step S30: Generate a classroom interaction quantification matrix based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and construct a classroom interaction feature tensor based on the classroom interaction quantification matrix.

[0097] It should be understood that after obtaining the verbal and nonverbal interaction characteristics of teachers and students in the classroom, these interaction characteristics can be graded and quantified. Specifically, for the verbal interaction characteristics, questions at different cognitive levels can be mapped to interaction depth levels. For example, the memory level corresponds to interaction depth level 1, the comprehension level to interaction depth level 2, the application level to interaction depth level 3, the analysis level to interaction depth level 4, and the evaluation and creation levels to interaction depth level 5. Furthermore, the overall classroom interaction depth is calculated using the average depth value of all question texts.

[0098] Understandably, after standardizing and quantifying the characteristics of verbal and nonverbal interactions between teachers and students in the classroom, a quantitative matrix of classroom interaction is constructed to provide a data foundation for subsequent clustering and comprehensive evaluation.

[0099] Step S40: Based on the target tensor clustering model, intelligent evaluation of classroom teacher-student interaction is performed according to the classroom interaction feature tensor to obtain a comprehensive evaluation report of classroom teacher-student interaction.

[0100] It should be noted that tensor clustering, as an important method in multidimensional data mining, can simultaneously process multidimensional information such as time, space, individuals, and features, enabling joint modeling and cluster analysis of complex behavioral patterns. Compared with traditional two-dimensional matrix clustering, tensor clustering can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure, making it suitable for multimodal fusion analysis of classroom teacher-student interactions.

[0101] Understandably, the target tensor clustering model can aggregate and analyze multiple modalities of data, such as teacher behavior, student feedback, semantic interaction, and emotional changes, in the same space. This enables intelligent evaluation of classroom teacher-student interaction and generates a comprehensive evaluation report, including but not limited to classroom interaction depth index, teacher-student emotional interaction index, and student participation index, thus achieving intelligent and quantitative evaluation of classroom teacher-student interaction. The target tensor clustering model can be a high-order Tucker decomposition or CP decomposition model.

[0102] This embodiment performs speech recognition on the classroom teaching video to be evaluated and determines the verbal interaction features between teachers and students based on the recognition results; it then performs target detection on the video and determines the nonverbal interaction features between teachers and students based on the detection results; a quantitative matrix of classroom interaction is generated based on the verbal and nonverbal interaction features, and a classroom interaction feature tensor is constructed based on this matrix; and intelligent evaluation of teacher-student interaction is performed based on the target tensor clustering model and the classroom interaction feature tensor, resulting in a comprehensive evaluation report of teacher-student interaction. By determining the verbal and nonverbal interaction features between teachers and students based on the video to be evaluated, and then performing intelligent evaluation based on a target tensor clustering model that reveals high-dimensional interaction relationships while maintaining data structure integrity, the accuracy and fairness of evaluating teacher-student interaction can be effectively improved.

[0103] In one specific implementation, this application provides steps for intelligent evaluation of classroom teacher-student interaction. Please refer to... Figure 2 , Figure 2 This is the second flowchart illustrating the intelligent evaluation method for classroom teacher-student interaction provided in this application embodiment. Step S40 includes steps S401 to S404:

[0104] Step S401: Determine the target core feature tensor based on the classroom interaction feature tensor.

[0105] It should be noted that the target core feature tensor refers to a low-dimensional, concise feature tensor, which no longer contains the original high-dimensional data such as audio and video. Instead, it uses the core weights that best represent the classroom interaction mode to describe it again. This classroom interaction mode can also be called the global correlation tensor.

[0106] Step S402: Calculate the target similarity matrix based on the target core feature tensor, the classroom interaction feature tensor, and the feature space selection vector.

[0107] Understandably, when calculating the target similarity matrix based on the target core feature tensor, classroom interaction feature tensor, and vector selection in the feature space, the distance vector concept can be used to calculate the difference vector between two lessons in the core feature space. The target similarity matrix is ​​the foundation for achieving multi-angle, refined clustering and can determine the clustering results. Classrooms that are similar from the target perspective will be grouped into the same category in subsequent clustering analysis. For example, the similarity matrix calculated from the perspective of "student participation" can divide classrooms into "high participation" and "low participation" clusters.

[0108] Step S403: Perform cluster analysis on the target similarity matrix based on the target clustering algorithm to obtain the clustering results of different classrooms in the dimension of interaction features.

[0109] It should be understood that the target clustering algorithm can be the Affinity Propagation Clustering (AP) algorithm. The core idea of ​​the AP clustering algorithm is to vote among data points through "message passing" to elect the most representative "representative point" as the cluster center. At this time, the target similarity matrix can be used as the input of the target clustering algorithm. Each target similarity matrix is ​​independently input into the target clustering algorithm, and cluster analysis is performed based on the target clustering algorithm to obtain the clustering results of different classrooms in terms of interaction feature dimensions.

[0110] Step S404: Based on the target tensor clustering model, intelligent evaluation of classroom teacher-student interaction is performed according to the clustering results of different classrooms in the interaction feature dimension, and a comprehensive evaluation report of classroom teacher-student interaction is obtained.

[0111] Understandably, after obtaining the clustering results of different classrooms in terms of interaction characteristics, the results are input into the target tensor clustering model. At this point, the classroom teacher-student interaction can be intelligently evaluated based on the target tensor clustering model. The output comprehensive evaluation report of classroom teacher-student interaction includes, but is not limited to, the classroom interaction depth index, the teacher-student emotional interaction index, and the student participation index, thereby realizing the intelligent and quantitative evaluation of classroom teacher-student interaction.

[0112] Further, the step of determining the target core feature tensor based on the classroom interaction feature tensor includes: fusing the classroom interaction feature tensor to obtain a global correlation tensor; normalizing the global correlation tensor to obtain a transition probability tensor; calculating a weight ranking vector based on the transition probability tensor and the model parameters of the target tensor clustering model using a target iteration algorithm; and determining the target core feature tensor based on the weight ranking vector and the feature space selection vector.

[0113] It should be understood that after obtaining the classroom interaction feature tensor, a tensor modeling strategy can be adopted to perform multimodal fusion of the different dimensions of the classroom interaction feature tensor to capture the overall classroom interaction pattern, which is the global correlation tensor. The transition probability tensor reveals the intrinsic correlation and dependency between different interaction features and quantifies the implicit and dynamic rules in classroom teacher-student interaction.

[0114] Understandably, after obtaining the transition probability tensor, the weight ranking vector can be calculated by combining the model parameters of the target tensor clustering model. That is, a stable "importance" score is calculated for each feature through the target iterative algorithm, representing the relative importance weight of each feature in the classroom interaction mode. In other words, when the iteration meets the upper limit or the optimization condition, the target core feature tensor can be determined by combining the feature space selection vector, that is, the feature space selection vector and the corresponding weight ranking vector are multiplied by tensor. Conversely, when the above conditions are not met, the inertial weight is adaptively adjusted, the particle state is updated, and the particles are mutated.

[0115] This embodiment determines the target core feature tensor based on the classroom interaction feature tensor; calculates the target similarity matrix based on the target core feature tensor, the classroom interaction feature tensor, and the feature space selection vector; performs cluster analysis on the target similarity matrix based on the target clustering algorithm to obtain clustering results for different classrooms in the interaction feature dimension; and performs intelligent evaluation of classroom teacher-student interaction based on the target tensor clustering model and the clustering results for different classrooms in the interaction feature dimension to obtain a comprehensive evaluation report of classroom teacher-student interaction. Through the above method, the classroom interaction feature tensor is converted into a low-dimensional, simplified target core feature tensor. Then, after calculating the target similarity matrix by combining the classroom interaction feature tensor and the feature space selection vector, the target similarity matrix is ​​used as input to the target clustering algorithm. After outputting the clustering results for different classrooms in the interaction feature dimension, intelligent evaluation is performed based on the target tensor clustering model, which can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure. This effectively improves the accuracy of evaluating teacher-student interaction.

[0116] The intelligent evaluation device for classroom teacher-student interaction provided in this application is described below. The intelligent evaluation device for classroom teacher-student interaction described below corresponds to and can be referred to in relation to the intelligent evaluation method for classroom teacher-student interaction described above. Please refer to... Figure 3 , Figure 3 This is a schematic diagram of the module structure of the intelligent evaluation device for classroom teacher-student interaction provided in the embodiments of this application, including:

[0117] The T10 module is used to perform speech recognition on the classroom teaching video to be evaluated and to determine the characteristics of the verbal interaction between teachers and students in the classroom based on the recognition results.

[0118] The determining module T10 is also used to perform target detection on the classroom teaching video to be evaluated, and determine the non-verbal interaction features between teachers and students in the classroom based on the detection results.

[0119] The construction module T20 is used to generate a classroom interaction quantization matrix based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and to construct a classroom interaction feature tensor based on the classroom interaction quantization matrix.

[0120] The evaluation module T30 is used to intelligently evaluate classroom teacher-student interaction based on the target tensor clustering model and the classroom interaction feature tensor, and obtain a comprehensive evaluation report of classroom teacher-student interaction.

[0121] This embodiment performs speech recognition on the classroom teaching video to be evaluated and determines the verbal interaction features between teachers and students based on the recognition results; it then performs target detection on the video and determines the nonverbal interaction features between teachers and students based on the detection results; a quantitative matrix of classroom interaction is generated based on the verbal and nonverbal interaction features, and a classroom interaction feature tensor is constructed based on this matrix; and intelligent evaluation of teacher-student interaction is performed based on the target tensor clustering model and the classroom interaction feature tensor, resulting in a comprehensive evaluation report of teacher-student interaction. By determining the verbal and nonverbal interaction features between teachers and students based on the video to be evaluated, and then performing intelligent evaluation based on a target tensor clustering model that reveals high-dimensional interaction relationships while maintaining data structure integrity, the accuracy and fairness of evaluating teacher-student interaction can be effectively improved.

[0122] It is understood that the detailed functional implementation of each of the above modules can be found in the description of the aforementioned method embodiments, and will not be repeated here.

[0123] It should be understood that the above-described device is used to execute the methods in the above embodiments. The implementation principle and technical effect of the corresponding program modules in the device are similar to those described in the above methods. The working process of the device can be referred to the corresponding process in the above methods, and will not be repeated here.

[0124] Based on the methods in the above embodiments, this application provides an electronic device, please refer to... Figure 4 , Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.

[0125] It should be noted that the system may include: a processor 10, a communications interface 20, a memory 30, and a communication bus 40. The processor 10, communications interface 20, and memory 30 communicate with each other via the communication bus 40. The processor 10 can invoke logical instructions stored in the memory 30 to execute the methods described in the above embodiments.

[0126] Furthermore, the logical instructions in the aforementioned memory 30 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0127] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0128] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0129] It is understood that the processor in the embodiments of this application can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0130] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disks, portable hard disks, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor.

[0131] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. Those skilled in the art will readily understand that the above descriptions are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An intelligent evaluation method for classroom teacher-student interaction, characterized in that, include: Speech recognition is performed on the classroom teaching videos to be evaluated, and the characteristics of teacher-student verbal interaction in the classroom are determined based on the recognition results; The classroom teaching video to be evaluated is subjected to target detection, and the non-verbal interaction characteristics of teachers and students in the classroom are determined based on the detection results; A classroom interaction quantification matrix is ​​generated based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and a classroom interaction feature tensor is constructed based on the classroom interaction quantification matrix. Based on the target tensor clustering model, the classroom teacher-student interaction is intelligently evaluated according to the classroom interaction feature tensor, and a comprehensive evaluation report of classroom teacher-student interaction is obtained. The steps of performing speech recognition on the classroom teaching video to be evaluated and determining the characteristics of teacher-student verbal interaction based on the recognition results include: Obtain the classroom teaching videos to be evaluated, and perform data cleaning on the classroom teaching videos to be evaluated; Extracting classroom audio data from cleaned classroom teaching videos; The classroom teaching audio data is transcribed, and the transcribed classroom teaching audio data is subjected to speech recognition using a target speaker recognition model to obtain speech segments of the teacher and students; The speech segments of the teacher and students are detected using a target speech detection algorithm, and semantic recognition is performed on the detected classroom text. The semantically recognized classroom texts are identified, and the cognitive levels are classified according to the identified question texts. The characteristics of classroom teacher-student verbal interaction were determined based on the results of cognitive level classification. The steps of performing target detection on the classroom teaching video to be evaluated and determining the nonverbal interaction features of teachers and students in the classroom based on the detection results include: Extract the teacher-facing and student-facing classroom video footage data from the classroom teaching videos to be evaluated; Target detection is performed on the classroom video footage data for teachers, and the teacher's face region is determined based on the detection results of the first frame. The image data of the teacher's face region is used to perform facial expression recognition, and the proportion of each type of teacher's facial expression recognized in the class time is determined. Target detection is performed on the classroom video footage data facing students, and the student face regions are determined based on the detection results of the second screen. Perform facial expression recognition on the image data of the student's face area and statistically analyze the distribution ratio of the recognized student facial expressions; Using time series analysis, the duration of each student's facial expression and the trend of the distribution ratio change are determined based on the distribution ratio. The characteristics of nonverbal interaction between teachers and students in the classroom are obtained based on the proportion of each type of teacher's facial expression in the class duration, the duration of each student's facial expression, and the trend of the distribution ratio. The steps of performing target detection on the classroom teaching video to be evaluated and determining the nonverbal interaction features of teachers and students in the classroom based on the detection results include: Extract the target classroom video frame data from the classroom teaching video to be evaluated; Target detection is performed on the target classroom video frame data, and the teacher's body area bounding box is determined based on the third detection result; The proportion of time a teacher spends leaving the podium and entering the student area is determined based on the spatial overlap between the teacher's body area frame and the podium area frame. Target detection is performed on the target classroom video data, and the attitude recognition is performed on the third detection result by the head posture estimation algorithm to obtain the pitch angle and yaw angle of all students. The pitch angle and yaw angle are used to determine the head-up ratio of all students within a unit of time, and the head-up ratio is used to determine the head-up rate of students in the classroom. The characteristics of nonverbal interaction between teachers and students in the classroom are determined based on the proportion of time the teacher spends away from the podium and enters the student area and the rate at which students look up in the classroom.

2. The method of claim 1, wherein, The steps of using the target tensor clustering model to intelligently evaluate classroom teacher-student interaction based on the classroom interaction feature tensor, and obtaining a comprehensive evaluation report of classroom teacher-student interaction, include: Determine the target core feature tensor based on the classroom interaction feature tensor; Calculate the target similarity matrix based on the target core feature tensor, the classroom interaction feature tensor, and the feature space selection vector; Cluster analysis of the target similarity matrix is ​​performed based on the target clustering algorithm to obtain clustering results of different classrooms in terms of interaction feature dimension; Based on the target tensor clustering model, the classroom teacher-student interaction is intelligently evaluated according to the clustering results of different classrooms in the dimension of interaction features, and a comprehensive evaluation report of classroom teacher-student interaction is obtained.

3. The method of claim 2, wherein, The step of determining the target core feature tensor based on the classroom interaction feature tensor includes: The classroom interaction feature tensors are fused to obtain a global correlation tensor; The global correlation tensor is normalized to obtain the transition probability tensor; The weight ranking vector is calculated using the target iteration algorithm based on the model parameters of the transition probability tensor and the target tensor clustering model. The target core feature tensor is determined based on the weight ranking vector and the feature space selection vector.

4. An intelligent evaluation device for classroom teacher-student interaction, characterized in that, include: The determination module is used to perform speech recognition on the classroom teaching videos to be evaluated, and to determine the characteristics of the verbal interaction between teachers and students in the classroom based on the recognition results; The determining module is also used to perform target detection on the classroom teaching video to be evaluated, and determine the non-verbal interaction features between teachers and students in the classroom based on the detection results; A construction module is used to generate a classroom interaction quantification matrix based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and to construct a classroom interaction feature tensor based on the classroom interaction quantification matrix; The evaluation module is used to intelligently evaluate classroom teacher-student interaction based on the target tensor clustering model and the classroom interaction feature tensor, and obtain a comprehensive evaluation report of classroom teacher-student interaction. The determining module is further configured to acquire the classroom teaching video to be evaluated and perform data cleaning on the video; extract classroom teaching audio data from the cleaned video; transcribe the audio data and perform speech recognition on the transcribed audio data using a target speaker recognition model to obtain speech segments of the teacher and students; detect the speech segments of the teacher and students using a target speech detection algorithm and perform semantic recognition on the detected classroom text; identify the semantically recognized text and classify it according to the cognitive level based on the identified question text; and determine the characteristics of classroom teacher-student verbal interaction based on the cognitive level classification results. The determining module is further configured to extract classroom video frame data facing the teacher and classroom video frame data facing the student from the classroom teaching video to be evaluated; perform target detection on the classroom video frame data facing the teacher, and determine the teacher's face region based on the first frame detection result; perform expression recognition on the frame data of the teacher's face region, and determine the proportion of the recognized teacher expressions in the class duration. Target detection is performed on the classroom video footage data facing students, and student face regions are determined based on the second image detection results; facial expression recognition is performed on the image data of the student face regions, and the distribution ratio of the recognized student facial expressions is statistically analyzed; through time series analysis, the duration of each student's facial expression and the trend of the distribution ratio are determined based on the distribution ratio; nonverbal interaction features between teachers and students in the classroom are obtained based on the proportion of each type of teacher's facial expression in the class duration, the duration of each student's facial expression, and the trend of the distribution ratio. The determining module is further configured to extract target classroom video frame data from the classroom teaching video to be evaluated; perform target detection on the target classroom video frame data, and determine the teacher's body area bounding box based on the third detection result; determine the duration ratio of the teacher leaving the podium and entering the student area based on the spatial overlap result of the teacher's body area bounding box and the podium area bounding box; perform target detection on the target classroom video frame data, and perform posture recognition on the third detection result through a head posture estimation algorithm to obtain the pitch angle and yaw angle of all students; determine the head-up ratio of all students per unit time based on the pitch angle and the yaw angle, and determine the student head-up rate in the teaching classroom based on the head-up ratio; and determine the non-verbal interaction features between teachers and students in the classroom based on the duration ratio of the teacher leaving the podium and entering the student area and the student head-up rate in the teaching classroom.

5. An electronic device, characterized in that, include: At least one memory for storing computer programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-3.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is run on the processor, it causes the processor to perform the method as described in any one of claims 1-3.

7. A computer program product, characterized in that, When the computer program product is run on a processor, the processor causes the processor to perform the method as described in any one of claims 1-3.