Intelligent evaluation method, device and equipment for classroom teacher-student interaction and storage medium
By performing speech recognition and target detection on classroom teaching videos, classroom teacher-student interaction features are generated. Using tensor clustering models for intelligent evaluation, the inaccuracy of evaluation caused by the neglect of multimodal information and the subjectivity of human observation in existing technologies are solved, and more accurate and fair teacher-student interaction evaluation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies neglect multimodal information in classroom teacher-student interaction assessments, resulting in lower accuracy and fairness in assessments, and human observation is greatly affected by subjective factors.
By performing speech recognition and target detection on classroom teaching videos, the verbal and non-verbal interaction features of teachers and students in the classroom are determined, an interaction quantification matrix is generated, and an intelligent evaluation is performed using a target tensor clustering model to generate a comprehensive evaluation report.
It improves the accuracy and fairness of classroom teacher-student interaction evaluation, overcomes the shortcomings of traditional methods such as human subjectivity and high operational difficulty, and provides a scientific basis for teaching quality monitoring and classroom behavior research.
Smart Images

Figure CN121640156A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of educational informatization, and more particularly, relates to an intelligent evaluation method, device and equipment for classroom teacher-student interaction and a storage medium. BACKGROUND
[0002] Classroom teacher-student interaction is not only the core of the teaching process, but also the embodiment of teacher-student communication, which has an important influence on the participation of students and the quality of classroom teaching. In the 1960s, Flanders proposed the Flanders' Interaction Analysis System (FIAS), marking the beginning of the systematic and scientific stage of teaching interaction research. FIAS observes and records the verbal interaction behavior of teachers and students through timed sampling, and then analyzes the data in a matrix table. The analysis results can be used to evaluate teaching quality, find teaching patterns, and guide teaching improvement. Since then, the FIAS system has been widely used and has been adapted and expanded in different disciplines, such as the Information Technology-based Interaction Analysis System (ITIAS), the Improved Flanders' Interaction Analysis System, and the Interactive Electronic Double Board-based Interaction Analysis System.
[0003] Classroom teacher-student interaction is a process of generating, transmitting and receiving multi-modal information, while the above-mentioned analysis systems directly ignore multi-modal information, resulting in incomplete and inaccurate analysis. Moreover, manual observation is involved in recording verbal interaction behavior. Obviously, manual observation is greatly affected by subjective factors and is difficult to operate. Therefore, the above-mentioned methods have low accuracy and fairness in evaluating teacher-student interaction. SUMMARY
[0004] In view of the defects of the prior art, the present application aims to provide an intelligent evaluation method, device, equipment and storage medium for classroom teacher-student interaction, which aims to solve the problem of low accuracy and fairness in evaluating teacher-student interaction due to the direct neglect of multi-modal information and the involvement of manual observation which is greatly affected by subjective factors and difficult to operate.
[0005] To achieve the above-mentioned purpose, in a first aspect, the present application provides an intelligent evaluation method for classroom teacher-student interaction, comprising: performing speech recognition on a to-be-evaluated classroom teaching video, and determining classroom teacher-student verbal interaction features according to the recognition result; performing target detection on the to-be-evaluated classroom teaching video, and determining classroom teacher-student non-verbal interaction features according to the detection result; generate a classroom interaction quantization matrix according to the classroom teacher-student verbal interaction features and the classroom teacher-student non-verbal interaction features, and construct a classroom interaction feature tensor according to the classroom interaction quantization matrix; based on a target tensor clustering model, intelligently evaluate classroom teacher-student interaction according to the classroom interaction feature tensor, and obtain a classroom teacher-student interaction comprehensive evaluation report.
[0006] In an embodiment, the step of performing speech recognition on the classroom teaching video to be evaluated and determining classroom teacher-student verbal interaction features according to the recognition result comprises: obtaining a classroom teaching video to be evaluated, and performing data cleaning on the classroom teaching video to be evaluated; extracting classroom teaching audio data in the data-cleaned classroom teaching video; transcribing the classroom teaching audio data, and performing speech recognition on the transcribed classroom teaching audio data through a target speaker recognition model to obtain teacher and student speech segments; detecting the teacher and student speech segments through a target speech detection algorithm, and performing semantic recognition on the detected classroom corpus text; recognizing the classroom corpus text after semantic recognition, and classifying the cognitive level according to the recognized question text; determining classroom teacher-student verbal interaction features according to the cognitive level classification result.
[0007] In an embodiment, the step of performing target detection on the classroom teaching video to be evaluated and determining classroom teacher-student non-verbal interaction features according to the detection result comprises: extracting teacher-oriented classroom video picture data and student-oriented classroom video picture data in the classroom teaching video to be evaluated; performing target detection on the teacher-oriented classroom video picture data, and determining a teacher face area according to a first picture detection result; performing expression recognition on the picture data of the teacher face area, and determining the proportion of each type of teacher expression in the classroom duration; performing target detection on the student-oriented classroom video picture data, and determining a student face area according to a second picture detection result; performing expression recognition on the picture data of the student face area, and counting the distribution proportion of the recognized student expressions; determining the duration and distribution proportion change trend of each student expression according to the distribution proportion through a time series analysis strategy; obtaining classroom teacher-student non-verbal interaction features according to the proportion of each type of teacher expression in the classroom duration, the duration of each student expression, and the distribution proportion change trend.
[0008] In an embodiment, the step of performing target detection on the classroom teaching video to be evaluated and determining a teacher-student non-verbal interaction feature according to a detection result comprises: extracting target classroom video picture data in the classroom teaching video to be evaluated; performing target detection on the target classroom video picture data and determining a teacher body region frame according to a third detection result; determining a proportion of time that a teacher leaves a podium and enters a student area according to a spatial overlap result of the teacher body region frame and the podium region frame; performing target detection on the target classroom video picture data and performing pose recognition on the third detection result through a head pose estimation algorithm to obtain a pitch angle and a yaw angle of all students; determining a proportion of students who look up in a unit of time according to the pitch angle and the yaw angle, and determining a student head-up rate in a teaching classroom according to the proportion of students who look up; determining a teacher-student non-verbal interaction feature according to the proportion of time that the teacher leaves the podium and enters the student area and the student head-up rate in the teaching classroom.
[0009] In an embodiment, the step of performing intelligent evaluation on teacher-student interaction in a classroom based on a target tensor clustering model according to the classroom interaction feature tensor to obtain a comprehensive evaluation report of teacher-student interaction in a classroom comprises: determining a target core feature tensor according to the classroom interaction feature tensor; calculating a target similarity matrix according to the target core feature tensor, the classroom interaction feature tensor, and a feature space selection vector; performing clustering analysis on the target similarity matrix based on a target clustering algorithm to obtain a clustering result of different classrooms in an interaction feature dimension; performing intelligent evaluation on teacher-student interaction in a classroom based on a target tensor clustering model according to the clustering result of different classrooms in an interaction feature dimension to obtain a comprehensive evaluation report of teacher-student interaction in a classroom.
[0010] In an embodiment, the step of determining a target core feature tensor according to the classroom interaction feature tensor comprises: fusing the classroom interaction feature tensor to obtain a global correlation tensor; performing normalization processing on the global correlation tensor to obtain a transition probability tensor; calculating a weight ranking vector according to the transition probability tensor and model parameters of a target tensor clustering model through a target iterative algorithm; determining a target core feature tensor according to the weight ranking vector and the feature space selection vector.
[0011] In a second aspect, the present application provides an intelligent evaluation device for classroom teacher-student interaction, comprising: a determination module configured to perform voice recognition on a to-be-evaluated classroom teaching video, and determine a classroom teacher-student verbal interaction feature according to a recognition result; The determination module is further configured to perform target detection on the to-be-evaluated classroom teaching video, and determine a classroom teacher-student non-verbal interaction feature according to a detection result; a construction module configured to generate a classroom interaction quantization matrix according to the classroom teacher-student verbal interaction feature and the classroom teacher-student non-verbal interaction feature, and construct a classroom interaction feature tensor according to the classroom interaction quantization matrix; an evaluation module configured to perform intelligent evaluation on classroom teacher-student interaction according to the classroom interaction feature tensor based on a target tensor clustering model, and obtain a classroom teacher-student interaction comprehensive evaluation report.
[0012] In a third aspect, the present application provides an electronic device, comprising: at least one memory configured to store a program; and at least one processor configured to execute the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0013] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program runs on a processor, the processor is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0014] In a fifth aspect, the present application provides a computer program product, and when the computer program product runs on a processor, the processor is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0015] It can be understood that the beneficial effects of the above-mentioned second aspect to fifth aspect can be referred to the related description in the first aspect, and will not be repeated here.
[0016] Overall, compared with the prior art, the above technical solutions conceived by the present application have the following beneficial effects: (1) In determining the multi-dimensional interaction features, the application will use the object detection strategy, which can accurately identify and locate the characters, objects and positions in the video pictures. In the classroom teaching scene, the object detection strategy can realize the identity recognition, position tracking and behavior capture of teachers and students, and provide data support for subsequent posture estimation and non-verbal behavior analysis. Compared with the traditional method based on background difference or motion detection, the object detection based on deep learning convolutional neural network model can maintain stable recognition performance in complex lighting, occlusion and multi-target scene, effectively improve the accuracy of detection in real classroom environment, and further improve the accuracy of evaluating teacher-student interaction.
[0017] (2) In determining the non-verbal interaction features of teachers and students in the classroom, the application will also use the facial expression recognition (FER) strategy, which uses deep convolutional neural network to analyze face images or video sequences to determine the individual emotional state. By statistically analyzing the time distribution and proportion of changes in the expressions of teachers and students, the classroom atmosphere, teaching emotions and student participation level, and other non-verbal interaction features of teachers and students in the classroom can be effectively reflected. Compared with single posture or voice signal, facial expression recognition plays an irreplaceable role in capturing emotional dimension and interaction intention.
[0018] (3) The application is based on the target tensor clustering model which can reveal high-dimensional interaction relationship while maintaining the integrity of data structure for intelligent evaluation. Tensor clustering, as an important method of multi-dimensional data mining, can process multi-dimensional information such as time, space, individual and features at the same time, realize joint modeling and clustering analysis of complex behavior patterns. Compared with traditional two-dimensional matrix clustering, tensor clustering can reveal high-dimensional interaction relationship while maintaining the integrity of data structure, which is suitable for multi-modal fusion analysis of classroom teacher-student interaction. By constructing a classroom interaction tensor model, multiple modal data can be aggregated and analyzed in the same space to complete intelligent evaluation of classroom teacher-student interaction, thereby effectively improving the accuracy and fairness of evaluating teacher-student interaction, overcoming the defects of strong artificial subjectivity and difficulty in traditional methods, and providing scientific basis and technical support for teaching quality monitoring and classroom behavior research.
[0019] In summary, the speech recognition is performed on the to-be-evaluated classroom teaching video, and the teacher-student verbal interaction features are determined according to the recognition result; the target detection is performed on the to-be-evaluated classroom teaching video, and the teacher-student non-verbal interaction features are determined according to the detection result; the classroom interaction quantization matrix is generated according to the teacher-student verbal interaction features and the teacher-student non-verbal interaction features, and the classroom interaction feature tensor is constructed according to the classroom interaction quantization matrix; the intelligent evaluation of the teacher-student interaction is performed according to the classroom interaction feature tensor based on the target tensor clustering model, and the comprehensive evaluation report of the teacher-student interaction is obtained. In the above manner, the teacher-student verbal interaction features and the teacher-student non-verbal interaction features are determined according to the to-be-evaluated classroom teaching video, and then the intelligent evaluation is performed based on the target tensor clustering model which can maintain the integrity of the data structure while revealing the high-dimensional interaction relationship, so that the accuracy and fairness of the evaluation of the teacher-student interaction can be effectively improved. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 FIG. 1 is one of flow diagrams of the intelligent evaluation method of the teacher-student interaction provided by the embodiments of the present application; Figure 2 FIG. 2 is another of flow diagrams of the intelligent evaluation method of the teacher-student interaction provided by the embodiments of the present application; Figure 3 FIG. 3 is a module structure diagram of the intelligent evaluation device of the teacher-student interaction provided by the embodiments of the present application; Figure 4 FIG. 4 is a structure diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0022] The term "and / or" in this paper is a description of the association relationship between the associated objects, which means that there may be three kinds of relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. The symbol " / " in this paper represents the relationship of or in the associated objects, for example, A / B represents A or B.
[0023] The terms "first" and "second" and the like in the specification and claims herein are used to distinguish different objects, and are not used to describe the specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, and are not used to describe the specific order of the response messages.
[0024] In the embodiments of the present application, the words such as "exemplary" or "for example" are used to mean serving as an example, instance, or illustration, and should not be necessarily construed as a preference or a benefit. Rather, use of the words such as "exemplary" or "for example" is intended to present concepts in a concrete manner.
[0025] Based on this, the embodiments of the present application provide an intelligent evaluation method for classroom teacher-student interaction, referring to Figure 1 , Figure 1 is one of the process schematic diagrams of the intelligent evaluation method for classroom teacher-student interaction provided by the embodiments of the present application. In the embodiments, the intelligent evaluation method for classroom teacher-student interaction includes steps S10 to S40: Step S10, performing speech recognition on a to-be-evaluated classroom teaching video, and determining a classroom teacher-student verbal interaction feature according to a recognition result.
[0026] It should be noted that the to-be-evaluated classroom teaching video refers to a teaching time frequency in which a teacher gives a lecture to students in a classroom. The to-be-evaluated classroom teaching video can be collected by using a microphone and a high-definition camera in a teaching classroom. The classroom teacher-student verbal interaction feature refers to an interaction feature in a verbal dimension between the teacher and the students, for example, a number of question and answer rounds between the teacher and the students.
[0027] Further, step S10 includes: acquiring a to-be-evaluated classroom teaching video, and performing data cleaning on the to-be-evaluated classroom teaching video; extracting classroom teaching audio data in the data-cleaned classroom teaching video; transcribing the classroom teaching audio data, and performing speech recognition on the transcribed classroom teaching audio data by using a target speaker recognition model to obtain a voice segment of the teacher and the students; detecting the voice segment of the teacher and the students by using a target voice detection algorithm, and performing semantic recognition on detected classroom corpus text; performing recognition on the classroom corpus text after semantic recognition, and performing cognitive level classification according to recognized question text; and determining the classroom teacher-student verbal interaction feature according to a cognitive level classification result.
[0028] It can be understood that, in order to effectively improve the quality and integrity of the classroom teaching video, after the to-be-evaluated classroom teaching video is acquired, data cleaning needs to be performed on the to-be-evaluated classroom teaching video. The data cleaning operation includes but is not limited to picture stabilization processing, audio denoising, face detection, format unification, time length pruning, and audio-video synchronization, and can provide standardized input data for subsequent operations.
[0029] It should be understood that in order to provide a semantic basis for classroom speech interaction analysis, after extracting the classroom teaching audio data from the classroom teaching video after data cleaning, the classroom teaching audio data needs to be transcribed. Since voiceprint is similar to biological information such as fingerprints, facial features, walking posture, and pupil iris, it can be used to identify the physiological attributes of a person's identity, so the key teacher and student speech segments can be identified through a target speaker recognition model. The classroom corpus text includes teacher question corpus text and student answer corpus text. After identifying the classroom corpus text after semantic recognition, semantic features and keywords can be used to classify the cognitive level according to the identified question text. Specifically, if the question text contains keywords such as "what", "who", "when", "point out", and "define", it is identified as the memory level; if the question text contains keywords such as "why", "reason", "how to explain", and "what does it represent", it is identified as the understanding level; if the question text contains keywords such as "give an example", "how to use", "how to apply", and "use... to", it is identified as the application level; if the question text contains keywords such as "compare", "distinguish", "difference", "connection", and "cause", it is identified as the analysis level; if the question text contains keywords such as "good or not", "whether reasonable", and "what do you think", it is identified as the evaluation level; if the question text contains keywords such as "design", "propose a solution", "think of a method", and "predict the result", it is identified as the creation level.
[0030] In step S20, target detection is performed on the classroom teaching video to be evaluated, and classroom teacher-student non-verbal interaction features are determined based on the detection results.
[0031] It can be understood that the target detection (Object Detection) strategy is based on a convolutional neural network model of deep learning, which can accurately identify and locate people, objects, and positions in video images. In the classroom teaching scenario, the target detection strategy can realize the identification of teachers and students, location tracking, and behavior capture, providing data support for subsequent pose estimation and non-verbal behavior analysis. Compared with traditional methods based on background difference or motion detection, the target detection based on the convolutional neural network model of deep learning can still maintain stable recognition performance in complex lighting, occlusion, and multi-target scenarios, and is more suitable for application in real classroom environments. Classroom teacher-student non-verbal interaction features refer to the interaction features of teachers and students in the non-verbal dimension, such as the proportion of various teacher expressions in the classroom duration, the duration and distribution trend of various student expressions, etc.
[0032] It should be noted that the facial expression recognition (FER) strategy is a strategy of using a deep convolutional neural network to analyze a face image or a video sequence to determine the emotional state of an individual. Common expression categories are positive, neutral, and negative. By statistically analyzing the time distribution and proportion of changes in the expressions of teachers and students, the classroom atmosphere, teaching emotions, and student participation can be effectively reflected. Compared with single posture or voice signal, facial expression recognition plays an irreplaceable role in capturing emotional dimensions and interactive intentions.
[0033] Further, step S20 comprises: extracting teacher-oriented classroom video frame data and student-oriented classroom video frame data in the classroom teaching video to be evaluated; performing target detection on the teacher-oriented classroom video frame data, and determining a teacher face region according to a first frame detection result; performing facial expression recognition on frame data of the teacher face region, and determining a proportion of each type of teacher expression in a classroom time length; performing target detection on the student-oriented classroom video frame data, and determining a student face region according to a second frame detection result; performing facial expression recognition on frame data of the student face region, and statistically analyzing a distribution proportion of the recognized student expressions; determining a duration and a distribution proportion change trend of each student expression according to the distribution proportion through a time series analysis strategy; and obtaining classroom teacher-student non-verbal interaction features according to the proportion of each type of teacher expression in the classroom time length, the duration of each student expression, and the distribution proportion change trend.
[0034] It should be noted that for teachers, the best indicator reflecting classroom interactive emotional investment can be a parameter strongly associated with expressions, such as the proportion of each type of teacher expression in the classroom time length. Therefore, after extracting teacher-oriented classroom video frame data from the classroom teaching video to be evaluated, the teacher face region needs to be determined, and the frame data of the teacher face region needs to be subjected to facial expression recognition. In the entire teaching process, the teacher's expression is different, so the proportion of each type of teacher expression in the classroom time length can be determined to form a teacher emotion change curve, which is used to reflect the teacher's classroom interactive emotional investment.
[0035] It can be understood that the teacher expression score can be calculated according to the grading rules of the teacher's expression in the classroom time period, specifically:
[0036] wherein, teacher expression score, positive teacher expression, negative teacher expression.
[0037] It should be emphasized that after calculating the teacher's expression score using the above teacher grading rules, the teacher's expression score also needs to be normalized to a five-level scoring range, which can be [1, 5].
[0038] It should be understood that, for students, the best indicators reflecting their overall emotional state and classroom participation are parameters strongly correlated with facial expressions, such as the duration and distribution trend of facial expressions. Therefore, after extracting classroom video footage data from the teaching videos to be evaluated, it is necessary to perform face detection and facial expression recognition on each frame, and then statistically analyze the distribution of student facial expressions, such as the distribution of positive, neutral, and negative expressions. Then, time series analysis strategies should be used to determine the duration and distribution trend of each student's facial expressions to reflect the students' overall emotional state and classroom participation.
[0039] It should be understood that, during classroom time, the average percentage of positive, neutral, and negative facial expressions can be expressed as:
[0040] in, This indicates the average percentage of students displaying positive facial expressions. Indicates class time period, The number of students displaying positive expressions. This indicates the total number of people tested. This indicates the average percentage of students with neutral facial expressions. The number of students who expressed neutral expressions. This indicates the average percentage of students with negative facial expressions. The number of students displaying negative expressions.
[0041] Understandably, after obtaining the average percentage of each of the above-mentioned student expressions, the student expression score can be calculated using the following formula:
[0042] in, The score is based on the student's facial expression. The weighting coefficients representing students' positive facial expressions This represents the weighting coefficient of students' neutral facial expressions. The weighting coefficient representing the negative facial expressions of students.
[0043] It should be noted that the weighting coefficients for the students' positive facial expressions mentioned above... The weighting factor for neutral student expressions can be preferably set to 1.0, and the weighting factor for negative student expressions can be preferably set to 0.5. It can be preferably set to 1.0. In addition, after calculating the birth expression score, the teacher expression score also needs to be normalized to a five-level rating.
[0044] Further, step S20 includes: extracting target classroom video frame data from the classroom teaching video to be evaluated; performing target detection on the target classroom video frame data and determining the teacher's body area bounding box based on the third detection result; determining the duration ratio of the teacher leaving the podium and entering the student area based on the spatial overlap result of the teacher's body area bounding box and the podium area bounding box; performing target detection on the target classroom video frame data and performing posture recognition on the third detection result through a head posture estimation algorithm to obtain the pitch angle and yaw angle of all students; determining the head-up ratio of all students per unit time based on the pitch angle and the yaw angle, and determining the student head-up rate in the teaching classroom based on the head-up ratio; and determining the non-verbal interaction features between teachers and students in the classroom based on the duration ratio of the teacher leaving the podium and entering the student area and the student head-up rate in the teaching classroom.
[0045] Understandably, to effectively improve the accuracy of determining the proportion of time a teacher leaves the podium to enter the student area, it is necessary to perform target detection on the target classroom video footage data and determine the teacher's body area bounding box based on the third detection result. Then, it is necessary to judge in real time whether there is spatial overlap between the teacher's body area bounding box and the podium area bounding box. If so, it indicates that the teacher has not left the podium; otherwise, it indicates that the teacher has left the podium and entered the student area, which is considered close proximity. At this point, the proportion of time the teacher leaves the podium to enter the student area can be determined based on the spatial overlap result between the teacher's body area bounding box and the podium area bounding box. The higher this proportion, the closer the teacher-student interaction distance, thereby assessing the degree of spatial interaction and classroom affinity between teachers and students. For example, through... Indicates the duration ratio, in At that time, it is divided into level 1, in At that time, it is divided into two levels, in At that time, it was divided into 3 levels. At that time, it was divided into 4 levels. The time is divided into 4 levels.
[0046] It should be understood that, in order to accurately measure the level of student attention during interaction, target detection is required on the target classroom video data, and posture recognition is performed using a head pose estimation algorithm. The pitch angle describes the angle of the head nodding motion; specifically, when the pitch angle is within a preset angle range and the yaw angle is less than a preset angle threshold, it indicates that the student is looking up. Through this method, it is possible to determine whether a student is looking up, and also to count the number of frames in which each student is looking up per unit time. Combined with the total number of detected frames, the proportion of each student looking up per unit time can be calculated.
[0047] in, This indicates the percentage of students who look up within a given unit of time. This represents the number of frames each student is in a head-up position within a unit of time. This indicates the total number of detected frames.
[0048] Understandably, after obtaining the head-up ratio, the student head-up rate in the classroom can be determined by averaging the results. Specifically:
[0049] in, This indicates the rate at which students look up in the classroom. This indicates the percentage of students who look up within a given unit of time. This indicates the total number of students.
[0050] It should be noted that the student head-up rate in the classroom can represent the students' attention level. After obtaining the student head-up rate in the classroom, the numerical range of the student head-up rate can be normalized and scored, and the score results can be divided into five levels. Among them, the higher the student head-up rate, the higher the corresponding score.
[0051] Step S30: Generate a classroom interaction quantification matrix based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and construct a classroom interaction feature tensor based on the classroom interaction quantification matrix.
[0052] It should be understood that after obtaining the verbal and nonverbal interaction characteristics of teachers and students in the classroom, these interaction characteristics can be graded and quantified. Specifically, for the verbal interaction characteristics, questions at different cognitive levels can be mapped to interaction depth levels. For example, the memory level corresponds to interaction depth level 1, the comprehension level to interaction depth level 2, the application level to interaction depth level 3, the analysis level to interaction depth level 4, and the evaluation and creation levels to interaction depth level 5. Furthermore, the overall classroom interaction depth is calculated using the average depth value of all question texts.
[0053] Understandably, after standardizing and quantifying the characteristics of verbal and nonverbal interactions between teachers and students in the classroom, a quantitative matrix of classroom interaction is constructed to provide a data foundation for subsequent clustering and comprehensive evaluation.
[0054] Step S40: Based on the target tensor clustering model, intelligent evaluation of classroom teacher-student interaction is performed according to the classroom interaction feature tensor to obtain a comprehensive evaluation report of classroom teacher-student interaction.
[0055] It should be noted that tensor clustering, as an important method in multidimensional data mining, can simultaneously process multidimensional information such as time, space, individuals, and features, enabling joint modeling and cluster analysis of complex behavioral patterns. Compared with traditional two-dimensional matrix clustering, tensor clustering can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure, making it suitable for multimodal fusion analysis of classroom teacher-student interactions.
[0056] Understandably, the target tensor clustering model can aggregate and analyze multiple modalities of data, such as teacher behavior, student feedback, semantic interaction, and emotional changes, in the same space. This enables intelligent evaluation of classroom teacher-student interaction and generates a comprehensive evaluation report, including but not limited to classroom interaction depth index, teacher-student emotional interaction index, and student participation index, thus achieving intelligent and quantitative evaluation of classroom teacher-student interaction. The target tensor clustering model can be a high-order Tucker decomposition or CP decomposition model.
[0057] This embodiment performs speech recognition on the classroom teaching video to be evaluated and determines the verbal interaction features between teachers and students based on the recognition results; it then performs target detection on the video and determines the nonverbal interaction features between teachers and students based on the detection results; a quantitative matrix of classroom interaction is generated based on the verbal and nonverbal interaction features, and a classroom interaction feature tensor is constructed based on this matrix; and intelligent evaluation of teacher-student interaction is performed based on the target tensor clustering model and the classroom interaction feature tensor, resulting in a comprehensive evaluation report of teacher-student interaction. By determining the verbal and nonverbal interaction features between teachers and students based on the video to be evaluated, and then performing intelligent evaluation based on a target tensor clustering model that reveals high-dimensional interaction relationships while maintaining data structure integrity, the accuracy and fairness of evaluating teacher-student interaction can be effectively improved.
[0058] In one specific implementation, this application provides steps for intelligent evaluation of classroom teacher-student interaction. Please refer to... Figure 2 , Figure 2 This is the second flowchart illustrating the intelligent evaluation method for classroom teacher-student interaction provided in this application embodiment. Step S40 includes steps S401 to S404: Step S401: Determine the target core feature tensor based on the classroom interaction feature tensor.
[0059] It should be noted that the target core feature tensor refers to a low-dimensional, concise feature tensor, which no longer contains the original high-dimensional data such as audio and video. Instead, it uses the core weights that best represent the classroom interaction mode to describe it again. This classroom interaction mode can also be called the global correlation tensor.
[0060] Step S402: Calculate the target similarity matrix based on the target core feature tensor, the classroom interaction feature tensor, and the feature space selection vector.
[0061] Understandably, when calculating the target similarity matrix based on the target core feature tensor, classroom interaction feature tensor, and vector selection in the feature space, the distance vector concept can be used to calculate the difference vector between two lessons in the core feature space. The target similarity matrix is the foundation for achieving multi-angle, refined clustering and can determine the clustering results. Classrooms that are similar from the target perspective will be grouped into the same category in subsequent clustering analysis. For example, the similarity matrix calculated from the perspective of "student participation" can divide classrooms into "high participation" and "low participation" clusters.
[0062] Step S403: Perform cluster analysis on the target similarity matrix based on the target clustering algorithm to obtain the clustering results of different classrooms in the dimension of interaction features.
[0063] It should be understood that the target clustering algorithm can be the Affinity Propagation Clustering (AP) algorithm. The core idea of the AP clustering algorithm is to vote among data points through "message passing" to elect the most representative "representative point" as the cluster center. At this time, the target similarity matrix can be used as the input of the target clustering algorithm. Each target similarity matrix is independently input into the target clustering algorithm, and cluster analysis is performed based on the target clustering algorithm to obtain the clustering results of different classrooms in terms of interaction feature dimensions.
[0064] Step S404: Based on the target tensor clustering model, intelligent evaluation of classroom teacher-student interaction is performed according to the clustering results of different classrooms in the interaction feature dimension, and a comprehensive evaluation report of classroom teacher-student interaction is obtained.
[0065] Understandably, after obtaining the clustering results of different classrooms in terms of interaction characteristics, the results are input into the target tensor clustering model. At this point, the classroom teacher-student interaction can be intelligently evaluated based on the target tensor clustering model. The output comprehensive evaluation report of classroom teacher-student interaction includes, but is not limited to, the classroom interaction depth index, the teacher-student emotional interaction index, and the student participation index, thereby realizing the intelligent and quantitative evaluation of classroom teacher-student interaction.
[0066] Further, the step of determining the target core feature tensor based on the classroom interaction feature tensor includes: fusing the classroom interaction feature tensor to obtain a global correlation tensor; normalizing the global correlation tensor to obtain a transition probability tensor; calculating a weight ranking vector based on the transition probability tensor and the model parameters of the target tensor clustering model using a target iteration algorithm; and determining the target core feature tensor based on the weight ranking vector and the feature space selection vector.
[0067] It should be understood that after obtaining the classroom interaction feature tensor, a tensor modeling strategy can be adopted to perform multimodal fusion of the different dimensions of the classroom interaction feature tensor to capture the overall classroom interaction pattern, which is the global correlation tensor. The transition probability tensor reveals the intrinsic correlation and dependency between different interaction features and quantifies the implicit and dynamic rules in classroom teacher-student interaction.
[0068] Understandably, after obtaining the transition probability tensor, the weight ranking vector can be calculated by combining the model parameters of the target tensor clustering model. That is, a stable "importance" score is calculated for each feature through the target iterative algorithm, representing the relative importance weight of each feature in the classroom interaction mode. In other words, when the iteration meets the upper limit or the optimization condition, the target core feature tensor can be determined by combining the feature space selection vector, that is, the feature space selection vector and the corresponding weight ranking vector are multiplied by tensor. Conversely, when the above conditions are not met, the inertial weight is adaptively adjusted, the particle state is updated, and the particles are mutated.
[0069] This embodiment determines the target core feature tensor based on the classroom interaction feature tensor; calculates the target similarity matrix based on the target core feature tensor, the classroom interaction feature tensor, and the feature space selection vector; performs cluster analysis on the target similarity matrix based on the target clustering algorithm to obtain clustering results for different classrooms in the interaction feature dimension; and performs intelligent evaluation of classroom teacher-student interaction based on the target tensor clustering model and the clustering results for different classrooms in the interaction feature dimension to obtain a comprehensive evaluation report of classroom teacher-student interaction. Through the above method, the classroom interaction feature tensor is converted into a low-dimensional, simplified target core feature tensor. Then, after calculating the target similarity matrix by combining the classroom interaction feature tensor and the feature space selection vector, the target similarity matrix is used as input to the target clustering algorithm. After outputting the clustering results for different classrooms in the interaction feature dimension, intelligent evaluation is performed based on the target tensor clustering model, which can reveal high-dimensional interaction relationships while maintaining the integrity of the data structure. This effectively improves the accuracy of evaluating teacher-student interaction.
[0070] The intelligent evaluation device for classroom teacher-student interaction provided in this application is described below. The intelligent evaluation device for classroom teacher-student interaction described below corresponds to and can be referred to in relation to the intelligent evaluation method for classroom teacher-student interaction described above. Please refer to... Figure 3 , Figure 3 This is a schematic diagram of the module structure of the intelligent evaluation device for classroom teacher-student interaction provided in the embodiments of this application, including: The T10 module is used to perform speech recognition on the classroom teaching video to be evaluated and to determine the characteristics of the verbal interaction between teachers and students in the classroom based on the recognition results.
[0071] The determining module T10 is also used to perform target detection on the classroom teaching video to be evaluated, and determine the non-verbal interaction features between teachers and students in the classroom based on the detection results.
[0072] The construction module T20 is used to generate a classroom interaction quantization matrix based on the classroom teacher-student verbal interaction features and the classroom teacher-student nonverbal interaction features, and to construct a classroom interaction feature tensor based on the classroom interaction quantization matrix.
[0073] The evaluation module T30 is used to intelligently evaluate classroom teacher-student interaction based on the target tensor clustering model and the classroom interaction feature tensor, and obtain a comprehensive evaluation report of classroom teacher-student interaction.
[0074] This embodiment performs speech recognition on the classroom teaching video to be evaluated and determines the verbal interaction features between teachers and students based on the recognition results; it then performs target detection on the video and determines the nonverbal interaction features between teachers and students based on the detection results; a quantitative matrix of classroom interaction is generated based on the verbal and nonverbal interaction features, and a classroom interaction feature tensor is constructed based on this matrix; and intelligent evaluation of teacher-student interaction is performed based on the target tensor clustering model and the classroom interaction feature tensor, resulting in a comprehensive evaluation report of teacher-student interaction. By determining the verbal and nonverbal interaction features between teachers and students based on the video to be evaluated, and then performing intelligent evaluation based on a target tensor clustering model that reveals high-dimensional interaction relationships while maintaining data structure integrity, the accuracy and fairness of evaluating teacher-student interaction can be effectively improved.
[0075] It is understood that the detailed functional implementation of each of the above modules can be found in the description of the aforementioned method embodiments, and will not be repeated here.
[0076] It should be understood that the above-described device is used to execute the methods in the above embodiments. The implementation principle and technical effect of the corresponding program modules in the device are similar to those described in the above methods. The working process of the device can be referred to the corresponding process in the above methods, and will not be repeated here.
[0077] Based on the methods in the above embodiments, this application provides an electronic device, please refer to... Figure 4 , Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.
[0078] It should be noted that the system may include: a processor 10, a communications interface 20, a memory 30, and a communication bus 40. The processor 10, communications interface 20, and memory 30 communicate with each other via the communication bus 40. The processor 10 can invoke logical instructions stored in the memory 30 to execute the methods described in the above embodiments.
[0079] Furthermore, the logical instructions in the aforementioned memory 30 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0080] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0081] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0082] It is understood that the processor in the embodiments of this application can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.
[0083] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disks, portable hard disks, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor.
[0084] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. Those skilled in the art will readily understand that the above descriptions are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An intelligent evaluation method for classroom teacher-student interaction, characterized in that, The method comprises the following steps: performing speech recognition on a to-be-evaluated classroom teaching video, and determining classroom teacher-student verbal interaction features according to the recognition result; performing target detection on the to-be-evaluated classroom teaching video, and determining classroom teacher-student non-verbal interaction features according to the detection result; generating a classroom interaction quantization matrix according to the classroom teacher-student verbal interaction features and the classroom teacher-student non-verbal interaction features, and constructing a classroom interaction feature tensor according to the classroom interaction quantization matrix; based on a target tensor clustering model, intelligently evaluating classroom teacher-student interaction according to the classroom interaction feature tensor to obtain a classroom teacher-student interaction comprehensive evaluation report.
2. The method of claim 1, wherein, The step of performing speech recognition on a to-be-evaluated classroom teaching video, and determining classroom teacher-student verbal interaction features according to the recognition result, comprises the following steps: obtaining a to-be-evaluated classroom teaching video, and performing data cleaning on the to-be-evaluated classroom teaching video; extracting classroom teaching audio data in the data-cleaned classroom teaching video; transcribing the classroom teaching audio data, and performing speech recognition on the transcribed classroom teaching audio data through a target speaker recognition model to obtain teacher and student voice segments; detecting the teacher and student voice segments through a target voice detection algorithm, and performing semantic recognition on the detected classroom corpus text; performing recognition on the classroom corpus text after semantic recognition, and classifying the recognized question texts according to cognitive levels; determining classroom teacher-student verbal interaction features according to the cognitive level classification result.
3. The method of claim 1, wherein, The step of performing target detection on the to-be-evaluated classroom teaching video, and determining classroom teacher-student non-verbal interaction features according to the detection result, comprises the following steps: extracting teacher-oriented classroom video picture data and student-oriented classroom video picture data in the to-be-evaluated classroom teaching video; performing target detection on the teacher-oriented classroom video picture data, and determining a teacher face area according to a first picture detection result; performing expression recognition on picture data of the teacher face area, and determining the proportion of each type of teacher expression in the classroom time; performing target detection on the student-oriented classroom video picture data, and determining a student face area according to a second picture detection result; performing expression recognition on picture data of the student face area, and counting the distribution proportion of the recognized student expressions; determining the duration and distribution proportion change trend of each student expression according to the distribution proportion through a time series analysis strategy; obtaining classroom teacher-student non-verbal interaction features according to the proportion of each type of teacher expression in the classroom time, the duration of each student expression, and the distribution proportion change trend.
4. The method of claim 1, wherein, The step of performing target detection on the to-be-evaluated classroom teaching video, and determining classroom teacher-student non-verbal interaction features according to the detection result, comprises the following steps: extracting target classroom video picture data in the to-be-evaluated classroom teaching video; performing target detection on the target classroom video picture data, and determining a teacher body area frame according to a third detection result; determining the proportion of the time when the teacher leaves the podium and enters the student area according to the spatial overlap result of the teacher body area frame and the podium area frame; The target classroom video picture data is subjected to target detection, and a head pose estimation algorithm is used to perform pose recognition on the third detection result to obtain the pitch angle and the yaw angle of all the students; The pitch angle and the yaw angle are used to determine the head-raising proportion of the students in a unit time, and the head-raising rate of the students in the teaching classroom is determined according to the head-raising proportion; The head-raising rate of the students in the teaching classroom and the time length proportion of the teacher leaving the platform and entering the student area are used to determine the classroom teacher-student non-verbal interaction feature.
5. The method of any one of claims 1 to 4, wherein, The step of intelligently evaluating the classroom teacher-student interaction according to the classroom interaction feature tensor based on the target tensor clustering model to obtain a classroom teacher-student interaction comprehensive evaluation report comprises: A target core feature tensor is determined according to the classroom interaction feature tensor; A target similarity matrix is calculated according to the target core feature tensor, the classroom interaction feature tensor, and a feature space selection vector; The target similarity matrix is subjected to clustering analysis based on a target clustering algorithm to obtain a clustering result of different classrooms in the interaction feature dimension; The classroom teacher-student interaction is intelligently evaluated according to the clustering result of different classrooms in the interaction feature dimension based on the target tensor clustering model to obtain a classroom teacher-student interaction comprehensive evaluation report.
6. The method of claim 5, wherein, The step of determining a target core feature tensor according to the classroom interaction feature tensor comprises: A global correlation tensor is obtained by fusing the classroom interaction feature tensor; A transition probability tensor is obtained by normalizing the global correlation tensor; A weight ranking vector is calculated according to the transition probability tensor and the model parameter of the target tensor clustering model by a target iteration algorithm; A target core feature tensor is determined according to the weight ranking vector and the feature space selection vector.
7. An intelligent evaluation device for classroom teacher-student interaction, characterized in that, Comprise: A determination module is configured to perform speech recognition on a classroom teaching video to be evaluated, and determine classroom teacher-student verbal interaction features according to the recognition result; The determination module is further configured to perform target detection on the classroom teaching video to be evaluated, and determine classroom teacher-student non-verbal interaction features according to the detection result; A construction module is configured to generate a classroom interaction quantization matrix according to the classroom teacher-student verbal interaction features and the classroom teacher-student non-verbal interaction features, and construct a classroom interaction feature tensor according to the classroom interaction quantization matrix; An evaluation module is configured to intelligently evaluate the classroom teacher-student interaction according to the classroom interaction feature tensor based on a target tensor clustering model to obtain a classroom teacher-student interaction comprehensive evaluation report.
8. An electronic device, comprising: Comprise: At least one memory for storing a computer program; At least one processor for executing the program stored in the memory, when the program stored in the memory is executed, the processor is used to execute the method as claimed in any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. When the computer program runs on the processor, the processor is caused to execute the method as claimed in any one of claims 1-6.
10. A computer program product, characterised in that, When the computer program product runs on the processor, the processor is caused to execute the method as claimed in any one of claims 1-6.
Citation Information
Patent Citations
Multi-dimensional classroom quantification system and method based on computer vision
CN110334610A
Classroom teaching effect evaluation system based on voice multi-feature progressive embedding
CN118782096A
Independent component repeatability analysis method based on correlation tensor clustering
CN120929858A
Learning situation analysis method, electronic device, and storage medium
US20220254158A1