Intelligent teaching evaluation and diagnosis system and method based on multi-mode audio and video analysis

By acquiring multimodal audio and video data and analyzing it using MTES technology, the problems of modal fragmentation and fixed-segment interference in teaching evaluation were solved. This enabled dynamic segmentation and accurate diagnosis of teaching events, improving the accuracy of teaching evaluation and the completeness of segmented segments.

CN120997010AActive Publication Date: 2025-11-21NANJING HEARING TECH CO LTD

Patent Information

Application Number
CN202511517111.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing teaching evaluation technologies process video, audio, and screen content independently, failing to capture cross-modal correlation phenomena, and the fixed-time slicing method leads to interference between teaching evaluation and diagnostic results.

Method used

By deploying camera arrays, directional microphone groups, and screen capture devices, time-stamp-aligned video streams of teacher and student behavior, audio streams, and teaching screen content streams are collected simultaneously to construct multimodal audio and video data. Combined with MTES technology, teaching status features are extracted and time-series model analysis is performed to achieve dynamic segmentation and diagnostic assessment.

Benefits of technology

It effectively solves the problem of cross-modal spatiotemporal mismatch, improves the accuracy and precision of teaching evaluation, and ensures the integrity of segmented segments and the reliability of diagnostic reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997010A_ABST
    Figure CN120997010A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of teaching evaluation and diagnosis management, in particular to an intelligent teaching evaluation and diagnosis system and method based on multi-mode audio and video analysis. Carrying out classroom teaching event identification on the teaching state feature set extracted at the current time; in combination with an MTES technology, classroom teaching event evolution process analysis based on a time sequence model is carried out, and classroom audios and videos are dynamically segmented based on an obtained classroom teaching event evolution process analysis result. According to the method, a mode of comprehensively quantifying the index set bound with each classroom teaching event based on the MTES technology can effectively solve a modal splitting phenomenon among modal acquisition data in a teaching evaluation and diagnosis process; according to the invention, the dynamic segmentation points of the classroom audio and video are screened, locked and calibrated from multiple angles, and the accuracy of the classroom audio and video segmentation segments is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of teaching evaluation, diagnosis and management technology, specifically to an intelligent teaching evaluation and diagnosis system and method based on multimodal audio and video analysis. Background Technology

[0002] Teaching evaluation is a core component for improving education quality, promoting teacher professional development, and ensuring student learning outcomes. Traditional teaching evaluation mainly relies on experts or peers to conduct evaluations through on-site classroom observations and manual completion of evaluation questionnaires.

[0003] With the development of computer technology and smart education, classroom teaching evaluation technology has shifted from manual observation to automated analysis. Existing technologies, such as analyzing student focus using a single camera or assessing teacher instruction quality based on speech recognition, have some effectiveness but also suffer from several bottlenecks. For example, existing technologies often process video, audio, and screen content independently, failing to capture cross-modal correlations such as students looking down at courseware while the teacher asks a question, resulting in significant fragmentation between data collected from different modalities. Furthermore, current technologies often use fixed time slices (e.g., 5 minutes) for teaching evaluation and diagnosis, which can lead to incorrect segmentation of transitional stages like "questioning → discussion," significantly interfering with subsequent evaluation and diagnostic results. Therefore, to address these technological bottlenecks, a new intelligent teaching evaluation and diagnostic system and method based on multimodal audio and video analysis is urgently needed. Summary of the Invention

[0004] The purpose of this invention is to provide an intelligent teaching evaluation and diagnosis system and method based on multimodal audio and video analysis to solve the problems mentioned in the background art.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: an intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis, the method comprising: S1. By deploying camera arrays, directional microphone groups and screen capture devices in the classroom, synchronously collect time-stamped video streams of teacher and student behavior, audio streams and teaching screen content streams to construct multimodal audio and video data; S2. Perform content parsing on the constructed multimodal audio and video data, and extract the corresponding teaching status feature set based on the content parsing results; S3. Combining teaching status feature data from historical data, identify classroom teaching events from the teaching status feature set extracted at the current time; and combine MTES technology to analyze the evolution process of classroom teaching events based on time series models. Based on the analysis results of the classroom teaching event evolution process, dynamically segment the classroom audio and video. S4. Based on the dynamic segmentation results of classroom audio and video, generate a diagnostic evaluation report corresponding to each segmented segment of classroom teaching audio and video.

[0006] This invention extracts classroom camera, microphone, and screen recording information to construct multimodal audio and video data (including teacher and student behavior, expressions, voice, teaching content, etc.), performs spatiotemporal alignment and feature extraction, and integrates the features of each modality to identify teaching events. Furthermore, it combines MTES technology to analyze the evolution process of classroom teaching events based on a time-series model (quantifying teaching events and the evolution trend of quantified indicators), dynamically segments classroom audio and video, and generates diagnostic evaluation reports (each segment is bound to a diagnostic evaluation report, each segment consists of one or more consecutive classroom teaching events, and the diagnostic evaluation report is a comprehensive analysis result of each classroom teaching event corresponding to each segment).

[0007] Furthermore, the camera array includes a panoramic camera covering the teacher's overall dynamics, a teacher tracking camera capturing the teacher's body language and blackboard writing, and a student expression camera collecting student facial response data.

[0008] This invention effectively solves the problem of spatiotemporal mismatch in data acquisition across modalities by using a multimodal data synchronous acquisition method. Furthermore, this invention achieves cross-modal correlation analysis between behavior, voice, and content through a camera array, directional microphone group, and screen capture device. The hierarchical deployment of the camera array, which achieves combined coverage of panoramic and close-up views, is intended to improve the detection accuracy of teacher movement trajectory and the recognition accuracy of student facial expressions.

[0009] Furthermore, the teaching status feature set in S2 includes teacher and student behavior features based on video stream parsing, voice interaction features based on audio stream parsing, and teaching content structure features based on screen content parsing. The teacher-student behavioral characteristics are composed of teacher movement trajectory density and student head-up rate. The teacher movement trajectory density at time t is denoted as Dt. Where Nt represents the number of times the teacher's position is updated within a preset unit time interval before time t in the acquired video stream; T represents the preset sampling time window length; and A represents the classroom area. The student head-up rate at time t is equal to the ratio of the number of students in the captured video stream who are looking at the teacher / PPT playback screen at time t to the total number of students. This invention uses quantitative formulas corresponding to various parameters in the process of acquiring teacher and student behavioral characteristics to effectively transform subjective teaching behaviors into objective indicator characteristics, which facilitates comprehensive analysis of various teaching state characteristics in subsequent steps. In the process of acquiring speech interaction features based on audio stream parsing, the frame length of the acquired audio stream is obtained, where the audio frame length represents the number of audio sampling points contained in the corresponding audio frame. The audio frame corresponding to any acquisition time and the midpoint of each sampling point within the corresponding audio frame are acquired. The short-time energy of the speech corresponding to the corresponding audio frame is obtained, serving as the speech interaction feature of the acquisition time corresponding to the midpoint of each sampling point within the corresponding audio frame. The short-time energy of the speech corresponding to the nth frame is denoted as En. Where K represents the number of preset audio sampling points contained in the audio frame; This represents the square of the amplitude (instantaneous power) corresponding to the k-th sampling point within the n-th frame. In the process of obtaining the structured features of teaching content based on screen content parsing, the screen recording text is recognized using OCR technology. Preset keyword types corresponding to the teaching content at collection time t are extracted, along with the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords. The teaching themes corresponding to the preset keyword types extracted from the database's preset forms are queried. Based on the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords, a teaching theme bias coefficient at collection time t is calculated. This teaching theme bias coefficient at collection time t is equal to the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords multiplied by the sum of the weight factors of the corresponding preset keyword types based on the corresponding teaching themes in the database's preset forms. The obtained teaching theme bias coefficient and the corresponding teaching theme are used as the first and second parameters in the structured features of the teaching content, respectively.

[0010] Furthermore, in the process of identifying classroom teaching events using the teaching state feature set extracted at the current time in step S3, a preset teaching event sequence is obtained, and the event probability of each element in the preset teaching event sequence based on the teaching state feature data at collection time t is calculated. The element with the highest event probability based on the teaching state feature data at collection time t in the preset teaching event sequence is taken as the classroom teaching event identification result at collection time t. The calculation formulas involved are as follows: Where FGt represents the teaching state characteristic data collected at time t; P(E|FGt) represents the event probability of any element E in the preset teaching event sequence based on the teaching state characteristic data FGt collected at time t; g mBoth gh and gh represent state feature functions, which are functions with base e and exponents calculated by subtracting the square of the difference between the average of all values ​​of the same teaching state feature type corresponding to the first parameter and the teaching state feature data corresponding to the second parameter in historical data. H represents a transition feature function. If the event state data pair formed by the first and second parameters in the transition feature function belongs to a preset effective transition set of event states, then the function value of the corresponding transition feature function is determined to be 1; otherwise, the function value of the corresponding transition feature function is determined to be 0. Each element in the effective transition set of event states corresponds to an event state data pair. The summary set of the first parameters in teacher movement trajectory density, student head-up rate, voice interaction features, and teaching content structured features in FGt is obtained and denoted as the first analysis set. m E represents the m-th parameter value in the first analysis set. t-1 The most recently identified teaching event is represented before the acquisition time t; W(FGt, m, E) represents the event adaptation evaluation coefficient of any element E in the preset teaching event sequence based on the teaching state feature data FGt at acquisition time t. Z(FGt) represents the sum of each W(FGt, m, E) corresponding to each element in the preset teaching event sequence; This represents the preset state feature weight corresponding to the m-th parameter in the first analysis set; The preset state feature weights corresponding to the teaching state evaluation coefficients are represented by μ; μ represents the transition feature weights, where the value of μ is equal to the ratio of the frequency of the event state data pair formed by the first and second parameters in the corresponding transition feature function appearing in the historical data to the sum of the frequencies of all event state data pairs in the historical data that have the same first parameter and the first parameter of the corresponding transition feature function; max{} represents the maximum value function. Where Ft represents the teaching status evaluation coefficient at acquisition time t; Lt represents the student head-up rate at acquisition time t; E(t) represents the short-time energy of the corresponding audio frame at acquisition time t; Xt represents the first parameter in the structured features of the teaching content at acquisition time t; r1, r2 and r3 are the preset first evaluation weight, second evaluation weight and third evaluation weight, respectively, and r1+r2+r3=1.

[0011] Furthermore, in the process of analyzing the evolution of classroom teaching events based on a time-series model using MTES technology, an array consisting of the identification results of each classroom teaching event is obtained sequentially according to time, denoted as the classroom teaching event evolution analysis array. In this array, before a new classroom teaching event is identified, the original classroom teaching identification results remain unchanged, and each time point corresponds to a unique classroom teaching event identification result. Based on MTES technology, the indicator set bound to each classroom teaching event in the classroom teaching time evolution analysis array is calculated. The indicator set includes interaction density and cognitive load index. The calculation method for interaction density is as follows: Where QN represents the number of valid questions asked by the teacher within the time period corresponding to the corresponding classroom teaching event identification result; TS represents the interval length of the time period corresponding to the corresponding classroom teaching event identification result; NSUM represents the sum of the number of students who participated in each teacher questioning interaction within the time period corresponding to the corresponding classroom teaching event identification result; and NA represents the total number of students in the classroom. The cognitive load index is calculated as follows: Where CH represents the cognitive load index corresponding to the identification results of the corresponding classroom teaching event; YP tb tb represents the average facial confusion level of each student at any time point tb within the time period corresponding to the corresponding classroom teaching event recognition result; the facial confusion level of a student is equal to the quotient of the number of wrinkles between the eyebrows obtained through image recognition divided by the area of ​​the area between the eyebrows; tb1 represents the minimum time point within the time period corresponding to the corresponding classroom teaching event recognition result; VS tb The VS represents the speech rate change coefficient at any time point tb within the time period corresponding to the classroom teaching event recognition result. tb The value is equal to the difference between the most recent monitored speech rate of the audio stream after time point tb in the historical data and the most recent monitored speech rate of the audio stream before time point tb in the historical data, divided by the quotient of the interval between two adjacent speech rate monitoring.

[0012] This invention uses MTES technology to calculate the interaction density and cognitive load index of each classroom teaching event in the classroom teaching time evolution analysis array. Combined with the constructed multimodal audio and video data, it can effectively solve the modal fragmentation phenomenon between the data collected from different modalities in the teaching evaluation and diagnosis process.

[0013] Furthermore, based on the analysis results of the evolution of classroom teaching events, during the dynamic segmentation of classroom audio and video, when the absolute value of the difference between the cognitive load indices corresponding to two adjacent classroom teaching events is greater than or equal to a preset value, and the time interval between the time boundary point between two adjacent classroom teaching events and the previous time segmentation point is greater than or equal to a preset duration, then the time boundary point between the corresponding two adjacent classroom teaching events is determined as a dynamic segmentation reserve point for the corresponding classroom audio and video; otherwise, it is determined that there is no dynamic segmentation point for the corresponding classroom audio and video between the corresponding two adjacent classroom teaching events. Based on the teaching themes corresponding to the second parameter in the structured features of teaching content at different times, each dynamic segmentation reserve point is calibrated to obtain classroom audio-visual segments between two adjacent dynamic segmentation reserve points, which are recorded as segments to be analyzed. The minimum absolute value of the difference in cognitive load index between each classroom teaching event within the previous preset duration in the segment to be analyzed and the previous classroom teaching event of the first dynamic segmentation storage point in the two adjacent dynamic segmentation reserve points is extracted and recorded as the cognitive load index deviation reference value. If the teaching themes of the two classroom teaching events corresponding to the cognitive load index deviation reference value, as well as the teaching themes of the two classroom teaching events corresponding to the first dynamic segmentation storage point in the two adjacent dynamic segmentation reserve points, are the same, then the first dynamic segmentation storage point in the two adjacent dynamic segmentation reserve points is deleted. Each calibrated dynamic segmentation reserve point is used as a dynamic segmentation point of the classroom audio-visual video to obtain each segment of the classroom audio-visual video. The diagnostic assessment report for each segment of classroom teaching audio and video includes each segment and its corresponding score. The calculation formulas involved in scoring the corresponding classroom audio and video segment are as follows: Wherein, GS represents the score of the corresponding classroom audio-visual segment; DCP represents the average interaction density of each classroom teaching event within the corresponding classroom audio-visual segment; CHP represents the average cognitive load index of each classroom teaching event within the corresponding classroom audio-visual segment; and φ1 represents the preset scoring factor.

[0014] This invention comprehensively filters dynamic segmentation reserve points for classroom audio and video based on the cognitive load index and the time interval between the time boundary point of two adjacent classroom teaching events and the previous time segmentation point; and based on the teaching theme corresponding to the second parameter in the structured features of teaching content, it realizes secondary calibration of each dynamic segmentation reserve point of the acquired classroom audio and video (taking into account the synergy of teaching themes), effectively eliminating useless dynamic segmentation reserve points at the teaching theme level, ensuring the integrity of classroom audio and video segmentation segments (reducing the frequency of the same teaching theme being divided into different segmentation segments and corresponding diagnostic assessment reports showing large differences).

[0015] An intelligent teaching evaluation and diagnosis system based on multimodal audio and video analysis, the system comprising: A multimodal audio and video data construction module, which constructs multimodal audio and video data by simultaneously collecting time-stamp-aligned video streams of teacher and student behavior, audio streams of speech, and content streams of teaching screens through a camera array, directional microphone group, and screen capture device deployed in the classroom; The teaching status feature extraction module performs content parsing on the constructed multimodal audio and video data and extracts the corresponding teaching status feature set based on the content parsing results. The audio and video dynamic segmentation module combines teaching status feature data from historical data to identify classroom teaching events based on the teaching status feature set extracted at the current time; and combines MTES technology to analyze the evolution process of classroom teaching events based on a time-series model. Based on the analysis results of the classroom teaching event evolution process, the module dynamically segments the classroom audio and video. The teaching evaluation and diagnosis management module generates a diagnostic evaluation report for each segment of classroom teaching audio and video based on the dynamic segmentation results of the classroom audio and video.

[0016] Furthermore, the audio and video dynamic segmentation module includes a teaching event recognition unit, an event evolution process analysis unit, and a segmentation point dynamic management unit; The teaching event identification unit combines teaching status feature data from historical data to identify classroom teaching events based on the teaching status feature set extracted at the current time. The event evolution process analysis unit combines MTES technology to analyze the evolution process of classroom teaching events based on a time-series model; The segmentation point dynamic management unit dynamically segments classroom audio and video based on the analysis results of the evolution process of classroom teaching events.

[0017] Compared with the prior art, the beneficial effects achieved by the present invention are: (1) The present invention can effectively solve the problem of spatiotemporal mismatch of corresponding data acquisition when cross-modal data is acquired through multimodal data synchronous acquisition; (2) The present invention is based on the comprehensive quantitative method of the indicator set bound to each classroom teaching event in the classroom teaching event evolution analysis array using MTES technology. Combined with the constructed multimodal audio and video data, it can effectively solve the modal fragmentation phenomenon between the modal data collected in the teaching evaluation and diagnosis process. (3) This invention filters and locks the dynamic segmentation points of classroom audio and video from multiple angles, and performs secondary calibration of the filtered and locked dynamic segmentation points through teaching themes to ensure the accuracy of classroom audio and video segmentation. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the intelligent teaching evaluation and diagnosis system based on multimodal audio and video analysis of the present invention; Figure 2 This is a flowchart illustrating the intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis of this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figures 1-2 The present invention provides a technical solution: such as Figure 1 As shown, this embodiment provides an intelligent teaching evaluation and diagnosis system based on multimodal audio and video analysis. The system includes: A multimodal audio and video data construction module, which constructs multimodal audio and video data by simultaneously collecting time-stamp-aligned video streams of teacher and student behavior, audio streams of speech, and content streams of teaching screens through a camera array, directional microphone group, and screen capture device deployed in the classroom; The teaching status feature extraction module performs content parsing on the constructed multimodal audio and video data and extracts the corresponding teaching status feature set based on the content parsing results. The audio and video dynamic segmentation module includes a teaching event recognition unit, an event evolution process analysis unit, and a segmentation point dynamic management unit. The teaching event identification unit combines teaching status feature data from historical data to identify classroom teaching events based on the teaching status feature set extracted at the current time. The event evolution process analysis unit combines MTES technology to analyze the evolution process of classroom teaching events based on a time-series model; The segmentation point dynamic management unit dynamically segments classroom audio and video based on the analysis results of the evolution process of classroom teaching events. The teaching evaluation and diagnosis management module generates a diagnostic evaluation report for each segment of classroom teaching audio and video based on the dynamic segmentation results of the classroom audio and video.

[0021] like Figure 2 As shown in the example, this paper provides an intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis. The method includes: S1. By deploying camera arrays, directional microphone groups and screen capture devices in the classroom, synchronously collect time-stamped video streams of teacher and student behavior, audio streams and teaching screen content streams to construct multimodal audio and video data; The camera array includes a panoramic camera covering the teacher's overall movements, a teacher tracking camera capturing the teacher's body language and blackboard writing, and a student expression camera collecting student facial response data.

[0022] This example uses a physics classroom in the second year of high school. The classroom is 60 square meters and there are 45 students. Two panoramic cameras, model DS-2CD3T86WD, are deployed diagonally. The teacher tracking camera is a PTZ camera (30x optical zoom). Six Logitech C930e cameras are used to capture student expressions. Each camera is placed in each row of student seats in the classroom. The directional microphone array uses Shure MXA910 (ceiling array, 8 nodes); The screen capture device is a computer with screen recording software installed; In this embodiment, the IEEE 1588 PTP protocol is used to realize the device clock synchronization, and the monitored synchronization error is less than or equal to ±0.1ms (the measured standard deviation σ=0.03ms).

[0023] S2. Perform content parsing on the constructed multimodal audio and video data, and extract the corresponding teaching status feature set based on the content parsing results; The teaching status feature set in S2 includes teacher and student behavior features based on video stream parsing, voice interaction features based on audio stream parsing, and teaching content structure features based on screen content parsing. The teacher-student behavioral characteristics are composed of teacher movement trajectory density and student head-up rate. The teacher movement trajectory density at time t is denoted as Dt. Where Nt represents the number of times the teacher's position is updated within a preset unit time interval before time t in the acquired video stream; T represents the preset sampling time window length; and A represents the classroom area. The student head-up rate at time t is equal to the ratio of the number of students in the captured video stream who are looking at the teacher / PPT playback screen at time t to the total number of students. In the process of acquiring speech interaction features based on audio stream parsing, the frame length of the acquired audio stream is obtained, where the audio frame length represents the number of audio sampling points contained in the corresponding audio frame. The audio frame corresponding to any acquisition time and the midpoint of each sampling point within the corresponding audio frame are acquired. The short-time energy of the speech corresponding to the corresponding audio frame is obtained, serving as the speech interaction feature of the acquisition time corresponding to the midpoint of each sampling point within the corresponding audio frame. The short-time energy of the speech corresponding to the nth frame is denoted as En. Where K represents the number of preset audio sampling points contained in the audio frame; This represents the square of the amplitude (instantaneous power) corresponding to the k-th sampling point within the n-th frame. In the process of obtaining the structured features of teaching content based on screen content parsing, the screen recording text is recognized using OCR technology. Preset keyword types corresponding to the teaching content at collection time t are extracted, along with the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords. The teaching themes corresponding to the preset keyword types extracted from the database's preset forms are queried. Based on the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords, a teaching theme bias coefficient at collection time t is calculated. This teaching theme bias coefficient at collection time t is equal to the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords multiplied by the sum of the weight factors of the corresponding preset keyword types based on the corresponding teaching themes in the database's preset forms. The obtained teaching theme bias coefficient and the corresponding teaching theme are used as the first and second parameters in the structured features of the teaching content, respectively.

[0024] S3. Combining teaching status feature data from historical data, identify classroom teaching events from the teaching status feature set extracted at the current time; and combine MTES technology to analyze the evolution process of classroom teaching events based on time series models. Based on the analysis results of the classroom teaching event evolution process, dynamically segment the classroom audio and video. In this embodiment, MTES stands for Multimodal Teaching Evaluation System. This system combines multiple modalities (such as text, audio, video, and images) to comprehensively evaluate and analyze the teaching process. Utilizing advanced technologies such as artificial intelligence, machine learning, and big data analytics, it captures, records, and evaluates various elements of teaching activities to provide more comprehensive, objective, and detailed teaching feedback. It can generate targeted classroom teaching evaluation reports for teachers of different levels, helping them improve their classroom teaching and promoting the transformation from knowledge-based classrooms to competency-based classrooms.

[0025] In the process of S3 identifying classroom teaching events from the teaching state feature set extracted at the current time, a preset teaching event sequence is obtained, and the event probability of each element in the preset teaching event sequence based on the teaching state feature data at collection time t is calculated. The element with the highest event probability based on the teaching state feature data at collection time t in the preset teaching event sequence is taken as the classroom teaching event identification result at collection time t. The calculation formulas involved are as follows: Where FGt represents the teaching state characteristic data collected at time t; P(E|FGt) represents the event probability of any element E in the preset teaching event sequence based on the teaching state characteristic data FGt collected at time t; g m Both gh and gh represent state feature functions, which are functions with base e and exponents calculated by subtracting the square of the difference between the average of all values ​​of the same teaching state feature type corresponding to the first parameter and the teaching state feature data corresponding to the second parameter in historical data. H represents a transition feature function. If the event state data pair formed by the first and second parameters in the transition feature function belongs to a preset effective transition set of event states, then the function value of the corresponding transition feature function is determined to be 1; otherwise, the function value of the corresponding transition feature function is determined to be 0. Each element in the effective transition set of event states corresponds to an event state data pair. The summary set of the first parameters in teacher movement trajectory density, student head-up rate, voice interaction features, and teaching content structured features in FGt is obtained and denoted as the first analysis set. m E represents the m-th parameter value in the first analysis set. t-1 The most recently identified teaching event is represented before the acquisition time t; W(FGt, m, E) represents the event adaptation evaluation coefficient of any element E in the preset teaching event sequence based on the teaching state feature data FGt at acquisition time t. Z(FGt) represents the sum of each W(FGt, m, E) corresponding to each element in the preset teaching event sequence; This represents the preset state feature weight corresponding to the m-th parameter in the first analysis set; The preset state feature weights corresponding to the teaching state evaluation coefficients are represented by μ; μ represents the transition feature weights, where the value of μ is equal to the ratio of the frequency of the event state data pair formed by the first and second parameters in the corresponding transition feature function appearing in the historical data to the sum of the frequencies of all event state data pairs in the historical data that have the same first parameter and the first parameter of the corresponding transition feature function; max{} represents the maximum value function. Where Ft represents the teaching status evaluation coefficient at acquisition time t; Lt represents the student head-up rate at acquisition time t; E(t) represents the short-time energy of the corresponding audio frame at acquisition time t; Xt represents the first parameter in the structured features of the teaching content at acquisition time t; r1, r2 and r3 are the preset first evaluation weight, second evaluation weight and third evaluation weight, respectively, and r1+r2+r3=1.

[0026] In the process of analyzing the evolution of classroom teaching events based on a time-series model using MTES technology, an array consisting of the identification results of each classroom teaching event is obtained sequentially according to time, denoted as the classroom teaching event evolution analysis array. In this array, before a new classroom teaching event is identified, the original classroom teaching identification results remain unchanged, and each time point corresponds to a unique classroom teaching event identification result. Based on MTES technology, a set of indicators bound to each classroom teaching event in the classroom teaching time evolution analysis array is calculated. This set of indicators includes interaction density and cognitive load index. The calculation method for interaction density is as follows: Where QN represents the number of valid questions asked by the teacher within the time period corresponding to the corresponding classroom teaching event identification result; TS represents the interval length of the time period corresponding to the corresponding classroom teaching event identification result; NSUM represents the sum of the number of students who participated in each teacher questioning interaction within the time period corresponding to the corresponding classroom teaching event identification result; and NA represents the total number of students in the classroom. The cognitive load index is calculated as follows: Where CH represents the cognitive load index corresponding to the identification results of the corresponding classroom teaching event; YP tbtb represents the average facial confusion level of each student at any time point tb within the time period corresponding to the corresponding classroom teaching event recognition result; the facial confusion level of a student is equal to the quotient of the number of wrinkles between the eyebrows obtained through image recognition divided by the area of ​​the area between the eyebrows; tb1 represents the minimum time point within the time period corresponding to the corresponding classroom teaching event recognition result; VS tb The VS represents the speech rate change coefficient at any time point tb within the time period corresponding to the classroom teaching event recognition result. tb The value is equal to the difference between the most recent monitored speech rate of the audio stream after time point tb in the historical data and the most recent monitored speech rate of the audio stream before time point tb in the historical data, divided by the quotient of the interval between two adjacent speech rate monitoring.

[0027] Based on the analysis results of the evolution of classroom teaching events, during the dynamic segmentation of classroom audio and video, when the absolute value of the difference between the cognitive load indices corresponding to two adjacent classroom teaching events is greater than or equal to a preset value, and the time interval between the time boundary point between two adjacent classroom teaching events and the previous time segmentation point is greater than or equal to a preset time, then the time boundary point between the corresponding two adjacent classroom teaching events is determined as a dynamic segmentation reserve point for the corresponding classroom audio and video; otherwise, it is determined that there is no dynamic segmentation point for the corresponding classroom audio and video between the corresponding two adjacent classroom teaching events. Based on the teaching themes corresponding to the second parameter in the structured features of teaching content at different times, each dynamic segmentation reserve point is calibrated to obtain classroom audio-visual segments between two adjacent dynamic segmentation reserve points, which are recorded as segments to be analyzed. The minimum absolute value of the difference in cognitive load index between each classroom teaching event within the previous preset duration in the segment to be analyzed and the previous classroom teaching event of the first dynamic segmentation storage point in the two adjacent dynamic segmentation reserve points is extracted and recorded as the cognitive load index deviation reference value. If the teaching themes of the two classroom teaching events corresponding to the cognitive load index deviation reference value, as well as the teaching themes of the two classroom teaching events corresponding to the first dynamic segmentation storage point in the two adjacent dynamic segmentation reserve points, are the same, then the first dynamic segmentation storage point in the two adjacent dynamic segmentation reserve points is deleted. Each calibrated dynamic segmentation reserve point is used as a dynamic segmentation point of the classroom audio-visual video to obtain each segment of the classroom audio-visual video. S4. Based on the dynamic segmentation results of classroom audio and video, generate a diagnostic evaluation report corresponding to each segmented segment of classroom teaching audio and video.

[0028] The diagnostic assessment report for each segment of classroom teaching audio and video includes each segment and its corresponding score. The calculation formulas involved in scoring the corresponding classroom audio and video segment are as follows: Wherein, GS represents the score of the corresponding classroom audio-visual segment; DCP represents the average interaction density of each classroom teaching event within the corresponding classroom audio-visual segment; CHP represents the average cognitive load index of each classroom teaching event within the corresponding classroom audio-visual segment; and φ1 represents the preset scoring factor.

[0029] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0030] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis, characterized in that, The method includes: S1. By deploying camera arrays, directional microphone groups and screen capture devices in the classroom, synchronously collect time-stamped video streams of teacher and student behavior, audio streams and teaching screen content streams to construct multimodal audio and video data; S2. Perform content parsing on the constructed multimodal audio and video data, and extract the corresponding teaching status feature set based on the content parsing results; S3. Combining teaching status feature data from historical data, identify classroom teaching events from the teaching status feature set extracted at the current time; and combine MTES technology to analyze the evolution process of classroom teaching events based on time series models. Based on the analysis results of the classroom teaching event evolution process, dynamically segment the classroom audio and video. S4. Based on the dynamic segmentation results of classroom audio and video, generate a diagnostic evaluation report corresponding to each segmented segment of classroom teaching audio and video.

2. The intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis according to claim 1, characterized in that: The camera array includes a panoramic camera covering the teacher's overall movements, a teacher tracking camera capturing the teacher's body language and blackboard writing, and a student expression camera collecting student facial response data.

3. The intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis according to claim 1, characterized in that: The teaching status feature set in S2 includes teacher and student behavior features based on video stream parsing, voice interaction features based on audio stream parsing, and teaching content structure features based on screen content parsing. The teacher-student behavioral characteristics are composed of teacher movement trajectory density and student head-up rate. The teacher movement trajectory density at time t is denoted as Dt. Where Nt represents the number of times the teacher's position is updated within a preset unit time interval before time t in the acquired video stream; T represents the preset sampling time window length; and A represents the classroom area. The student head-up rate at time t is equal to the ratio of the number of students in the captured video stream who are looking at the teacher / PPT playback screen at time t to the total number of students. In the process of acquiring speech interaction features based on audio stream parsing, the frame length of the acquired audio stream is obtained, where the audio frame length represents the number of audio sampling points contained in the corresponding audio frame. The audio frame corresponding to any acquisition time and the midpoint of each sampling point within the corresponding audio frame are acquired. The short-time energy of the speech corresponding to the corresponding audio frame is obtained, serving as the speech interaction feature of the acquisition time corresponding to the midpoint of each sampling point within the corresponding audio frame. The short-time energy of the speech corresponding to the nth frame is denoted as En. Where K represents the number of preset audio sampling points contained in the audio frame; This represents the squared amplitude corresponding to the k-th sampling point within the n-th frame; In the process of obtaining the structured features of teaching content based on screen content parsing, the screen recording text is recognized using OCR technology. Preset keyword types corresponding to the teaching content at collection time t are extracted, along with the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords. The teaching themes corresponding to the preset keyword types extracted from the database's preset forms are queried. Based on the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords, a teaching theme bias coefficient at collection time t is calculated. This teaching theme bias coefficient at collection time t is equal to the ratio of the frequency of each preset keyword to the total frequency of extracted preset keywords multiplied by the sum of the weight factors of the corresponding preset keyword types based on the corresponding teaching themes in the database's preset forms. The obtained teaching theme bias coefficient and the corresponding teaching theme are used as the first and second parameters in the structured features of the teaching content, respectively.

4. The intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis according to claim 3, characterized in that: In the process of S3 identifying classroom teaching events from the teaching state feature set extracted at the current time, a preset teaching event sequence is obtained, and the event probability of each element in the preset teaching event sequence based on the teaching state feature data at collection time t is calculated. The element with the highest event probability based on the teaching state feature data at collection time t in the preset teaching event sequence is taken as the classroom teaching event identification result at collection time t. The calculation formulas involved are as follows: Where FGt represents the teaching state characteristic data collected at time t; P(E|FGt) represents the event probability of any element E in the preset teaching event sequence based on the teaching state characteristic data FGt collected at time t; g m Both gh and gh represent state feature functions, which are functions with base e and exponents calculated by subtracting the square of the difference between the average of all values ​​of the same teaching state feature type corresponding to the first parameter and the teaching state feature data corresponding to the second parameter in historical data. H represents a transition feature function. If the event state data pair formed by the first and second parameters in the transition feature function belongs to a preset effective transition set of event states, then the function value of the corresponding transition feature function is determined to be 1; otherwise, the function value of the corresponding transition feature function is determined to be 0. Each element in the effective transition set of event states corresponds to an event state data pair. The summary set of the first parameters in teacher movement trajectory density, student head-up rate, voice interaction features, and teaching content structured features in FGt is obtained and denoted as the first analysis set. m E represents the m-th parameter value in the first analysis set. t-1 The most recently identified teaching event is represented before the acquisition time t; W(FGt, m, E) represents the event adaptation evaluation coefficient of any element E in the preset teaching event sequence based on the teaching state feature data FGt at acquisition time t. Z(FGt) represents the sum of each W(FGt, m, E) corresponding to each element in the preset teaching event sequence; This represents the preset state feature weight corresponding to the m-th parameter in the first analysis set; The preset state feature weights corresponding to the teaching state evaluation coefficients are represented by μ; μ represents the transition feature weights, where the value of μ is equal to the ratio of the frequency of the event state data pair formed by the first and second parameters in the corresponding transition feature function appearing in the historical data to the sum of the frequencies of all event state data pairs in the historical data that have the same first parameter and the first parameter of the corresponding transition feature function; max{} represents the maximum value function. Where Ft represents the teaching status evaluation coefficient at acquisition time t; Lt represents the student head-up rate at acquisition time t; E(t) represents the short-time energy of the corresponding audio frame at acquisition time t; Xt represents the first parameter in the structured features of the teaching content at acquisition time t; r1, r2 and r3 are the preset first evaluation weight, second evaluation weight and third evaluation weight, respectively, and r1+r2+r3=1.

5. The intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis according to claim 1, characterized in that: In the process of analyzing the evolution of classroom teaching events based on a time-series model using MTES technology, an array consisting of the identification results of each classroom teaching event is obtained sequentially according to time, denoted as the classroom teaching event evolution analysis array. In this array, before a new classroom teaching event is identified, the original classroom teaching identification results remain unchanged, and each time point corresponds to a unique classroom teaching event identification result. Based on MTES technology, a set of indicators bound to each classroom teaching event in the classroom teaching time evolution analysis array is calculated. This set of indicators includes interaction density and cognitive load index. The calculation method for interaction density is as follows: Where QN represents the number of valid questions asked by the teacher within the time period corresponding to the corresponding classroom teaching event identification result; TS represents the interval length of the time period corresponding to the corresponding classroom teaching event identification result; NSUM represents the sum of the number of students who participated in each teacher questioning interaction within the time period corresponding to the corresponding classroom teaching event identification result; and NA represents the total number of students in the classroom. The cognitive load index is calculated as follows: Where CH represents the cognitive load index corresponding to the identification results of the corresponding classroom teaching event; YP tb tb represents the average facial confusion level of each student at any time point tb within the time period corresponding to the corresponding classroom teaching event recognition result; the facial confusion level of a student is equal to the quotient of the number of wrinkles between the eyebrows obtained through image recognition divided by the area of ​​the area between the eyebrows; tb1 represents the minimum time point within the time period corresponding to the corresponding classroom teaching event recognition result; VS tb The VS represents the speech rate change coefficient at any time point tb within the time period corresponding to the classroom teaching event recognition result. tb The value is equal to the difference between the most recent monitored speech rate of the audio stream after time point tb in the historical data and the most recent monitored speech rate of the audio stream before time point tb in the historical data, divided by the quotient of the interval between two adjacent speech rate monitoring.

6. The intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis according to claim 5, characterized in that: Based on the analysis results of the evolution of classroom teaching events, during the dynamic segmentation of classroom audio and video, when the absolute value of the difference between the cognitive load indices corresponding to two adjacent classroom teaching events is greater than or equal to the preset value and the time interval between the time boundary point between two adjacent classroom teaching events and the previous time segmentation point is greater than or equal to the preset time, the time boundary point between the corresponding two adjacent classroom teaching events is determined as a dynamic segmentation reserve point of the corresponding classroom audio and video. Conversely, it is determined that there is no dynamic segmentation point of corresponding classroom audio and video between two adjacent classroom teaching events; Based on the teaching theme corresponding to the second parameter in the structured features of teaching content at different times, each dynamic segmentation reserve point is calibrated, and the classroom audio and video segments between two adjacent dynamic segmentation reserve points are obtained and recorded as the segments to be analyzed. The minimum absolute value of the difference between the cognitive load index between each classroom teaching event corresponding to the previous classroom teaching event of the first dynamic segmentation storage point in the segments to be analyzed within the previous preset time period is extracted and recorded as the cognitive load index deviation reference value. If the teaching topics of the two classroom teaching events corresponding to the cognitive load index deviation reference value, and the teaching topics of the two classroom teaching events corresponding to the first dynamic segmentation storage point in two adjacent dynamic segmentation storage points are the same, then delete the first dynamic segmentation storage point in two adjacent dynamic segmentation storage points; use each calibrated dynamic segmentation storage point as a dynamic segmentation point of the classroom audio and video to obtain each segment of the classroom audio and video. The diagnostic assessment report for each segment of classroom teaching audio and video includes each segment and its corresponding score. The calculation formulas involved in scoring the corresponding classroom audio and video segment are as follows: Wherein, GS represents the score of the corresponding classroom audio-visual segment; DCP represents the average interaction density of each classroom teaching event within the corresponding classroom audio-visual segment; CHP represents the average cognitive load index of each classroom teaching event within the corresponding classroom audio-visual segment; and φ1 represents the preset scoring factor.

7. An intelligent teaching evaluation and diagnosis system based on multimodal audio and video analysis, employing the intelligent teaching evaluation and diagnosis method based on multimodal audio and video analysis as described in any one of claims 1-6, characterized in that, The system includes: A multimodal audio and video data construction module, which constructs multimodal audio and video data by simultaneously collecting time-stamp-aligned video streams of teacher and student behavior, audio streams of speech, and content streams of teaching screens through a camera array, directional microphone group, and screen capture device deployed in the classroom; The teaching status feature extraction module performs content parsing on the constructed multimodal audio and video data and extracts the corresponding teaching status feature set based on the content parsing results. The audio and video dynamic segmentation module combines teaching status feature data from historical data to identify classroom teaching events based on the teaching status feature set extracted at the current time; and combines MTES technology to analyze the evolution process of classroom teaching events based on a time-series model. Based on the analysis results of the classroom teaching event evolution process, the module dynamically segments the classroom audio and video. The teaching evaluation and diagnosis management module generates a diagnostic evaluation report for each segment of classroom teaching audio and video based on the dynamic segmentation results of the classroom audio and video.

8. The intelligent teaching evaluation and diagnosis system based on multimodal audio and video analysis according to claim 7, characterized in that: The audio and video dynamic segmentation module includes a teaching event recognition unit, an event evolution process analysis unit, and a segmentation point dynamic management unit. The teaching event identification unit combines teaching status feature data from historical data to identify classroom teaching events based on the teaching status feature set extracted at the current time. The event evolution process analysis unit combines MTES technology to analyze the evolution process of classroom teaching events based on a time-series model; The segmentation point dynamic management unit dynamically segments classroom audio and video based on the analysis results of the evolution process of classroom teaching events.

Citation Information

Patent Citations

  • Teaching state evaluation method and device and electronic equipment

    CN112308746A

  • Classroom student expression recognition and classroom state evaluation method and device

    CN113239914A

  • Teaching management quality evaluation system based on big data and Al technology

    CN114493129A

  • Classroom type intelligent evaluation system based on teaching video

    CN115170064A

  • Teaching behavior analysis system for teaching feature fusion and modeling based on knowledge base

    CN115239527A

Cited By

  • Multi-modal data real-time analysis and feedback method and system

    CN122045936A