Teaching behavior data classification method and device based on multiple modes, equipment and medium

By collecting video and audio information in the classroom, performing dynamic semantic segmentation and action analysis, and combining it with knowledge graph weight allocation, the teaching behavior classification is achieved on edge devices using knowledge distillation. This solves the problems of computational complexity and recognition accuracy in multimodal data processing, and improves the real-time performance and accuracy of teaching behavior analysis.

CN120997007APending Publication Date: 2025-11-21BEIJING POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511115100.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing methods for analyzing classroom teaching behavior suffer from problems such as missing information, insufficient recognition accuracy, and high computational complexity when processing multimodal data, making it difficult to achieve real-time classification on edge devices.

Method used

By collecting classroom video and audio information, performing preprocessing, dynamic semantic segmentation and action analysis, and combining the knowledge graph of teaching links to allocate cross-modal attention weights, the features are transferred to edge devices for teaching behavior classification using knowledge distillation.

Benefits of technology

It improves the accuracy and real-time performance of teaching behavior recognition, reduces computational complexity, is suitable for edge devices with limited computing power, and supports students to review and repeatedly study classroom content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997007A_ABST
    Figure CN120997007A_ABST
Patent Text Reader

Abstract

The invention discloses a teaching behavior data classification method and device based on multiple modes, equipment and a medium, and relates to the field of teaching behavior analysis, and the method comprises the steps: collecting video and voice information, and carrying out the preprocessing; dynamically cutting the voice data, and identifying an explanation statement behavior; performing action analysis on the video data, and identifying a display behavior and a guide behavior; questioning behaviors are recognized by recognizing doubt expression features and actions pointing to students through gestures and detecting doubt voices and intonations; carrying out cross-modal attention weight distribution by utilizing the knowledge graph and fusing to form a multi-modal fusion feature; and migrating to edge equipment by utilizing knowledge distillation and outputting a teaching behavior classification result. Voice segments with the same or similar semantics can be prevented from being segmented through dynamic segmentation; questioning behavior recognition is performed in combination with video information and voice information, so that the teaching behavior recognition accuracy is improved; improving the weights of different data sources according to teaching links by using the knowledge graph; and through knowledge distillation, the edge equipment can also meet the computing power requirement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of teaching behavior analysis, and in particular to a teaching behavior data classification method and device based on multi-modal, equipment and medium. BACKGROUND

[0002] When a teacher teaches in a classroom, there are various teaching behaviors, such as language behavior of making statements or explanations through language, display behavior of expressing through body movements or facial expressions, questioning behavior of interacting with students, and guiding behavior of guiding and discussing students' difficult problems. By analyzing the teaching behavior of the teacher in the classroom, teaching records can be generated and saved, which helps students to understand, trace, query and understand the teaching content in real time.

[0003] However, the existing classroom teaching behavior analysis often needs to rely on intelligent models, and the teaching behavior has great differences among the teaching habits of teachers. When recording the video data of the classroom teaching process and using intelligent models for processing, the single modal data often has information missing due to the rapid movement of the teacher, the movement of the students, the change of light, etc., which affects the recognition accuracy of the teaching behavior and even the board writing information, and the great differences in scientific knowledge and teaching habits among different subjects and different teachers lead to different forms of different behaviors, which easily leads to the lack of generalization ability and robustness of the recognition model. If multi-source data such as video, voice and board writing are used, the multi-modal data will be highly heterogeneous in the time and space dimensions, leading to exponential expansion of the parameters of the deep learning network, which is difficult to deploy on edge devices and other devices with limited computing performance, and cannot realize real-time teaching behavior classification, and the importance and influence of multi-source data on different teaching behaviors are also different, if the differences between multi-source data cannot be accurately used to distinguish the influence degree of different characteristics, it will also cause the problem of insufficient teaching behavior analysis accuracy. SUMMARY

[0004] The embodiments of the present application provide a teaching behavior data classification method, device, equipment and medium based on multi-modal, to solve the technical problem that it is difficult to accurately classify teaching behavior using multi-source data.

[0005] In a first aspect, the embodiments of the present application provide a teaching behavior data classification method based on multi-modal, comprising:

[0006] S101, collecting video information and voice information in the classroom and performing corresponding preprocessing to form teaching video data and teaching voice data;

[0007] S102, dynamically cutting the teaching voice data to form a dynamic semantic event unit, generating a dynamic semantic feature according to the dynamic semantic event unit, and using the dynamic semantic feature to identify the explanation and statement behavior; S103, extracting the board writing information from the teaching video data, and generating a board writing feature according to the board writing information, and using the board writing feature to identify the display behavior; S104, extracting the body movement information from the teaching video data, and generating a body movement feature according to the body movement information, and using the body movement feature to identify the display behavior; S105, extracting the facial expression information from the teaching video data, and generating a facial expression feature according to the facial expression information, and using the facial expression feature to identify the display behavior; S106, extracting the voice information from the teaching video data, and generating a voice feature according to the voice information, and using the voice feature to identify the display behavior; S107, extracting the student movement information from the teaching video data, and generating a student movement feature according to the student movement information, and using the student movement feature to identify the display behavior; S108, extracting the student voice information from the teaching video data, and generating a student voice feature according to the student voice information, and using the student voice feature to identify the display behavior; S109, extracting the student facial expression information from the teaching video data, and generating a student facial expression feature according to the student facial expression information, and using the student facial expression feature to identify the display behavior; S110, extracting the student body movement information from the teaching video data, and generating a student body movement feature according to the student body movement information, and using the student body movement feature to identify the display behavior; S111, extracting the student question information from the teaching voice data, and generating a student question feature according to the student question information, and using the student question feature to identify the questioning behavior; S112, extracting the teacher question information from the teaching voice data, and generating a teacher question feature according to the teacher question information, and using the teacher question feature to identify the questioning behavior; S113, extracting the teacher guiding information from the teaching voice data, and generating a teacher guiding feature according to the teacher guiding information, and using the teacher guiding feature to identify the guiding behavior; S114, extracting the student guiding information from the teaching voice data, and generating a student guiding feature according to the student guiding information, and using the student guiding feature to identify the guiding behavior.

[0008] S103, performing action analysis on the teaching video data, forming a teacher action feature according to a teacher action in the video stream, and used for identifying the demonstration behavior and the guidance behavior;

[0009] S104, performing teacher facial expression analysis and action analysis on the teaching video data, respectively identifying a teacher facial doubt expression and a gesture pointing to a student action, forming a facial doubt feature and a gesture pointing feature, performing voice intonation detection on a dynamic semantic event unit, identifying a teacher question sentence intonation to form a question intonation feature, and according to the facial doubt feature, the gesture pointing feature and the question intonation feature, identifying the questioning behavior;

[0010] S105, determining a teaching link using a teaching link knowledge graph, respectively assigning cross-modal attention weights to the dynamic semantic feature, the teacher action feature, the facial doubt feature, the gesture pointing feature and the question intonation feature according to the teaching link, and performing feature fusion according to the different weights to form a multi-modal fusion feature;

[0011] S106, migrating the multi-modal fusion feature to an edge device using knowledge distillation, and outputting a teaching behavior classification result on the edge device.

[0012] Further, the method further comprises:

[0013] When the output probabilities of the explanation statement behavior, the demonstration behavior, the guidance behavior and the questioning behavior are all lower than 0.3, and a pen trace is detected in the teaching video data, a visual recognition model is used to track the teacher's hand action in the teaching video data, to perform teacher pen trace recognition, and at the same time, the blackboard writing in the teaching video data is recognized, according to the pen trace recognition result and the text recognition result, a prediction model is used to complete the occluded part, and teaching blackboard writing data is generated.

[0014] Further, the S102 further comprises:

[0015] Based on the voice silence detection and the question sentence intonation recognition, the teaching voice data is dynamically cut to form a dynamic semantic event unit;

[0016] The dynamic semantic event unit is subjected to acoustic feature and text keyword joint recognition to generate a dynamic semantic feature, which is used for identifying the explanation statement behavior.

[0017] Further, the S103 comprises:

[0018] According to the teaching video data, a teaching aid is identified using a space-time attention, and a teaching aid motion feature is generated according to a motion track of the identified teaching aid, which is used for identifying the demonstration behavior;

[0019] According to the teaching video data, teacher bone key points are generated, and according to the continuous spatial association of the teacher bone key points and the student region, bone posture spatial features are generated, which are used for identifying the guiding behavior.

[0020] Further, the S106 includes:

[0021] The multi-modal fusion features are subjected to Logits distillation, feature distillation and structure distillation in sequence respectively to generate distilled multi-modal features.

[0022] The distilled multi-modal features are migrated to an edge device, the distilled multi-modal features are classified by the edge device to output a classification result.

[0023] Further, the S101 includes:

[0024] Video information and voice information on the classroom are collected;

[0025] The video information is subjected to spatial denoising and information enhancement processing to form teaching video data, and the voice information is subjected to environmental noise reduction, echo cancellation and voice separation processing to form teaching voice data.

[0026] Further, the collection of the video information and the voice information on the classroom includes:

[0027] The video information and the voice information of the teacher on the classroom are collected through a synchronous clock source and are subjected to timestamp alignment to generate time-aligned video information and voice information.

[0028] In a second aspect, an embodiment of the present application provides a multi-modal based teaching behavior data classification device, which includes:

[0029] An information collection module is configured to collect video information and voice information on a classroom and perform corresponding preprocessing to form teaching video data and teaching voice data;

[0030] A voice data recognition module is configured to perform dynamic cutting on the teaching voice data to form dynamic semantic event units, generate dynamic semantic features according to the dynamic semantic event units, and identify explanation and statement behaviors;

[0031] A video action analysis module is configured to perform action analysis on the teaching video data, form teacher action features according to teacher actions in a video stream, and identify display behaviors and guiding behaviors;

[0032] The question behavior recognition module is configured to analyze the teacher facial expression and the action of the teaching video data, recognize the teacher facial doubt expression and the gesture pointing student action respectively, form the facial doubt feature and the gesture pointing feature, detect the voice tone of the dynamic semantic event unit, recognize the teacher question sentence tone to form the question tone feature, and recognize the question behavior according to the facial doubt feature, the gesture pointing feature and the question tone feature.

[0033] The multi-modal attention feature fusion module is configured to determine the teaching link by using the teaching link knowledge graph, respectively distribute the cross-modal attention weights of the dynamic semantic feature, the teacher action feature, the facial doubt feature, the gesture pointing feature and the question tone feature according to the teaching link, fuse the features according to the different weights, and form the multi-modal fusion feature.

[0034] The teaching behavior classification module is configured to migrate the multi-modal fusion feature to the edge device by using the knowledge distillation, and output the teaching behavior classification result on the edge device.

[0035] In a third aspect, an electronic device is provided, including:

[0036] one or more processors;

[0037] a storage device configured to store one or more programs,

[0038] When the one or more programs are executed by the one or more processors, the one or more processors implement the above multi-modal based teaching behavior data classification method.

[0039] In a fourth aspect, a storage medium containing computer executable instructions is provided, which, when executed by a computer processor, is used to execute the above multi-modal based teaching behavior data classification method.

[0040] The embodiment of the present application provides a kind of based on multi-modal teaching behavior data classification method, device, equipment and medium, the method is by collecting the video data and speech data of teacher in classroom, by dynamic semantic segmentation to speech data, form the dynamic semantic features for identifying explanation statement behavior in speech stream;By action analysis and facial expression analysis to video data, form the teacher action features for identifying display behavior and guiding behavior in video stream, and combine the tone of voice detected in dynamic semantic analysis Question tone and the puzzled expression identified by facial expression analysis And the action of gesture pointing student space identified by action analysis, form the facial puzzled features, gesture pointing features and question tone features for identifying teacher questioning behavior;Again, the teaching link knowledge graph is used to determine the teaching link, and the different teaching behaviors in the teaching link are correspondingly weighted and fused into multi-modal fusion features, and finally the fused features are migrated to the edge device using knowledge distillation, so that the limited computing capacity of the edge device can also generate teaching behavior classification results in real time. By dynamic semantic segmentation, the same or similar semantic speech segments are avoided to be segmented, the semantic segmentation accuracy is improved, and the speech recognition accuracy is improved, so that the explanation statement behavior in speech stream can be more accurately identified;Questioning behavior can be identified by multi-source data joint, video information and speech information can be combined, and the identification accuracy of questioning behavior is improved;Different weights are allocated by using knowledge graph to distinguish teaching links, which can improve the weight of different data sources in different behavior identification, and is more conducive to the performance of data features, and further improves the accuracy of teaching behavior identification and classification;Through knowledge distillation, the device with limited computing power can also obtain teaching behavior classification results in real time, and classroom information can be recorded by occupying small storage space, which is especially convenient for students to review and repeatedly learn classroom content, and improves the understanding and mastery of students to teaching content. BRIEF DESCRIPTION OF DRAWINGS

[0041] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of the present application and their

[0042] Figure 1 A flowchart of a multi-modal based teaching behavior data classification method according to the first embodiment of the present application;

[0043] Figure 2 A flowchart of a multi-modal based teaching behavior data classification method according to the second embodiment of the present application;

[0044] Figure 3 A flowchart of a multi-modal based teaching behavior data classification method according to the third embodiment of the present application;

[0045] Figure 4A structural schematic diagram of a multi-modal based teaching behavior data classification device according to Embodiment Four of the present application;

[0046] Figure 5 A structural diagram of an electronic device according to Embodiment Five of the present application. DETAILED DESCRIPTION

[0047] The present application will be further described below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely intended for the purpose of interpretation of the present application and are not limiting of the present application. In addition, it should be noted that, for the purpose of description, only the parts related to the present application are shown in the accompanying drawings rather than all the structures.

[0048] When a teacher teaches in a classroom, the teacher will impart knowledge to students through language explanation, demonstration of teaching aids, action guidance, and interactive questioning. However, in a university classroom, especially for subjects with high professional degree, it is often difficult to master complex scientific knowledge at one time through classroom learning, and it is necessary to review, rewatch, or repeatedly review after class to consolidate the mastery of scientific knowledge. Even some difficult scientific knowledge may be difficult to understand when it is taught in the classroom. It is more conducive to rapid learning of the scientific knowledge taught by using an intelligent model to analyze the teaching behavior of the teacher and providing corresponding prompts in real time on the student terminal in the classroom. By recognizing the video stream, the teaching behavior of the teacher can be recognized through the actions and blackboard writing of the teacher. By recognizing the speech stream, the behaviors such as statement, explanation, guidance, and questioning of the teacher can be recognized through the language of the teacher. However, single-mode behavior recognition has limitations in scene and recognition ability. If multi-source data is combined for behavior recognition, although the behavior analysis accuracy can be improved to a certain extent, the computational amount and computational complexity of the recognition model will be greatly improved. The devices held by the students are mostly edge devices with limited computing power, which are difficult to meet the complex computing needs of the intelligent recognition model. The feature importance of different data forms in multi-modal data under different teaching behaviors also has certain differences. Due to the diversity of teaching behaviors and the differences in teaching habits among teachers, the adaptability and generalization ability of the model for multi-source data or multi-modal recognition are poor, and it is difficult to achieve high-precision teaching behavior analysis.

[0049] Embodiment One

[0050] Figure 1 A flowchart of a multi-modal based teaching behavior data classification method according to Embodiment One of the present application. In this embodiment, the video information and speech information collected in the classroom are jointly analyzed, the importance of different data sources to different teaching behaviors is utilized, different weights are assigned when different types of teaching behaviors are recognized, the multi-modal data are fused, and the edge device can also meet the computing needs of generating the teaching behavior classification result through knowledge distillation. The specific steps include the following steps:

[0051] S101, collect video information and voice information on the classroom and perform corresponding preprocessing to form teaching video data and teaching voice data.

[0052] The video information of the teacher on the classroom is collected by a camera, and the information collected by the camera is preprocessed, such as motion blur compensation, light adaptation, background filtering, etc. to repair the blur caused by the teacher's fast writing or gesture action, eliminate the reflection interference of blackboard, projector and other equipment, and reduce the interference of student movement in the background on the analysis of the teacher's behavior, to form teaching video data; the voice information of the teacher on the classroom is collected by a microphone, and the collected voice information is preprocessed, such as noise reduction, echo cancellation, and human separation, to eliminate environmental noise, microphone howling and other interference elements, and reduce the influence of student voice on teacher voice, to form teaching voice data. Exemplarily, the camera is set at a perspective facing the blackboard, so that the camera can collect the entire blackboard and the range of the podium, and the behavior of the teacher is collected by the camera. The microphone can be set on the podium or the blackboard, or the voice can be collected by the teacher wearing the microphone. In order to facilitate data transmission, the microphone worn by the teacher can be in a passive form, and the voice data can be transmitted in a wireless form such as Bluetooth or WiFi.

[0053] S102, dynamically cutting the teaching voice data to form dynamic semantic event units, generating dynamic semantic features from the dynamic semantic event units for identifying the explanation and statement behavior.

[0054] According to the semantic events, the teaching voice data is dynamically cut, and the voice stream is cut into dynamic semantic event units of different sizes according to different semantics. The current semantic state can be identified according to the interruption, pause, tone, etc. of the language, and the voice paragraphs of different lengths and different semantics are dynamically cut according to the semantic state, and each semantic unit has a relatively complete and independent semantic, and the dynamic semantic features can be generated according to the meaning of each semantic unit to express the language behavior of the teacher in different semantic paragraphs, and the teaching behavior of the teacher can be determined, and the explanation and statement behavior of the teacher to students can be identified.

[0055] S103, action analysis on the teaching video data to form teacher action features from the teacher actions in the video stream for identifying the demonstration behavior and the guidance behavior.

[0056] The body posture of the teacher in the teaching video data is used to identify the action formed by the teacher to form a teacher action feature, and the teacher action feature can be used to analyze the teaching behavior being performed by the teacher. In some professional classrooms, in order to facilitate students to understand professional scientific knowledge, the teacher often uses some entity teaching aids to demonstrate, uses the entity teaching aids to tell the students scientific knowledge, captures and analyzes the moving condition of the teaching aids, and combines the teacher action feature to identify the demonstration behavior of the teacher using the teaching aids to teach scientific knowledge. In some classrooms, the teacher will also arrange practical operation tasks to improve the students' mastery of scientific knowledge. When the students perform the practical operation tasks, the teacher will observe the task execution process of the students and guide the execution process. According to the teacher action feature and the space interaction of one side of the student, the guiding behavior of the teacher guiding the practical operation task of the student can be identified.

[0057] In S104, the teacher facial expression analysis and action analysis are performed on the teaching video data, the teacher facial doubt expression and gesture pointing student action are identified respectively, the facial doubt feature and gesture pointing feature are formed, the voice tone detection is performed on the dynamic semantic event unit, the teacher question sentence tone is identified to form the question tone feature, and the facial doubt feature, the gesture pointing feature and the question tone feature are used to identify the questioning behavior.

[0058] In the classroom, the teacher will also ask questions according to the scientific achievements made by the students, the functional tasks achieved, etc. Through the students' answers, the specific quality and effect of the achievements or the implementation of the functions are distinguished, and the questioning behavior is made. The teacher may also ask questions to the students according to the students' statements of the knowledge they have mastered, to understand the students' mastery of knowledge or whether the understanding is correct. When the teacher makes the questioning behavior, it is mainly to understand the situation of the students and express their own doubts, so there will be a doubtful expression on the face, a gesture of pointing to the side of the student, and a sentence with a questioning tone in the language expression. Therefore, when identifying the questioning behavior, not only the questioning sentence in the language should be paid attention to, but also the facial expression and action behavior of the teacher should be paid attention to, so as to improve the identification accuracy of the questioning behavior. Due to the differences in teaching habits, some teachers will use the form of continuous questioning to convey information to students, but in this case, the teacher's face will not appear a doubtful expression, so by combining the teacher's facial expression and the space interaction with the student side, the questioning behavior can be more accurately identified. For example, the facial doubtful expression of the teacher is identified by using a facial recognition model (MobileViT model, classifying the expression into neutral, smile, doubt, surprise, serious, and other 6 categories, and outputting the doubtful expression probability by Sigmoid activation function), forming a facial doubtful feature. The teacher's body movement is identified by using a visual recognition model (YOLOv8 model, capturing the key points of the teacher's hand), and the interactive gesture or body movement of the teacher with the student side is distinguished, forming a gesture pointing feature. The questioning tone in the teacher's language is identified by voice tone detection, forming a questioning tone feature. Then, a decision tree model is used to determine the questioning behavior according to the three features. The weights in the decision tree can adopt the strategy of the weight of the questioning tone +0.6, the weight of the facial doubtful expression +0.3, and the weight of the gesture pointing to the student +0.4. Finally, whether it is a questioning behavior is confirmed according to the determination threshold (≥0.8), and whether the behavior made by the teacher is a questioning behavior can be identified accordingly.

[0059] S105, determine the teaching link by using the teaching link knowledge graph, respectively distribute cross-modal attention weights to the dynamic semantic feature, the teacher action feature, the facial doubtful feature, the gesture pointing feature and the questioning tone feature according to the teaching link, and fuse the features according to the different weights to form a multi-modal fusion feature.

[0060] The teaching link knowledge graph records different teaching links existing in the classroom. Teachers will use different teaching behaviors to teach corresponding scientific knowledge in different teaching links. The teaching link knowledge graph is used to identify the current teaching stage, and different weights are assigned to the data features of different modalities according to the teaching stage. For example, in the demonstration teaching link, teachers will use language statements, definition explanations, and demonstration of teaching aids to teach, so the teaching behaviors in this teaching link can be classified as explanation and statement behaviors, display behaviors, and guidance behaviors. In the questioning teaching link, teachers will ask questions to students through language, and use gestures or body movements to select or guide students to answer, so the teaching behaviors in this teaching link can be classified as questioning behaviors. Some disciplines also have an experimental link, that is, scientific experiments are conducted in the classroom to enable students to master scientific knowledge through actual experimental operations. However, in different teaching links, the importance of different forms of multi-source data to different teacher behaviors is different. In the demonstration teaching link, the importance of the teacher's speech content to the explanation and statement behavior is higher, so the weight of the dynamic semantic feature can be increased. The importance of the teacher's actions and the change state of the teaching aids to the display behavior is higher, so the weight of the teacher's action feature can be increased. The importance of the teacher's facial expression, action, and interactive action with students to the questioning behavior is higher, so the weights of the facial doubt feature, gesture pointing feature, and questioning tone feature can be increased. Therefore, for different teaching links, different weights can be assigned to the features obtained from different data sources, which can improve the accuracy of teaching behavior analysis. By consulting the knowledge graph, the current teaching link is determined, and the corresponding data features are assigned higher weights. In subsequent teaching behavior recognition analysis using multi-source data, important features can be better captured, and more accurate teaching behavior recognition analysis results can be generated.

[0061] In S106, the multi-modal fusion features are migrated to the edge device using knowledge distillation, and the teaching behavior classification results are output on the edge device.

[0062] Since the teacher behavior classification utilizes multi-source data composed of video data and speech data, but the analysis and recognition methods of data from different sources are different, the teacher end with higher computing power deployed in the university computer room can be used to identify features of multi-source data. However, in order to facilitate students' understanding and mastery of teaching behaviors, it is necessary to generate teaching behavior classification results on the student end. Since the devices held by students are mostly edge devices with limited computing power, knowledge distillation is needed to perform data dimensionality reduction on the multi-modal fusion features, while preserving important features, so that they can be migrated to the student end and real-time output teaching behavior classification results on edge devices with limited computing power.

[0063] The embodiment collects video data and voice data of a teacher in a classroom, forms a dynamic semantic feature for identifying an explanation and statement behavior in a voice stream by dynamically cutting the voice data, forms a teacher action feature for identifying a demonstration behavior and a guidance behavior in a video stream by analyzing the video data, and forms a facial doubt feature, a gesture pointing feature and a question tone feature for identifying a question behavior of the teacher by combining a detected question tone in the dynamic semantic analysis, a detected doubt expression in the facial expression analysis and a detected gesture pointing to a student space in the action analysis. Then, a teaching link knowledge graph is used to determine a teaching link, different teaching behaviors in the teaching link are distributed with corresponding weights and fused into a multi-modal fusion feature, and finally, the fused feature is migrated to an edge device by using knowledge distillation, so that the limited computing power of the edge device can also generate a teaching behavior classification result in real time. By dynamic semantic cutting, the same or similar semantic voice segments are avoided from being cut, the semantic cutting accuracy is improved, and the voice recognition accuracy is improved, so that the explanation and statement behavior in the voice stream can be more accurately identified. By jointly identifying the question behavior from multiple sources of data, the video information and the voice information can be combined, and the identification accuracy of the question behavior is improved. By using the knowledge graph to distinguish the teaching link and distribute different weights, the weights of different data sources can be improved in different behavior recognitions, which is more conducive to the performance of data features and further improves the accuracy of teaching behavior recognition and classification. By knowledge distillation, the device with limited computing power can also obtain the teaching behavior classification result in real time, and the classroom information can be recorded by occupying a small storage space, which is especially convenient for students to review and repeatedly learn the classroom content, and improves the understanding and mastery of the teaching content by the students.

[0064] In an optional implementation of the embodiment, the method further includes:

[0065] When the output probabilities of the explanation and statement behavior, the demonstration behavior, the guidance behavior and the question behavior are all lower than 0.3, and a pen trace is detected in the teaching video data, a visual recognition model is used to track the teacher's hand action in the teaching video data, the teacher's pen trace is recognized, the blackboard writing in the teaching video data is recognized, the occluded part is completed by using a prediction model according to the recognition results of the pen trace and the text, and the teaching blackboard data is generated.

[0066] In addition to teaching through language and actions, teachers also teach students scientific knowledge through blackboard writing in class. Blackboard writing is also an important part of classroom teaching behavior. When the output probability of explanation and statement behavior, demonstration behavior, guidance behavior, and questioning behavior is less than 0.3 in the recognition result obtained by recognizing teaching behavior through video data and voice data, it indicates that the teacher is less likely to perform the above behaviors. At this time, the teacher is most likely to be writing on the blackboard. At this time, handwriting detection can be performed on the teaching video data to determine whether handwriting is formed, and whether the teacher is writing on the blackboard can be determined. The output probability is the confidence degree of the model outputting the classification result of each teaching behavior when recognizing various teaching behaviors using video or audio data. The explanation and statement behavior is identified by combining acoustic features and text keywords using dynamic semantic features formed by dynamic semantic event units, and is identified by matching speech speed, tone, and keywords. The demonstration behavior is identified by a visual recognition model (YOLO) that identifies the consistency of the movement of the teacher's action, gaze, and teaching aids in the video data. The guidance behavior is identified by capturing the teacher's human skeleton key points and using a visual recognition model (YOLO) to identify the interaction between the hand key points and the student space. The questioning behavior is identified by a facial expression recognition model and a visual model that identify the teacher's doubtful expression and body movements, respectively, and use a decision tree model to identify the questioning tone identified by the dynamic semantic feature. When each of the above behaviors is identified using the respective model, if the probability (confidence) of the recognition result of the corresponding teaching behavior output by the model is less than 0.3, it can be considered that the teacher is less likely to perform the corresponding behavior. When it is confirmed that the teacher is writing on the blackboard, the hand movement of the teacher is tracked using a visual recognition model, the handwriting of the written blackboard is recognized using the movement trajectory of the teacher's hand, and the text of the blackboard is recognized. Since the teacher's standing position often causes some occlusion of the blackboard, the handwriting of the occluded part is predicted using a prediction model based on the defects in the text recognition and the handwriting formed by the teacher's hand movement, and the occluded text part is completed to generate teaching blackboard data. Students can use the teaching blackboard data for learning, which can avoid missing other teaching behaviors of the teacher by manually copying the blackboard, and also facilitates review after class, improves learning efficiency, and is more conducive to mastering the knowledge taught. For example, when it is confirmed that the teacher is writing on the blackboard, the YOLOv8 model is used to dynamically crop the hand expansion area, which is a rectangular area formed by using 2 times the width and 1.5 times the length of the hand. By performing OCR recognition in this area, local blackboard data can be gradually generated following the movement of the teacher's hand.If no hand or hand movement is detected in a continuous period of time, global EAST text detection is started using the blackboard part image in the video data, global text data is generated, and the occluded defect part is identified. The local board writing data of the defect part is used, and the LSTM network is used to predict the board writing content of the defect part according to the teacher hand movement track of the defect part, to complete the board writing and generate complete board writing data. The input data of the LSTM network can be the past 10 frame hand displacement vector sequence before the occluded part, and the displacement of the future 5 frames is predicted through the double-layer LSTM network structure. Whether the blackboard has pen marks can be generated by identifying the teaching video data, for example, when the teacher hand moves, there is a line on the blackboard that does not exist before the teacher hand passes, which is significantly different from the color of the blackboard, and the trajectory is consistent with the teacher hand trajectory, that is, it is considered that the blackboard has pen marks.

[0067] Embodiment Two

[0068] Figure 2 A flowchart of a multi-modal based teaching behavior data classification method according to Embodiment Two of the present application, which is optimized based on the above-mentioned embodiments. In this embodiment, S102 is specifically optimized as follows:

[0069] Based on speech silence detection and question tone recognition, the teaching speech data is dynamically cut to form dynamic semantic event units.

[0070] The dynamic semantic event units are subjected to acoustic feature and text keyword joint recognition to generate dynamic semantic features, which are used to identify the explanation statement behavior.

[0071] Correspondingly, the multi-modal based teaching behavior data classification method provided in the present embodiment specifically includes:

[0072] S201, collect video information and speech information in the classroom and perform corresponding preprocessing to form teaching video data and teaching speech data.

[0073] S202, based on speech silence detection and question tone recognition, the teaching speech data is dynamically cut to form dynamic semantic event units.

[0074] The speech silence condition is detected by using a double-threshold detection based on RMS energy. The condition that the energy threshold is less than -40 dBFS is regarded as the speech silence condition. The condition that the speech silence duration is greater than 1 second is identified. Meanwhile, the pitch rise rate of the last 100 ms of a sentence is calculated based on the PYIN (Probability YIN) algorithm. The pitch with a pitch rise rate greater than 20% is determined as the intonation of a question sentence. According to the speech silence condition and the intonation of a question sentence, the teaching speech data is dynamically cut. The interval of the speech silence is used to determine the condition of sentence segmentation. The speech data is dynamically divided in length to form a dynamic semantic event unit. If the pitch rise rate of the last 100 ms of a sentence is greater than 20%, the dynamic semantic event unit expresses the semantics of a question sentence.

[0075] In S203, the dynamic semantic event unit is subjected to joint recognition of acoustic features and text keywords to generate dynamic semantic features, which are used to recognize the explaining and stating behaviors.

[0076] Since the teacher describes the scientific knowledge in a relatively flat and stable speed and intonation when stating or explaining the knowledge to facilitate the understanding of the students, the speed of the dynamic semantic event unit can be calculated based on the ASR text segmentation statistics. The fundamental frequency variance of the dynamic semantic unit can be calculated based on the F0 standard deviation. The energy envelope of the dynamic semantic unit can be analyzed based on the Hilbert transform to extract the acoustic features of the dynamic semantic unit. The acoustic features with a speed less than 5 words per second, a fundamental frequency variance less than 20 Hz, and an energy envelope fluctuation less than 3 dB can be regarded as the acoustic features representing the relatively flat and stable stating and explaining. The dynamic semantic unit is subjected to keyword matching by using the BERT fine-tuning model to identify the words commonly used in stating or explaining sentences, such as “therefore”, “because”, “for example”, “that is”, and “that is to say”. The explaining and stating behaviors can be recognized based on the acoustic features representing the stating and explaining and the identified keywords.

[0077] In S204, the teaching video data is subjected to action analysis to form teacher action features based on the teacher actions in the video stream, which are used to recognize the demonstrating and guiding behaviors.

[0078] Specifically, the teaching aid is identified by using the spatiotemporal attention based on the teaching video data. The teaching aid motion features are generated based on the motion trajectories of the identified teaching aid, which are used to recognize the demonstrating behavior.

[0079] In order to make it easier for students to master the scientific knowledge with high professional degree or complex technology, teachers often demonstrate the entity teaching aid in the classroom, and teach scientific knowledge combined with the entity teaching aid. When demonstrating with the teaching aid, the motion state of the teaching aid can be tracked and identified according to the teaching video data using the space-time attention, and the teaching aid motion feature is formed. Whether the teacher is performing a demonstration behavior can be identified according to the motion of the teaching aid. For example, first, the YOLOv5 model is used to detect the teaching aid to obtain the bounding box of the teaching aid, and then the motion trajectory of the teaching aid is analyzed. The optical flow of the teaching aid region is calculated (using the Farneback optical flow algorithm), the optical flow histogram (HOFO) of the teaching aid region is calculated, the motion amplitude and direction are counted, and if the optical flow amplitude of the teaching aid region is continuously > 30 pixels / second (i.e. more than 5 consecutive frames), the demonstration behavior candidate is triggered. Then the teacher's gaze is confirmed, the teacher's gaze direction is estimated by the Gaze360 model, and if the teacher's gaze is consistent with the motion direction of the teaching aid (included angle < 30 degrees), it can be considered that the teacher's gaze falls on the teaching aid, and the teacher's demonstration behavior can be identified at this time.

[0080] According to the teaching video data, the teacher's skeleton key points are generated, and the skeleton posture space feature is generated according to the continuous spatial association between the teacher's skeleton key points and the student region, which is used to identify the guidance behavior.

[0081] When teaching part of the professional knowledge, it is also necessary to pass through the practical experiment link to make the students better master the scientific knowledge of the related subjects. In the experimental classroom, the teacher often arranges practical or experimental tasks for the students to perform, and the teacher guides the students appropriately by observing the students' operation execution process to improve the students' understanding and mastery of scientific knowledge. The guidance behavior often produces spatial interaction with the student side on the teacher's body movement, and a visual recognition model is used to recognize the teacher's body to form human skeleton key points, and the human skeleton key points are respectively bound to the corresponding parts of the human body. When the teacher makes a movement, the visual recognition model can identify the movement made by the teacher according to the movement of the bound key points through the change of the posture of the skeleton, and the guidance behavior can be identified according to the interaction of the teacher's movement. For example, first, the HRNet model is used to extract the skeleton key points, 17 skeleton key points of the teacher are extracted from the video frame, and then the YOLOv3 is used to detect the student region in the video frame. The spatial relationship between the hand skeleton key points (left and right wrist) of the teacher and the student region is calculated, if the teacher's hand key points point to the student region for more than 3 seconds, i.e. the Euclidean distance between the teacher's hand and the student region is < 50 pixels, and the direction vector included angle is < 45 degrees, it can be considered that the teacher is guiding the student's practical behavior, and the teacher's guidance behavior can be identified at this time. The calculation of the shoulder key points can also be added to judge whether the teacher's body is facing the student region as an additional basis for judging the guidance behavior.

[0082] S205, performing teacher facial expression analysis and action analysis on the teaching video data, respectively identifying a teacher facial doubt expression and a gesture pointing to a student action, forming a facial doubt feature and a gesture pointing feature, performing voice intonation detection on the dynamic semantic event unit, identifying a teacher question sentence intonation to form a question intonation feature, and using the facial doubt feature, the gesture pointing feature and the question intonation feature to identify a questioning behavior.

[0083] S206, determining a teaching link using a teaching link knowledge graph, respectively performing cross-modal attention weight distribution on the dynamic semantic feature, the teacher action feature, the facial doubt feature, the gesture pointing feature and the question intonation feature according to the teaching link, and performing feature fusion according to the different weights to form a multi-modal fusion feature.

[0084] S207, migrating the multi-modal fusion feature to an edge device using knowledge distillation, and outputting a teaching behavior classification result on the edge device.

[0085] The dynamic semantic event unit formed by dynamic cutting based on voice silence detection and question sentence intonation recognition, and the acoustic feature and text keyword joint recognition, generates a dynamic semantic feature, which can preserve the semantic integrity of the sentence paragraph, avoid splitting the same semantic sentence paragraph, improve the accuracy of the semantic feature, and further improve the recognition accuracy of the explanation statement behavior. By using the spatio-temporal attention to recognize the teaching aids in the teaching video data, the motion trajectory of the teaching aids is formed, and according to the motion and interaction between the teacher and the teaching aids, the teacher's demonstration behavior can be accurately recognized. By generating teacher skeleton key points and tracking the key points, analyzing the spatial relationship between the hand skeleton key points and the student area, and identifying the interaction between the teacher and the students, the teacher's guidance behavior is identified. By classifying the demonstration behavior and the guidance behavior according to the teacher's action feature, the teaching aid motion and the interaction with the student space, the accuracy of the behavior classification is improved, and the probability of behavior misjudgment is reduced.

[0086] Optionally, the S207 comprises:

[0087] The multi-modal fusion feature is sequentially subjected to Logits distillation, feature distillation and structure distillation to generate a distilled multi-modal feature.

[0088] Because the classification of teacher behavior uses multi-source data, extracts a large number of features, and assigns different weights, it requires significant computing power, often necessitating deployment on school-level teacher terminals to meet the necessary computational resources. Therefore, knowledge distillation can reduce the number of model parameters, lowering computational complexity while retaining important features to save computing resources, enabling edge devices with limited computing power to meet the required requirements. First, Logits distillation softens the output probability, and KL divergence aligns the output distribution, preserving decision boundary information. Then, the Hint layer (6th layer activation value) from feature distillation is used to match intermediate features, constraining the similarity of intermediate representations. Finally, structural distillation modifies the network layers, replacing fully connected layers with grouped convolutions, achieving hardware-friendly deployment.

[0089] The distillation multimodal features are transferred to an edge device, and the edge device is used to classify the teaching behavior based on the distillation multimodal features, and the classification results are output.

[0090] Simplified distilled multimodal features are transferred to edge devices, which can classify teaching behaviors and output classification results with fewer parameters while retaining important features. For example, the edge device can use the lightweight MobileViT model, performing feature dimensionality reduction through grouped convolutions, followed by channel attention weighting. Then, a three-layer MobileViT block is input, and multiple parallel classification heads are used to identify teaching behaviors. The lecturing / declarative behavior classification head can use the Softmax activation function, while the demonstration, guidance, and questioning behaviors can use the Sigmoid activation function. Finally, the output layer outputs the behavior classification probability according to its respective weight. Temporal smoothing can be applied to the output results, and confidence filtering is used to output the classification results for behaviors with a probability greater than 0.5, thus obtaining the teaching behavior classification result. Edge devices are often handheld devices on the student's end, such as tablets and laptops, which have limited computing resources. Knowledge distillation allows students to generate teaching behavior classifications in real time with limited computing power, aiding in understanding teachers' classroom lectures and scientific knowledge, and also for after-class review and repeated learning.

[0091] Example 3

[0092] Figure 3 This is a flowchart of a multimodal teaching behavior data classification method according to Embodiment 3 of the present invention. This embodiment is an optimization based on the above embodiment. In this embodiment, S101 includes:

[0093] Collect video and audio information from the classroom;

[0094] The video information is subjected to spatial denoising and information enhancement processing to form teaching video data, and the speech information is subjected to environmental noise reduction, echo cancellation and human voice separation processing to form teaching speech data.

[0095] Correspondingly, the multi-modal based teaching behavior data classification method provided in the embodiment specifically includes the following steps.

[0096] S301, video information and speech information on a classroom are collected.

[0097] A camera is used to record the actions of a teacher and a blackboard on the classroom to form video data and board writing data, and a microphone is used to record the language of the teacher to form speech information. In order to ensure the acquisition ability of the board writing, which is used for later identification and storage of the content of the board writing, the camera needs to be erected at a visual angle directly opposite the blackboard, and the best position is the center directly opposite the blackboard. The camera records the actions of the teacher and the blackboard at the visual angle directly opposite the blackboard. The video recorded by the camera can be used to analyze the actions of the teacher and to perform character recognition on the board writing on the blackboard to extract the content of the board writing. The microphone can be worn on the body of the teacher and transmitted in a wireless manner, or can be arranged at a fixed position on the platform or the blackboard to facilitate the collection of the language information of the teacher with high quality. Exemplarily, the camera can adopt 1080P-30fps and a wide-angle lens to ensure the clarity and frame rate continuity of the video recording and the comprehensive coverage of the platform and the blackboard. The microphone can adopt a directional array microphone and can be arranged on the platform or the blackboard, with a signal-to-noise ratio > 60dB to reduce the interference of irrelevant noise and ensure the quality of speech recording and later identification.

[0098] Optionally, the video information and the speech information of the teacher on the classroom are collected by using a synchronous clock source and are time-stamped to be aligned to generate time-aligned video information and speech information.

[0099] When the video information and the speech information are collected, a synchronous clock source of IEEE 1588 PTP protocol is used to align the time stamps of the video information and the speech information, so that the master clock synchronization error of the two is less than 10ms, and the later analysis of the speech data and the video data can be jointly performed to ensure the accuracy of the analysis result. The video frame can use a PTS display time stamp (accurate to ms), and the audio stream can use an RTP message sequence number.

[0100] S302, the video information is subjected to spatial denoising and information enhancement processing to form teaching video data, and the speech information is subjected to environmental noise reduction, echo cancellation and human voice separation processing to form teaching speech data.

[0101] Video information is mainly used to identify the teacher's actions, facial expressions and board content, and speech data is mainly used to identify the teacher's statement tone and question tone, so appropriate preprocessing of video information and speech information is needed to eliminate the interference of irrelevant factors, so that video data and speech data can be used for subsequent feature extraction and teaching behavior analysis. For example, first, motion blur compensation is performed on the video data to repair the blur caused by the teacher's fast writing and gestures. A non-local mean deblurring algorithm is used to extract a blur kernel from three consecutive frames, and then a Wiener filter is used to restore the clear edges. Then, light adaptive processing is performed on the video data to solve the problem of blackboard reflection and projector interference. The video data is processed in different regions. The CLAHE (Clip Limit = 2.0) is used for the blackboard area, and the Retinex multi-scale light correction is used for the teacher area. Then, background filtering is performed to remove student movement interference. The DeepLabV3+ semantic segmentation model can be used to extract the mask of the podium area (pixel-level accuracy) and only keep the teacher and blackboard areas. Frame rate unification can also be used to compatible different camera equipment. The video data is interpolated to 30fps through bilinear interpolation, ensuring the timing consistency of action recognition. Key frame extraction can also be used to reduce redundant calculations. Based on motion saliency, the algorithm can only process frames with a motion intensity greater than 15 pixels / second, reducing 70% of invalid calculations, lightening the model, and saving computing resources. The YOLOv4-Tiny model can be used for teacher detection, and the Affine transformation can be used for pose alignment to a unified front view, solving the deformation problem when the teacher writes on the side. Environmental noise reduction processing is performed on the speech data to eliminate fan noise, table and chair movement, etc. The RNNoise deep learning model can be used to screen 32ms frames or 10ms frame shifts, and joint time-frequency mask estimation is performed to improve the signal-to-noise ratio by 25dB. Then, echo cancellation is used to suppress microphone howling. The adaptive filtering (NLMS) algorithm is used to attenuate the echo by 40dB. Then, human voice separation is performed using the multi-channel beamforming (MVDR) algorithm to filter out student interference speech. An 8-microphone array can be used during acquisition, with the main lobe pointing in the teacher's direction ± 15°, achieving a teacher speech retention rate of > 90%. Volume equalization processing can also be used to solve the volume fluctuations caused by distance changes. Dynamic range compression (DRC) is used to make the target volume reach -23LUFS, with an attack / release time of 50ms / 300ms, making the volume fluctuation < ± 3dB. Sampling rate unification can also be used to compatible different equipment. Resampling to 16kHz (using the SoX tool) meets the ASR input requirements. Silent section marking can also be used to identify invalid speech segments based on energy threshold (RMS < -40dBFS) and zero-crossing rate, improving the efficiency of semantic event segmentation, and normalizing the speech parameters.

[0102] S303, dynamically cutting the teaching voice data to form a dynamic semantic event unit, and identifying the explanation and statement behavior according to the dynamic semantic event unit.

[0103] S304, performing action analysis on the teaching video data, and identifying the display behavior and the guidance behavior according to the teacher action in the video stream.

[0104] S305, performing teacher facial expression analysis and action analysis on the teaching video data, respectively identifying the teacher facial confusion expression feature and the gesture pointing to the student feature, performing voice intonation detection on the dynamic semantic event unit, identifying the teacher question intonation feature, and identifying the questioning behavior according to the teacher facial confusion expression feature, the gesture pointing to the student feature and the teacher question intonation feature.

[0105] S306, using the teaching link knowledge graph to respectively perform cross-modal attention weight distribution on the identified explanation and statement behavior, display behavior, guidance behavior and questioning behavior, and perform feature fusion according to the different weights distributed, to form a multi-modal fusion feature.

[0106] S307, migrating the multi-modal fusion feature to the edge device by using knowledge distillation, and outputting the teaching behavior classification result on the edge device.

[0107] The embodiment synchronizes the clock source to collect the video information and the voice information of the teacher in the classroom and performs timestamp alignment, ensures the timing consistency of the video data and the voice data, and guarantees the association consistency of the video data and the voice data when performing joint feature extraction and analysis on the multi-source data. The video data and the voice data are also preprocessed, the video information is subjected to spatial denoising and information enhancement processing, the voice information is subjected to environmental noise reduction, echo cancellation and human voice separation processing, the OCR recognition accuracy of the blackboard writing is improved, the action delay is reduced, the recognition accuracy of the facial expression is guaranteed, the voice signal-to-noise ratio is reduced, the question sentence detection recall rate can be effectively controlled, and the ASR word error rate is reduced. The teaching video data and the teaching voice data formed can reduce the interference information on the premise of retaining important features, and guarantee the accuracy of the teaching behavior recognition.

[0108] Embodiment four

[0109] Figure 4 A structural schematic diagram of a teaching behavior data classification device based on multiple modes according to the fourth embodiment of the application, in the embodiment, the teaching behavior data classification device based on multiple modes comprises:

[0110] The information collection module 810 is configured to collect the video information and the voice information in the classroom, and perform corresponding preprocessing to form teaching video data and teaching voice data.

[0111] The voice data recognition module 820 is configured to perform dynamic cutting on the teaching voice data to form dynamic semantic event units, generate dynamic semantic features according to the dynamic semantic event units, and recognize the explanation statement behavior;

[0112] The video action analysis module 830 is configured to perform action analysis on the teaching video data, form teacher action features according to the teacher actions in the video stream, and recognize the display behavior and the guidance behavior;

[0113] The question behavior recognition module 840 is configured to perform teacher facial expression analysis and action analysis on the teaching video data, recognize the teacher facial confusion expression and the gesture pointing student action respectively to form facial confusion features and gesture pointing features, perform voice intonation detection on the dynamic semantic event units, recognize the teacher question sentence intonation to form a question intonation feature, and recognize the question behavior according to the facial confusion features, the gesture pointing features and the question intonation feature.

[0114] The multi-modal attention feature fusion module 850 is configured to determine the teaching link by using a teaching link knowledge graph, respectively perform cross-modal attention weight distribution on the dynamic semantic features, the teacher action features, the facial confusion features, the gesture pointing features and the question intonation features according to the teaching link, perform feature fusion according to the different weights distributed, and form multi-modal fusion features.

[0115] The teaching behavior classification module 860 is configured to migrate the multi-modal fusion features to an edge device by using knowledge distillation, and output a teaching behavior classification result on the edge device.

[0116] The embodiment collects video information and voice information on the classroom through the information collection module and performs corresponding preprocessing, cuts the teaching voice data dynamically to form a dynamic semantic event unit and extracts dynamic semantic features through the voice data recognition module, which is used for the recognition of explanation and statement behaviors, analyzes the teaching video data through the video action analysis module to form the teacher action features, which is used for the recognition of display behaviors and guidance behaviors, detects the teacher's facial confusion expression and gesture pointing to the student features in the teaching video data through the question behavior recognition module, combines the question intonation features in the dynamic semantic features, which is used for the recognition of question behaviors, determines the teaching link through the multi-modal attention feature fusion module, and performs cross-modal attention weight distribution on the features of different data sources, and fuses to form multi-modal fusion features, and the teaching behavior classification module is used for knowledge distillation of the multi-modal fusion features and migration to the edge device to output the teaching behavior classification result. Through dynamic semantic cutting, the same or similar semantic voice fragments are avoided to be cut, the semantic cutting accuracy is improved, and the voice recognition accuracy is improved, so that the explanation and statement behaviors in the voice stream can be more accurately recognized. The question behaviors are recognized through multi-source data joint recognition, the video information and the voice information are combined, and the recognition accuracy of the question behaviors is improved. The knowledge graph is used for distinguishing the teaching link and distributing different weights, which can improve the weight of different data sources in different behavior recognition, is more conducive to the performance of data features, and further improves the accuracy of teaching behavior recognition and classification. Through knowledge distillation, the device with limited computing power can also obtain the teaching behavior classification result in real time, and the classroom information can be recorded through a small storage space occupation, which is especially convenient for students to review and repeatedly learn the classroom content, and improves the understanding and mastery of the teaching content by the students.

[0117] The multi-modal based teaching behavior data classification device provided by the embodiment of the application can execute the multi-modal based teaching behavior data classification method provided by any embodiment of the application, and has the corresponding function modules and beneficial effects of executing the method.

[0118] Embodiment five

[0119] Figure 5 A structural diagram of an electronic device according to the fifth embodiment of the application, Figure 5 A block diagram of an exemplary electronic device 12 suitable for implementing embodiments of the application is shown. Figure 5 The electronic device 12 shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the application.

[0120] As Figure 5As shown, the electronic device 12 is in the form of a general- purpose computer. Components of the electronic device 12 can include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including the system memory 28 to the processing unit 16.

[0121] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics bus (e.g., an Accelerated Graphics Port, or AGP bus) and a processor or local bus using any of a variety of bus architectures. By way of example, these architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0122] The electronic device 12 typically includes a variety of computer system readable media. Such media can be any available media that is located either internally or externally to the electronic device 12, including both volatile and nonvolatile media, removable and non-removable media.

[0123] The system memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 can be provided for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Figure 5 Not shown, a removable / non-removable interface can also be provided and can include at least one drive or port that enables the removal and / or attachment of external media storage devices, such as a magnetic floppy disk drive, a magnetic hard drive, or optical super media drive (e.g., DVD-ROM drive, CD-RW drive, etc.) to the electronic device 12. Although not specifically shown, such computing-based devices can also include other peripheral output devices such as an audio device, which can include a speaker, a microphone, etc. Such computing-based devices can further include a display device, which can be a cathode ray tube (CRT), a liquid crystal display (LCD), or any other display device suitable for creating a graphical interface. Figure 5 The system memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 can be provided for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive").

[0124] A program / utility 40 having a set (at least one) of program modules 42 can be stored in, for example, system memory 28 by way of example, such program modules 42 include an operating system, one or more application programs, other program modules, and program data, each of which or a combination can include implementation of a network environment. Program modules 42 generally carry out the functions and / or methodologies of embodiments of the present application as described herein.

[0125] The electronic device 12 can also be in communication with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; can also be in communication with one or more devices that enable a user to interact with the electronic device 12 / server / computer; and / or can be in communication with any devices (such as a network card, a modem, etc.) that enable the electronic device 12 to communicate with one or more other computing devices. Such communication can be facilitated by an Input / Output (I / O) interface 22. Still yet, the electronic device 12 can be in communication with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or the Internet) through a network adapter 20. As Figure 5 illustrated, the network adapter 20 is in communication with the other components of the electronic device 12 through a bus 18. It should be understood that although not shown, other hardware and / or software components that are ​ described in connection with the electronic device 12 can be used in connection with the electronic device 12. Furthermore, the embodiments of the present application can be implemented by hardware only or combinations of both software and / or firmware and hardware, although the present description is generally presented in terms of software or symbolic representations of hardware as a lending grammatical convenience. Examples of such hardware can include devices, apparatus, etc. for implementing software in accordance with this disclosure, including inter alia: application-specific integrated circuits (ASICs); field-programmable gate arrays (FPGAs); servers, computers, etc. for hosting or executing processes and programs in accordance with the present disclosure; memory, etc. for storing these processes and programs; and the like.

[0126] The processing unit 16 performs various functions described above as provided by embodiments of the present application by executing programs stored in the system memory 28.

[0127] Embodiment Six

[0128] Embodiment Six of the present application also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, perform a multi-modal based teaching behavior data classification method as provided by the above embodiments.

[0129] The computer storage media of the embodiments of the present application can be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired computer program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, or twisted pair, then the coaxial cable, fiber optic cable, or twisted pair are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-Ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0130] A computer readable signal medium can include a propagated data signal with computer executable code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport programming code.

[0131] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0132] Computer program code for carrying out operations for aspects of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0133] Note that the foregoing are merely examples of the preferred embodiments of the present application and the principles of the technology employed. It will be understood by those skilled in the art that the present application is not limited to the specific embodiments described herein, and that various obvious changes, modifications and substitutions can be made to the present application without departing from the scope of the present application. Therefore, although the present application has been described in detail with reference to the foregoing embodiments, the present application is not limited to the foregoing embodiments, and can include other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.

Claims

1. A method for classifying teaching behavior data based on multi-modal, characterized in that, The method comprises: S101, collecting video information and voice information on the classroom and performing corresponding preprocessing to form teaching video data and teaching voice data; S102, performing dynamic cutting on the teaching voice data to form a dynamic semantic event unit, generating a dynamic semantic feature according to the dynamic semantic event unit, and using the dynamic semantic feature to identify the explanation statement behavior; S103, performing action analysis on the teaching video data, forming a teacher action feature according to the teacher action in the video stream, and using the teacher action feature to identify the display behavior and the guidance behavior; S104, performing teacher facial expression analysis and action analysis on the teaching video data, respectively identifying the teacher facial confusion expression and the gesture pointing to the student action to form a facial confusion feature and a gesture pointing feature, performing voice intonation detection on the dynamic semantic event unit, identifying the teacher question sentence intonation to form a question intonation feature, and using the facial confusion feature, the gesture pointing feature and the question intonation feature to identify the questioning behavior; S105, determining the teaching link using a teaching link knowledge graph, respectively assigning cross-modal attention weights to the dynamic semantic feature, the teacher action feature, the facial confusion feature, the gesture pointing feature and the question intonation feature according to the teaching link, and performing feature fusion according to the different weights to form a multi-modal fusion feature; S106, migrating the multi-modal fusion feature to an edge device using knowledge distillation, and outputting a teaching behavior classification result on the edge device.

2. The method of claim 1, wherein, The method further comprises: When the output probabilities of the explanation statement behavior, the display behavior, the guidance behavior and the questioning behavior are all lower than 0.3, and the pen trace is detected in the teaching video data, a visual recognition model is used to track the teacher's hand action in the teaching video data, the teacher's pen trace is recognized, the handwriting in the teaching video data is recognized, the occluded part is completed using a prediction model according to the recognition results of the pen trace and the handwriting, and teaching handwriting data is generated.

3. The method of claim 1, wherein, The S102 comprises: Based on voice silence detection and question sentence intonation recognition, the teaching voice data is dynamically cut to form a dynamic semantic event unit; The dynamic semantic event unit is subjected to joint recognition of acoustic features and text keywords to generate a dynamic semantic feature for identifying the explanation statement behavior.

4. The method of claim 1, wherein, The S103 comprises: According to the teaching video data, a teaching aid is identified using a space-time attention, a teaching aid motion feature is generated according to the motion track of the identified teaching aid, and the teaching aid motion feature is used to identify the display behavior; According to the teaching video data, a teacher skeleton key point is generated, a skeleton posture space feature is generated according to the continuous spatial association between the teacher skeleton key point and the student area, and the skeleton posture space feature is used to identify the guidance behavior.

5. The method of claim 1, wherein, The S106 comprises: The multi-modal fusion feature is subjected to Logits distillation, feature distillation and structure distillation in sequence to generate a distilled multi-modal feature; The distilled multi-modal feature is migrated to an edge device, the edge device is used to classify the teaching behavior of the distilled multi-modal feature, and a classification result is output.

6. The method of claim 1, wherein, The S101 comprises: Collecting video information and voice information on the classroom; The video information is subjected to spatial denoising and information enhancement processing to form teaching video data, and the voice information is subjected to environmental noise reduction, echo cancellation and voice separation processing to form teaching voice data.

7. The method of claim 1, wherein, The video information and voice information on the classroom are collected, including: The video information and voice information of the teacher on the classroom are collected through a synchronous clock source, and time stamp alignment is performed to generate time-aligned video information and voice information. 8.A multi-modal based teaching behavior data classification apparatus, characterized in that, Including: An information collection module is configured to collect video information and voice information on the classroom and perform corresponding preprocessing to form teaching video data and teaching voice data; A voice data recognition module is configured to perform dynamic cutting on the teaching voice data to form a dynamic semantic event unit, generate dynamic semantic features based on the dynamic semantic event unit, and recognize explanation and statement behaviors; A video action analysis module is configured to perform action analysis on the teaching video data, form teacher action features based on the teacher actions in the video stream, and recognize display behaviors and guidance behaviors; A questioning behavior recognition module is configured to perform teacher facial expression analysis and action analysis on the teaching video data, recognize teacher facial doubt expressions and hand gesture pointing actions to students, form facial doubt features and hand gesture pointing features, perform voice intonation detection on the dynamic semantic event unit, recognize teacher question sentence intonation to form question intonation features, and recognize questioning behaviors based on the facial doubt features, hand gesture pointing features and question intonation features; A multi-modal attention feature fusion module is configured to determine a teaching link based on a teaching link knowledge graph, perform cross-modal attention weight distribution on the dynamic semantic features, teacher action features, facial doubt features, hand gesture pointing features and question intonation features based on the teaching link, perform feature fusion based on the different weights, and form multi-modal fusion features; A teaching behavior classification module is configured to migrate the multi-modal fusion features to an edge device using knowledge distillation, and output teaching behavior classification results on the edge device.

9. An electronic device, comprising: The electronic device includes: One or more processors; A storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-modal based teaching behavior data classification method of any one of claims 1-7.

10. A storage medium containing computer executable instructions for performing the multi-modal based teaching behavior data classification method of any one of claims 1-7 when executed by a computer processor.