Multi-modal analysis and evaluation method for classroom teaching skill video

By constructing a systematic evaluation system through multimodal data collaborative analysis, the problems of subjectivity and high cost in traditional classroom teaching evaluation are solved. It realizes full-process automation and multi-dimensional teaching skills analysis and evaluation, improves the objectivity and practicality of evaluation, and supports the digital transformation of education.

CN121684718APending Publication Date: 2026-03-17GUANGXI NORMAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511880068.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Traditional classroom teaching evaluation suffers from problems such as subjectivity, high cost, one-sided information capture, fragmented system, and insufficient practical value, and cannot meet the needs of large-scale campus group assessment and regional education quality survey.

Method used

Employing a multimodal data collaborative analysis method, and utilizing audio, video, and text processing technologies, a systematic evaluation system is constructed across eight dimensions, including teaching objectives, content, methods, behaviors, and atmosphere. This system generates quantitative scores and targeted improvement suggestions, enabling fully automated evaluation throughout the entire process.

Benefits of technology

It improves the objectivity and comparability of evaluation, significantly reduces labor costs, comprehensively captures diverse classroom information, constructs a scientific and comprehensive evaluation system, outputs practical teaching improvement guidelines, activates the value of classroom data, and supports teaching and research decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121684718A_ABST
    Figure CN121684718A_ABST
Patent Text Reader

Abstract

The invention belongs to a multi-modal analysis and evaluation method for a classroom teaching skill video, and belongs to the crossing field of education technology and artificial intelligence, and the method comprises the following steps: uploading a video and filling metadata; automatically preprocessing the video; teaching targets and characteristic innovation are analyzed; analyzing the teaching content and the teaching method; analyzing teacher behaviors; analyzing student performance; analyzing the classroom atmosphere; analyzing teaching evaluation; analyzing characteristic innovation; comprehensive evaluation is given; and generating an analysis report. According to the method, a multi-view quantitative evaluation system is constructed by fusing a multi-AI technology, full-process automation is realized, content score and improvement suggestion reports are output, evaluation objectivity and efficiency are improved, and accurate improvement of teaching quality and professional development of teachers are supported.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of educational technology and artificial intelligence, and in particular to a multi-modal analysis and evaluation method for classroom teaching skill videos. BACKGROUND

[0002] Classroom teaching skill evaluation is the core support for ensuring the quality of education and promoting the professional development of teachers, and is widely used in school research and supervision, regional education quality assessment, and teacher self-reflection under normal circumstances. With the acceleration of digital transformation in education, the traditional evaluation mode centered on manual evaluation has gradually exposed obvious limitations: this mode relies on expert observation, peer review or student feedback, and requires evaluators to participate in the classroom or repeatedly review teaching videos, which is extremely costly in terms of manpower and time. Moreover, the evaluation results are easily influenced by the evaluators' teaching experience, professional background and subjective preferences, and there is a lack of unified quantitative standards for judging the core indicators such as the degree of achievement of teaching goals and the effectiveness of classroom interaction, resulting in insufficient objectivity and poor cross-scene comparability of the evaluation, making it difficult to meet the needs of large-scale campus group evaluation or regional education quality survey.

[0003] To alleviate the shortcomings of manual evaluation, there are only some preliminary surface-level technical attempts in the industry, which have significant gaps with the current digital and systematic evaluation needs and have not broken through the core limitations of traditional modes: such solutions mostly rely on mechanical data processing such as video duration statistics and text keyword matching, or are limited to fragmented evaluation of a single modality, and the technical level only stays at simple data extraction without real intelligent analysis capabilities. The core defects are as follows: first, the data analysis dimension is lacking, only the surface extraction of single sensory information, which cannot capture the core dynamics such as teacher-student interaction and student concentration, resulting in serious distortion of the evaluation; second, the degree of automation is insufficient, and a large amount of manual intervention is still required in data screening and indicator determination, which has not achieved full-process closed-loop automation, and the efficiency has not been substantially improved; third, the evaluation system lacks systematization, and there is no complete framework covering teaching goals, methods and atmosphere, and no special rules are designed for different subjects, which has poor adaptability and only outputs formal results.

[0004] At the same time, the in-depth promotion of education informatization has enabled schools at all levels to accumulate a large amount of classroom teaching audio and video data, but these data are mostly stored in raw format and have not been converted into structured and quantifiable evaluation indicators and decision-making basis, and the data value has been in a "sleeping" state for a long time. Current education evaluation is gradually shifting from experience-driven to data-driven, and there is an increasing demand for standardized and scaled evaluation data for school research and management and regional education quality analysis, but the traditional evaluation mode and existing technical tools cannot meet this transformational demand, making it difficult to improve the scientificity, accuracy and timeliness of education evaluation. Under this background, developing a full-process automated, multi-modal data collaborative analysis, and multi-dimensional systematic classroom teaching skill analysis and evaluation method has become a key problem to be solved in the field of educational technology. SUMMARY

[0005] The present application aims to solve the problems of subjectivity, high cost, one-sided information capture, scattered system, and insufficient practical value in traditional classroom teaching evaluation, to realize full-process automatic evaluation to reduce human time cost, to fully mine multi-dimensional information in the classroom through multi-modal data collaborative analysis, to build a systematic evaluation system containing eight dimensions to adapt to the needs of nine disciplines, to output quantitative scores and targeted improvement suggestions to help teachers optimize teaching accurately, and to activate the value of classroom data to support teaching research decisions and promote the transformation of education evaluation to data-driven.

[0006] To achieve the above purpose, the present application provides a multi-modal analysis and evaluation method for classroom teaching skill video, characterized in that it comprises the following steps:

[0007] S1, video uploading and metadata filling: the user uploads the classroom teaching skill teaching video to the system, and adds the course name, teaching topic, school name, location, grade metadata information corresponding to the video in the system;

[0008] S2, automatic video preprocessing: extract the audio file in the video, and generate text and subtitle files after standardization processing after model transcription, synchronously generate word cloud diagram, and count video length and language density;

[0009] S3, analysis of teaching objectives and characteristic innovation: for the teaching process, the target clarity of the teaching objectives and the teaching design innovation of the characteristic innovation are analyzed;

[0010] S4, analysis of teaching content and teaching method: for the teaching process, the scientificity and accuracy, appropriateness and relevance of the teaching content, and the technical integration, diversity and appropriateness of the teaching method are analyzed;

[0011] S5, analysis of teacher behavior: for the teaching process, the classroom control ability and interaction guidance of the teacher behavior are analyzed;

[0012] S6, analysis of student performance: for the teaching process, the thinking depth and innovation, participation and concentration of student performance are analyzed;

[0013] S7, analysis of classroom atmosphere: for the teaching process, the democracy and inclusiveness, learning enthusiasm of the classroom atmosphere are analyzed;

[0014] S8, analysis of teaching evaluation: for the teaching process, the evaluation method and feedback effectiveness of the teaching evaluation are analyzed;

[0015] S9, give comprehensive evaluation: after completing the analysis in steps S3 to S10 in parallel, according to the analysis results of each dimension, give the overall evaluation;

[0016] S10. Generate an analysis report: Summarize and organize the analysis results of each dimension obtained in steps S3 to S11 to generate an analysis report.

[0017] Preferably, step S2, automated video preprocessing, specifically includes the following steps:

[0018] S2-1. Use the audio and video processing tool ffmpeg to extract the audio from the lecture video to obtain an independent audio file; use the Faster-Whisper speech recognition model to process the audio file into speech and generate initial text data.

[0019] S2-2. Standardize and clean the initial text data and automatically add punctuation marks that conform to grammatical rules to generate standardized subtitle files and text files; generate a visual word cloud based on the text files using a word cloud algorithm;

[0020] S2-3. Record the total duration of the lecture video, perform word count per minute on the text file, and form a quantitative language density index.

[0021] Preferably, in step S3, the analysis of teaching objectives and distinctive innovations is conducted through the following two aspects:

[0022] S3-1. Based on the multimodal data preprocessed in step S2, for the dimension of "clarity of teaching objectives," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the teaching objectives conform to the curriculum standards, whether they conform to students' cognitive levels, whether they are observable, and whether they are achievable." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the performance of teaching objectives corresponding to matching cognition, conforming to standards, and goal completion, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the quantitative scores of the sub-items to obtain a comprehensive score for the dimension of "clarity of teaching objectives." The two-level evaluation model includes requesting the large language model to generate a performance description of "clarity of teaching objectives" based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of about 120 words based on the comprehensive score and the content detection conclusion.

[0023] S3-2. Based on the multimodal data preprocessed in step S2, for the dimension of "innovative instructional design," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether there are unique teaching ideas, activity designs, or resource integration methods." Specifically, the multi-dimensional scoring mechanism involves calling a large language model to professionally score the uniqueness of resource integration, teaching ideas, and activity designs, and outputting quantitative scores for each sub-item. Simultaneously, based on preset rules, a weighted algorithm is used to fuse the two types of scores to obtain a comprehensive score for the "innovative instructional design" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection function based on the teaching text to describe the performance of "innovative instructional design," and generating a comprehensive review of approximately 120 words of teaching suggestions based on the score evaluation results and content detection conclusions.

[0024] Preferably, in step S4, the analysis of teaching content and teaching methods is conducted through the following four aspects:

[0025] S4-1. Scientific Rigor and Accuracy: Based on the multimodal data preprocessed in step S2, a comprehensive analysis is conducted on the "scientific rigor and accuracy" dimension, focusing on the evaluation criteria of "whether the knowledge explanation is accurate and error-free, whether the logic is rigorous, and whether it reflects the forefront of the discipline." A multi-dimensional scoring mechanism and a two-level evaluation model are used for comprehensive analysis. The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching content performance corresponding to the accuracy of knowledge, the rigor of logic, and the relevance to the forefront of the discipline, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "scientific rigor and accuracy" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "scientific rigor and accuracy" performance based on the teaching text, and generating a comprehensive evaluation of approximately 120 words of teaching suggestions based on the comprehensive score and the content detection conclusions.

[0026] S4-2, Appropriateness and Relevance: Based on the multimodal data preprocessed in step S2, for the dimension of "appropriateness and relevance," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the content difficulty matches the students' level, whether it connects with prior and subsequent knowledge, and whether it relates to real-life situations." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching content's performance in terms of its relevance to real-world situations, appropriate difficulty, and knowledge connection, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "appropriateness and relevance" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "appropriateness and relevance" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of approximately 120 words based on the comprehensive score and content detection conclusions.

[0027] S4-3, Technology Integration Degree: Based on the multimodal data preprocessed in step S2, intelligent analysis is conducted on the "technology integration degree" dimension, focusing on the evaluation criteria of "whether information technology effectively assists teaching and whether there is excessive reliance on it." Classroom videos are processed using computer vision technology, with VideoCapture used to read basic information such as video frame rate and total number of frames, employing a 2-second frame sampling strategy. A pre-trained YOLO object detection model is loaded to detect and identify teaching items in the video frames. The detection results are quantitatively analyzed, statistically analyzing the frequency of appearance, duration, and continuous use of various teaching aids, calculating the usage ratio, and generating a comprehensive score based on preset evaluation criteria. Simultaneously, a large language model is requested to generate a comprehensive review and improvement suggestions for teaching based on the detection data, statistical results, and classroom scene information regarding the "technology integration degree."

[0028] S4-4. Diversity and Appropriateness: Based on the multimodal data preprocessed in step S2, for the dimension of "diversity and appropriateness," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the methods are flexible and diverse, and whether they conform to the characteristics of the subject and the teaching content." Specifically, the multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching methods corresponding to flexibility and diversity and conformity to requirements, and outputting quantitative scores for each sub-item. Simultaneously, based on preset rules, a weighted algorithm is used to fuse the two types of scores to obtain a comprehensive score for the "diversity and appropriateness" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "diversity and appropriateness" performance based on the teaching text, and generating a comprehensive review of approximately 120 words of teaching suggestions based on the score evaluation results and content detection conclusions.

[0029] Preferably, in step S5, the analysis of teacher behavior is conducted through the following two aspects:

[0030] S5-1, Classroom Management Ability: Based on the multimodal data preprocessed in step S2, for the dimension of "classroom management ability," deep analysis of classroom audio is performed using audio processing technology. The silence function of the pydub library is used to detect silent segments in the audio, calculate the percentage of silent duration, and generate a basic classroom management score by combining a preset silence percentage threshold and evaluation rules for the rationality of silent period distribution. The teacher's classroom language content is extracted from the preprocessed text data, and the accuracy of their use of professional terminology, clarity of instructions, and stability of speech rate are analyzed. Based on the silence analysis results, language expression statistics, and a large language model for classroom scene information requests, a comprehensive review and improvement suggestions for teaching recommendations regarding "classroom management ability" are generated.

[0031] S5-2, Interactive Guidance: Based on the multimodal data preprocessed in step S2, for the "interactive guidance" dimension, a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether an interactive scenario is created and whether students can be effectively guided to think, ask questions, and discuss." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teacher's behavioral performance corresponding to the interactive scenario, teaching questions, teaching reflections, and teaching discussions, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "interactive guidance" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "interactive guidance" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of approximately 120 words based on the comprehensive score and the content detection conclusions.

[0032] Preferably, in step S6, the analysis of student performance is conducted from the following two aspects:

[0033] S6-1, Depth of Thinking and Innovation: Based on the multimodal data preprocessed in step S2, a comprehensive analysis is conducted on the dimension of "depth of thinking and innovation," focusing on the evaluation criteria of "whether students can raise questions, express unique insights, and apply knowledge to solve complex problems." A multi-dimensional scoring mechanism and a two-level evaluation model are used. The multi-dimensional scoring mechanism involves calling a large language model to professionally score students' knowledge application and complex problem-solving abilities, as well as their questioning and unique insight abilities, and outputting quantitative scores for each sub-item. A weighted algorithm is then used to fuse these scores based on preset rules to obtain a comprehensive score for the "depth of thinking and innovation" dimension. The two-level evaluation model includes requesting the large language model to generate a content description of the "depth of thinking and innovation" performance based on the teaching text, and generating a comprehensive evaluation of approximately 120 words of teaching suggestions based on the comprehensive score and the content detection conclusions.

[0034] S6-2, Engagement and Attention: Based on the multimodal data preprocessed in step S2, for the "engagement and attention" dimension, the deployed 3D-Speaker speech processing technology is invoked to perform speech activity detection to identify speech segments and filter out silent parts. Speaker embeddings are extracted from the speech segments and clustered to identify different speakers. The number of speakers and speaking duration information are counted. Based on preset rules, a comprehensive score for the "engagement and attention" dimension is obtained. The content detection results obtained by the large language model based on speech processing technology are requested. Combining the comprehensive score and content detection conclusion, a comprehensive review of teaching suggestions of about 120 words is generated.

[0035] Preferably, step S7, analyzing the classroom atmosphere, specifically includes the following steps:

[0036] S7-1, Democracy and Inclusivity: Based on the multimodal data preprocessed in step S2, a comprehensive analysis is conducted on the "democracy and inclusiveness" dimension, focusing on the evaluation criteria of "whether the classroom reflects equal dialogue, whether it accommodates different viewpoints, and whether the teacher-student relationship is harmonious." A multi-dimensional scoring mechanism and a two-level evaluation model are used. The multi-dimensional scoring mechanism involves calling a large language model to professionally score the classroom atmosphere corresponding to teacher-student relationships, equal dialogue, and tolerance of viewpoints, and outputting quantitative scores for each sub-item. A weighted algorithm is then used to fuse these scores based on preset rules to obtain a comprehensive score for the "democracy and inclusiveness" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection function based on the teaching text to describe the performance of "democracy and inclusiveness," and generating a comprehensive evaluation of approximately 120 words of teaching suggestions based on the comprehensive score and the content detection conclusions.

[0037] S7-2, Learning Motivation: Based on the multimodal data preprocessed in step S2, for the dimension of "learning motivation," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether students demonstrate a thirst for knowledge and curiosity, and whether they are willing to cooperate and share." The multi-dimensional scoring mechanism involves calling a large language model to analyze the text and identify and statistically analyze all "problem" sentences for professional scoring. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "learning motivation" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "learning motivation" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of about 120 words based on the comprehensive score and the content detection conclusion.

[0038] Preferably, step S8, analyzing teaching evaluation, specifically includes the following steps:

[0039] S8-1, Evaluation Method: Based on the multimodal data preprocessed in step S2, for the "evaluation method" dimension, focusing on the evaluation point of "whether to adopt multi-dimensional evaluation", the corresponding subject evaluation method is selected from the nine preset subject evaluation methods according to the teaching theme. The deployed RAGFlow service is called to conduct professional evaluation. At the same time, combined with preset keywords and rules, a weighted algorithm is used to obtain the comprehensive score of the "evaluation method" dimension. The two-level evaluation mode includes requesting the large language model to generate a performance description of the "evaluation method" based on the teaching text for content detection, and generating a comprehensive evaluation of teaching suggestions of about 120 words based on the comprehensive score and content detection conclusion.

[0040] S8-2, Feedback Effectiveness: Based on the multimodal data preprocessed in step S2, for the "Feedback Effectiveness" dimension, a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the feedback is timely and specific, and whether it can guide students to improve their learning." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching performance corresponding to timely and specific feedback, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "Feedback Effectiveness" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "Feedback Effectiveness" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of approximately 120 words based on the comprehensive score and the content detection conclusion.

[0041] Preferably, in step S9, the specific operation process for providing a comprehensive evaluation is as follows: after completing the analysis in parallel based on steps S3 to S8, the evaluation criteria, evaluation scores, suggestions, and analysis results of each dimension are summarized, and the large language model is requested to conduct a comprehensive quantitative evaluation and improvement suggestions for the teaching skills of the classroom.

[0042] Preferably, the specific operation process of generating the analysis report in step S10 is as follows: the analysis results of each dimension from steps S3 to S9 are classified, summarized and logically integrated. The analysis results include the quantitative scores of each dimension sub-item, the comprehensive score, the content detection conclusion, and the specific teaching suggestions; at the same time, the comprehensive quantitative evaluation conclusion of step S9 is integrated and presented in the form of a structured analysis report. The report includes video basic metadata, multi-dimensional evaluation details from different perspectives, and a comprehensive conclusion module, and simultaneously integrates the visualized word cloud and language density index quantitative data generated in step S2.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. Addressing the core pain points of traditional evaluation and improving its objectivity and comparability: By abandoning the traditional model that relies on subjective expert evaluation, and through standardized data analysis processes and quantitative scoring systems, the interference caused by evaluators' personal preferences and professional backgrounds is eliminated, so that the evaluation results have a unified measurement standard, significantly improving objectivity and cross-scenario comparability, and filling the gap in traditional evaluation that lacks quantitative basis.

[0045] 2. Achieve full-process automation, significantly reduce evaluation costs and adapt to large-scale needs: From video preprocessing and multimodal data analysis to multi-dimensional evaluation and report generation, no deep human intervention is required throughout the process, greatly reducing manpower and time costs. It can efficiently cover a large-scale campus group and meet the routine needs of teachers' self-reflection, school teaching and research supervision and regional education quality assessment.

[0046] 3. Multimodal data collaborative analysis to fully capture diverse classroom information: Integrating multi-dimensional data such as audio, text, and video, and using technologies such as audio extraction, speech-to-text conversion, target detection, and speech activity recognition, we comprehensively explore key information in classroom teaching, such as language interaction, use of teaching aids, and teacher and student participation. This avoids the one-sidedness of information caused by single data analysis and provides more comprehensive data support for evaluation.

[0047] 4. Construct a systematic and multi-dimensional evaluation system to significantly improve the scientific nature and comprehensiveness of the evaluation: Focusing on the core teaching links, eight dimensions and detailed evaluation points are set up, including teaching objectives, teaching content, and teaching methods. It is adapted to the exclusive assessment needs of nine disciplines, covering not only surface indicators such as classroom process and language expression, but also in-depth dimensions such as teaching interaction, thinking cultivation, and classroom management, so as to achieve a panoramic evaluation of teaching skills.

[0048] 5. Provide targeted improvement guidance to enhance the practical value of evaluation: Each evaluation dimension provides quantitative scores and targeted teaching suggestions simultaneously. The comprehensive evaluation clearly identifies strengths and areas for improvement, avoiding the evaluation results from becoming merely a formality. It provides teachers with actionable guidelines for accurately identifying problems and optimizing teaching strategies, helping to improve teaching skills efficiently.

[0049] 6. Activate the value of classroom data to support educational research and decision-making: Transform classroom audio and video data into structured and quantifiable evaluation indicators and decision-making basis, which not only serves individual teacher improvement, but also provides data support for school teaching and research management and regional education quality analysis, promotes the transformation of education evaluation from experience-driven to data-driven, and helps the digital and intelligent development of education.

[0050] 7. Adaptable to specific campus scenarios, with strong applicability and scalability: The evaluation system is designed closely with school curriculum standards and students' cognitive characteristics, supports diverse teaching scenarios and nine major subjects, and its technical architecture is compatible with mainstream audio and video formats and AI models. It can flexibly expand evaluation dimensions and indicators according to the needs of teaching reform and has continuous iteration capabilities. Attached Figure Description

[0051] Figure 1 This is a flowchart of a multimodal analysis and evaluation method for classroom teaching skills videos according to the present invention;

[0052] Figure 2 This is an example image of a classroom teaching skills video provided in the embodiment. Detailed Implementation

[0053] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings and examples. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0054] Example:

[0055] Please see Figure 1 This invention provides a multimodal analysis and evaluation method for classroom teaching skills videos, comprising the following steps:

[0056] S1. Video Upload and Metadata Entry: Users upload classroom teaching skills lecture videos to the system and add metadata information such as the course name, teaching topic, school name, location, and grade level corresponding to the video. To ensure the accuracy of the analysis, the video should ideally have good clarity, and the instructor should always be within the shooting range. The image area of ​​the submitted video should be as follows: Figure 2 As shown;

[0057] S2. Automated preprocessing of video: Extract audio files from the video, transcribe them into text using a model, and then standardize the text and subtitle files to generate text and subtitle files. Simultaneously generate word cloud images and calculate video duration and language density.

[0058] S3. Analysis of teaching objectives and innovative features: Analyze the clarity of teaching objectives and the innovative features of the teaching design in the teaching process;

[0059] S4. Analyze teaching content and methods: Analyze the scientific accuracy, appropriateness and relevance of the teaching content, and the technical integration, diversity and appropriateness of the teaching methods in the teaching process.

[0060] S5. Analyze Teacher Behavior: Analyze the teacher's classroom management skills and interactive guidance during the teaching process;

[0061] S6. Analyze student performance: Analyze the depth and creativity of students' thinking, participation and concentration during the teaching process;

[0062] S7. Analyze the classroom atmosphere: Analyze the democracy and inclusiveness of the classroom atmosphere and the students' enthusiasm for learning during the teaching process;

[0063] S8. Analyze teaching evaluation: Analyze the evaluation methods and feedback effectiveness of the teaching process.

[0064] S9. Provide a comprehensive evaluation: After completing the analysis in parallel from steps S3 to S10, provide an overall evaluation based on the analysis results of each dimension;

[0065] S10. Generate an analysis report: Summarize and organize the analysis results of each dimension obtained in steps S3 to S11 to generate an analysis report.

[0066] Preferably, step S2, automated video preprocessing, specifically includes the following steps:

[0067] S2-1. Use the audio and video processing tool ffmpeg to extract the audio from the lecture video to obtain an independent audio file; use the Faster-Whisper speech recognition model to process the audio file into speech and generate initial text data.

[0068] S2-2. Standardize and clean the initial text data and automatically add punctuation marks that conform to grammatical rules to generate standardized subtitle files and text files; generate a visual word cloud based on the text files using a word cloud algorithm;

[0069] S2-3. Record the total duration of the lecture video, perform word count per minute on the text file, and form a quantitative language density index.

[0070] Preferably, in step S3, the analysis of teaching objectives and distinctive innovations is conducted through the following two aspects:

[0071] S3-1. Based on the multimodal data preprocessed in step S2, for the dimension of "clarity of teaching objectives," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the teaching objectives conform to the curriculum standards, whether they conform to students' cognitive levels, whether they are observable, and whether they are achievable." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the performance of teaching objectives corresponding to matching cognition, conforming to standards, and goal completion, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the quantitative scores of the sub-items to obtain a comprehensive score for the dimension of "clarity of teaching objectives." The two-level evaluation model includes requesting the large language model to generate a performance description of "clarity of teaching objectives" based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of approximately 120 words based on the comprehensive score and the content detection conclusion.

[0072] S3-2. Based on the multimodal data preprocessed in step S2, for the dimension of "innovative instructional design," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether there are unique teaching ideas, activity designs, or resource integration methods." Specifically, the multi-dimensional scoring mechanism involves calling a large language model to professionally score the uniqueness of resource integration, teaching ideas, and activity designs, and outputting quantitative scores for each sub-item. Simultaneously, based on preset rules, a weighted algorithm is used to fuse the two types of scores to obtain a comprehensive score for the "innovative instructional design" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection function based on the teaching text to describe the performance of "innovative instructional design," and generating a comprehensive review of approximately 120 words of teaching suggestions based on the score evaluation results and content detection conclusions.

[0073] Preferably, in step S4, the analysis of teaching content and teaching methods is conducted through the following four aspects:

[0074] S4-1. Scientific Rigor and Accuracy: Based on the multimodal data preprocessed in step S2, a comprehensive analysis is conducted on the "scientific rigor and accuracy" dimension, focusing on the evaluation criteria of "whether the knowledge explanation is accurate and error-free, whether the logic is rigorous, and whether it reflects the forefront of the discipline." A multi-dimensional scoring mechanism and a two-level evaluation model are used for comprehensive analysis. The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching content performance corresponding to the accuracy of knowledge, the rigor of logic, and the relevance to the forefront of the discipline, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "scientific rigor and accuracy" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "scientific rigor and accuracy" performance based on the teaching text, and generating a comprehensive evaluation of approximately 120 words of teaching suggestions based on the comprehensive score and the content detection conclusions.

[0075] S4-2, Appropriateness and Relevance: Based on the multimodal data preprocessed in step S2, for the dimension of "appropriateness and relevance," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the content difficulty matches the students' level, whether it connects with prior and subsequent knowledge, and whether it relates to real-life situations." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching content's performance in terms of its relevance to real-world situations, appropriate difficulty, and knowledge connection, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "appropriateness and relevance" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "appropriateness and relevance" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of approximately 120 words based on the comprehensive score and content detection conclusions.

[0076] S4-3, Technology Integration Degree: Based on the multimodal data preprocessed in step S2, intelligent analysis is conducted on the "technology integration degree" dimension, focusing on the evaluation criteria of "whether information technology effectively assists teaching and whether there is excessive reliance on it." Classroom videos are processed using computer vision technology, with VideoCapture used to read basic information such as video frame rate and total number of frames, employing a 2-second frame sampling strategy. A pre-trained YOLO object detection model is loaded to detect and identify teaching items in the video frames. The detection results are quantitatively analyzed, statistically analyzing the frequency of appearance, duration, and continuous use of various teaching aids, calculating the usage ratio, and generating a comprehensive score based on preset evaluation criteria. Simultaneously, a large language model is requested to generate a comprehensive review and improvement suggestions for teaching based on the detection data, statistical results, and classroom scene information regarding the "technology integration degree."

[0077] S4-4. Diversity and Appropriateness: Based on the multimodal data preprocessed in step S2, for the dimension of "diversity and appropriateness," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the methods are flexible and diverse, and whether they conform to the characteristics of the subject and the teaching content." Specifically, the multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching methods corresponding to flexibility and diversity and conformity to requirements, and outputting quantitative scores for each sub-item. Simultaneously, based on preset rules, a weighted algorithm is used to fuse the two types of scores to obtain a comprehensive score for the "diversity and appropriateness" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "diversity and appropriateness" performance based on the teaching text, and generating a comprehensive review of approximately 120 words of teaching suggestions based on the score evaluation results and content detection conclusions.

[0078] Preferably, in step S5, the analysis of teacher behavior is conducted through the following two aspects:

[0079] S5-1, Classroom Management Ability: Based on the multimodal data preprocessed in step S2, for the dimension of "classroom management ability," deep analysis of classroom audio is performed using audio processing technology. The silence function of the pydub library is used to detect silent segments in the audio, calculate the percentage of silent duration, and generate a basic classroom management score by combining a preset silence percentage threshold and evaluation rules for the rationality of silent period distribution. The teacher's classroom language content is extracted from the preprocessed text data, and the accuracy of their use of professional terminology, clarity of instructions, and stability of speech rate are analyzed. Based on the silence analysis results, language expression statistics, and a large language model for classroom scene information requests, a comprehensive review and improvement suggestions for teaching recommendations regarding "classroom management ability" are generated.

[0080] S5-2, Interactive Guidance: Based on the multimodal data preprocessed in step S2, for the "interactive guidance" dimension, a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether an interactive scenario is created and whether students can be effectively guided to think, ask questions, and discuss." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teacher's behavioral performance corresponding to the interactive scenario, teaching questions, teaching reflections, and teaching discussions, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "interactive guidance" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "interactive guidance" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of approximately 120 words based on the comprehensive score and the content detection conclusions.

[0081] Preferably, in step S6, the analysis of student performance is conducted from the following two aspects:

[0082] S6-1, Depth of Thinking and Innovation: Based on the multimodal data preprocessed in step S2, a comprehensive analysis is conducted on the dimension of "depth of thinking and innovation," focusing on the evaluation criteria of "whether students can raise questions, express unique insights, and apply knowledge to solve complex problems." A multi-dimensional scoring mechanism and a two-level evaluation model are used. The multi-dimensional scoring mechanism involves calling a large language model to professionally score students' knowledge application and complex problem-solving abilities, as well as their questioning and unique insight abilities, and outputting quantitative scores for each sub-item. A weighted algorithm is then used to fuse these scores based on preset rules to obtain a comprehensive score for the "depth of thinking and innovation" dimension. The two-level evaluation model includes requesting the large language model to generate a content description of the "depth of thinking and innovation" performance based on the teaching text, and generating a comprehensive evaluation of approximately 120 words of teaching suggestions based on the comprehensive score and the content detection conclusions.

[0083] S6-2, Engagement and Attention: Based on the multimodal data preprocessed in step S2, for the "engagement and attention" dimension, the deployed 3D-Speaker speech processing technology is invoked to perform speech activity detection to identify speech segments and filter out silent parts. Speaker embeddings are extracted from the speech segments and clustered to identify different speakers. The number of speakers and speaking duration information are counted. Based on preset rules, a comprehensive score for the "engagement and attention" dimension is obtained. The content detection results obtained by the large language model based on speech processing technology are requested. Combining the comprehensive score and content detection conclusion, a comprehensive review of teaching suggestions of about 120 words is generated.

[0084] Preferably, step S7, analyzing the classroom atmosphere, specifically includes the following steps:

[0085] S7-1, Democracy and Inclusivity: Based on the multimodal data preprocessed in step S2, a comprehensive analysis is conducted on the "democracy and inclusiveness" dimension, focusing on the evaluation criteria of "whether the classroom reflects equal dialogue, whether it accommodates different viewpoints, and whether the teacher-student relationship is harmonious." A multi-dimensional scoring mechanism and a two-level evaluation model are used. The multi-dimensional scoring mechanism involves calling a large language model to professionally score the classroom atmosphere corresponding to teacher-student relationships, equal dialogue, and tolerance of viewpoints, and outputting quantitative scores for each sub-item. A weighted algorithm is then used to fuse these scores based on preset rules to obtain a comprehensive score for the "democracy and inclusiveness" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection function based on the teaching text to describe the performance of "democracy and inclusiveness," and generating a comprehensive evaluation of approximately 120 words of teaching suggestions based on the comprehensive score and the content detection conclusions.

[0086] S7-2, Learning Motivation: Based on the multimodal data preprocessed in step S2, for the dimension of "learning motivation," a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether students demonstrate a thirst for knowledge and curiosity, and whether they are willing to cooperate and share." The multi-dimensional scoring mechanism involves calling a large language model to analyze the text and identify and statistically analyze all "problem" sentences for professional scoring. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "learning motivation" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "learning motivation" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of about 120 words based on the comprehensive score and the content detection conclusion.

[0087] Preferably, step S8, analyzing teaching evaluation, specifically includes the following steps:

[0088] S8-1, Evaluation Method: Based on the multimodal data preprocessed in step S2, for the "evaluation method" dimension, focusing on the evaluation point of "whether to adopt multi-dimensional evaluation", the corresponding subject evaluation method is selected from the nine preset subject evaluation methods according to the teaching theme. The deployed RAGFlow service is called to conduct professional evaluation. At the same time, combined with preset keywords and rules, a weighted algorithm is used to obtain the comprehensive score of the "evaluation method" dimension. The two-level evaluation mode includes requesting the large language model to generate a performance description of the "evaluation method" based on the teaching text for content detection, and generating a comprehensive evaluation of teaching suggestions of about 120 words based on the comprehensive score and content detection conclusion.

[0089] S8-2, Feedback Effectiveness: Based on the multimodal data preprocessed in step S2, for the "Feedback Effectiveness" dimension, a comprehensive analysis is conducted using a multi-dimensional scoring mechanism and a two-level evaluation model, focusing on the evaluation criteria of "whether the feedback is timely and specific, and whether it can guide students to improve their learning." The multi-dimensional scoring mechanism involves calling a large language model to professionally score the teaching performance corresponding to timely and specific feedback, and outputting quantitative scores for each sub-item. Based on preset rules, a weighted algorithm is used to fuse the scores to obtain a comprehensive score for the "Feedback Effectiveness" dimension. The two-level evaluation model includes requesting the large language model to generate a content detection description of the "Feedback Effectiveness" performance based on the teaching text, and generating a comprehensive evaluation of teaching suggestions of approximately 120 words based on the comprehensive score and the content detection conclusion.

[0090] Preferably, in step S9, the specific operation process for providing a comprehensive evaluation is as follows: after completing the analysis in parallel based on steps S3 to S8, the evaluation criteria, evaluation scores, suggestions, and analysis results of each dimension are summarized, and the large language model is requested to conduct a comprehensive quantitative evaluation and improvement suggestions for the teaching skills of the classroom.

[0091] Preferably, the specific operation process of generating the analysis report in step S10 is as follows: the analysis results of each dimension from steps S3 to S9 are classified, summarized and logically integrated. The analysis results include the quantitative scores of each dimension sub-item, the comprehensive score, the content detection conclusion, and the specific teaching suggestions; at the same time, the comprehensive quantitative evaluation conclusion of step S9 is integrated and presented in the form of a structured analysis report. The report includes video basic metadata, multi-dimensional evaluation details from different perspectives, and a comprehensive conclusion module, and simultaneously integrates the visualized word cloud and language density index quantitative data generated in step S2.

Claims

1. A method for multi-modal analysis and evaluation of classroom teaching skills video, characterized in that, Comprising the following steps: S1, video uploading and metadata filling: the user uploads a classroom teaching skill lecture video to the system, and adds the course name, teaching topic, school name, location, grade metadata information corresponding to the video in the system; S2, automatic pre-processing video: extracting the audio file in the video, normalizing the text and subtitle file after model transcription, synchronously generating word cloud diagram, and counting video length and language density; S3, analyzing teaching objectives and innovative features: for the teaching process, the target clarity of teaching objectives and the teaching design innovation of innovative features are analyzed; S4, analyzing teaching content and teaching method: for the teaching process, the scientificity and accuracy, appropriateness and relevance of teaching content, and the technical integration, diversity and appropriateness of teaching method are analyzed; S5, analyzing teacher behavior: for the teaching process, the classroom control ability and interaction guidance of teacher behavior are analyzed; S6, analyzing student performance: for the teaching process, the thinking depth and innovation of student performance, participation and concentration are analyzed; S7, analyzing classroom atmosphere: for the teaching process, the democracy and inclusiveness of classroom atmosphere, learning enthusiasm are analyzed; S8, analyzing teaching evaluation: for the teaching process, the evaluation method and feedback effectiveness of teaching evaluation are analyzed; S9, giving a comprehensive evaluation: after completing the analysis in steps S3 to S10 in parallel, according to the analysis results of each dimension, the overall evaluation is given; S10, generating an analysis report: the analysis results of each dimension obtained in steps S3 to S11 are summarized and arranged to generate an analysis report.

2. The method for multi-modal analysis and evaluation of classroom teaching skill video according to claim 1, characterized in that, The step S2 of automatically pre-processing the video specifically comprises the following steps: S2-1: calling the audio and video processing tool ffmpeg to extract the audio in the lecture video to obtain an independent audio file; performing voice-to-text processing on the audio file through the faster-whisper speech recognition model to generate initial text data; S2-2: normalizing and cleaning the initial text data and automatically adding punctuation symbols conforming to grammatical rules to generate standardized subtitle files and text files; generating a visual word cloud diagram based on the text file through a word cloud algorithm; S2-3: recording the total length of the lecture video, and performing per-minute word count on the text file to form a quantitative language density index.

3. The method for multi-modal analysis and evaluation of classroom teaching skill video as claimed in claim 1 wherein, In the step S3 of analyzing teaching objectives and innovative features, the analysis is performed through the following two aspects: S3-1: Based on the multi-modal data pre-processed in step S2, for the "teaching goal clarity" dimension, around the evaluation points of "whether the teaching goal meets the curriculum standards, whether it meets the students' cognitive level, whether it has observability, and whether it has attainability", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis; the multi-dimensional scoring mechanism is to call a large language model to score the teaching goal performance corresponding to the matching cognition, standard fit, and target completion degree, and output the quantitative scores of each sub-item, and based on the preset rules, the weighted algorithm is used to fuse the sub-item quantitative scores to obtain the comprehensive score of the "teaching goal clarity" dimension; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "teaching goal clarity" based on teaching text, and generating a 120-word or so comprehensive review of teaching suggestions based on the comprehensive score and content detection conclusion; S3-2: Based on the multi-modal data pre-processed in step S2, for the "teaching design innovation" dimension, around the evaluation points of "whether there is a unique teaching idea, activity design or resource integration method", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis, wherein the multi-dimensional scoring mechanism specifically calls a large language model to score the unique characteristics of resource integration, teaching idea, and activity design, and outputs the quantitative scores of each sub-item, and based on the preset rules, the weighted algorithm is used to fuse the two types of scores to obtain the comprehensive score of the "teaching design innovation" dimension; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "teaching design innovation" based on teaching text, and generating a 120-word or so comprehensive review of teaching suggestions based on the score evaluation result and the content detection conclusion.

4. The method for multi-modal analysis and evaluation of classroom teaching skills video as claimed in claim 1 wherein, In the step S4 of analyzing the teaching content and teaching method, analysis is performed in the following four aspects: S4-1: Scientificity and accuracy: based on the multi-modal data pre-processed in step S2, for the "scientificity and accuracy" dimension, around the evaluation points of "whether the knowledge explanation is accurate and correct, whether the logic is rigorous, and whether it can reflect the discipline frontier", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis; the multi-dimensional scoring mechanism is to call a large language model to score the teaching content performance corresponding to the knowledge accuracy, logic rigor, and discipline frontier correlation, and output the quantitative scores of each sub-item, and based on the preset rules, the weighted algorithm is used to fuse to obtain the comprehensive score of the "scientificity and accuracy" dimension; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "scientificity and accuracy" based on teaching text, and generating a 120-word or so comprehensive review of teaching suggestions based on the comprehensive score and content detection conclusion; S4-2: Relevance and relevance: based on the pre-processed multi-modal data in step S2, for the "relevance and relevance" dimension, around the evaluation points of "whether the content difficulty is consistent with the student level, whether the knowledge is connected, and whether it is related to the actual life", a multi-dimensional scoring mechanism and a two-level evaluation mode are used to carry out comprehensive analysis; the multi-dimensional scoring mechanism is to call a large language model to score the teaching content performance corresponding to the combination of reality, appropriate difficulty and knowledge connection, and output the quantitative scores of each sub-item, and the comprehensive score of the "relevance and relevance" dimension is obtained by weighting algorithm based on the preset rules; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "relevance and relevance" based on teaching text, and generating a comprehensive review of about 120 words of teaching suggestions based on the comprehensive score and content detection conclusion; S4-3: Technology integration: based on the pre-processed multi-modal data in step S2, for the "technology integration" dimension, around the evaluation points of "whether information technology effectively assists teaching, whether it is over-reliant", intelligent analysis is carried out; the classroom video is processed by computer vision technology, the VideoCapture is used to read the video frame rate and total frame number basic information, and the frame sampling strategy of 2 seconds interval is adopted; the pre-trained YOLO target detection model is loaded to detect and identify the teaching materials in the video frame; the detection results are quantitatively analyzed, the appearance frequency, duration and continuous use of each type of teaching aid are counted, the use ratio is calculated and the comprehensive score is generated combined with the preset evaluation standard; at the same time, a large language model is requested to generate a comprehensive review of teaching suggestions and improvement suggestions for "technology integration" based on detection data, statistical results and classroom scene information; S4-4: Diversity and appropriateness: based on the pre-processed multi-modal data in step S2, for the "diversity and appropriateness" dimension, around the evaluation points of "whether the method is flexible and diverse, whether it is consistent with the characteristics of the subject and the teaching content", a multi-dimensional scoring mechanism and a two-level evaluation mode are used to carry out comprehensive analysis, wherein the multi-dimensional scoring mechanism specifically calls a large language model to score the teaching method performance corresponding to the method flexible and diverse and the method meets the requirements and outputs the quantitative scores of each sub-item, and at the same time, based on the preset rules, the two types of scores are fused to obtain the comprehensive score of the "diversity and appropriateness" dimension by weighting algorithm; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "diversity and appropriateness" based on teaching text, and generating a comprehensive review of about 120 words of teaching suggestions based on the score evaluation result and the content detection conclusion.

5. The method for multi-modal analysis and evaluation of classroom teaching skill video as claimed in claim 1 wherein, In the step S5 of analyzing the teacher's behavior, the following two aspects are analyzed: S5-1: Classroom regulation ability: Based on the multi-modal data pre-processed in step S2, for the "classroom regulation ability" dimension, the audio processing technology is used to deeply analyze the classroom audio, the silence in the audio is detected by using the silence of the pydub library, the proportion of the silence time is calculated, and the preset silence proportion threshold and the silence period distribution rationality evaluation rule are combined to generate the classroom management basic score; the teacher's classroom language content is extracted from the pre-processed text data, and the accuracy of the professional term use, the clarity of the instruction and the stability of the language speed are analyzed; Based on the silence analysis result, the language expression statistical data and the classroom scene information, the large language model is requested to generate the teaching suggestion comprehensive comment and the improvement suggestion for "classroom regulation ability"; S5-2: Interaction guidance: Based on the multi-modal data pre-processed in step S2, for the "interaction guidance" dimension, around the evaluation points of "whether to create an interactive situation, and whether to effectively guide students to think, ask questions and discuss", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis; the multi-dimensional scoring mechanism is to call the large language model to score the teacher's behavior performance in interactive situation, teaching questions, teaching reflection and teaching discussion, and output the quantitative scores of each sub-item, and the comprehensive score of "interaction guidance" dimension is obtained by weighting algorithm based on the preset rules; the two-level evaluation mode includes requesting the large language model to generate performance description content detection for "interaction guidance" based on the teaching text, and generating a teaching suggestion comprehensive comment of about 120 words based on the comprehensive score and the content detection conclusion.

6. The method for multi-modal analysis and evaluation of classroom instruction skills video as claimed in claim 1 wherein, In the step S6, the student performance is analyzed by the following two aspects: S6-1: Thinking depth and innovation: Based on the multi-modal data pre-processed in step S2, for the "thinking depth and innovation" dimension, around the evaluation points of "whether students can ask questions and express unique insights, and whether they can use knowledge to solve complex problems", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis; the multi-dimensional scoring mechanism is to call the large language model to score the student's performance in knowledge application and complex problem solving ability, and the student's questioning and unique insight ability, and output the quantitative scores of each sub-item, and the comprehensive score of "thinking depth and innovation" dimension is obtained by weighting algorithm based on the preset rules; the two-level evaluation mode includes requesting the large language model to generate performance description content detection for "thinking depth and innovation" based on the teaching text, and generating a teaching suggestion comprehensive comment of about 120 words based on the comprehensive score and the content detection conclusion; S6-2: Participation and concentration: Based on the multi-modal data pre-processed in step S2, for the "participation and concentration" dimension, the 3D-Speaker speech processing technology is called to perform speech activity detection to identify speech segments and filter silence, extract speaker embeddings from the speech segments and cluster the embeddings to identify different speakers, count the number of speakers and the speaking time information, and obtain the comprehensive score of "participation and concentration" dimension based on the preset rules; The request large language model obtains the content detection result based on the voice processing technology, combines the comprehensive score and the content detection conclusion, and generates a teaching suggestion comprehensive review of about 120 words.

7. The method for multi-modal analysis and evaluation of classroom instruction skills video as claimed in claim 1 wherein, The step S7 analyzes the classroom atmosphere, specifically including the following steps: S7-1: Democracy and inclusiveness: based on the multi-modal data pre-processed in step S2, for the "democracy and inclusiveness" dimension, around the evaluation points of "whether the classroom embodies equal dialogue, whether it is inclusive of different opinions, and whether the teacher-student relationship is harmonious", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis; the multi-dimensional scoring mechanism is to call a large language model to score the corresponding classroom atmosphere of teacher-student relationship, equal dialogue, and opinion inclusiveness, and output the quantitative scores of each sub-item, and based on the preset rules, the comprehensive score of the "democracy and inclusiveness" dimension is obtained by weighting algorithm fusion; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "democracy and inclusiveness" based on the teaching text, and generating a teaching suggestion comprehensive review of about 120 words based on the comprehensive score and the content detection conclusion; S7-2: Learning enthusiasm: based on the multi-modal data pre-processed in step S2, for the "learning enthusiasm" dimension, around the evaluation points of "whether students show curiosity, whether they are willing to cooperate and share", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis; the multi-dimensional scoring mechanism is to call a large language model to analyze the text and identify all "question" sentences for professional scoring, and based on the preset rules, the comprehensive score of the "learning enthusiasm" dimension is obtained by weighting algorithm fusion; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "learning enthusiasm" based on the teaching text, and generating a teaching suggestion comprehensive review of about 120 words based on the comprehensive score and the content detection conclusion.

8. The method for multi-modal analysis and evaluation of classroom instruction skills video as claimed in claim 1 wherein, The step S8 analyzes the teaching evaluation, specifically including the following steps: S8-1: Evaluation method: based on the multi-modal data pre-processed in step S2, for the "evaluation method" dimension, around the evaluation point of "whether to use multi-element evaluation", according to the teaching theme, select the corresponding subject evaluation method from the preset nine disciplines evaluation methods, call the deployed RAGFlow service for professional evaluation, and at the same time, combine the preset keywords and rules, and get the comprehensive score of the "evaluation method" dimension by weighting algorithm fusion; the two-level evaluation mode includes requesting a large language model to generate performance description content detection for "evaluation method" based on the teaching text, and generating a teaching suggestion comprehensive review of about 120 words based on the comprehensive score and the content detection conclusion. S8-2: Feedback effectiveness: based on the pre-processed multi-modal data in step S2, for the dimension of "feedback effectiveness", around the evaluation points of "whether the feedback is timely, specific, and can guide students to improve learning", a multi-dimensional scoring mechanism and a two-level evaluation mode are used for comprehensive analysis; the multi-dimensional scoring mechanism is to call a large language model to score the feedback and the corresponding teaching performance, and output the quantitative scores of each sub-item, and based on the preset rules, the comprehensive score of the "feedback effectiveness" dimension is obtained by weighted algorithm; the two-level evaluation mode includes requesting a large language model to generate performance description content for "feedback effectiveness" based on teaching text detection, and generating a 120-word teaching suggestion comprehensive review based on the comprehensive score and content detection conclusion.

9. The method for multi-modal analysis and evaluation of classroom instruction skills video as claimed in claim 1 wherein, The specific operation process of the step S9 is: based on the analysis completed in steps S3 to S8, the evaluation criteria, evaluation scores, suggestion content, and analysis results of each dimension are summarized, and a large language model is requested to make a comprehensive quantitative evaluation and improvement suggestion on the teaching skills of the classroom.

10. The method for multi-modal analysis and evaluation of classroom instruction skills video as claimed in claim 1 wherein, The specific operation process of the step S10 is: classify and integrate the analysis results of each dimension in steps S3 to S9, including the quantitative scores of each dimension sub-item, the comprehensive score, the content detection conclusion, and the special teaching suggestion; At the same time, the comprehensive quantitative evaluation conclusion of step S9 is integrated, and a structured analysis report is generated, which includes video basic metadata, multi-dimensional evaluation details from different perspectives, a comprehensive conclusion module, and is synchronized with the visual word cloud diagram and language density index quantitative data generated in step S2.