Classroom teaching analysis method based on large language model and multi-modal information fusion
Through the classroom teaching analysis method that integrates large language model and multimodal information, the accuracy of classroom speech transcription and behavior analysis is solved, efficient multimodal data fusion and behavior recognition are achieved, and the accuracy and intelligent application of teaching analysis are improved.
Patent Information
- Application Number
- CN202510461604.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The existing classroom speech transcription technology has high error rates in complex environments and lacks targeted optimization. The behavioral analysis technology fails to make full use of audio and video information in multimodal data fusion, resulting in insufficient reliability and integrity of the analysis results.
The method of fusion of large language model and multimodal information is adopted to generate an N-best candidate list through ASR, filter high confidence candidates, and multimodal period division and role inference analysis are performed by combining visual model and visual large model, and multimodal data analysis is performed by combining classroom video frame extraction to obtain behavioral time periods.
It significantly improves the quality of speech transcription and behavior recognition accuracy in classroom teaching scenarios, provides more comprehensive data support, and provides a foundation for teaching analysis and intelligent feedback.
Smart Images

Figure CN120337145A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of classroom teaching analysis, and specifically relates to a classroom teaching analysis method based on the fusion of large language models and multimodal information. Background Art
[0002] With the development of information technology and artificial intelligence, the application of artificial intelligence in the field of education has become increasingly widespread. However, in actual classroom scenarios, the results of speech transcription are often affected by factors such as environmental noise, unclear sound collection, a large number of speakers, and dialect accents, resulting in problems such as a high error rate and incoherent semantics in the transcribed content. In addition, existing speech transcription technologies lack targeted optimization in specific classroom scenarios and cannot effectively cope with the interference of the above complex factors, limiting the actual application effect of classroom speech transcription technology.
[0003] On the other hand, behavior analysis in the classroom (such as teachers' patrols, writing on the blackboard, lectures, and students' reading, writing, raising hands, listening to lectures, etc.) is of great significance for understanding the teaching process and evaluating teaching quality. However, existing behavior analysis technologies mostly rely on a single visual model and lack adaptability and accuracy in complex classroom scenarios. Especially in the aspect of multimodal data fusion, existing technologies have not fully utilized the complementarity of audio and video information, resulting in insufficient reliability and integrity of the analysis results. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a classroom teaching analysis method based on the fusion of large language models and multimodal information, which improves the accurate analysis ability of classroom teaching.
[0005] The present invention adopts the following technical solutions to achieve the above purpose. The present invention provides a classroom teaching analysis method based on the fusion of large language models and multimodal information, including:
[0006] S1. Generate an N-best candidate list;
[0007] S2. Correct the N-best candidate list to obtain a corrected SRT;
[0008] S3. Multimodal time period division, divide the corrected SRT document into effective speech transcription time periods, noisy time periods, and quiet time periods;
[0009] S4. Multimodal role inference analysis;
[0010] S5. Perform frame extraction on the classroom video to obtain the time periods of student behavior and teacher behavior;
[0011] S6. Use the time periods of student behavior and teacher behavior as an aid for SRT analysis to perform multimodal data analysis.
[0012] Furthermore, step S1 specifically includes:
[0013] Process the classroom audio through an ASR (Automatic Speech Recognition) model to output an N-best candidate list containing timestamp, speech content, and speaker information. Each candidate list contains a confidence score and time alignment information. Filter out the low-confidence lists among the N-best candidates by setting a confidence threshold, and only retain the high-confidence candidates.
[0014] Furthermore, step S2 specifically includes:
[0015] Adopt the zero-shot or few-shot learning ability of the generative large language model, and guide the large language model to complete the task of correcting the filtered N-best candidate list through a constraint rule prompt template to obtain the corrected SRT. The input of the prompt template is the filtered N-best candidate text, and the output is the integrated optimal transcription result.
[0016] Furthermore, the noisy period recognition includes:
[0017] Use the large language model to detect the noisy period, identify students' free discussions, readings, or noises, and mark the noisy period;
[0018] The quiet period includes:
[0019] Measure the time interval between the end time of each SRT record and the start time of the next record. If the interval exceeds the set time, mark it as a quiet period and record the start and end points of this time period;
[0020] The effective speech transcription period includes:
[0021] After excluding the noisy and quiet periods, mark the remaining periods as the effective speech transcription period.
[0022] Furthermore, step S4 specifically includes:
[0023] Multimodal role reasoning includes S-T analysis of the effective speech transcription period, S-T analysis of the noisy period, and S-T analysis of the quiet period;
[0024] S-T analysis of the effective speech transcription period:
[0025] For the effective speech transcription period, directly perform S-T determination according to the speaker's identity. Mark the teacher's speech as T and the student's speech as S;
[0026] S-T analysis of the noisy period and the quiet period:
[0027] Collect multiple classroom scene pictures and label the behaviors of students and teachers in the classroom;
[0028] Train the YOLOv12 model with the labeled pictures;
[0029] Use the trained YOLOv12 model to detect the behaviors of teachers and students in the classroom scene, and then analyze and judge the S-T roles based on the detected behaviors. Determine whether it belongs to students or teachers at the current time period according to the set priority rules.
[0030] Furthermore, the priority rules are as follows:
[0031] Teacher behavior priority: Lecturing > Writing on the blackboard > Patrolling > Teacher-student interaction > Others;
[0032] Student behavior priority: Discussing > Reading aloud > Writing > Reading > Raising hands > Responding > Listening > Others.
[0033] Furthermore, step S5 specifically includes:
[0034] Perform frame extraction on the classroom video to obtain classroom scene pictures containing the behaviors of teachers and students, and label the teacher behaviors and student behaviors separately. Teacher behaviors include patrolling, teacher-student interaction, lecturing, writing on the blackboard, and others. Student behaviors include raising hands, responding, reading aloud, reading, writing, discussing, listening, and others;
[0035] Train the visual model and the visual large model with the labeled picture data, and optimize the behavior detection of the visual large model by designing prompt words;
[0036] Detect the behaviors in the classroom scene video frames with the trained visual model and visual large model to obtain the time periods of student and teacher behaviors.
[0037] Furthermore, designing prompt words specifically includes:
[0038] Define the characteristics of teacher behaviors and student behaviors in the form of natural language descriptions;
[0039] Provide context information for behavior detection to enhance the model's ability to recognize behaviors;
[0040] Use prompt words to guide the visual large model to recognize and distinguish the behavior patterns of teachers and students.
[0041] Furthermore, step S6 specifically includes:
[0042] Use the time periods of student and teacher behaviors as an aid for SRT analysis, and combine audio transcription and visual behavior detection results to perform multimodal data analysis, so as to accurately distinguish and label the role identities of teachers and students.
[0043] The beneficial effects of the present invention are as follows:
[0044] Through the synergistic effects of automatic speech recognition, visual behavior analysis, and video frame extraction and behavior detection, the present invention significantly improves the accuracy and efficiency of S-T analysis in classroom teaching scenarios. This method not only optimizes the quality of classroom speech transcription but also further enhances the accurate recognition of teachers' and students' behaviors through video frame extraction and behavior annotation. In addition, the behavior detection technology combining traditional visual models and large visual models provides more comprehensive data support for S-T role analysis. Description of the Drawings
[0045] Figure 1 It is a flowchart of a classroom teaching analysis method based on the fusion of large language models and multi-modal information provided by an embodiment of the present invention. Detailed Embodiments
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0047] The present invention provides a classroom teaching analysis method based on the fusion of large language models and multi-modal information, as Figure 1 shown, specifically including:
[0048] S1. Generate an N-best candidate list;
[0049] Process the classroom audio through the ASR model and output an N-best candidate list containing timestamp, speech content, and speaker information, where each candidate hypothesis contains a confidence score and time alignment information.
[0050] In a complex classroom environment, speech transcription is affected by various factors such as environmental noise, unclear sound collection, a large number of speakers, and dialect accents. A single transcription result often has a high error rate. The role of the N-best candidate list is to provide multiple possible transcription hypotheses for the subsequent large language model. The large language model can comprehensively analyze and judge from different hypotheses based on these multiple candidate inputs, fully utilize the effective information in each candidate hypothesis, thereby improving the accuracy of the final transcription result. It can avoid deviations in subsequent analysis caused by inaccurate single transcription results and provide richer and more reliable basic data for subsequent precise correction.
[0051] To reduce the interference of low-confidence content on the large model, the present invention screens the low-confidence list in the N-best candidates by setting a threshold. Specifically, a confidence threshold (for example, 0.7) is set, and only the candidate hypotheses with a confidence greater than or equal to this threshold are retained.
[0052] The ASR model is set to give 3 candidate hypotheses for each audio segment (N = 3), and provides a confidence score and time alignment information for each candidate hypothesis. The following is an example of an N-best candidate list (assuming N = 3):
[0053] Table 1 N-best candidate list
[0054]
[0055] S2. Revise the N-best candidate list to obtain the revised SRT;
[0056] Utilize the zero-shot or few-shot learning ability of the generative large language model. Through the constraint rule prompt template, guide the large language model to complete the task of revising the filtered N-best candidate list, and obtain the revised SRT. The input of the prompt template is the filtered N-best candidate text, and the output is the integrated optimal transcription result.
[0057] The prompt template is as follows:
[0058] "Please generate the most accurate transcription from these ASR hypotheses:
[0059] <hypothesis1>{Text 1}… <hypothesisn>{Text N}< / hypothesisn>
[0060] Rule: Disable new words, ignore punctuation, output without explanation, etc.
[0061] The SRT corrected by the large model rules: After being corrected by the large language model constraints, an SRT document containing start time, end time, speech content, and speaker is obtained. The example is as follows:
[0062] Table 2 Voice transcription results of classroom teaching videos
[0063]
[0064]
[0065] S3. Multimodal time period division, dividing the corrected SRT document into effective speech transcription time periods, noisy time periods, and quiet time periods;
[0066] Noisy time period identification:
[0067] Use the large language model to detect noisy time periods, identify students' free discussions, readings, or noises, and mark the noisy time periods;
[0068] The following are the prompt words for the large language model to identify and detect noisy time periods:
[0069] The following is a detailed description of the task:
[0070] 1. Background and objectives
[0071] The video transcription content is from classroom recordings and includes conversations between students and teachers. Find the "chaotic" time periods in the content and return the results in the specified format.
[0072] 2. Definition of chaotic content
[0073] 1) It is mainly manifested as noisy and unthemed conversations when students have free discussions, readings, or noises.
[0074] 2) The speakers are mostly students, and a small number may involve teachers, but the teachers' words do not constitute the main part.
[0075] 3) Paragraphs with short conversations and clear themes do not belong to chaotic content.
[0076] 3. Definition of normal content
[0077] 1) There is a clear theme, and the conversations between students and teachers are logical and coherent.
[0078] 2) The conversations are clear and related to classroom activities.
[0079] The video transcription content is the transcription of the classroom language between students and teachers, which may contain some disorganized content. This part is usually caused by the noisy voices during students' free reading or discussion.
[0080] 4. Other rules
[0081] 1) Ignore the speaker numbers (students 1, student 2, etc. are all regarded as one entity).
[0082] 2) The disorganized time period should be at least one and a half minutes. Time periods that are too short do not need to be marked separately.
[0083] 3) The disorganized time period should generally not exceed 8 minutes.
[0084] 4) If there are multiple disorganized time periods, merge adjacent time periods according to continuity.
[0085] 5. Input and output
[0086] 1) Input: The complete classroom video transcription content.
[0087] 2) Output: Return all identified disorganized time periods in JSON format. Do not output other content and strictly follow the following structure:
[0088]
[0089]
[0090] Please output the disorganized conversation time periods in the following video transcription content based on the following video transcription content.
[0091] The following is the video transcription content:
[0092] {{Video transcription content}}
[0093] The output sample result is as follows:
[0094]
[0095] Then mark the noisy time periods with a fixed string as follows:
[0096] Table 3 Noisy time periods identified in SRT
[0097] 00:28:59 00:28:59 Still want to eat Student 1 00:29:00 00:29:01 (The sound is noisy during this period) Student 1 00:29:01 00:29:03 (The sound is noisy during this period) Teacher
[0098] Quiet time periods:
[0099] The quiet time period should follow the noisy process. The judgment logic for the quiet time period is as follows: Check whether there is an interval between the end time of each record and the start time of the next record in chronological order. If the interval exceeds 5 seconds, it is considered a quiet time period, and the start and end points of this time period are recorded.
[0100] Examples of the output blank time periods are as follows:
[0101] Blank time: 00:03:04 to 00:03:14
[0102] Blank time: 00:04:03 to 00:04:11
[0103] Blank time: 00:04:26 to 00:04:56
[0104] Blank time: 00:04:59 to 00:05:05
[0105] Blank time: 00:05:29 to 00:05:37
[0106] Blank time: 00:06:00 to 00:06:08 ...
[0108] Blank time: 00:22:52 to 00:23:02
[0109] Blank time: 00:26:32 to 00:26:41
[0110] Blank time: 00:30:51 to 00:30:57
[0111] Blank time: 00:31:14 to 00:31:19
[0112] Blank time: 00:39:22 to 00:39:40
[0113] Blank time: 00:39:43 to 00:39:58
[0114] Blank time: 00:42:49 to 00:42:57
[0115] There is no audio information in this part of the blank time period, and the behavior of teachers and students (such as writing on the blackboard, patrolling, students reading and writing) is filled by using a visual model (YOLOv12 model).
[0116] The effective speech transcription time periods include:
[0117] After excluding the noisy and quiet time periods, the remaining time periods are marked as effective speech transcription time periods.
[0118] S4. Multimodal role reasoning and analysis;
[0119] Multimodal role reasoning includes: effective speech transcription period S-T analysis logic, noisy period S-T analysis logic, and quiet period S-T analysis logic. The effective speech transcription period refers to the period in the speech recognition system when clear speech content can be accurately and stably extracted from the audio. The noisy period refers to the period in the audio when the environmental noise is large or there are interferences from multiple speech sources, resulting in a decrease in the accuracy of the speech transcription system. The quiet period refers to the period in the audio when there is almost no recognizable speech and the speech recognition system cannot extract any information. Perform S-T role analysis on the effective period, noisy period, and quiet period, and make a comprehensive judgment in combination with the multimodal visual model.
[0120] Effective speech transcription period S-T analysis logic:
[0121] During effective speech periods such as teacher lectures, student answers, and unison reading, the speech recognition system can accurately judge S-T without relying on other auxiliary technologies. Therefore, for the effective speech transcription period, S-T determination can be directly performed according to the speaker's identity: teacher speech is marked as T, and student speech is marked as S.
[0122] Noisy and quiet period S-T analysis logic:
[0123] Noisy periods are usually caused by the following factors: such as student discussions, student free reading, teacher playing videos, etc. Quiet periods are mainly caused by teacher blackboard writing, teacher patrol, students reading and writing, the blank periods during student-teacher Q&A, and the low voice of students when answering questions.
[0124] In order to be able to recognize the above teacher and student behaviors, the present invention collected 7428 classroom scene pictures and labeled 106830 behaviors, and then trained the YOLOv12 model with the labeled data;
[0125] Detect the teacher and student behaviors in the classroom scene through the trained YOLOv12 model, and then analyze and judge the S-T role based on the detected behaviors, and judge whether the current period belongs to the student or the teacher according to the set priority rules. Specifically, the larger the priority number, the greater the impact of the behavior on the role determination. As shown in Table 4 below, when a student is detected to be discussing, its priority is 5; while the priority of the teacher's patrol behavior is 2. Therefore, the S-T determination for this period is the student (S). On the contrary, when a student is detected to be listening (priority 1), and the teacher is writing on the blackboard (priority 4), the S-T determination for the period is the teacher (T).
[0126] Table 4 Student-Teacher Behavior Priorities
[0127] Priority (S - Student) Student behavior Priority (T - Teacher) Teacher behavior 3 Raise hand 2 Lecture 3 Respond 2 Patrol 5 Read aloud 4 Write on the blackboard 3 Read 2 Teacher-student interaction 3 Write 0 Others 5 Discuss — — 1 Listen — — 1 Others — —
[0128] S5. Perform frame extraction on the classroom video to obtain the time periods of student and teacher behaviors;
[0129] Perform frame extraction on the classroom video to obtain classroom scene pictures containing teacher and student behaviors, and label the teacher and student behaviors respectively. Teacher behaviors include patrol, teacher-student interaction, lecture, writing on the blackboard, and others; student behaviors include raising hands, answering, reading aloud, reading, writing, discussing, listening, and others.
[0130] Train the visual model and the visual large model with the labeled picture data, and optimize the behavior detection of the visual large model by designing prompt words;
[0131] Detect the behaviors in the classroom scene video frames with the trained visual model and visual large model to obtain the time periods of student and teacher behaviors.
[0132] Among them, the design of prompt words specifically includes:
[0133] Define the characteristics of teacher and student behaviors in the form of natural language description;
[0134] Provide context information for behavior detection to enhance the model's ability to recognize behaviors;
[0135] Guide the visual large model to recognize and distinguish the behavior patterns of teachers and students through prompt words.
[0136] S6. Use the time periods of student and teacher behaviors as an aid for SRT analysis to perform multi-modal data analysis.
[0137] Use the time periods of student and teacher behaviors as an aid for SRT analysis, combine the audio transcription and the visual behavior detection results to perform multi-modal data analysis, so as to accurately distinguish and label the role identities of teachers and students.
[0138] The method proposed by the present invention integrates advanced large language models and multi-modal information fusion analysis technologies. Through the coordinated actions of automatic speech recognition, visual behavior analysis, and video frame extraction and behavior detection, the accuracy and efficiency of S-T analysis in the classroom teaching scenario are significantly improved. This method not only optimizes the quality of classroom speech transcription, but also further enhances the accurate recognition of teacher and student behaviors through video frame extraction and behavior annotation. In addition, combining the behavior detection technologies of traditional visual models and visual large models provides more comprehensive data support for S-T role analysis. Finally, the present invention provides a solid foundation for subsequent teaching analysis, intelligent feedback, and education quality evaluation applications, and promotes the intelligent application of educational technology in classroom teaching.
[0139] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in the relevant field. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A classroom teaching analysis method based on the integration of large language models and multimodal information, characterized in that, Including: S1. Generate an N-best candidate list; S2. Revise the N-best candidate list to obtain the revised SRT; S3. Multimodal time segment division, dividing the revised SRT document into effective speech transcription segments, noisy segments, and quiet segments; S4. Multimodal role inference analysis; S5. Perform frame extraction on the classroom video to obtain the time periods of student behavior and teacher behavior; S6. Use the time periods of student behavior and teacher behavior as an auxiliary for SRT analysis to perform multimodal data analysis.
2. The classroom teaching analysis method based on the integration of large language models and multimodal information according to claim 1, wherein Step S1 specifically includes: Process the classroom audio through an ASR model to output an N-best candidate list containing timestamp, speech content, and speaker information. Each candidate list contains a confidence score and time alignment information. Filter out the low-confidence lists in the N-best candidates by setting a confidence threshold, and only retain the high-confidence candidates.
3. The classroom teaching analysis method based on the fusion of large language models and multi-modal information according to claim 1, wherein, Step S2 specifically includes: Adopt the zero-shot or few-shot learning ability of the generative large language model, and use the constraint rule prompt template to guide the large language model to complete the task of revising the filtered N-best candidate list to obtain the revised SRT. The input of the prompt template is the filtered N-best candidate text, and the output is the integrated optimal transcription result.
4. The classroom teaching analysis method based on the fusion of large language models and multimodal information according to claim 1, wherein Noisy segment identification includes: Use the large language model to detect noisy segments, identify students' free discussions, readings, or noises, and mark the noisy segments; Quiet segments: Measure the time interval between the end time of each SRT record and the start time of the next record. If the interval exceeds the set time, mark it as a quiet segment and record the start and end points of this time period; Effective speech transcription segments include: After excluding the noisy and quiet segments, mark the remaining segments as effective speech transcription segments.
5. The classroom teaching analysis method based on the integration of large language models and multimodal information according to claim 1, wherein, Step S4 specifically includes: Multimodal role inference includes S-T analysis of effective speech transcription segments, S-T analysis of noisy segments, and S-T analysis of quiet segments; S-T analysis of effective speech transcription segments: For effective speech transcription segments, directly perform S-T determination according to the speaker's identity. Mark the teacher's speech as T and the student's speech as S; S-T analysis of noisy segments and quiet segments: Collect multiple classroom scene pictures and annotate the student behavior and teacher behavior in the classroom; Train the YOLOv12 model with the annotated pictures; Use the trained YOLOv12 model to detect the teacher and student behavior in the classroom scene, and then analyze and judge the S-T role according to the detected behavior, and judge whether the current segment belongs to the student or the teacher according to the set priority rules.
6. The classroom teaching analysis method based on the integration of large language models and multimodal information according to claim 5, characterized in that The priority rules are as follows: Teacher behavior priority: teaching > writing on the blackboard > patrolling > teacher-student interaction > others; Student behavior priority: discussion > reading aloud > writing > reading > raising hands > answering > listening > others.
7. The classroom teaching analysis method based on the fusion of large language models and multimodal information according to claim 1, characterized in that, Step S5 specifically includes: Perform frame extraction on the classroom video to obtain classroom scene pictures containing teacher and student behavior, and annotate the teacher behavior and student behavior respectively. Teacher behavior includes patrolling, teacher-student interaction, teaching, writing on the blackboard, and others. Student behavior includes raising hands, answering, reading aloud, reading, writing, discussion, listening, and others; Train the visual model and the large visual model with the labeled image data, and optimize the behavior detection of the large visual model by designing prompts; Detect the behaviors in the classroom scene video frames with the trained visual model and the large visual model to obtain the behavior time periods of the students and teachers.
8. The classroom teaching analysis method based on the fusion of large language models and multimodal information according to claim 5, characterized in that, The designed prompts specifically include: Define the characteristics of teachers' behaviors and students' behaviors in the form of natural language descriptions; Provide the context information for behavior detection to enhance the model's ability to recognize behaviors; Guide the large visual model to recognize and distinguish the behavior patterns of teachers and students through prompts.
9. The classroom teaching analysis method based on the integration of large language models and multimodal information according to claim 1, wherein Step S6 specifically includes: Use the behavior time periods of the students and teachers as an aid for SRT analysis, and combine the audio transcription and the visual behavior detection results to perform multi-modal data analysis, so as to accurately distinguish and label the role identities of the teachers and students.
Citation Information
Cited By
Learning activity identification method and device based on memory module
CN121191212A
Teaching interaction quality evaluation method and system based on large language model
CN121810129A
A teaching interaction quality evaluation method and system based on a large language model
CN121810129B