An ai-driven multi-modal classroom interaction data analysis method and platform
By using multimodal data analysis methods and platforms to comprehensively process voice and image data, the problem of insufficient teaching effectiveness caused by the diversity of data types in the teaching process is solved, and the teaching process is optimized in real time and its quality is improved.
Patent Information
- Application Number
- CN202511010390.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing technologies are insufficient for comprehensively analyzing various types of teaching data to improve teaching effectiveness, resulting in the inability to adjust and optimize the teaching process in a timely manner.
By acquiring multimodal data (speech and image data) of teaching courses, visual tracking and speech interaction analysis are performed to form teaching interaction data. Combined with the feature information of image and speech data, teaching behavior adjustment analysis is conducted.
It enables real-time feedback and adjustment of the teaching process, ensuring that students can fully adapt to the teaching situation and improve teaching quality.
Smart Images

Figure CN120807240B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal teaching analysis, in particular to an AI-driven multi-modal classroom interaction data analysis method and platform. BACKGROUND
[0002] With the progress of society and the development of science, the teaching method is gradually diversified, and adaptive teaching modes are formed for different teaching contents, further improving the teaching effect.
[0003] With the development of artificial intelligence, the integration of artificial intelligence in teaching can assist in the perfection of the teaching process to improve the situation that affects the teaching effect in the teaching process. Since the data obtained for the teaching process is mainly processed and analyzed by artificial intelligence to form a reference result, but the data that can be obtained for the teaching process is diverse, and the results and angles formed by analyzing different data types are also different. If different types of data are analyzed and processed, the teaching process can be further improved.
[0004] Therefore, an AI-driven multi-modal classroom interaction data analysis method and platform are designed, which collects and compares multi-modal data of teaching classrooms to provide real-time and reasonable reference data for improving teaching effect, further improving teaching quality, which is a problem to be solved at present. SUMMARY
[0005] The purpose of the present application is to provide an AI-driven multi-modal classroom interaction data analysis method, which obtains multi-modal teaching interaction data including voice and image data of different classrooms of the same teaching course, analyzes and extracts interactive data of different data types, compares interactive data corresponding to different classrooms of the same type, and further determines adjustment data based on different types of data reference. On the one hand, the multi-modal data can be used to analyze and process the classroom teaching condition from multiple data types to provide teaching guidance from different aspects, and on the other hand, the multi-modal intelligent data analysis can output responsive analysis data in real time to provide timely teaching effect feedback for the teaching process, ensure that the teaching can be adjusted in time, so that the students in the classroom can fully adapt to the teaching situation, ensure that the students in the classroom have high teaching effect, and effectively ensure the teaching quality of the classroom.
[0006] The application also aims to provide an AI-driven multi-modal classroom interaction data analysis platform, which forms multi-modal classroom teaching data through the collection of image data and voice data, realizes the extraction and comparative analysis of feature data under the processing of a comprehensive comparative analysis unit, and further realizes the reasonable analysis and processing of multi-modal data of classroom teaching, thereby providing important reference data for the adjustment and analysis of teaching behavior and being an important material basis for realizing reasonable teaching behavior adjustment.
[0007] In a first aspect, the application provides an AI-driven multi-modal classroom interaction data analysis method, which comprises: acquiring multi-modal teaching interaction data of the same teaching course, performing interactive data extraction analysis, forming classroom interaction type data corresponding to different teaching courses; performing comprehensive comparative analysis on different classroom interaction type data to form teaching interaction comparative result data; performing classroom teaching adjustment analysis according to the teaching interaction comparative result data to form classroom teaching adjustment reference data.
[0008] In the application, the method acquires multi-modal teaching interaction data including voice and image data of different classrooms of the same teaching course, performs interactive data extraction analysis on different data types, and compares the interactive data of different classrooms of the same type, and further comprehensively determines adjustment data based on different types of data reference. On the one hand, the multi-modal data can be used to analyze and process the classroom teaching condition from multiple data types to provide teaching guidance in different aspects, and on the other hand, the multi-modal intelligent data analysis can output responsive analysis data in real time to provide timely teaching effect feedback for the teaching process, ensure that the teaching can be adjusted in time, and make the students in the classroom fully adapt to the teaching situation, ensure that the students in the classroom have high teaching effect, and effectively ensure the teaching quality of the classroom.
[0009] As a possible implementation manner, the multi-modal teaching interaction data of the same teaching course are acquired, interactive data extraction analysis is performed, and classroom interaction type data corresponding to different teaching courses are formed, which comprises: performing visual tracking analysis based on image data on different multi-modal teaching interaction data of the same teaching course to form teaching visual tracking interaction data; performing sound interaction analysis based on voice data on different multi-modal teaching interaction data of the same teaching course to form teaching voice tracking interaction data; and combining the teaching visual tracking interaction data and the teaching voice tracking interaction data to form classroom interaction type data corresponding to the teaching course.
[0010] In the present application, the collection of multi-modal teaching interaction data in the classroom process of teaching courses mainly considers two types of data, one is image data, and the other is voice data. For image data, it mainly reflects the action change of teachers and students in the teaching process. The present application considers whether the students can well absorb the teaching content. The performance in the image data is mainly the concentration degree of the students to the teaching content of the teacher. Therefore, the image data obtained is mainly analyzed to form the feature data reflecting the situation of the students absorbing the teaching content. For voice data, it mainly reflects the language information interaction of teachers and students in the teaching process. The present application considers that the students will interact with the teacher in the process of absorbing the teaching content, such as asking questions, answering and other voice interaction modes. Through analyzing and extracting the voice interaction features, data information reflecting the situation of the students absorbing the teaching content from different angles and directions can be formed. The feature information extracted based on the image data and the feature data extracted based on the voice data are combined to form multi-modal teaching interaction data displayed in the teaching process of the teaching classroom.
[0011] As a possible implementation manner, for different multi-modal teaching interaction data of the same teaching course, visual tracking analysis based on image data is performed to form teaching visual tracking interaction data, including: extracting classroom interaction image data in the different multi-modal teaching interaction data of the same teaching course, performing teacher teaching position analysis based on spatial position to form corresponding classroom teacher teaching position tracking data; performing student visual angle tracking analysis based on spatial position according to the corresponding classroom interaction image data to form corresponding classroom student visual angle tracking data; and performing teaching position coordination analysis according to the classroom teacher teaching position tracking data and the classroom student visual angle tracking data to form corresponding teaching visual tracking interaction data.
[0012] In the present application, the teaching interaction data extraction target of the image data based on visual tracking is to reflect whether the students can keep coordination with the teaching speed of the teacher when absorbing the teaching content. If the coordination cannot be kept, the teacher will enter the explanation of the next knowledge point teaching content before the previous teaching content is absorbed, and the students cannot apply the previous knowledge point to the next knowledge point to understand and absorb the current knowledge point. Such teaching speed leads to the disconnection of the students absorbing the teaching content, which affects the teaching effect. Therefore, in order to ensure whether the absorption of the students can keep coordination with the teaching speed of the teacher, the mapping position of the blackboard writing or display position of the teaching content of the teacher and the visual angle position of the students in the display area is analyzed to accurately determine whether the students have obvious disconnection.
[0013] As one possible approach, classroom interaction image data is extracted from different multimodal teaching interaction data of the same course, and spatial location-based teacher teaching position analysis is performed to form corresponding classroom teacher teaching position tracking data. This includes: establishing a classroom spatial coordinate system based on the classroom interaction image data; and determining the change data of the teacher's teaching gesture pointing position in the classroom spatial coordinate system over time based on the classroom interaction image data. 'n' represents the number of different multimodal teaching interaction data; an effective distance for teaching pointing is defined, with the teacher's teaching gesture pointing position as the center and the effective distance for teaching pointing as the diameter, based on the corresponding teaching gesture pointing position change data. On a plane parallel to the teaching display, determine the corresponding changes in the pointing area of the teaching gestures. .
[0014] In this invention, to track the student's perspective position based on image data to determine whether it aligns with the teacher's teaching pace, it is first necessary to determine the changes in the display position of the teaching content during the teaching process. To accurately determine the display position of the teaching content, spatial coordinates are established based on image data to achieve accurate positional data acquisition. Here, the display position of the teaching content is determined by the teacher's teaching gestures in the classroom. Whether writing on the blackboard or displaying on a projection screen, the teacher will indicate the content being taught. By recognizing and analyzing the direction of the gestures through image data, the specific position of the gestures on teaching planes such as blackboards and projection screens can be accurately determined. It should be noted that for teaching gestures that are essentially fixed, including finger pointing and dogma pointing, the position can be determined through feature-based deep analysis of the image data. The extraction method can be image recognition processing methods such as big data, artificial intelligence, and neural networks. It should also be noted that the direction of the teaching gesture indicates a point on the teaching plane, while the teaching content is usually displayed as regional information such as paragraphs, formulas, and pictures. Therefore, by determining a reasonable range of teaching content areas centered on the point pointed to by the gesture, the change data of the teaching gesture pointing area can be formed. The effective distance of the teaching gesture can be set according to the teaching content or determined based on big data analysis. As long as most of the teaching content can be mapped within the area, it is acceptable.
[0015] As one possible implementation, based on the corresponding classroom interaction image data, student perspective tracking analysis based on spatial location is performed to form corresponding classroom student perspective tracking data, including: determining the perspective center position of each student in the classroom spatial coordinate system based on the classroom interaction image data; and determining the perspective center angle change data of the student's perspective center as a function of time parameters, with the perspective center position as the origin. k represents the ID of a different student in the multimodal teaching interaction data with ID n; the effective viewing angle is set, with time parameter as reference, the student's center viewpoint as the center line, and the effective viewing angle as the cone angle, based on the data of changes in the center viewpoint angle. Determine the data on changes in the student's perspective within the area indicated by the teaching gesture. Data on changes in the student's perspective tracking area formed on the plane .
[0016] In this invention, after determining the location changes of the teacher's teaching content area, to reflect whether the students' absorption speed of the teaching content is coordinated with the teacher's teaching speed, it is also necessary to determine the area that the students' perspective is focused on on the teaching plane to demonstrate the positional relationship between the area of teaching instructions and the area that the students are focused on. Considering that students will definitely acquire content and capture teaching data such as the teacher's body language when absorbing teaching information, determining the perspective can accurately determine whether the students are coordinating with the teaching speed. After all, there are connections between knowledge points in the teaching content, and if the current knowledge point is related to the previous knowledge point, then when learning the current knowledge point, students will constantly review the previous knowledge points, thus shifting their gaze to acquire the previous knowledge point. The acquisition of the student's perspective mainly involves the extraction of the visual center line, which can be obtained through image data processing. Considering the distance between students' seats and the teaching surface, the viewing angle forms a visual area after reaching the teaching surface. However, while the range of visual angles is relatively large for the human body, the range of visual angles that are memorable or easily identifiable is relatively small. This range can be limited by the effective viewing angle, forming data on the effective area of the viewing angle projected onto the teaching surface. The effective viewing angle can be set based on factors such as the distance between students and the teaching surface, the actual teaching content, or it can be determined based on big data analysis.
[0017] As one possible approach, based on classroom teacher position tracking data and student perspective tracking data, a collaborative analysis of teaching positions is conducted to generate corresponding visual tracking and interaction data, including: for different students, tracking area change data based on their respective student perspectives. Data on changes in the area pointed to by teaching gestures The data on the change of non-overlapping area of the student's viewpoint region and the area pointed to by the teaching gesture in the time dimension were determined. Data on changes in non-overlapping areas from different students' perspectives Determine the corresponding non-overlapping total area from the student's perspective. ,in, , The class teaching duration corresponds to the multimodal teaching interaction data numbered n; for all students, the non-overlapping total area is calculated based on the corresponding student's perspective. Determine the cumulative non-overlapping area of the entire classroom. ,in, .
[0018] In this invention, whether students' absorption of teaching content is synchronized with the teacher's teaching pace can be characterized by whether the student's gaze area overlaps with the content area pointed to by the teacher's teaching gestures over time. Considering the individual differences among students in different classrooms, the overlap data of a single student cannot be analyzed individually. Instead, the overlap of all students in the classroom is statistically analyzed to characterize the synergy between student absorption speed and teaching pace in the corresponding classroom. After all, a low overlap rate for one student while a high overlap rate for others may be due to students not paying attention. Using data from the student group as a reference avoids the impact of factors such as inattentiveness or accidental gaze on the synergy analysis.
[0019] As one possible approach, voice interaction analysis based on speech data is performed on different multimodal teaching interaction data of the same course to form teaching speech tracking interaction data. This includes: extracting corresponding classroom speech data from different multimodal teaching interaction data of the same course; and determining the duration of each teacher's single speech instruction based on the classroom speech data. Based on the duration of single-speech teaching and corresponding classroom teaching time Determine the corresponding percentage of interaction time. ,in, .
[0020] In this invention, the analysis of audio data in classroom teaching primarily aims to determine whether there is excessive interaction time between teachers and students. If the interaction time is excessively long in some classes within the same curriculum, it indicates lower teaching quality, as it's impossible to conduct continuous instruction without providing a quiet classroom environment. Here, the proportion of interaction time is determined by the duration of individual teacher-led audio instruction. It should be noted that individual teacher-led audio instruction includes the teacher's voice, the voice information of the music played, and the voice information of the video, among other teaching content. Considering the variability in the total duration of classroom teaching, providing reference data in the form of a proportion is more accurate and reasonable.
[0021] As one possible approach, a comprehensive comparative analysis of data from different types of classroom interactions is conducted to generate comparative data on teaching interactions, including: data based on cumulative non-overlapping areas. Determine the corresponding cumulative area per unit. ,in, Establish an area numerical axis and accumulate the area of different units. The area center value is determined by arbitrarily selecting a unit cumulative area, and the area is marked on the area numerical axis. This is marked as the preset center point of the area; with the preset center point as the center and the area envelope distance as the radius, the radius value is gradually increased to determine the cumulative area of all other units. The cumulative increase in unit area per unit distance during each increase in the area enveloped into the circular region. The quantity, and fitted to form the corresponding area preset center envelope change function. , i represents the index of the sequentially increasing area envelope distance; iterate through all the unit cumulative areas Preset the center envelope transformation function for all areas The cumulative area per unit corresponding to the preset center point with the largest average rate of change The central cumulative unit area was determined; a time-based numerical axis was established, and the proportions of different interaction durations were assigned. The value is calibrated on the duration axis, and the duration center value is determined in the following way: arbitrarily select a percentage of the interaction duration. This is marked as the preset center point of the duration; with the preset center point of the duration as the center and the duration envelope distance as the radius, the radius value is gradually increased to determine the proportion of the duration of all other interactions. The percentage of additional interaction time added by each increment of the envelope distance during the process of enveloping the circular range. The number of values is determined, and a corresponding time-preset center envelope variation function is fitted to form it. 'm' represents the number of times the duration envelope distance is sequentially increased; iterate through all interaction duration percentages. All duration preset center envelope change functions The percentage of interaction time corresponding to the preset center point with the largest average rate of change The proportion of interaction time at the center was determined.
[0022] In this invention, different classrooms teaching the same content can affect students' absorption of the content due to differences in teachers' teaching styles. Therefore, comparative analysis of visual tracking interaction data and audio tracking interaction data between different classrooms can determine the impact of different teachers on teaching effectiveness. Furthermore, by comparing similar data, appropriate adjustments to teaching behaviors for different student groups can be made. For the comparison of visual tracking interaction data, it is first necessary to determine a reference data benchmark. Visual tracking interaction data from different classrooms exhibits clustering characteristics under large datasets; therefore, using the cluster center data as a reference standard is reasonable. This application determines the center data by calculating the rate of change of the data envelope based on the step size of the cumulative non-overlapping area corresponding to the classroom. It can be understood that if the cumulative non-overlapping area is the center data, then the sum of the numerical distances from other cumulative non-overlapping areas to it is the smallest. Therefore, the rate of change of the data envelope on the step size will be the largest, even for the average rate of change. Similarly, the comparative analysis of teaching voice tracking interaction data also uses the rate of change of the data envelope based on the step size on the numerical axis to determine the central data.
[0023] As one possible approach, classroom teaching adjustments can be analyzed based on comparative data of teaching interactions to generate reference data for adjustments. This includes: setting visual deviation thresholds and combining this with the cumulative unit area to form a cumulative unit area threshold range; and adjusting the cumulative unit area... For lessons that do not fall within the cumulative unit area threshold range, adjust and calibrate the teaching speed; set a speech deviation threshold, and combine it with the proportion of central interaction time to form an interaction time proportion threshold range; and adjust the interaction time proportion. For teaching sessions that do not fall within the threshold range for interaction time percentage, the interaction time should be adjusted and calibrated.
[0024] In this invention, by obtaining benchmark data on visual and audio interaction tracking data from different classrooms teaching the same content, classrooms with poor teaching effectiveness can be identified through reasonable scope. The adjustments needed for teaching behavior differ depending on the type of comparative data. For example, the comparative analysis of visual interaction tracking data indicates that students' absorption of the teaching content cannot keep pace with the teacher's pace, requiring teachers to adjust their teaching speed by simplifying the explanation process or further refining the content. Similarly, the comparative analysis of audio interaction tracking data suggests that the teaching content provided by the teacher cannot meet students' understanding and learning needs, which can be addressed by improving the teaching presentation content and expanding on knowledge points. These comparative analysis results provide clear and reasonable guidance for adjusting teaching behavior, offering real-time assistance to improve teaching effectiveness and aid students' learning absorption in the current lesson.
[0025] Secondly, this invention provides an AI-driven multimodal classroom interaction data analysis platform, comprising: an image data acquisition unit for acquiring classroom interaction image data; an audio data acquisition unit for acquiring classroom audio data; and a comprehensive comparison and analysis unit for acquiring the classroom interaction image data acquired by the image data acquisition unit and the classroom audio data acquired by the audio data acquisition unit, performing comprehensive comparison and analysis to form teaching interaction comparison result data, and performing classroom teaching adjustment analysis based on the teaching interaction comparison result data to form classroom teaching adjustment reference data.
[0026] This platform generates multimodal classroom teaching data by collecting image and voice data. Through the processing of the comprehensive comparative analysis unit, it extracts and compares feature data, thereby enabling reasonable analysis and processing of multimodal classroom teaching data. This provides important reference data for the analysis and adjustment of teaching behavior and is an important material basis for achieving reasonable adjustment of teaching behavior.
[0027] The beneficial effects of the AI-driven multimodal classroom interaction data analysis method and platform provided by this invention are as follows:
[0028] This method acquires multimodal teaching interaction data, including voice and image data, from different classrooms teaching the same curriculum. It then extracts and analyzes interaction data of different data types and compares the interaction data from different classrooms within the same data type. This comprehensive approach determines adjustment data based on different data types. On one hand, it leverages multimodal data to analyze classroom teaching from multiple data types, providing diverse teaching guidance. On the other hand, intelligent multimodal data analysis can output real-time analytical results, providing timely feedback on teaching effectiveness. This ensures timely adjustments to the teaching process, allowing students to fully adapt to the learning environment and achieving high learning outcomes, thus effectively guaranteeing the quality of the lesson.
[0029] This platform generates multimodal classroom teaching data by collecting image and voice data. Through the processing of the comprehensive comparative analysis unit, it extracts and compares feature data, thereby enabling reasonable analysis and processing of multimodal classroom teaching data. This provides important reference data for the analysis and adjustment of teaching behavior and is an important material basis for achieving reasonable adjustment of teaching behavior. Attached Figure Description
[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0031] Fig. 1 A flowchart illustrating the steps of an AI-driven multimodal classroom interaction data analysis method provided in an embodiment of the present invention;
[0032] Fig. 2 This is a schematic diagram of the structure of the AI-driven multimodal classroom interaction data analysis platform provided in an embodiment of the present invention. Detailed Implementation
[0033] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.
[0034] With social progress and scientific development, teaching methods have gradually diversified, forming adaptive teaching models for different teaching content, which further improves teaching effectiveness.
[0035] With the development of artificial intelligence (AI), integrating AI into teaching can assist in improving the teaching process and mitigating factors that affect teaching effectiveness. Since the data acquired in the teaching process is primarily processed and analyzed by AI to generate referential results, the data is diverse, and different data types produce different results and perspectives. Therefore, by comprehensively analyzing and processing different types of data, the teaching process can be further improved.
[0036] refer to Figs. 1-2 This invention provides an AI-driven multimodal classroom interaction data analysis method. This method acquires multimodal teaching interaction data, including voice and image data, from different classrooms teaching the same curriculum. It then extracts and analyzes interaction data based on different data types, compares the interaction data from different classrooms with data of the same type, and comprehensively determines adjustment data based on different data types. On the one hand, it leverages multimodal data to analyze and process classroom teaching from multiple data types, providing different aspects of teaching guidance. On the other hand, the intelligent multimodal data analysis can output real-time analysis results, providing timely feedback on teaching effectiveness, ensuring timely adjustments to teaching, and enabling students to fully adapt to the teaching situation, thus ensuring high teaching effectiveness and effectively guaranteeing the teaching quality of the lesson.
[0037] The AI-driven multimodal classroom interaction data analysis method specifically includes the following steps:
[0038] S1: Obtain multimodal teaching interaction data for the same course, extract and analyze the interaction data, and form classroom interaction type data corresponding to different courses.
[0039] Acquire multimodal teaching interaction data for the same course, extract and analyze the interaction data to form classroom interaction type data corresponding to different courses, including: performing visual tracking analysis based on image data on different multimodal teaching interaction data for the same course to form teaching visual tracking interaction data; performing audio interaction analysis based on speech data on different multimodal teaching interaction data for the same course to form teaching audio tracking interaction data; and combining teaching visual tracking interaction data and teaching audio tracking interaction data to form classroom interaction type data corresponding to the course.
[0040] This study collects multimodal teaching interaction data during classroom instruction, primarily considering two types of data: image data and audio data. Image data mainly reflects the changes in the actions of teachers and students during the teaching process. This application considers whether students can effectively absorb the teaching content; the image data mainly shows their focus on the content being taught by the teacher. Therefore, the acquired image data is primarily analyzed using visual tracking of students to form feature data reflecting their absorption of the teaching content. Audio data mainly reflects the linguistic information interaction between teachers and students during the teaching process. This application considers the interaction between students and teachers during the absorption of teaching content, such as asking and answering questions via audio. By analyzing and extracting the interactive features of audio, data information reflecting students' absorption of the teaching content can be formed from a different perspective than that of image data. Combining the feature information extracted from image data and the feature data extracted from audio data, multimodal teaching interaction data demonstrating the teaching process in the classroom is formed.
[0041] For different multimodal teaching interaction data of the same course, visual tracking analysis based on image data is performed to form teaching visual tracking interaction data. This includes: extracting classroom interaction image data from different multimodal teaching interaction data of the same course, performing teacher teaching position analysis based on spatial location to form corresponding classroom teacher teaching position tracking data; performing student perspective tracking analysis based on spatial location based on the corresponding classroom interaction image data to form corresponding classroom student perspective tracking data; and performing teaching position synergy analysis based on classroom teacher teaching position tracking data and classroom student perspective tracking data to form corresponding teaching visual tracking interaction data.
[0042] The goal of visual tracking-based teaching interaction data extraction from image data is to reflect whether students' absorption of teaching content is in sync with the teacher's pace. If this sync fails, the teacher moves on to the next topic before the previous content has been absorbed, preventing students from applying previous knowledge to the next and thus hindering their understanding and absorption of the current material. This disjointed teaching pace negatively impacts teaching effectiveness. Therefore, to ensure that student absorption is synchronized with the teacher's pace, a collaborative analysis is conducted to determine the mapping between the teacher's blackboard writing or display position and the students' viewpoint within the display area. This analysis accurately identifies any significant gaps in student engagement.
[0043] Classroom interaction image data is extracted from different multimodal teaching interaction data of the same teaching course. Spatial location-based teacher teaching position analysis is performed to form corresponding classroom teacher teaching position tracking data, including: establishing a classroom spatial coordinate system based on the classroom interaction image data; and determining the change of the teacher's teaching gesture pointing position in the classroom spatial coordinate system over time based on the classroom interaction image data. 'n' represents the number of different multimodal teaching interaction data; an effective distance for teaching pointing is defined, with the teacher's teaching gesture pointing position as the center and the effective distance for teaching pointing as the diameter, based on the corresponding teaching gesture pointing position change data. On a plane parallel to the teaching display, determine the corresponding changes in the pointing area of the teaching gestures. .
[0044] To track students' perspective positions based on image data and determine whether they align with the teacher's teaching pace, it's first necessary to determine the changes in the teacher's presentation position of the teaching content during the lesson. Accurately determining the presentation position requires establishing spatial coordinates based on image data to obtain precise location data. Here, the teacher's presentation position is determined through classroom gestures. Whether writing on the blackboard or displaying on a projection screen, the teacher will indicate the content being taught. By recognizing and analyzing the direction of these gestures using image data, the specific location of the gestures on teaching surfaces such as the blackboard or projection screen can be accurately determined. It should be noted that for relatively fixed gestures, including finger pointing and pointing with pointers, the position can be determined through feature-based deep analysis of the image data. This extraction can be achieved using image recognition processing methods such as big data, artificial intelligence, and neural networks. It should also be noted that the direction of the teaching gesture indicates a point on the teaching plane, while the teaching content is usually displayed as regional information such as paragraphs, formulas, and pictures. Therefore, by determining a reasonable range of teaching content areas centered on the point pointed to by the gesture, the change data of the teaching gesture pointing area can be formed. The effective distance of the teaching gesture can be set according to the teaching content or determined based on big data analysis. As long as most of the teaching content can be mapped within the area, it is acceptable.
[0045] Based on the corresponding classroom interaction image data, a student perspective tracking analysis based on spatial location is performed to generate corresponding classroom student perspective tracking data, including: determining the perspective center position of each student in the classroom spatial coordinate system based on the classroom interaction image data; and determining the perspective center angle change data of the student's perspective center as a function of time parameters, with the perspective center position as the origin. k represents the ID of a different student in the multimodal teaching interaction data with ID n; the effective viewing angle is set, with time parameter as reference, the student's center viewpoint as the center line, and the effective viewing angle as the cone angle, based on the data of changes in the center viewpoint angle. Determine the data on changes in the student's perspective within the area indicated by the teaching gesture. Data on changes in the student's perspective tracking area formed on the plane .
[0046] After determining the data on the changes in the teacher's teaching content location, to reflect whether the students' absorption speed of the teaching content is coordinated with the teacher's teaching speed, it is also necessary to determine the area of the students' gaze on the teaching plane to show the positional relationship between the area of teaching instructions and the area of students' gaze. Considering that students will definitely acquire content and capture teaching data such as teacher's body language when absorbing teaching, determining the perspective can accurately determine whether students are coordinating with the teaching speed. After all, there are connections between knowledge points in the teaching content, and if the current knowledge point is related to the previous knowledge point, then when learning the current knowledge point, students will constantly review the previous knowledge point, thus shifting their gaze to acquire the previous knowledge point. The acquisition of students' perspective mainly involves the extraction of the visual center line, which can be obtained through image data processing. Considering the distance between the student's seat and the teaching plane, the perspective will form a perspective area after reaching the teaching plane. However, for the human body, although the perspective range is relatively large, the perspective range with memory or obvious recognition is relatively small. The effective perspective range angle can be used to limit the range, forming the effective range area data of the perspective projection onto the teaching plane. The effective viewing angle can be set based on factors such as the distance between students and the teaching plane, the actual teaching content, or it can be determined based on big data analysis.
[0047] Based on classroom teacher positioning tracking data and student perspective tracking data, a collaborative analysis of teaching positions is conducted to generate corresponding visual tracking and interaction data, including: data on changes in the tracking area based on the student's perspective for different students. Data on changes in the area pointed to by teaching gestures The data on the change of non-overlapping area of the student's viewpoint region and the area pointed to by the teaching gesture in the time dimension were determined. Data on changes in non-overlapping areas from different students' perspectives Determine the corresponding non-overlapping total area from the student's perspective. ,in, , The class teaching duration corresponds to the multimodal teaching interaction data numbered n; for all students, the non-overlapping total area is calculated based on the corresponding student's perspective. Determine the cumulative non-overlapping area of the entire classroom. ,in, .
[0048] Whether students' absorption of teaching content is synchronized with the teacher's teaching pace can be characterized by whether the students' gaze areas overlap with the content areas pointed to by the teacher's teaching gestures over time. Considering the individual differences among students in different classrooms, it's not feasible to analyze the overlap data of a single student alone. Instead, the overlap of all students in the classroom should be statistically analyzed to represent the synergy between student absorption speed and teaching pace in the corresponding classroom. After all, a low overlap rate for one student while a high overlap rate for others might be due to students not paying attention. Using data from the student group as a reference can avoid the influence of factors such as inattentiveness or accidental gaze on the synergy analysis.
[0049] For different multimodal teaching interaction data of the same course, voice interaction analysis based on speech data is performed to form teaching speech tracking interaction data, including: extracting corresponding classroom speech data from different multimodal teaching interaction data of the same course; and determining the duration of each teacher's single speech instruction based on the classroom speech data. Based on the duration of single-speech teaching and corresponding classroom teaching time Determine the corresponding percentage of interaction time. ,in, .
[0050] The analysis of audio data in classrooms primarily aims to determine whether there is excessive interaction time between teachers and students. If the interaction time is excessively long in some classes within the same curriculum, it indicates lower teaching quality, as it's impossible to conduct continuous instruction without providing a quiet classroom environment. Here, we determine the proportion of interaction time by analyzing the duration of teacher-led individual audio instruction. It should be noted that teacher-led individual audio instruction includes the teacher's voice, the audio of the music played, and the audio of the video, among other teaching content. Considering the variability in total classroom duration, providing reference data in the form of proportions is more accurate and reasonable.
[0051] S2: Conduct a comprehensive comparative analysis of data on different types of classroom interaction to generate comparative data on teaching interaction results.
[0052] A comprehensive comparative analysis of data from different types of classroom interactions was conducted to generate comparative data on teaching interactions, including: data based on cumulative non-overlapping areas. Determine the corresponding cumulative area per unit. ,in, Establish an area numerical axis and accumulate the area of different units. The area center value is determined by arbitrarily selecting a unit cumulative area, and the area is marked on the area numerical axis. This is marked as the preset center point of the area; with the preset center point as the center and the area envelope distance as the radius, the radius value is gradually increased to determine the cumulative area of all other units. The cumulative increase in unit area per unit distance during each increase in the area enveloped into the circular region. The quantity, and fitted to form the corresponding area preset center envelope change function. , i represents the index of the sequentially increasing area envelope distance; iterate through all the unit cumulative areas Preset the center envelope transformation function for all areas The cumulative area per unit corresponding to the preset center point with the largest average rate of change The central cumulative unit area was determined; a time-based numerical axis was established, and the proportions of different interaction durations were assigned. The value is calibrated on the duration axis, and the duration center value is determined in the following way: arbitrarily select a percentage of the interaction duration. This is marked as the preset center point of the duration; with the preset center point of the duration as the center and the duration envelope distance as the radius, the radius value is gradually increased to determine the proportion of the duration of all other interactions. The percentage of additional interaction time added by each increment of the envelope distance during the process of enveloping the circular range. The number of values is determined, and a corresponding time-preset center envelope variation function is fitted to form it. 'm' represents the number of times the duration envelope distance is sequentially increased; iterate through all interaction duration percentages. All duration preset center envelope change functions The percentage of interaction time corresponding to the preset center point with the largest average rate of change The proportion of interaction time at the center was determined.
[0053] Different classrooms teaching the same content can have varying impacts on student absorption due to differences in teaching styles. Therefore, comparative analysis of visual tracking interaction data and audio tracking interaction data across different classrooms can determine the influence of different teachers on teaching effectiveness and allow for adjustments to teaching behaviors for different student groups based on the comparison of similar data. For comparing visual tracking interaction data, it's first necessary to establish a reference data benchmark. Visual tracking interaction data from different classrooms exhibits clustering characteristics under large datasets; therefore, using the cluster's central data as a reference standard is reasonable. This application determines the central data by calculating the rate of change of the data envelope along the numerical axis based on the step size of the cumulative non-overlapping area corresponding to each classroom. It's understood that if the cumulative non-overlapping area is the central data, then the sum of the numerical distances from other cumulative non-overlapping areas to it is minimized. Therefore, the rate of change of the data envelope along the step size will be the largest, even for the average rate of change. Similarly, the comparative analysis of teaching voice tracking interaction data also uses the rate of change of the data envelope based on the step size on the numerical axis to determine the central data.
[0054] S3: Based on the comparison data of teaching interactions, conduct classroom teaching adjustment analysis to form reference data for classroom teaching adjustment.
[0055] Based on the comparative data of teaching interactions, classroom teaching adjustments and analyses were conducted to generate reference data for these adjustments. This included: setting visual deviation thresholds and combining this with the cumulative unit area to form a cumulative unit area threshold range; and defining the cumulative unit area... For lessons that do not fall within the cumulative unit area threshold range, adjust and calibrate the teaching speed; set a speech deviation threshold, and combine it with the proportion of central interaction time to form an interaction time proportion threshold range; and adjust the interaction time proportion. For teaching sessions that do not fall within the threshold range for interaction time percentage, the interaction time should be adjusted and calibrated.
[0056] By obtaining benchmark data on visual and audio interaction tracking data from different classrooms teaching the same content, we can identify classrooms with poor teaching effectiveness through reasonable scope. The adjustments needed for different types of comparative data differ. For example, the comparative analysis of visual interaction tracking data indicates that students' absorption of the content is not keeping pace with the teacher's pace, requiring teachers to adjust their teaching speed by simplifying the explanation process or further refining the content. Similarly, the comparative analysis of audio interaction tracking data suggests that the content provided by the teacher is insufficient for students' comprehension and learning needs, which can be addressed by improving the presentation content and expanding on key knowledge points. These comparative analyses provide clear and reasonable guidance for adjusting teaching behaviors, offering real-time assistance to improve teaching effectiveness and aid students' learning absorption in the current lesson.
[0057] This application also provides an AI-driven multimodal classroom interaction data analysis platform, which includes: an image data acquisition unit for acquiring classroom interaction image data; an audio data acquisition unit for acquiring classroom audio data; and a comprehensive comparison and analysis unit for acquiring the classroom interaction image data acquired by the image data acquisition unit and the classroom audio data acquired by the audio data acquisition unit, performing comprehensive comparison and analysis to form teaching interaction comparison result data, and analyzing classroom teaching adjustments based on the teaching interaction comparison result data to form classroom teaching adjustment reference data.
[0058] This platform generates multimodal classroom teaching data by collecting image and voice data. Through the processing of the comprehensive comparative analysis unit, it extracts and compares feature data, thereby enabling reasonable analysis and processing of multimodal classroom teaching data. This provides important reference data for the analysis and adjustment of teaching behavior and is an important material basis for achieving reasonable adjustment of teaching behavior.
[0059] In summary, the beneficial effects of the AI-driven multimodal classroom interaction data analysis method and platform provided in this embodiment of the invention are as follows:
[0060] This method acquires multimodal teaching interaction data, including voice and image data, from different classrooms teaching the same curriculum. It then extracts and analyzes interaction data of different data types and compares the interaction data from different classrooms within the same data type. This comprehensive approach determines adjustment data based on different data types. On one hand, it leverages multimodal data to analyze classroom teaching from multiple data types, providing diverse teaching guidance. On the other hand, intelligent multimodal data analysis can output real-time analytical results, providing timely feedback on teaching effectiveness. This ensures timely adjustments to the teaching process, allowing students to fully adapt to the learning environment and achieving high learning outcomes, thus effectively guaranteeing the quality of the lesson.
[0061] This platform generates multimodal classroom teaching data by collecting image and voice data. Through the processing of the comprehensive comparative analysis unit, it extracts and compares feature data, thereby enabling reasonable analysis and processing of multimodal classroom teaching data. This provides important reference data for the analysis and adjustment of teaching behavior and is an important material basis for achieving reasonable adjustment of teaching behavior.
[0062] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information is called the information to be instructed. In the specific implementation process, there are many ways to instruct the information to be instructed, such as, but not limited to, directly instructing the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly instruct the information to be instructed by instructing other information, where there is a relationship between the other information and the information to be instructed. It can also instruct only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing instruction overhead to some extent. At the same time, common parts of various pieces of information can be identified and uniformly indicated to reduce the instruction overhead caused by individually indicating the same information.
[0063] Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be repeated here. As described above, for example, when multiple pieces of information of the same type need to be indicated, the indication methods for different pieces of information may differ. In the specific implementation process, the required indication method can be selected according to specific needs. This application embodiment does not limit the selected indication method; therefore, the indication methods involved in this application embodiment should be understood to cover various methods that enable the party to be indicated to obtain the information to be indicated.
[0064] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information messages sent separately, and the sending period and / or timing of these sub-information messages can be the same or different. The specific sending method is not limited in this application embodiment. The sending period and / or timing of these sub-information messages can be predefined, for example, according to a protocol, or configured by the sending device by sending configuration information to the receiving device.
[0065] "Predefined" or "pre-configured" can be achieved by pre-saving corresponding codes, tables, or other means that can be used to indicate relevant information in the device. This application does not limit the specific implementation method. "Saving" can refer to saving in one or more memories. These memories can be separate installations or integrated into the encoder, decoder, processor, or communication device. Alternatively, some memories can be separately installed, while others are integrated into the decoder, processor, or communication device. The type of memory can be any form of storage medium, and this application does not limit this.
[0066] The “protocol” mentioned in the embodiments of this application may refer to a protocol family in the field of communication, a standard protocol with a similar protocol family frame structure, or a related protocol applied to future communication systems. The embodiments of this application do not specifically limit this.
[0067] In the embodiments of this application, descriptions such as "when," "under the circumstances," "if," and "if" all refer to the device making corresponding processing under certain objective circumstances, and are not limited to a specific time. They do not require the device to make a judgment action during implementation, nor do they imply any other limitations.
[0068] In the description of the embodiments of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in the embodiments of this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of the embodiments of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Additionally, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or order of execution, and that "first," "second," etc., are not necessarily different. Furthermore, in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.
[0069] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0070] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0071] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0072] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0073] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0074] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0075] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0076] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0077] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0079] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0080] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0081] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An AI-driven method for analyzing multimodal classroom interaction data, characterized in that, Configured as: Acquire multimodal teaching interaction data for the same course, extract and analyze the interaction data, and form classroom interaction type data corresponding to different courses; A comprehensive comparative analysis of the data from different types of classroom interaction is conducted to generate comparative data on teaching interaction. Based on the comparative results of the teaching interactions, classroom teaching adjustments and analyses are conducted to generate reference data for classroom teaching adjustments. This involves acquiring multimodal teaching interaction data for the same course, extracting and analyzing the interaction data to form classroom interaction type data corresponding to different courses, including: Visual tracking analysis based on image data is performed on different multimodal teaching interaction data of the same teaching course to form teaching visual tracking interaction data; For different multimodal teaching interaction data of the same teaching course, conduct voice interaction analysis based on voice data to form teaching voice tracking interaction data; The teaching visual tracking interaction data and the teaching voice tracking interaction data are combined to form the classroom interaction type data corresponding to the teaching course; Visual tracking analysis based on image data is performed on different multimodal teaching interaction data of the same teaching course to form teaching visual tracking interaction data, including: Classroom interaction image data are extracted from different multimodal teaching interaction data of the same teaching course, and teacher teaching position analysis is performed based on spatial location to form corresponding classroom teacher teaching position tracking data. Based on the corresponding classroom interaction image data, a student perspective tracking analysis based on spatial location is performed to form corresponding classroom student perspective tracking data. Based on the classroom teacher's teaching position tracking data and the classroom student's perspective tracking data, a teaching position synergy analysis is performed to generate corresponding teaching visual tracking interaction data. Classroom interaction image data is extracted from different multimodal teaching interaction data of the same teaching course, and teacher teaching position analysis is performed based on spatial location to form corresponding classroom teacher teaching position tracking data, including: Based on the classroom interaction image data, establish a classroom spatial coordinate system; Based on the classroom interaction image data, the data on the change of the teacher's teaching gesture pointing position in the classroom spatial coordinate system over time were determined. , where n represents the number of the different multimodal teaching interaction data; Define an effective distance for the teaching gesture, using the teacher's gesture pointing position as the center and the effective distance as the diameter, based on the corresponding changes in the gesture pointing position. On a plane parallel to the teaching display, determine the corresponding changes in the pointing area of the teaching gestures. ; Based on the corresponding classroom interaction image data, a student perspective tracking analysis based on spatial location is performed to form corresponding classroom student perspective tracking data, including: Based on the classroom interaction image data, the position of the center of view of each student in the classroom spatial coordinate system is determined; Using the aforementioned center of view as the origin, determine the data on the change of the student's center of view angle as a function of time parameters. k represents the ID of a different student in the multimodal teaching interaction data numbered n; Define an effective viewing angle range, using time parameters as a reference, with the student's center viewpoint as the center line, and the effective viewing angle range as the cone angle, based on the change data of the center viewpoint angle. Determine the data on changes in the student's perspective within the area indicated by the teaching gesture. Data on changes in the student's perspective tracking area formed on the plane ; Based on the classroom teacher's teaching position tracking data and the classroom student's perspective tracking data, a teaching position synergy analysis is performed to generate corresponding teaching visual tracking interaction data, including: For different students, track regional change data according to the corresponding student's perspective. and the change data of the teaching gesture pointing area The data on the change of non-overlapping area of the student's viewpoint region and the area pointed to by the teaching gesture in the time dimension were determined. ; Data on changes in non-overlapping areas based on different students' perspectives Determine the corresponding non-overlapping total area from the student's perspective. ,in, , The class teaching duration corresponding to the multimodal teaching interaction data numbered n; For all students, the total non-overlapping area from the corresponding student's perspective. Determine the cumulative non-overlapping area of the entire classroom. ,in, .
2. The AI-driven multimodal classroom interaction data analysis method according to claim 1, characterized in that, The process of performing voice interaction analysis on different multimodal teaching interaction data for the same teaching course to form teaching voice tracking interaction data includes: For different multimodal teaching interaction data of the same course, extract the corresponding classroom voice data; Based on the classroom audio data, the duration of each teacher's single-voice teaching session was determined. ; Based on the single-voice teaching duration and the corresponding classroom teaching duration Determine the corresponding percentage of interaction time. ,in, .
3. The AI-driven multimodal classroom interaction data analysis method according to claim 2, characterized in that, The comprehensive comparative analysis of data from different types of classroom interactions to form comparative teaching interaction results data includes: Based on the cumulative non-overlapping area Determine the corresponding cumulative area per unit. ,in, ; Establish an area numerical axis and accumulate the area of different units. The area center value is determined by aligning it on the area numerical axis and performing the following steps: Arbitrarily select one of the aforementioned cumulative areas , and mark it as the preset center point of the area; Using the preset center point of the area as the center, and the area envelope distance as the radius, the radius value is gradually increased to determine the cumulative area of all other units. During the process of enveloping the circular area, the cumulative area per unit is increased by the enveloping distance each time. The quantity, and fitted to form the corresponding area preset center envelope change function. , i represents the number of times the area envelope distance is sequentially increased; Traverse all the cumulative areas of the units mentioned The area preset center envelope transformation function The cumulative area per unit corresponding to the preset center point of the area with the largest average rate of change The area is determined as the central cumulative unit area; Establish a duration numerical axis and assign the proportion of different interaction durations. The duration center value is determined by aligning it on the duration value axis and performing the following steps: Arbitrarily select one of the interaction duration percentages Mark it as the preset center point of the duration; Using the preset center point of the duration as the center, and the duration envelope distance as the radius, the radius value is gradually increased to determine the proportion of all other interaction durations. The percentage of the interaction time added by each increment of the duration envelope distance during the process of enveloping the circular range. The number of values is determined, and a corresponding time-preset center envelope variation function is fitted to form it. m represents the number of times the duration envelope distance is sequentially incremented; Iterate through all the interaction duration percentages. All the time-preset center envelope change functions The percentage of interaction time corresponding to the preset center point with the largest average rate of change The proportion of interaction time at the center was determined.
4. The AI-driven multimodal classroom interaction data analysis method according to claim 3, characterized in that, The step of analyzing and adjusting classroom teaching based on the comparison results of the teaching interactions to generate reference data for classroom teaching adjustments includes: A visual deviation threshold is set, and the cumulative unit area threshold range is formed by combining the cumulative unit area of the center. For the unit cumulative area For teaching classes that do not fall within the aforementioned cumulative unit area threshold range, the teaching speed will be adjusted and calibrated. Set a voice deviation threshold, and combine it with the central interaction duration percentage to form an interaction duration percentage threshold range; Percentage of the interaction time For teaching classes that do not fall within the threshold range for the percentage of interaction time, the interaction time will be adjusted and calibrated.
5. An AI-driven multimodal classroom interaction data analysis platform, characterized in that, The AI-driven multimodal classroom interaction data analysis method according to any one of claims 1-4 includes: The image data acquisition unit is used to acquire classroom interaction image data during classroom teaching. The voice data acquisition unit is used to acquire classroom voice data for classroom teaching; The comprehensive comparison and analysis unit is used to acquire classroom interaction image data collected by the image data acquisition unit and classroom audio data collected by the audio data acquisition unit, perform comprehensive comparison and analysis to form teaching interaction comparison result data, and analyze classroom teaching adjustments based on the teaching interaction comparison result data to form classroom teaching adjustment reference data.
Citation Information
Patent Citations
Course quality evaluation and improvement method based on student visual attention and teacher behaviors
CN113506027A
Interaction analysis method and system based on cognition and social network mining
CN118733994A
Ideological and political classroom interaction analysis method and system based on multi-modal fusion
CN119478525A