AI-driven multi-mode classroom interaction data analysis method and platform
Through multimodal data analysis methods and platforms, the voice and image data in the teaching process are comprehensively processed, which solves the problem of poor teaching results caused by the diversity of data types in the teaching process and realizes real-time optimization and quality improvement of the teaching process.
Patent Information
- Application Number
- CN202511010390.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing technologies make it difficult to effectively and comprehensively analyze multiple teaching data types to improve teaching effectiveness, resulting in the inability to adjust and optimize the teaching process in a timely manner.
By acquiring multimodal interactive data (voice and image data) of teaching courses, visual tracking and voice interaction analysis are performed to form teaching visual tracking and voice tracking interactive data, and comprehensive comparative analysis is performed to provide reference data for teaching adjustments.
Real-time feedback and adjustment of the teaching process are achieved to ensure that students can fully adapt to the teaching situation and improve the quality and effectiveness of teaching.
Smart Images

Figure CN120807240A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal teaching analysis, in particular to an AI-driven multi-modal classroom interaction data analysis method and platform. BACKGROUND
[0002] With the progress of society and the development of science, the teaching method is gradually diversified, and adaptive teaching modes are formed for different teaching contents, further improving the teaching effect.
[0003] With the development of artificial intelligence, the integration of artificial intelligence in teaching can assist in the perfection of the teaching process to improve the situation that affects the teaching effect in the teaching process. Since the data obtained for the teaching process is mainly processed and analyzed by artificial intelligence to form a reference result, but the data that can be obtained for the teaching process is diverse, and the results and angles formed by analyzing different data types are also different. If different types of data are analyzed and processed, the teaching process can be further improved.
[0004] Therefore, an AI-driven multi-modal classroom interaction data analysis method and platform are designed, which collects and compares multi-modal data of teaching classrooms to provide real-time and reasonable reference data for improving teaching effect, further improving teaching quality, which is a problem to be solved at present. SUMMARY
[0005] The purpose of the present application is to provide an AI-driven multi-modal classroom interaction data analysis method, which obtains multi-modal teaching interaction data including voice and image data of different classrooms of the same teaching course, analyzes and extracts interactive data of different data types, compares interactive data corresponding to different classrooms of the same type, and further determines adjustment data based on different types of data reference. On the one hand, the multi-modal data can be used to analyze and process the classroom teaching condition from multiple data types to provide teaching guidance from different aspects, and on the other hand, the multi-modal intelligent data analysis can output responsive analysis data in real time to provide timely teaching effect feedback for the teaching process, ensure that the teaching can be adjusted in time, so that the students in the classroom can fully adapt to the teaching situation, ensure that the students in the classroom have high teaching effect, and effectively ensure the teaching quality of the classroom.
[0006] The application also aims to provide an AI-driven multi-modal classroom interaction data analysis platform, which forms multi-modal classroom teaching data through the collection of image data and voice data, realizes the extraction and comparative analysis of feature data under the processing of a comprehensive comparative analysis unit, and further realizes the reasonable analysis and processing of multi-modal data of classroom teaching, thereby providing important reference data for the adjustment and analysis of teaching behavior and being an important material basis for realizing reasonable teaching behavior adjustment.
[0007] In a first aspect, the application provides an AI-driven multi-modal classroom interaction data analysis method, which comprises: acquiring multi-modal teaching interaction data of the same teaching course, performing interactive data extraction analysis, forming classroom interaction type data corresponding to different teaching courses; performing comprehensive comparative analysis on different classroom interaction type data to form teaching interaction comparative result data; performing classroom teaching adjustment analysis according to the teaching interaction comparative result data to form classroom teaching adjustment reference data.
[0008] In the application, the method acquires multi-modal teaching interaction data including voice and image data of different classrooms of the same teaching course, performs interactive data extraction analysis on different data types, and compares the interactive data of different classrooms of the same type, and further comprehensively determines adjustment data based on different types of data reference. On the one hand, the multi-modal data can be used to analyze and process the classroom teaching condition from multiple data types to provide teaching guidance in different aspects, and on the other hand, the multi-modal intelligent data analysis can output responsive analysis data in real time to provide timely teaching effect feedback for the teaching process, ensure that the teaching can be adjusted in time, and make the students in the classroom fully adapt to the teaching situation, ensure that the students in the classroom have high teaching effect, and effectively ensure the teaching quality of the classroom.
[0009] As a possible implementation manner, the multi-modal teaching interaction data of the same teaching course is acquired, interactive data extraction analysis is performed, and classroom interaction type data corresponding to different teaching courses is formed, which comprises: performing visual tracking analysis based on image data on different multi-modal teaching interaction data of the same teaching course to form teaching visual tracking interaction data; performing sound interaction analysis based on voice data on different multi-modal teaching interaction data of the same teaching course to form teaching voice tracking interaction data; and combining the teaching visual tracking interaction data and the teaching voice tracking interaction data to form classroom interaction type data corresponding to the teaching course.
[0010] In the present application, the collection of multi-modal teaching interaction data in the classroom process of teaching courses mainly considers two types of data, one is image data, and the other is voice data. For image data, it mainly reflects the action change of teachers and students in the teaching process. The present application considers whether the students can well absorb the teaching content. The performance in the image data is mainly the concentration degree of the students to the teaching content of the teacher. Therefore, the image data obtained is mainly analyzed to form the feature data reflecting the situation of the students absorbing the teaching content. For voice data, it mainly reflects the language information interaction of teachers and students in the teaching process. The present application considers that the students will interact with the teacher in the process of absorbing the teaching content, such as asking questions, answering and other voice interaction modes. Through analyzing and extracting the voice interaction features, data information reflecting the situation of the students absorbing the teaching content from different angles and directions can be formed. The feature information extracted based on the image data and the feature data extracted based on the voice data are combined to form multi-modal teaching interaction data displayed in the teaching process of the teaching classroom.
[0011] As a possible implementation manner, for different multi-modal teaching interaction data of the same teaching course, visual tracking analysis based on image data is performed to form teaching visual tracking interaction data, including: extracting classroom interaction image data in the different multi-modal teaching interaction data of the same teaching course, performing teacher teaching position analysis based on spatial position to form corresponding classroom teacher teaching position tracking data; performing student visual angle tracking analysis based on spatial position according to the corresponding classroom interaction image data to form corresponding classroom student visual angle tracking data; and performing teaching position coordination analysis according to the classroom teacher teaching position tracking data and the classroom student visual angle tracking data to form corresponding teaching visual tracking interaction data.
[0012] In the present application, the teaching interaction data extraction target of the image data based on visual tracking is to reflect whether the students can keep coordination with the teaching speed of the teacher when absorbing the teaching content. If the coordination cannot be kept, the teacher will enter the explanation of the next knowledge point teaching content before the previous teaching content is absorbed, and the students cannot apply the previous knowledge point to the next knowledge point to understand and absorb the current knowledge point. Such teaching speed leads to the disconnection of the students absorbing the teaching content, which affects the teaching effect. Therefore, in order to ensure whether the absorption of the students can keep coordination with the teaching speed of the teacher, the mapping position of the blackboard writing or display position of the teaching content of the teacher and the visual angle position of the students in the display area is analyzed to accurately determine whether the students have obvious disconnection.
[0013] As a possible implementation, the classroom interaction image data in different multi-modal teaching interaction data of the same teaching course is extracted, spatial position-based teacher teaching position analysis is performed, and corresponding classroom teacher teaching position tracking data is formed, including: establishing a classroom space coordinate system according to the classroom interaction image data; determining the teaching gesture pointing position change data of the teacher's teaching gesture pointing position in the classroom space coordinate system according to the classroom interaction image data , n represents the number of different multi-modal teaching interaction data; the teaching pointing effective distance is set, the teaching gesture pointing position of the teacher is taken as the center, and the teaching pointing effective distance is taken as the diameter, and the corresponding teaching gesture pointing area change data is determined in the plane parallel to the teaching plane according to the corresponding teaching gesture pointing position change data .
[0014] In the present application, in order to realize the tracking of the student's visual angle position based on the image data to determine whether it is coordinated with the teacher's teaching speed, it is necessary to first determine the change of the display position of the teaching content of the teacher in the teaching process, and to accurately determine the display position of the teaching content, the space coordinate is established based on the image data to realize accurate position data acquisition. Here, the display position of the teaching content of the teacher is determined by the teaching gesture of the teacher in the classroom, whether it is blackboard writing or projection screen display, the teacher will indicate the content being transferred when actually teaching, and the gesture pointing direction recognition and analysis of the pointing direction through image data can accurately determine the specific position information of the gesture pointing on the teaching plane such as blackboard, projection screen, etc. It should be noted that the teaching gesture pointing is basically fixed gesture pointing, including finger and teaching pointing, which can be determined by feature-based deep analysis extraction of image data, and the extraction method can be image recognition processing method of big data, artificial intelligence, neural network, etc. It should be noted that the pointing direction of the teaching gesture indicates a point on the teaching plane, and the teaching content usually displays regional information such as paragraphs, formulas, pictures, etc., so the change data of the teaching gesture pointing area is formed by determining the reasonable teaching content area range with the gesture pointing point as the center, and the teaching pointing effective distance can be set according to the teaching content, or determined based on big data analysis, as long as most of the teaching content can be mapped in the area.
[0015] As a possible implementation method, based on the corresponding classroom interactive image data, the student perspective tracking analysis based on spatial position is performed to form the corresponding classroom student perspective tracking data, including: determining the perspective center position of each student in the classroom space coordinate system based on the classroom interactive image data; using the perspective center position as the origin, determining the perspective center angle change data of the student's perspective center changing with time parameters , k represents the number of different students in the multimodal teaching interaction data numbered n; set the effective viewing angle range, take the time parameter as the reference, the student's central viewing angle as the center line, and the effective viewing angle range as the cone angle, and change the data according to the viewing angle center Determine the data on changes in students' perspectives in the area pointed by teaching gestures Student perspective tracking area change data formed on the plane .
[0016] In the present invention, after determining the position change data of the teacher's teaching content area, in order to reflect whether the speed at which students absorb the teaching content is coordinated with the teacher's teaching speed, it is also necessary to determine the area where the student's perspective is focused on the teaching plane in order to show the positional relationship between the area of teaching instructions and the area where the student is focused. Considering that students will definitely acquire the content through their eyes when absorbing the teaching and intercept the teacher's body display information and other teaching data, determining the perspective orientation can accurately determine whether the students are coordinated with the teaching speed. After all, there is a correlation between the knowledge points in the teaching content, and if the current knowledge point involves the previous knowledge point, then when the previous knowledge point is not absorbed, the student will constantly review the previous knowledge point when studying the current knowledge point, and then shift their sight to acquire the previous knowledge point. The acquisition of the student's perspective is mainly the extraction of the visual center line, which can be obtained by processing the image data. Considering the distance between the student's seat and the teaching plane, the viewing angle will form a viewing area after reaching the teaching plane. However, for the human body, although the viewing angle range is relatively large, the viewing angle range that is memorable or clearly recognizable is smaller. This range can be limited by the effective viewing angle range, forming the effective range area data of the viewing angle projected onto the teaching plane. The effective viewing angle range can be set based on factors such as the distance between the student and the teaching plane and the actual teaching content, or it can be determined based on big data analysis.
[0017] As a possible implementation method, based on the classroom teacher's teaching position tracking data and the classroom student perspective tracking data, the teaching position collaborative analysis is carried out to form the corresponding teaching visual tracking interaction data, including: for different students, the corresponding student perspective tracking area change data and teaching gesture pointing area change data determining the visual angle non-coincidence area change data of the student visual angle area not coinciding with the teaching gesture pointing area in the time dimension ; determining the corresponding student visual angle non-coincidence total area according to the visual angle non-coincidence area change data of different students , , is the corresponding classroom teaching duration of the multi-modal teaching interaction data numbered n; according to the corresponding student visual angle non-coincidence total area , the cumulative non-coincidence area of the entire classroom is determined for all students .
[0018] In the present application, whether the student's absorption of teaching content is coordinated with the teacher's teaching speed can be represented by whether the student's visual angle in the time dimension coincides with the content area pointed by the teacher's teaching gesture. Considering that different teaching classrooms have individual differences, the coincidence data of a single student cannot be analyzed alone, but the coincidence of all students in the classroom is counted to represent the coordination between the student's absorption speed and the teaching speed in the corresponding teaching classroom. After all, the low coincidence of a single student and the high coincidence of other students may be due to the student not listening carefully. Referring to the data of the student group can avoid the influence of such as not paying attention in class and accidental attention on the coordination analysis.
[0019] As a possible implementation manner, for different multi-modal teaching interaction data of the same teaching course, sound interaction analysis based on voice data is performed to form teaching voice tracking interaction data, including: for different multi-modal teaching interaction data of the same teaching course, corresponding classroom voice data is extracted; according to the classroom voice data, the single voice teaching duration of the teacher is determined ; according to the single voice teaching duration and the corresponding classroom teaching duration , the corresponding interaction duration proportion is determined, wherein .
[0020] In the present application, the analysis of the speech data in the teaching classroom is mainly to determine whether there is more interaction time between the teacher and the students. The interaction time is longer in some teaching classrooms under the same teaching course, which indicates that the teaching quality is not high. After all, it is impossible to only interact without providing a quiet classroom environment for the teacher to provide continuous content teaching. Here, the proportion of the interaction time is determined by the length of the teacher's single speech teaching. It should be noted that the teacher's single speech teaching includes the teacher's speech information, the speech information of the played music and video, and other speech information of the teaching content. Considering that the total length of the teaching classroom has certain differences, it is more accurate and reasonable to provide reference data in the form of proportion.
[0021] As a possible implementation, the different classroom interaction type data are comprehensively compared and analyzed to form teaching interaction comparison result data, including: determining the corresponding unit cumulative area according to the cumulative non-overlapping area , wherein ; establishing an area numerical axis, and marking different unit cumulative areas on the area numerical axis, and determining the area center value in the following manner: arbitrarily selecting a unit cumulative area , and marking it as an area preset center point; taking the area preset center point as the center, taking the area envelope distance as the radius incremental step, gradually increasing the radius value, determining the number of unit cumulative areas added each time the area envelope distance is increased in the process of enveloping all other unit cumulative areas into a circular range, and fitting to form a corresponding area preset center envelope change function , i represents the number of times of sequentially increasing the area envelope distance; traversing all unit cumulative areas , determining the unit cumulative area corresponding to the area preset center point with the largest average change rate of all area preset center envelope change functions as the center cumulative unit area; establishing a time length numerical axis, and marking different interaction time length proportions on the time length numerical axis, and determining the time length center value in the following manner: arbitrarily selecting an interaction time length proportion , and marking it as a time length preset center point; taking the time length preset center point as the center, taking the time length envelope distance as the radius incremental step, gradually increasing the radius value, determining the number of interaction time length proportions added each time the time length envelope distance is increased in the process of enveloping all other interaction time length proportions into a circular range, and fitting to form a corresponding time length preset center envelope change function , m represents the number of times of sequentially increasing duration envelope distance; traverse all interaction duration proportions , preset center envelope change function of all durations , the interaction duration proportion corresponding to the duration preset center point with the maximum average change rate is determined as the center interaction duration proportion.
[0022] In the application, different teaching classes of the same teaching content will also cause the influence on the students' teaching content absorption due to the different teaching styles of teachers, so the comparative analysis of teaching visual tracking interaction data and the comparative analysis of teaching voice tracking interaction data between different teaching classes can not only determine the influence of different teachers on teaching effect, but also compare different classes through the same type of data to adjust the reasonable teaching behavior of different student groups. The comparison of teaching visual tracking interaction data first needs to determine a reference data benchmark. The teaching visual tracking interaction data corresponding to different teaching classes has clustering under big data, so it is reasonable to take the clustering center data as the reference standard. The application determines the center data by the step-based data volume envelope change rate of the cumulative non-overlapping area corresponding to the teaching class on the numerical axis. It can be understood that if the cumulative non-overlapping area is the center data, the sum of the numerical distance differences of other cumulative non-overlapping areas to it is the smallest, so the envelope data change rate of the envelope change of the determined step is the largest, and even the average change rate is also the largest. Similarly, the comparative analysis of teaching voice tracking interaction data is also determined by the step-based data volume envelope change rate of the interaction duration proportion on the numerical axis.
[0023] As a possible implementation manner, according to the teaching interaction comparison result data, classroom teaching adjustment analysis is performed to form classroom teaching adjustment reference data, including: setting a visual deviation threshold, combining the center cumulative unit area to form a cumulative unit area threshold range; for the unit cumulative area that does not belong to the cumulative unit area threshold range, teaching speed adjustment calibration is performed; a voice deviation threshold is set, and the center interaction duration proportion is combined to form an interaction duration proportion threshold range; for the interaction duration proportion that does not belong to the interaction duration proportion threshold range, interaction duration adjustment calibration is performed.
[0024] In the present application, after the reference data on the teaching visual tracking interaction data and the teaching voice interaction tracking data in different teaching classes of the same teaching content is obtained, the teaching classes with poor teaching effect can be determined through reasonable range limitation, and the teaching behavior adjustment reflected by the comparison result data of different types is different. The comparison and analysis result of the teaching visual tracking interaction data reflects that the speed of the students to absorb the teaching content cannot be synchronized with the teaching speed of the teacher, so the teacher needs to adjust the teaching speed, and the visual line of the teacher through the aspect of explanation process or further detailed explanation of the content. The comparison and analysis result of the teaching voice tracking interaction data reflects that the teaching content provided by the teacher cannot meet the understanding and learning needs of the students, so the teaching display content can be improved, the knowledge point content can be expanded, and the like. These comparison and analysis results have obvious and reasonable teaching behavior adjustment guiding reference significance, can provide real-time teaching behavior adjustment help, are beneficial to the improvement of the teaching effect, and provide help for the learning and absorption of the students in the class.
[0025] In a second aspect, the present application provides an AI-driven multi-modal classroom interaction data analysis platform, comprising: an image data acquisition unit, configured to acquire classroom interaction image data of classroom teaching; a voice data acquisition unit, configured to acquire classroom voice data of classroom teaching; a comprehensive comparison and analysis unit, configured to acquire the classroom interaction image data collected by the image data acquisition unit and the classroom voice data collected by the voice data acquisition unit, perform comprehensive comparison and analysis to form teaching interaction comparison result data, and perform teaching adjustment analysis according to the teaching interaction comparison result data to form classroom teaching adjustment reference data.
[0026] The platform forms multi-modal classroom teaching data through the acquisition of image data and voice data, realizes the extraction and comparison and analysis of feature data under the processing of the comprehensive comparison and analysis unit, and further realizes the reasonable analysis and processing of the multi-modal data of classroom teaching, provides important reference data for the adjustment and analysis of teaching behavior, and is an important material basis for realizing reasonable teaching behavior adjustment.
[0027] The AI-driven multi-modal classroom interaction data analysis method and platform provided by the present application have the following beneficial effects:
[0028] The method can analyze the teaching situation from multiple data types by means of multi-modal data, provide teaching guidance in different aspects, and output responsive analysis data in real time by means of multi-modal intelligent data analysis, thereby providing timely teaching effect feedback for the teaching process, ensuring that the teaching can be adjusted in time, so that the students in the class can fully adapt to the teaching situation, and the students in the class have high teaching effect, thereby effectively ensuring the teaching quality of the class.
[0029] The platform forms multi-modal classroom teaching data by collecting image data and voice data, realizes feature data extraction and comparative analysis under the processing of the comprehensive comparison and analysis unit, and then realizes reasonable analysis and processing of multi-modal classroom teaching data, thereby providing important reference data for adjustment and analysis of teaching behavior, and being an important material basis for realizing reasonable teaching behavior adjustment. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0031] Fig. 1 The steps of the AI-driven multi-modal classroom interactive data analysis method provided by the embodiments of the present application are shown in the figure.
[0032] Fig. 2 The structure schematic diagram of the AI-driven multi-modal classroom interactive data analysis platform provided by the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0034] With the progress of society and the development of science, the methods of education and teaching are gradually diversified, and adaptive teaching modes are formed for different teaching contents, thereby further improving the teaching effect.
[0035] With the development of artificial intelligence, the integration of artificial intelligence in teaching can assist in the perfection of the teaching process to improve the situation that affects the teaching effect existing in the teaching process. Since the data obtained for the teaching process is mainly processed and analyzed by artificial intelligence to form a reference result, but since the data that can be obtained for the teaching process is diverse, the results and angles formed by analyzing different data types are also different. If different types of data are analyzed and processed, the teaching process can be further improved.
[0036] Reference Figs. 1-2 The embodiment of the application provides an AI-driven multi-modal classroom interaction data analysis method, which obtains multi-modal teaching interaction data including voice and image data of different classrooms of the same teaching course, analyzes and extracts the interaction data of different data types, compares the interaction data corresponding to different classrooms of the same type, and comprehensively determines the adjustment data based on different types of data. On the one hand, the multi-modal data can be used to analyze and process the classroom teaching situation from multiple data types to provide teaching guidance from different aspects, and on the other hand, the multi-modal intelligent data analysis can output responsive analysis data in real time to provide timely teaching effect feedback for the teaching process, ensure that the teaching can be adjusted in time, so that the students in the classroom can fully adapt to the teaching situation, ensure that the students in the classroom have high teaching effect, and effectively ensure the teaching quality of the classroom.
[0037] The AI-driven multi-modal classroom interaction data analysis method specifically includes the following steps:
[0038] S1: Obtain multi-modal teaching interaction data of the same teaching course, analyze and extract the interaction data, and form classroom interaction type data corresponding to different teaching courses.
[0039] The multi-modal teaching interaction data of the same teaching course is obtained, the interaction data is analyzed and extracted, and the classroom interaction type data corresponding to different teaching courses is formed, including: the different multi-modal teaching interaction data of the same teaching course is analyzed and extracted based on image data, visual tracking analysis is performed, and teaching visual tracking interaction data is formed; the different multi-modal teaching interaction data of the same teaching course is analyzed and extracted based on voice data, sound interaction analysis is performed, and teaching voice tracking interaction data is formed; the teaching visual tracking interaction data and the teaching voice tracking interaction data are combined to form the classroom interaction type data corresponding to the teaching course.
[0040] The collection of multi-modal teaching interaction data in the classroom process of teaching courses mainly considers two types of data, one is image data, and the other is voice data. For image data, the main consideration is the action change of teachers and students in the teaching process. The application considers whether the students can well absorb the teaching content. The performance in the image data is mainly the concentration degree of the students to the teaching content of the teacher. Therefore, the image data obtained is mainly the analysis of the visual tracking state of the students to form feature data that can reflect the situation of the students absorbing the teaching content. For voice data, the main consideration is the language information interaction of teachers and students in the teaching process. The application considers that the students will interact with the teacher in the process of absorbing the teaching content, such as asking questions, answering, and other voice interaction modes. By analyzing and extracting the voice interaction features, data information reflecting the situation of students absorbing teaching content from different angles and directions can be formed. The feature information extracted based on the image data and the feature data extracted based on the voice data are combined to form multi-modal teaching interaction data displayed in the teaching process of the teaching classroom.
[0041] The same teaching course different multi-modal teaching interaction data is analyzed based on image data visual tracking to form teaching visual tracking interaction data, including: extracting classroom interaction image data in the same teaching course different multi-modal teaching interaction data, performing teacher teaching position analysis based on spatial position to form corresponding classroom teacher teaching position tracking data; based on the corresponding classroom interaction image data, performing student visual angle tracking analysis based on spatial position to form corresponding classroom student visual angle tracking data; based on the classroom teacher teaching position tracking data and the classroom student visual angle tracking data, performing teaching position coordination analysis to form corresponding teaching visual tracking interaction data.
[0042] The teaching interaction data extraction target of the image data based on visual tracking is to reflect whether the students can keep pace with the teaching speed of the teacher when absorbing the teaching content. If not, it will lead to the teacher entering the explanation of the next knowledge point teaching content before the previous teaching content is absorbed, and the students cannot apply the previous knowledge point to the next knowledge point to understand and absorb the current knowledge point. Such teaching speed leads to the disconnection of the students' absorption of the teaching content, which affects the teaching effect. Therefore, in order to ensure whether the students' absorption can keep pace with the teaching speed of the teacher, the mapping position of the teacher's teaching content board or display position and the student's visual angle position in the display area is analyzed to accurately determine whether the students have obvious disconnection.
[0043] The classroom interaction image data in different multi-modal teaching interaction data of the same teaching course is extracted, the teacher teaching position analysis based on the spatial position is carried out, and the corresponding classroom teacher teaching position tracking data is formed, including: establishing a classroom spatial coordinate system according to the classroom interaction image data; determining the teaching gesture pointing position change data of the teacher's teaching gesture pointing position in the classroom spatial coordinate system according to the classroom interaction image data , n represents the number of different multi-modal teaching interaction data; the teaching pointing effective distance is set, the teaching gesture pointing position of the teacher is taken as the center, and the teaching pointing effective distance is taken as the diameter, and the corresponding teaching gesture pointing area change data is determined in the plane parallel to the teaching version .
[0044] To realize the tracking of the student's visual angle position based on the image data to determine whether to cooperate with the teaching speed of the teacher, first, the change of the display position of the teaching content of the teacher in the teaching process needs to be determined, and to accurately determine the display position of the teaching content, the spatial coordinate is established based on the image data to realize the accurate position data acquisition. Here, the display position of the teaching content of the teacher is determined by the teaching gesture of the teacher in the classroom, whether it is blackboard writing on the blackboard or display on the projection screen, the teacher will indicate the content being transferred when actually teaching, and the gesture pointing direction recognition and analysis of the pointing direction based on the image data can accurately determine the specific position information of the gesture pointing on the teaching plane such as the blackboard and the projection screen. It should be noted that the teaching gesture pointing is basically a fixed gesture pointing, including finger and teaching pointing, which can be determined by feature-based deep analysis extraction of image data, and the extraction method can be image recognition processing method of big data, artificial intelligence, neural network, etc. It should be noted that the pointing direction of the teaching gesture indicates a point on the teaching plane, and the teaching content usually displays regional information such as paragraphs, formulas and pictures, so the reasonable teaching content area range is determined based on the point of the gesture pointing to form the teaching gesture pointing area change data, and the teaching pointing effective distance can be set according to the teaching content, or it can be determined based on big data analysis, as long as most of the teaching content can be mapped in the area.
[0045] According to the corresponding classroom interaction image data, the student visual angle tracking analysis based on the spatial position is carried out, and the corresponding classroom student visual angle tracking data is formed, including: determining the visual angle center position of each student in the classroom spatial coordinate system according to the classroom interaction image data; determining the visual angle center angle change data of the visual angle center of the student with the time parameter change with the visual angle center position as the origin , k represents the number of different students in the multimodal teaching interactive data numbered n; set the effective visual angle range angle, take the time parameter as the reference, take the central visual angle of the student as the center line, take the effective visual angle range angle as the cone angle, according to the visual angle center angle change data determine the change data of the student's visual angle in the teaching gesture pointing area formed on the plane where the student's visual angle tracking area changes
[0046] After determining the teaching content area position change data of the teacher, it is necessary to reflect whether the student's absorption speed of the teaching content is coordinated with the teacher's teaching speed, and the student's visual angle in the area observed by the teaching plane needs to be determined to show the positional relationship between the teaching instruction area and the area observed by the student. Considering that the student will definitely acquire the content through the observation of the eyes and intercept the body display information of the teacher and other teaching data when absorbing the teaching, the determination of the visual angle direction can accurately determine whether the student is coordinated with the teaching speed. After all, there is a correlation between the knowledge points in the teaching content, and if the current knowledge point is related to the previous knowledge point, the student will constantly review the previous knowledge point when learning the current knowledge point without absorbing the previous knowledge point, and then shift the line of sight to acquire the previous knowledge point. The collection of the student's visual angle is mainly the extraction of the visual center line, which can be obtained through the processing of image data. Considering that there is a distance between the student's seat and the teaching plane, the visual angle will form a visual angle area after reaching the teaching plane. However, for the human body, although the visual angle range is relatively large, the visual angle range with memory or obvious recognition is relatively small, which can be limited by the effective visual angle range angle to form the effective range area data of the visual angle projection to the teaching plane. The effective visual angle range angle can be set according to the distance between the student and the teaching plane, the actual teaching content, etc., or determined according to big data analysis.
[0047] According to the classroom teacher teaching position tracking data and the classroom student visual angle tracking data, the teaching position coordination analysis is performed to form the corresponding teaching visual tracking interactive data, including: for different students, according to the corresponding student visual angle tracking area change data and teaching gesture pointing area change data , determine the visual angle non-coincidence area change data that the student's visual angle area does not coincide with the teaching gesture pointing area in the time dimension; according to the visual angle non-coincidence area change data of different students, determine the corresponding student visual angle non-coincidence total area , wherein , is the teaching time length corresponding to the multimodal teaching interaction data numbered n; for all students, according to the corresponding student perspective non-overlapping total area , determine the cumulative non-overlapping area of the entire classroom , wherein .
[0048] Whether the student's absorption of the teaching content is coordinated with the teacher's teaching speed can be represented by whether the student's perspective gaze area in the time dimension coincides with the content area pointed by the teacher's teaching gesture. Considering that different teaching classes have individual differences among students, the coincidence data of individual students cannot be analyzed alone, but the coincidence of all students in the classroom is statistically analyzed to represent the coordination between the student's absorption speed and the teaching speed in the corresponding teaching class. After all, if a single student has a low coincidence rate while other students have a high coincidence rate, it may be because the student is not paying attention to the class. Referring to the data of the student group can avoid the influence of such as not paying attention to the class and accidental gaze on the coordination analysis.
[0049] For different multimodal teaching interaction data of the same teaching course, sound interaction analysis based on voice data is performed to form teaching voice tracking interaction data, including: for different multimodal teaching interaction data of the same teaching course, extracting corresponding classroom voice data; according to the classroom voice data, determining the single voice teaching time length of the teacher ; according to the single voice teaching time length and the corresponding classroom teaching time length , determine the interaction time length ratio , wherein .
[0050] The analysis of the voice data in the teaching class is mainly to determine whether there is more interaction time between the teacher and the student. The interaction time is longer in some teaching classes under the same teaching course, indicating that the teaching quality is not high. After all, it is impossible to only interact without providing a quiet classroom environment for the teacher to continue teaching content. Here, the interaction time length ratio is determined by the time length of the teacher's single voice teaching. It should be noted that the teacher's single voice teaching includes the teacher's voice information, the voice information of the played music and video, and the voice information of the teaching content. Considering that the total time length of the teaching class has certain differences, it is more accurate and reasonable to provide reference data in the form of ratio.
[0051] S2: Comprehensive comparative analysis of different classroom interaction type data to form teaching interaction comparison result data.
[0052] Comprehensive comparative analysis of different classroom interaction type data to form teaching interaction comparison result data, including: according to the cumulative non-overlapping area , determine the corresponding unit cumulative area , wherein, ; establish an area numerical axis, and mark different unit cumulative areas on the area numerical axis, and determine the area center value in the following manner: arbitrarily select one unit cumulative area , mark it as the area preset center point; take the area preset center point as the center, take the area envelope distance as the radius, incrementally increase the radius value, determine the number of unit cumulative areas that are newly added each time the area envelope distance is incremented in the process of enveloping all other unit cumulative areas into a circular range, and fit the corresponding area preset center envelope change function , i represents the number of times of sequentially increasing the area envelope distance; traverse all unit cumulative areas , and determine the unit cumulative area corresponding to the area preset center point with the largest average change rate of all area preset center envelope change functions as the center cumulative unit area; establish a time length numerical axis, and mark different interactive time length proportions on the time length numerical axis, and determine the time length center value in the following manner: arbitrarily select one interactive time length proportion , mark it as the time length preset center point; take the time length preset center point as the center, take the time length envelope distance as the radius, incrementally increase the radius value, determine the number of interactive time length proportions that are newly added each time the time length envelope distance is incremented in the process of enveloping all other interactive time length proportions into a circular range, and fit the corresponding time length preset center envelope change function , m represents the number of times of sequentially increasing the time length envelope distance; traverse all interactive time length proportions , and determine the interactive time length proportion corresponding to the time length preset center point with the largest average change rate of all time length preset center envelope change functions as the center interactive time length proportion.
[0053] Different teaching classes of the same teaching content will also cause the influence on the students' teaching content absorption due to the different teaching styles of teachers, so the comparison and analysis of the teaching visual tracking interaction data and the comparison and analysis of the teaching voice tracking interaction data between different teaching classes can not only determine the influence of different teachers on the teaching effect, but also compare the different classes through the same type of data to adjust the reasonable teaching behavior of different student groups. The comparison of the teaching visual tracking interaction data first needs to determine a reference data benchmark, and the teaching visual tracking interaction data corresponding to different teaching classes has clustering under big data, so it is reasonable to take the center data of clustering as the reference standard. The application determines the center data by the step-based data volume envelope change rate of the cumulative non-overlapping area corresponding to the teaching class on the numerical axis. It can be understood that if the cumulative non-overlapping area is the center data, the sum of the numerical distance differences of other cumulative non-overlapping areas to it is the smallest, so the envelope data change rate of the envelope change of the determined step will be the largest, even if it is the average change rate. Similarly, the comparison and analysis of the teaching voice tracking interaction data is also determined by the step-based data volume envelope change rate of the interactive time length proportion on the numerical axis.
[0054] S3: According to the teaching interaction comparison result data, the classroom teaching adjustment analysis is carried out to form the classroom teaching adjustment reference data.
[0055] According to the teaching interaction comparison result data, the classroom teaching adjustment analysis is carried out to form the classroom teaching adjustment reference data, including: setting a visual deviation threshold, combining the center cumulative unit area to form a cumulative unit area threshold range; adjusting the unit cumulative area The teaching class not belonging to the cumulative unit area threshold range is adjusted for teaching speed; a voice deviation threshold is set, and an interactive time length proportion threshold range is formed by combining the center interactive time length proportion; the interactive time length proportion The teaching class not belonging to the interactive time length proportion threshold range is adjusted for the interactive time length.
[0056] After the reference data on the teaching visual tracking interaction data and the teaching voice interaction tracking data in different teaching classes with the same teaching content is obtained, the teaching classes with poor teaching effect can be determined through reasonable range limitation, and the teaching behavior adjustment reflected by the comparison result data of different types is different. The comparison analysis result of the teaching visual tracking interaction data reflects that the speed of the students to absorb the teaching content cannot be synchronized with the teaching speed of the teacher, so the teacher needs to adjust the teaching speed, and the comparison analysis result of the teaching voice tracking interaction data reflects that the teaching content provided by the teacher cannot meet the understanding and learning needs of the students, so the teaching display content can be improved and the knowledge point content can be expanded to achieve the adjustment. These comparison analysis results have obvious and reasonable teaching behavior adjustment guiding reference significance, and can provide real-time teaching behavior adjustment help, which is beneficial to the improvement of the teaching effect and the learning absorption of the students in the class.
[0057] The application also provides an AI-driven multi-modal classroom interaction data analysis platform, which comprises: an image data acquisition unit, configured to acquire classroom interaction image data of classroom teaching; a voice data acquisition unit, configured to acquire classroom voice data of classroom teaching; and a comprehensive comparison and analysis unit, configured to acquire the classroom interaction image data acquired by the image data acquisition unit and the classroom voice data acquired by the voice data acquisition unit, perform comprehensive comparison and analysis to form teaching interaction comparison result data, and perform teaching adjustment analysis according to the teaching interaction comparison result data to form classroom teaching adjustment reference data.
[0058] The platform forms multi-modal classroom teaching data through the acquisition of image data and voice data, realizes the extraction and comparison analysis of feature data under the processing of the comprehensive comparison and analysis unit, and further realizes the reasonable analysis and processing of the multi-modal data of classroom teaching, thereby providing important reference data for the adjustment analysis of teaching behavior and being an important material basis for realizing reasonable teaching behavior adjustment.
[0059] In summary, the AI-driven multi-modal classroom interaction data analysis method and platform provided by the embodiments of the application have the following beneficial effects:
[0060] The method can analyze the interactive data of different data types by acquiring multi-modal teaching interactive data including voice and image data of different classes of the same teaching course, and comparing the interactive data corresponding to different classes between the same type of data, and then comprehensively determining the adjustment data based on different types of data reference. On the one hand, the multi-modal data can be used to analyze and process the classroom teaching condition from multiple data types to provide teaching guidance in different aspects. On the other hand, the multi-modal intelligent data analysis can output responsive analysis data in real time to provide timely teaching effect feedback for the teaching process, ensure that the teaching can be adjusted in time, so that the students in the class can fully adapt to the teaching situation, ensure that the students in the class have high teaching effect, and effectively ensure the teaching quality of the class.
[0061] The platform forms multi-modal classroom teaching data by collecting image data and voice data, realizes feature data extraction and comparison analysis under the processing of the comprehensive comparison and analysis unit, and then realizes reasonable analysis and processing of multi-modal classroom teaching data, provides important reference data for adjustment and analysis of teaching behavior, and is an important material basis for realizing reasonable teaching behavior adjustment.
[0062] In the embodiments of the present application, "indication" can include direct indication and indirect indication, and can also include explicit indication and implicit indication. The information indicated by a certain information is referred to as to-be-indicated information, and in the specific implementation process, there are many ways to indicate the to-be-indicated information, for example, but not limited to, the to-be-indicated information can be directly indicated, such as the to-be-indicated information itself or the index of the to-be-indicated information. The to-be-indicated information can also be indirectly indicated by indicating other information, wherein the other information and the to-be-indicated information have an association relationship. The to-be-indicated information can also be indicated only by a part of the to-be-indicated information, and the other part of the to-be-indicated information is known or agreed in advance. For example, the indication of a specific information can also be realized by means of the arrangement order of each information agreed in advance (for example, the protocol stipulates), thereby reducing the indication overhead to a certain extent. At the same time, the common part of each information can be identified and uniformly indicated to reduce the indication overhead caused by separately indicating the same information.
[0063] In addition, the specific indication manner can also be various existing indication manners, for example but not limited to the indication manners described above and various combinations thereof. The specific details of various indication manners can refer to the prior art, and will not be described herein. As can be known from the above, for example, when multiple information of the same type needs to be indicated, the indication manners of different information can be different. In the specific implementation process, the required indication manner can be selected according to the specific needs, and the selected indication manner is not limited by the embodiments of the application. In this way, the indication manners involved in the embodiments of the application should be understood as covering various methods that can enable the indicating party to know the information to be indicated.
[0064] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information and sent separately, and the sending period and / or sending occasion of the sub-information can be the same or different. The specific sending method is not limited by the embodiments of the application. The sending period and / or sending occasion of the sub-information can be predefined, for example, predefined according to a protocol, or configured by the sending end device by sending configuration information to the receiving end device.
[0065] The "predefined" or "preconfigured" can be implemented by pre-storing corresponding codes, tables or other methods that can be used to indicate related information in the device, and the specific implementation manner is not limited by the embodiments of the application. The "storage" can mean storage in one or more memories. The one or more memories can be separately arranged or integrated in the encoder or decoder, processor or communication device. The one or more memories can be partially separately arranged and partially integrated in the decoder, processor or communication device. The type of memory can be any form of storage medium, and the embodiments of the application do not limit this.
[0066] The "protocol" involved in the embodiments of the application can refer to a protocol family in the communication field, a standard protocol similar to the protocol family frame structure, or a related protocol applied to a future communication system, and the embodiments of the application do not make specific limitations.
[0067] In the embodiments of the application, "when", "in the case of", "if" and "if" and the like all refer to the device making corresponding processing under certain objective conditions, and are not limited by time, and do not require the device to have a judgment action when implemented, nor does it mean that there are other limitations.
[0068] In the description of the embodiments of the present application, unless otherwise specified, " / " represents that the objects before and after the " / " are in an "or" relationship, for example, A / B can represent A or B; "and / or" in the embodiments of the present application is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, wherein A and B can be singular or plural. In addition, in the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or the like means any combination of the items, including any combination of single item or multiple items. For example, at least one of a, b or c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, wherein a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, "first", "second", and the like are used to distinguish the same items or similar items with basically the same function and role. Those skilled in the art can understand that "first", "second", and the like do not limit the quantity and execution order, and "first", "second", and the like do not necessarily mean different. At the same time, in the embodiments of the present application, "exemplary" or "for example" means to serve as an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, "exemplary" or "for example" is used to present the relevant concept in a specific manner, for understanding.
[0069] It should be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0070] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of random access memory (RAM) can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0071] The above-described embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center through a wired (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0072] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects, but can also represent an "and / or" relationship, which can be understood in the context before and after.
[0073] In this application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0074] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined by their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0075] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0076] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0077] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0078] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0079] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.
[0080] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0081] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An AI-driven multimodal classroom interaction data analysis method, characterized by: Configured to: Obtain multimodal teaching interaction data of the same teaching course, extract and analyze the interaction data, and form classroom interaction type data corresponding to different teaching courses; Conducting comprehensive comparative analysis on the different types of classroom interaction data to form teaching interaction comparison result data; Based on the teaching interaction comparison result data, classroom teaching adjustment analysis is conducted to form classroom teaching adjustment reference data.
2. The AI-driven multimodal classroom interaction data analysis method according to claim 1 is characterized in that: The method of obtaining multimodal teaching interaction data of the same teaching course, extracting and analyzing the interaction data, and forming classroom interaction type data corresponding to different teaching courses includes: Performing visual tracking analysis based on image data on different multimodal teaching interaction data of the same teaching course to form teaching visual tracking interaction data; Performing voice interaction analysis based on voice data on different multimodal teaching interaction data of the same teaching course to form teaching voice tracking interaction data; The teaching visual tracking interaction data and the teaching voice tracking interaction data are collected to form the classroom interaction type data corresponding to the teaching course.
3. The AI-driven multimodal classroom interaction data analysis method according to claim 2 is characterized in that: The visual tracking analysis based on image data is performed on the different multimodal teaching interaction data of the same teaching course to form teaching visual tracking interaction data, including: Extracting classroom interaction image data from different multimodal teaching interaction data of the same teaching course, performing teacher teaching position analysis based on spatial position, and forming corresponding classroom teacher teaching position tracking data; Performing student perspective tracking analysis based on spatial positions according to the corresponding classroom interactive image data to form corresponding classroom student perspective tracking data; Based on the classroom teacher's teaching position tracking data and the classroom student perspective tracking data, a teaching position coordination analysis is performed to form the corresponding teaching visual tracking interaction data.
4. The AI-driven multimodal classroom interaction data analysis method according to claim 3 is characterized in that: The extracting of classroom interaction image data from different multimodal teaching interaction data of the same teaching course, performing teacher teaching position analysis based on spatial position, and forming corresponding classroom teacher teaching position tracking data includes: Establishing a classroom space coordinate system according to the classroom interactive image data; According to the classroom interactive image data, determine the teacher's teaching gesture pointing position in the classroom space coordinate system with the time parameter teaching gesture pointing position change data , n represents the number of different multimodal teaching interaction data; Set the effective distance of teaching pointing, take the teacher's teaching gesture pointing position as the center of the circle, take the teaching pointing effective distance as the diameter, and change the data of the corresponding teaching gesture pointing position. , determine the corresponding teaching gesture pointing area change data on a plane parallel to the teaching page .
5. The AI-driven multimodal classroom interaction data analysis method according to claim 4 is characterized in that: The method of performing spatial position-based student perspective tracking analysis based on the corresponding classroom interactive image data to form corresponding classroom student perspective tracking data includes: Determining the viewing center position of each student in the classroom space coordinate system based on the classroom interactive image data; Taking the center of view as the origin, determine the change data of the student's view center as the time parameter changes. , k represents the number of different students in the multimodal teaching interaction data numbered n; Set the effective viewing angle range, take the time parameter as reference, take the student's central viewing angle as the center line, take the effective viewing angle range as the cone angle, and change the data according to the viewing angle center angle. Determine the data of the change of the student's perspective in the area pointed by the teaching gesture Student perspective tracking area change data formed on the plane .
6. The AI-driven multimodal classroom interaction data analysis method according to claim 5, characterized in that: The method of performing teaching position synergy analysis based on the classroom teacher teaching position tracking data and the classroom student perspective tracking data to form the corresponding teaching visual tracking interactive data includes: For different students, track regional change data based on the corresponding student perspective and the teaching gesture pointing area change data , determine the change data of the non-overlapping area of the student's perspective area that does not overlap with the teaching gesture pointing area in the time dimension ; Data on the change of non-overlapping areas according to the perspectives of different students , determine the corresponding total non-overlapping area of the student's perspective ,in, , The classroom teaching duration corresponding to the multimodal teaching interaction data numbered n; For all students, the total non-overlapping area according to the corresponding student perspective , determine the cumulative non-overlapping area of the entire classroom ,in, .
7. The AI-driven multimodal classroom interaction data analysis method according to claim 6, characterized in that: The method of performing voice interaction analysis based on voice data on different multimodal teaching interaction data of the same teaching course to form teaching voice tracking interaction data includes: Extracting corresponding classroom voice data for different multimodal teaching interaction data of the same teaching course; Determine the teacher's single voice teaching time based on the classroom voice data ; According to the single voice teaching time and the corresponding classroom teaching time , determine the corresponding interaction time ratio ,in, .
8. The AI-driven multimodal classroom interaction data analysis method according to claim 7, characterized in that: The comprehensive comparative analysis of the different classroom interaction type data to form teaching interaction comparison result data includes: According to the cumulative non-overlapping area , determine the corresponding unit cumulative area ,in, ; Create an area value axis and convert the cumulative area of different units into The area center is calibrated on the area value axis and the area center value is determined in the following manner: Arbitrarily select one of the unit cumulative areas , mark it as the preset center point of the area; With the preset center point of the area as the center of the circle and the area envelope distance as the radius increment step, the radius value is gradually increased to determine the cumulative area of all other units. The unit cumulative area added by each incremental increase of the area envelope distance during the process of enveloping into the circular range The number of, and fitting to form the corresponding area preset center envelope change function , i represents the number of times the area envelope distance is increased sequentially; Traverse all the unit cumulative areas , all the areas are preset with the center envelope change function The unit cumulative area corresponding to the preset center point of the area with the largest average change rate Determined as the central cumulative unit area; Establish a time value axis and divide the different interaction time ratios The axis of the duration value is marked and the duration center value is determined in the following manner: Randomly select one of the interaction time ratios , mark it as the preset center point of duration; With the preset center point of the duration as the center of the circle, and the duration envelope distance as the radius increment step, gradually increase the radius value to determine the proportion of all other interaction durations. The proportion of the interaction duration added each time the duration envelope distance is increased during the process of enveloping into the circular range The number of, and fitting to form the corresponding duration preset center envelope change function , m represents the number of times the duration envelope distance is sequentially increased; Traverse all the interaction duration ratios , all the time length preset center envelope change function The interaction duration ratio corresponding to the preset center point of the duration with the largest average change rate Determined as the proportion of center interaction time.
9. The AI-driven multimodal classroom interaction data analysis method according to claim 8, characterized in that: The method of performing classroom teaching adjustment analysis based on the teaching interaction comparison result data to form classroom teaching adjustment reference data includes: Setting a visual deviation threshold value, and combining the central cumulative unit area to form a cumulative unit area threshold range; For the unit cumulative area For teaching classrooms that do not fall within the cumulative unit area threshold range, the teaching speed is adjusted and calibrated; Setting a voice deviation threshold, and combining the center interaction duration ratio to form an interaction duration ratio threshold range; The proportion of interaction time For teaching classes that do not fall within the interaction time ratio threshold range, the interaction time will be adjusted and calibrated.
10. An AI-driven multimodal classroom interaction data analysis platform, characterized by: The AI-driven multimodal classroom interaction data analysis method according to any one of claims 1 to 9 comprises: An image data acquisition unit, used to acquire classroom interactive image data of classroom teaching; A voice data collection unit, used to obtain classroom voice data of classroom teaching; The comprehensive comparison and analysis unit is used to obtain the classroom interaction image data collected by the image data acquisition unit and the classroom voice data collected by the voice data acquisition unit, perform comprehensive comparison and analysis to form teaching interaction comparison result data, and perform classroom teaching adjustment analysis based on the teaching interaction comparison result data to form classroom teaching adjustment reference data.
Citation Information
Patent Citations
Smart classroom teaching behavior analysis method, storage medium and smart television
CN111311131A
Teacher-student behavior analysis system based on classroom videos
CN111709358A
Classroom teaching quality monitoring method and system, terminal and storage medium
CN112766130A
Course quality evaluation and improvement method based on student visual attention and teacher behaviors
CN113506027A
Interaction analysis method and system based on cognition and social network mining
CN118733994A