Double-teacher classroom intelligent teaching terminal based on AI teaching assistance
By combining multimodal data collection and intelligent decision-making algorithms with an AI teaching assistant system, the problem of difficulty in comprehensively understanding teaching pace, content difficulty and participation fairness in existing technologies has been solved. This has enabled efficient and low-cost optimization of teaching interaction and improved the teaching effectiveness of dual-teacher classrooms.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-10
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technical solutions, due to their high cost and complexity in single-dimensional behavioral analysis, cannot comprehensively understand teaching pace, content difficulty, and fairness of participation. This leads to a disconnect between decision-making recommendations and actual teaching needs, making it difficult to promote and apply them in educational practice.
The system employs a smart teaching terminal for dual-teacher classrooms based on AI teaching assistants. It acquires visual and audio data through a multimodal data acquisition module, and combines lightweight target detection and real-time speech recognition technologies to identify students' hand-raising behavior and key knowledge points taught by teachers. The system then uses an intelligent decision-making module to integrate algorithms based on time correlation, participation balance, and rhythm adaptation to generate interactive recommendation information.
It significantly enhances the classroom presence and interactive decision-making efficiency of online lecturers, reduces the workload of offline tutors, enables a deeper understanding of the teaching context and adaptive decision-making, and improves the quality of dual-teacher classroom interaction and teaching equity.
Smart Images

Figure CN121789526A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to intelligent teaching assistance technology, specifically to a smart teaching terminal for dual-teacher classrooms based on AI teaching assistants. Background Technology
[0002] In the context of the rapid development of educational informatization, the dual-teacher classroom model has emerged as an innovative teaching approach. This model cleverly integrates the high-quality resources of renowned online teachers and offline tutors, aiming to break through geographical limitations, significantly expand the reach of quality education, and enable more students to enjoy high-quality teaching services.
[0003] In existing technological applications, several technical solutions have been proposed and are being tested to help dual-teacher classrooms better achieve teaching objectives and improve teaching quality. Among these, computer vision and artificial intelligence technologies have become popular choices, aiming to enhance online instructors' ability to perceive the classroom status of students offline through technological means. For example, cameras can be used to capture student images, and then complex visual analysis models can be used to identify and analyze student postures, movements, and other behaviors, hoping to obtain information on student classroom participation and thus assist online teachers in adjusting their teaching strategies.
[0004] However, in practical applications, the computational complexity of the visual analysis models relied upon by existing solutions is high, leading to a significant increase in system deployment costs. This makes it unaffordable for many schools with limited educational resources. Moreover, in actual operation, existing solutions often focus only on a single dimension of student behavior, such as whether students raise their hands or maintain proper posture, lacking a comprehensive understanding of the multi-dimensional factors in the teaching context. For example, they do not fully consider key factors such as the pace of teaching, changes in the difficulty of teaching content, and the fairness of student participation. This results in a serious disconnect between the generated decision-making suggestions and real, dynamic teaching needs, failing to provide targeted and practical guidance for online teachers and hindering the improvement of the quality and fairness of teaching interaction in dual-teacher classrooms. Therefore, large-scale promotion and application in educational practice is difficult. Summary of the Invention
[0005] The purpose of this invention is to provide a smart teaching terminal for dual-teacher classrooms based on AI teaching assistants, in order to solve the problem that existing technical solutions, due to their high cost and complexity of single-dimensional behavioral analysis, cannot comprehensively understand the teaching pace, content difficulty and participation fairness, resulting in decision-making suggestions that are out of touch with real teaching needs and difficult to be practically promoted.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a smart teaching terminal for a dual-teacher classroom based on AI teaching assistants, comprising a hardware terminal deployed in an offline classroom, the hardware terminal including a multimodal data acquisition module for collecting visual data, audio data, and environmental data from the offline classroom, and further comprising:
[0007] The AI teaching assistant system communicates and connects with hardware terminals. The AI teaching assistant system includes:
[0008] The student behavior analysis module is used to identify students' hand-raising behavior based on video data using a lightweight object detection model.
[0009] The classroom content comprehension module is used to identify key knowledge points and questions from the lecturer based on audio data through real-time speech recognition and keyword matching technology, and to analyze the current teaching pace.
[0010] The intelligent decision-making module is used to generate interactive recommendation information and output it to the teacher's end based on the hand-raising behavior, questioning, and teaching pace. It uses a decision-making algorithm that integrates time correlation, participation balance, and pace adaptation.
[0011] Furthermore, the classroom content comprehension module quantifies the teaching pace and status in the following ways. :
[0012] ;
[0013] in:
[0014] To slide the time window The number of key knowledge points identified in the lesson;
[0015] when At that time, it is judged to be a fast-paced state;
[0016] when If the time is right, it is considered a slow-paced state; otherwise, it is a normal-paced state.
[0017] and This is the preset rhythm state threshold.
[0018] Furthermore, the student behavior analysis module identifies hand-raising behavior in the following ways:
[0019] Within the pre-defined student interaction area, the duration of the detected hand-raising target frame is monitored. Exceeding the time threshold When a hand-raising action is detected, its mathematical expression is:
[0020] .
[0021] Furthermore, the intelligent decision-making module calculates a dynamic priority score for each hand-raising action using the following formula. :
[0022] ;
[0023] in, The basic priority score is determined by the time difference between the time the hand-raising action occurs and the time the question is recognized.
[0024] A rhythm adaptation factor related to the current teaching pace.
[0025] A fairness-enhancing factor related to the number of times the student has interacted historically.
[0026] Furthermore, basic priority scores The calculation method is as follows:
[0027] ;
[0028] in, Based on the score;
[0029] , representing the time difference between the time the hand-raising action occurs and the time the question is recognized;
[0030] λ is the preset attenuation coefficient.
[0031] Furthermore, rhythm adaptation factor The value selection strategy is as follows:
[0032] In a fast-paced environment, higher scores are assigned to the top k students who react fastest when raising their hands. value;
[0033] In a slower-paced environment, higher allocations are made to students with lower history engagement. value;
[0034] Under normal rhythm conditions Take the baseline value as 1.0.
[0035] Furthermore, equity promotion factors Calculated using the following formula:
[0036] ;
[0037] in, This represents the total number of times the student interacted with history in this class.
[0038] This represents the highest number of student interactions so far in this lesson.
[0039] This is a preset fairness adjustment coefficient.
[0040] Furthermore, the classroom content comprehension module determines the expected cognitive difficulty L of the question through the following steps:
[0041] Extracting the short-time average energy of the question sentence speech signal and short-time average zero crossing rate ;
[0042] The extracted features are compared with a preset threshold to determine the state of the speech features.
[0043] Simultaneously, the text of the question is analyzed and matched with the basic cognitive keyword library and the advanced cognitive keyword library;
[0044] Based on the matching results of speech features and keywords, the expected cognitive difficulty level L is determined to be high, medium or low according to preset rules.
[0045] Furthermore, rhythm adaptation factor The value is also affected by the expected cognitive difficulty level L:
[0046] When L is at a high level, the same applies as the slow-paced state. Value retrieval strategy;
[0047] When L is at a low level, the same applies as in the fast-paced state. Value retrieval strategy.
[0048] Furthermore, the AI teaching assistant system also includes a teaching rhythm suggestion module, which generates rhythm adjustment suggestions to the main teacher when the teaching rhythm state R is detected to be continuously abnormal.
[0049] Compared with existing technologies, the AI-assisted smart teaching terminal for dual-teacher classrooms provided by this invention, by constructing a collaborative system integrating multimodal data acquisition, lightweight behavior recognition, deep understanding of teaching content, and intelligent decision-making, fundamentally overcomes the shortcomings of existing dual-teacher classrooms, such as blind interaction, unbalanced tutoring, and technical solutions detached from practical applications. It can not only significantly improve the classroom presence and interactive decision-making efficiency of online lecturers, but also effectively reduce the workload of offline tutors. By comprehensively considering teaching pace, content difficulty, and fairness of participation, it achieves a deep understanding of the teaching situation and adaptive decision-making, thereby promoting the overall improvement of the quality of dual-teacher classroom interaction and teaching fairness.
[0050] The lightweight target detection model used in the student behavior analysis module can accurately identify students' hand-raising behavior in a low-computational-cost and efficient manner, avoiding the high resource consumption and privacy risks of complex visual analysis algorithms. At the same time, by introducing a time-based verification mechanism, the behavior is only confirmed as valid when the detected hand-raising target box continuously exceeds a preset threshold. This enables highly robust interactive signal capture in noisy classroom environments, providing a reliable basis for subsequent decision-making.
[0051] Through the real-time speech recognition and keyword matching technology of the classroom content comprehension module, it can automatically extract key knowledge points and questions from the teacher's audio, and quantify the teaching rhythm based on the knowledge point density within the sliding time window. This enables the AI teaching assistant to accurately grasp the classroom process like an experienced teacher, providing dynamic contextual awareness support for intelligent decision-making.
[0052] By integrating decision-making algorithms based on time correlation, participation balance, and rhythm adaptation through the intelligent decision-making module, a dynamic priority score is calculated for each hand-raising behavior. The optimal interactive recommendation sequence is generated by comprehensively considering multiple factors such as hand-raising response time, teaching rhythm, problem cognition difficulty, and students' historical participation. This ensures response efficiency while proactively promoting classroom fairness and achieving personalized teaching assistance.
[0053] Through continuous monitoring and anomaly detection by the teaching rhythm suggestion module, suggestions for rhythm adjustment can be generated in a timely manner when the teaching rhythm status value deviates from the normal range. This helps teachers optimize classroom process management based on objective data, improve teaching fluency and adaptability, and further enhance the overall teaching effect of the dual-teacher classroom. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0055] Figure 1 The intelligent interactive decision-making logic diagram provided in the embodiments of the present invention;
[0056] Figure 2 A logic diagram for teaching rhythm analysis and cognitive difficulty judgment provided in this embodiment of the invention;
[0057] Figure 3 This is a logic diagram for recognizing student hand-raising behavior provided in an embodiment of the present invention;
[0058] Figure 4 The intelligent decision-making and recommendation generation logic diagram provided in the embodiments of the present invention. Detailed Implementation
[0059] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0060] In this embodiment of the invention, the process of realizing smart teaching in a dual-teacher classroom based on AI teaching assistants may include the following devices: a smart teaching terminal and an AI teaching assistant system. The smart teaching terminal is deployed in the offline classroom to collect multimodal data during the classroom teaching process, including video data, audio data, and environmental data, and transmits the collected data to the AI teaching assistant system. The AI teaching assistant system is used to analyze and process the received multimodal data to achieve student status recognition, understanding of teaching content, and teaching decision recommendations. In a specific implementation, the smart teaching terminal includes an image acquisition module, an audio acquisition module, a main control module, and a communication module. The image acquisition module uses a high-definition camera, deployed at the front of the classroom, to collect video of students' classroom behavior; the audio acquisition module uses a microphone array, deployed in the podium area, to collect teacher and student voice data; the main control module uses an embedded processor, responsible for preliminary data processing and coordinated control of various modules; the communication module is used to establish a network connection with the AI teaching assistant system. During dual-teacher classroom teaching, the smart teaching terminal continuously collects multimodal classroom data and transmits it to the AI teaching assistant system. After analyzing and processing the data, the AI teaching assistant system pushes the generated teaching decision information to both the online lead teacher and the offline tutor, achieving intelligent teaching assistance. In another specific implementation, the AI teaching assistant system includes a data fusion module, a student status analysis module, a classroom content understanding module, and an intelligent decision-making module. The data fusion module performs spatiotemporal alignment and feature fusion on the received multimodal data; the student status analysis module, based on a lightweight object detection model, identifies students' hand-raising behavior and listening status from video data; the classroom content understanding module extracts key teaching knowledge points from audio data and identifies question statements through speech recognition and natural language processing technology; and the intelligent decision-making module generates interactive recommendations based on student status information and teaching content information using a dynamic weighted decision-making algorithm. During classroom interaction, when the system recognizes a teacher's question and detects multiple hand-raising actions, the intelligent decision-making module will comprehensively consider factors such as hand-raising response time, teaching pace, difficulty of understanding the question, and students' historical participation. It will calculate a dynamic priority score for each hand-raising action, generate an optimal interaction recommendation sequence, and push the recommendation results to the main teacher in real time. At the same time, it will push information about students who need special attention to the tutors.
[0061] The inventors of this invention discovered that in the actual teaching process of the dual-teacher classroom model, due to the physical isolation of the online lecturer, it is impossible to perceive the real-time classroom status of offline students, making it difficult to adjust the teaching pace based on students' immediate feedback. Interactive questioning often becomes aimless, leading to low efficiency in teaching interaction. Meanwhile, offline tutors need to simultaneously undertake multiple tasks such as equipment operation, maintaining classroom discipline, and answering student questions, resulting in an excessive workload and making it difficult to provide personalized attention and guidance to each student. While existing technologies have attempted to improve this situation through computer vision or artificial intelligence, most solutions require complex visual analysis algorithms, consuming significant computational resources, incurring high deployment costs, and posing risks of infringing on student privacy. Furthermore, these solutions often focus only on a single dimension of student behavior, lacking comprehensive consideration of teaching pace, content difficulty, and fairness of classroom participation. They fail to truly understand the teaching context, resulting in a disconnect between the auxiliary decision-making they provide and actual teaching needs, making it difficult to effectively promote and apply in educational practice. Based on this, in one embodiment of the present invention provided by the inventors, a lightweight, low-cost smart teaching terminal is constructed. Ordinary cameras and microphone arrays are used as the sensing basis, abandoning computationally complex and privacy-sensitive visual emotion and behavior analysis. Instead, the focus is on identifying the most core teaching interaction signals: students raising their hands and the content of teachers' questions. On this basis, a teaching rhythm perception and cognitive difficulty assessment model is introduced to quantify the teacher's teaching style and the level of thinking in the questions. Combined with a historical data feedback mechanism aimed at promoting fair classroom participation, a dynamic weighted intelligent decision-making algorithm is constructed. This algorithm does not simply rely on the single dimension of "who raises their hand first," but comprehensively considers multiple factors such as students' reaction speed, the current teaching pace, the difficulty of the questions, and students' historical participation, ultimately generating an optimal interactive recommendation sequence. This enables a new paradigm of intelligent dual-teacher teaching where online expert teachers can accurately assess the classroom remotely, and offline tutors can provide targeted assistance, fundamentally improving interaction efficiency and teaching fairness.
[0062] See Figures 1 to 4 In one embodiment of the present invention, a smart teaching terminal for a dual-teacher classroom based on AI teaching assistants includes:
[0063] The hardware terminal is deployed in the offline classroom. The hardware terminal includes a multimodal data acquisition module, which is used to collect visual data, audio data and environmental data of the offline classroom.
[0064] The AI teaching assistant system communicates and connects with hardware terminals. The AI teaching assistant system includes:
[0065] The student behavior analysis module is used to identify students' hand-raising behavior based on video data using a lightweight object detection model.
[0066] The classroom content comprehension module is used to identify key knowledge points and questions from the lecturer based on audio data through real-time speech recognition and keyword matching technology, and to analyze the current teaching pace.
[0067] The intelligent decision-making module is used to generate interactive recommendation information and output it to the teacher's end based on the hand-raising behavior, questioning, and teaching pace. It uses a decision-making algorithm that integrates time correlation, participation balance, and pace adaptation.
[0068] In this embodiment, hardware terminals deployed in offline classrooms serve as the system's senses. A multimodal data acquisition module continuously collects visual, audio, and environmental data from the classroom, providing rich raw information for subsequent analysis. The AI teaching assistant system, working in conjunction with the hardware terminals, acts as the brain. The student behavior analysis module is specifically responsible for accurately identifying the key interactive signal of students raising their hands from the video stream. The classroom content understanding module extracts teaching knowledge points from the teacher's voice and identifies question statements in parallel, while quantifying the current teaching pace by calculating the density of knowledge points. Finally, the intelligent decision-making module, as the system's decision-making center, deeply integrates the above multi-source information and generates optimal interactive recommendations through an intelligent algorithm that comprehensively considers response timeliness, participation fairness, and teaching pace adaptability. This provides precise teaching assistance to online lecturers and offline tutors, achieving a complete closed loop from simple data collection to intelligent teaching decision-making.
[0069] In one embodiment of the present invention, the classroom content comprehension module quantifies the teaching pace state in the following manner. :
[0070] ;
[0071] in:
[0072] To slide the time window The number of key knowledge points identified in the lesson;
[0073] when At that time, it is judged to be a fast-paced state;
[0074] when If the time is right, it is considered a slow-paced state; otherwise, it is a normal-paced state.
[0075] and This is the preset rhythm state threshold.
[0076] Specifically, the quantification of teaching rhythm in the classroom content comprehension module is a core process that transforms the teacher's subjective feeling of "speed" into objective data. The system first tracks and counts the number of "key knowledge points" taught by the teacher within a continuously sliding time window (e.g., five minutes) using real-time speech recognition and natural language processing technology. Then, this number is divided by the length of the time window to calculate a quantified "teaching rhythm state value" R, which intuitively represents the knowledge output density per unit time. To classify continuous rhythm values, the system presets two key thresholds: (High rhythm threshold) and (Low-rhythm threshold); when the calculated R value is greater than When the R value is less than a certain value, it indicates that the teacher is rapidly outputting knowledge, and the system determines that the current state is "fast-paced"; when the R value is less than a certain value... When the R value is low, it indicates that the teaching progress is relatively slow, and the system classifies it as a "slow-paced state"; if the R value is between the two, it is considered a "normal-paced state". This method enables the AI teaching assistant to accurately grasp the overall progress of the classroom, much like an experienced teacher observing the lesson.
[0077] Assuming the system is set to a sliding time window The time limit is 5 minutes, and the high and low tempo thresholds are respectively: =2.0 (pieces / minute) and =0.5 (pieces / minute). In a middle school physics class, the teacher quickly explained three key knowledge points—"the definition of buoyancy," "Archimedes' principle," and "the formula for calculating buoyancy"—within a continuous 5-minute period. At this time, the system statistics showed... If the value is 3, then the teaching rhythm state value R = 3 / 5 = 0.6 (practice sessions / minute). Since 0.6 is greater than... (0.5) but less than (2.0) The system determines that it is currently in a "normal pace state". Based on this judgment, the intelligent decision-making module may adopt a neutral recommendation strategy when processing student hand-raising interactions. Conversely, if the teacher explains more than 10 knowledge points in the same 5 minutes (R=2.0 or higher), the system will enter a "fast pace state" and may trigger a strategy to prioritize recommending the fastest-responding students to match the tight teaching pace; if the teacher only analyzes 1 knowledge point in depth in 5 minutes (R=0.2), the system will enter a "slow pace state", and its strategy will tend to encourage more students to participate in in-depth thinking.
[0078] In one embodiment of the present invention, the student behavior analysis module identifies hand-raising behavior in the following manner:
[0079] Within the pre-defined student interaction area, the duration of the detected hand-raising target frame is monitored. Exceeding the time threshold When a hand-raising action is detected, its mathematical expression is:
[0080] .
[0081] Specifically, firstly, a specific interaction area is preset for each student in the captured video footage to limit the analysis scope, improve recognition efficiency, and reduce misjudgments. Next, based on a lightweight object detection model, the module detects in real time whether a target box matching the "raised hand" feature appears within that area. Most importantly, to avoid misidentifying brief, unintentional hand movements as valid hand-raising signals, the module introduces a time-duration judgment condition—only if the detected hand-raising target box persists for a certain period of time (…). Exceeding the preset time threshold ( Only when the hand-raising action is confirmed as valid will the system ultimately confirm it as a valid hand-raising action and set its flag to 1; otherwise, if the hand-raising state is too short, it is considered an invalid action, and the flag remains at 0. Through the dual judgment mechanism of spatial positioning and temporal verification, the system ensures that it will only trigger the subsequent decision-making process when it captures a student's clear and stable intention to raise their hand, thus achieving highly robust behavior recognition in noisy classroom environments.
[0082] For example, when student Zhang San raises his hand after hearing the teacher's question, the system detects the appearance of a hand-raising target box in his interaction area and starts timing; if Zhang San maintains the raised hand posture for a full 2 seconds (i.e., =2.0s> =1.5s), the system confirms that the hand-raising is valid, marks position 1 and passes it to the decision module as a valid interaction signal; conversely, if Zhang San only briefly raises his hand to fix his hair, the hand-raising target box disappears after only 1 second. =1.0s< If the action is invalid, the system will determine it as invalid and keep the flag at 0, thus effectively avoiding interference from common instantaneous actions in the classroom and ensuring the accuracy and reliability of behavior recognition.
[0083] To address the challenge of balancing teaching interaction efficiency and participation fairness in dual-teacher classrooms, in one embodiment of this invention, the intelligent decision-making module calculates a dynamic priority score for each hand-raising action using the following formula. :
[0084] ;
[0085] in, The basic priority score is determined by the time difference between the time the hand-raising action occurs and the time the question is recognized.
[0086] A rhythm adaptation factor related to the current teaching pace.
[0087] A fairness-enhancing factor related to the number of times the student has interacted historically.
[0088] Specifically, through a dynamic priority score fused from multiple factors. To quantify the recommendation priority of each hand-raising action, the score is composed of the product of three key factors: a base priority score. This reflects the student's response speed to questions, achieved through an exponential decay function, ensuring that students whose hand-raising time is closer to the time of being asked a question receive a higher base score; rhythm adaptation factor. This reflects the system's intelligent perception of the teaching context. During fast-paced lessons, it prioritizes students who respond quickly to maintain flow, while during slow-paced or challenging lessons, it tends to select students with lower participation to promote classroom equity. (Equity-promoting factor) By using a calculation method that is negatively correlated with the number of times a student has interacted in the past, weight compensation is proactively provided to students who have had fewer opportunities to speak.
[0089] For example, in a math class, after the teacher asks a question, student A (who has spoken 5 times in the past) raises their hand after 2 seconds, and student B (who has spoken 1 time in the past) raises their hand after 3 seconds. At this point, the system detects that the class is in a slow-paced state; when calculating student A's score, their quick hand-raising earns them a higher score. It's worth it, but the large number of past posts makes it... The value is low, and it is in a slow-paced state. The strategy also disadvantages students who speak frequently; when calculating student B's score, although their Pb value is lower due to raising their hand slightly slower, their lower historical speaking frequency results in a higher score. Compensation value, and in a slow-paced state The strategy effectively supports and encourages such students; ultimately, student B's overall priority score Pd surpasses that of student A, and the system then recommends student B to answer first to the teacher, reflecting the design goal of effectively promoting fairness in classroom participation while ensuring response efficiency.
[0090] In one embodiment of the invention, the basic priority score The calculation method is as follows:
[0091] ;
[0092] in, Based on the score;
[0093] , representing the time difference between the time the hand-raising action occurs and the time the question is recognized;
[0094] λ is the preset attenuation coefficient.
[0095] Specifically, an exponential decay function mathematical model is used to transform the qualitative characteristic of students' response speed into a precise quantitative score. This represents the benchmark for a student to receive full marks when they raise their hand in a timely manner, while the key variable... The difference between the time it takes to raise your hand and the time it takes to ask a question is then expressed as an exponential function. To adjust the final score, which means The value does not follow Instead of a linear decline, the system exhibits a sensitive characteristic where quick responders score highly, while slow responders experience a sharp drop in scores. This significantly widens the score gap between quick and slow responders, ensuring the system prioritizes students who are quick-thinking and keep up with the teacher's pace. Assume the system is set... The value is 100, the attenuation coefficient λ is 0.5, and when the teacher asks a question, student A quickly raises their hand after 1 second. =1), Student B raised their hand after 3 seconds ( =3), then student A's The value is 100×e (-0.5×1) ≈100 × 0.606 = 60.6, while student B's The value is 100×e (-0.5×3) ≈100×0.223=22.3; This calculation result clearly shows that although both students raised their hands, student A received a much higher basic priority score than student B due to their faster response. This lays the foundation for emphasizing response time in subsequent comprehensive decision-making.
[0096] In one embodiment of the invention, the rhythm adaptation factor The value selection strategy is as follows:
[0097] In a fast-paced environment, higher scores are assigned to the top k students who react fastest when raising their hands. value;
[0098] In a slower-paced environment, higher allocations are made to students with lower history engagement. value;
[0099] Under normal rhythm conditions Take the baseline value as 1.0.
[0100] Specifically, rhythm adaptation factor The value allocation strategy is dynamically adjusted according to different teaching paces: in a fast-paced state, the system prioritizes ensuring the smoothness of the teaching process, therefore assigning higher values to the top k students who raise their hands the fastest. In a slower pace, the focus shifts to deeper understanding and full participation, while the system prioritizes classroom equity, allocating higher scores to students with lower participation in history. It is valuable to proactively create opportunities to promote the integration of these students; while under normal circumstances, The baseline value is set to 1.0, which does not impose any additional influence on the base priority score.
[0101] For example, in a fast-paced English grammar class, the teacher is explaining multiple grammar points intensively. The system detects that the teaching pace is "fast-paced," and at this moment, student A and student B raise their hands almost simultaneously. The time taken was 1.2 seconds and 1.5 seconds respectively, while student C was slightly slower to raise their hand ( The time limit is 2.8 seconds, but historical participation is low; the system follows a fast-paced strategy, assigning higher priority to the two students with the fastest reactions (k=2), namely student A and student B. The system assigns a base value (e.g., 1.3) to each student, while student C, who did not rank in the top two, receives a base value of 1.0. This gives quick-thinking students an advantage in the final ranking. Conversely, in another slow-paced philosophy discussion class, the system detects a "slow-paced" state, where even if student D is slightly slower to raise their hand ( (The allotted time was 3.5 seconds), but because he hadn't spoken this semester, the system assigned him a higher time limit. Values (e.g., 1.4), while student E, who frequently speaks, raises his hand quickly ( (The actual time allotted was 1.0 seconds), but only the baseline value of 1.0 was obtained, thus effectively balancing classroom participation opportunities through algorithmic intervention.
[0102] In one embodiment of the invention, a fairness-promoting factor Calculated using the following formula:
[0103] ;
[0104] in, This represents the total number of times the student interacted with history in this class.
[0105] This represents the highest number of student interactions so far in this lesson.
[0106] This is a preset fairness adjustment coefficient.
[0107] Specifically, through a number of historical interactions with an individual student ( ) and the highest number of interactions in the whole class ( The function is negatively correlated with the relative ratio of students, dynamically providing additional score bonuses to students with low participation; the adjustment coefficient in the formula... The strength of this fairness intervention is determined by the formula structure. This ensures that the value of the factor ranges from 1 to (1+ Between ) for a student who has not yet spoken ( =0), he will receive the maximum bonus (1+ ), and for a student whose interaction count has reached the highest in the class ( = Its factor is a baseline value of 1, which neither punishes nor rewards, thus systematically tilting interaction opportunities toward students with low participation without harming the enthusiasm of active students.
[0108] For example, in a class, student A has spoken 5 times ( =5), which is the highest total number of interactions so far, while student B has not yet spoken. =0), the system sets the fairness adjustment coefficient α to 0.5; then student A's =1+0.5*(1-5 / 5)=1+0.5*0=1.0, while student B's =1+0.5*(1-0 / 5)=1+0.5*1=1.5; This means that, under the same conditions (such as response speed), student B's final priority score will be 1.5 times that of student A. As a result, the system will prioritize recommending student B to answer questions, thereby effectively improving classroom participation.
[0109] In one embodiment of the present invention, the classroom content comprehension module determines the expected cognitive difficulty L of the question by the following steps:
[0110] Extracting the short-time average energy of the question sentence speech signal and short-time average zero crossing rate ;
[0111] The extracted features are compared with a preset threshold to determine the state of the speech features.
[0112] Simultaneously, the text of the question is analyzed and matched with the basic cognitive keyword library and the advanced cognitive keyword library;
[0113] Based on the matching results of speech features and keywords, the expected cognitive difficulty level L is determined to be high, medium or low according to preset rules.
[0114] Specifically, the classroom content comprehension module determines the expected cognitive difficulty L of a question by integrating a dual evidence chain of speech signal features and text semantic analysis. First, it extracts physical features reflecting the teacher's tone and rhythm from the speech signal—short-time average energy. (Reflects sound intensity) and short-time average zero-crossing rate (Reflecting speech speed), speech features are categorized into high-energy / speech-speed or low-energy / speech-speed states by comparing them with preset thresholds; simultaneously, keyword matching is performed on the identified question text, comparing it with preset basic cognitive keyword databases (containing factual questions such as "what" and "where") and higher-order cognitive keyword databases (containing analytical questions such as "why" and "how to evaluate"); finally, based on the cross-validation of speech and text features, a comprehensive judgment is made according to preset rules. For example, when low-energy, slow-speed speech features are detected and matched with higher-order cognitive keywords, it is determined that the question requires deep thinking, and the expected cognitive difficulty L is marked as high level.
[0115] In one embodiment of the invention, the rhythm adaptation factor The value is also affected by the expected cognitive difficulty level L:
[0116] When L is at a high level, the same applies as the slow-paced state. Value retrieval strategy;
[0117] When L is at a low level, the same applies as in the fast-paced state. Value retrieval strategy.
[0118] Specifically, high-level cognitive difficulty problems are similar to slow-paced situations, requiring more time for students to think and encouraging participation at different levels of thinking. Therefore, the same approach as in slow-paced situations should be adopted. The selection strategy prioritizes students with lower historical participation to promote fairness in deep thinking; conversely, low cognitive difficulty questions (L for low) are like a fast-paced signal, usually involving factual recall or simple application, aiming for an active classroom atmosphere and efficient feedback. Therefore, the system adopts the same strategy as the fast-paced state, that is, prioritizing the fastest-reacting students to maintain smooth teaching.
[0119] For example, when a language arts teacher poses a cognitively challenging question, such as "Please analyze the symbolic meaning of the scene in *Dream of the Red Chamber* where Daiyu buries the flowers," even if the lesson is generally at a normal pace, the system, through analysis, determines that the expected cognitive difficulty L is high and will immediately activate the high-difficulty question approach. The strategy (i.e., the slow-paced strategy) allows a student A, who is usually quiet but deep in thought, to achieve a higher score even if they are a little slow to raise their hand. Value, thereby increasing its chances of being recommended and guiding classroom interaction towards depth and breadth; conversely, when a teacher asks a question of low cognitive difficulty, such as "In which chapter of 'Dream of the Red Chamber' does the story of Daiyu burying flowers take place?", the system determines L to be of low level and then uses methods for low-difficulty questions. The strategy (i.e., the fast-paced strategy) prioritizes rewarding student B who raises their hand first, using rapid question-and-answer sessions to reinforce knowledge and create a positive classroom atmosphere.
[0120] In one embodiment of the present invention, the AI teaching assistant system further includes a teaching rhythm suggestion module, which is used to generate rhythm adjustment suggestions to the main teacher when the teaching rhythm state R is detected to be continuously abnormal.
[0121] Specifically, the teaching rhythm suggestion module continuously tracks and analyzes the teaching rhythm status value R calculated by the classroom content comprehension module. When the system detects that the rhythm status value R deviates from the normal range continuously within multiple consecutive time windows, a continuous anomaly occurs. The teaching rhythm suggestion module will then proactively generate rhythm adjustment suggestions with clear operational guidance and push them to the online lecturer in real time. This helps teachers transcend their subjective feelings and optimize classroom management based on objective data.
[0122] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A smart teaching terminal for dual-teacher classrooms based on AI teaching assistants, comprising hardware terminals deployed in offline classrooms, characterized in that: The hardware terminal includes a multimodal data acquisition module for collecting visual, audio, and environmental data from offline classrooms, and also includes: The AI teaching assistant system communicates and connects with hardware terminals. The AI teaching assistant system includes: The student behavior analysis module is used to identify students' hand-raising behavior based on video data using a lightweight object detection model; The classroom content comprehension module is used to identify key knowledge points and questions from the lecturer based on audio data through real-time speech recognition and keyword matching technology, and to analyze the current teaching pace. The intelligent decision-making module is used to generate interactive recommendation information and output it to the teacher's end based on the hand-raising behavior, questioning, and teaching pace. It uses a decision-making algorithm that integrates time correlation, participation balance, and pace adaptation.
2. The AI-assisted dual-teacher smart teaching terminal according to claim 1, characterized in that, The classroom content comprehension module quantifies the teaching pace and status using the following methods. : ; in: To slide the time window The number of key knowledge points identified in the lesson; when At that time, it is judged to be a fast-paced state; when At this time, it is judged to be a slow-paced state; Otherwise, it is in a normal rhythm state; and This is the preset rhythm state threshold.
3. The AI-assisted dual-teacher smart teaching terminal according to claim 2, characterized in that, The student behavior analysis module identifies hand-raising behavior in the following ways: Within the pre-defined student interaction area, the duration of the detected hand-raising target frame is monitored. Exceeding the time threshold When a hand-raising action is detected, its mathematical expression is: 。 4. The AI-assisted dual-teacher classroom smart teaching terminal according to claim 1, characterized in that, The intelligent decision-making module calculates a dynamic priority score for each hand-raising action using the following formula. : ; in, The basic priority score is determined by the time difference between the time the hand-raising action occurs and the time the question is recognized. A rhythm adaptation factor related to the current teaching pace. A fairness-enhancing factor related to the number of times the student has interacted historically.
5. The AI-assisted dual-teacher classroom smart teaching terminal according to claim 4, characterized in that, Basic priority score The calculation method is as follows: ; in, Based on the score; , representing the time difference between the time the hand-raising action occurs and the time the question is recognized; λ is the preset attenuation coefficient.
6. The AI-assisted dual-teacher classroom smart teaching terminal according to claim 4, characterized in that, Rhythm adaptation factor The value selection strategy is as follows: In a fast-paced environment, the top k students with the fastest hand-raising response times are assigned higher scores. value; In a slower-paced environment, higher allocations are made to students with lower history engagement. value; Under normal rhythm conditions Take the baseline value as 1.
0.
7. The AI-assisted dual-teacher smart teaching terminal according to claim 4, characterized in that, Fairness Promotion Factor Calculated using the following formula: ; in, This represents the total number of times the student interacted with history in this class. This represents the highest number of student interactions so far in this lesson. This is a preset fairness adjustment coefficient.
8. The AI-assisted dual-teacher classroom smart teaching terminal according to claim 4, characterized in that, The classroom content comprehension module determines the expected cognitive difficulty L of the question by following these steps: Extracting the short-time average energy of the question sentence speech signal and short-time average zero crossing rate ; The extracted features are compared with a preset threshold to determine the state of the speech features. Simultaneously, the question text is analyzed and matched with the basic cognitive keyword library and the advanced cognitive keyword library; Based on the matching results of speech features and keywords, the expected cognitive difficulty level L is determined to be high, medium or low according to preset rules.
9. The AI-assisted dual-teacher smart teaching terminal according to claim 8, characterized in that, Rhythm adaptation factor The value is also affected by the expected cognitive difficulty level L: When L is at a high level, the same applies as the slow-paced state. Value retrieval strategy; When L is at a low level, the same applies as in the fast-paced state. Value retrieval strategy.
10. The AI-assisted dual-teacher smart teaching terminal according to claim 1, characterized in that, The AI teaching assistant system also includes a teaching rhythm suggestion module, which generates rhythm adjustment suggestions to the main teacher when the teaching rhythm state R is detected to be continuously abnormal.