Virtual double-teacher teaching dynamic scheduling method, device and system, electronic equipment and storage medium

By acquiring multimodal teaching data in real time to identify teaching scenarios, determining the operating mode of virtual teachers and controlling their behavior, the problem of insufficient perception and decision-making of virtual teachers in complex teaching situations is solved, and efficient personalized teaching assistance is achieved.

CN121300945APending Publication Date: 2026-01-09IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511524463.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing virtual teachers lack the ability to perceive and understand complex and dynamic teaching situations in real time. They are unable to independently judge the stage of teaching activities and adjust the timing of intervention, and cannot effectively assist individual practice, self-study and group discussion, resulting in a significant reduction in teaching efficiency and depth.

Method used

By acquiring multimodal teaching data in real time, identifying the target activity type of the current teaching scenario, determining the target operation mode of the virtual teacher, and calling the behavior strategy library to control its execution of teaching tasks, the virtual teacher's proactive perception and intelligent decision-making are realized.

Benefits of technology

Virtual teachers can autonomously intervene in teaching activities, reducing the need for intervention from real teachers, improving the precision and breadth of personalized teaching, achieving efficient collaboration with real teachers, and ensuring that every student receives timely attention and guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121300945A_ABST
    Figure CN121300945A_ABST
Patent Text Reader

Abstract

The invention provides a virtual double-teacher teaching dynamic scheduling method, device and system, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence education, and the method comprises the steps: obtaining multi-modal teaching data in a classroom teaching process of a teacher; identifying a target teaching activity type of the current teaching scene by using the multi-modal teaching data; determining a target operation mode of the virtual teacher by using the target teaching activity type; by calling the target behavior strategy library in the target operation mode, the virtual teacher can be controlled to execute the teaching task corresponding to the target teaching activity type. According to the method, the virtual teachers are endowed with the ability of sensing and understanding the classroom situation, so that the virtual teachers can get rid of dependence on teacher instructions, an intelligent and dynamic cooperative relationship with real teachers is formed according to natural circulation, autonomous decision making and active intervention of teaching activities, each student in the classroom can get attention, and the teaching efficiency is improved. And the teaching idea of virtual double teachers is really realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence in education (AI in Education) technology, and in particular to a method, device, system, electronic device and storage medium for dynamic scheduling of virtual dual-teacher teaching. Background Technology

[0002] With the development of artificial intelligence and big data technologies, smart classrooms and online education platforms are gradually introducing artificial intelligence (AI) teaching assistants or virtual teachers to support human teachers and improve teaching efficiency and personalization. These virtual teachers are typically presented as digital humans and can perform some pre-set teaching assistance tasks, such as announcing announcements, keeping time during class, playing pre-set multimedia content, or conducting simple question-and-answer sessions based on keyword matching.

[0003] However, existing virtual teachers exhibit significant limitations in practical applications: they lack the ability to perceive and understand complex and dynamic teaching situations in real time, and cannot accurately determine the current stage of a teaching activity, such as whether the teacher is lecturing on core concepts, students are engaged in group discussions, or individual students are encountering difficulties in independent practice. Moreover, the behavioral logic of virtual teachers usually relies on the proactive instructions of real teachers or pre-set fixed scripts, and cannot autonomously adjust their intervention timing, role, and support strategies according to the real classroom atmosphere, teaching pace, and teacher-student dynamics.

[0004] Furthermore, modern instructional design includes numerous components that require students to explore independently or engage in collaborative learning. When students are engaged in individual practice, independent reading, or online research, it is difficult for a live teacher to simultaneously monitor the progress and status of each student. In such situations, virtual teachers, lacking proactive awareness and decision-making capabilities, cannot identify students who are stalled due to difficulties or who have strayed from their learning tasks due to inattention, and therefore cannot provide timely intervention and guidance.

[0005] When students engage in group discussions or project-based learning, live teachers cannot simultaneously delve into each group's discussion to understand its progress, atmosphere, and challenges. Similarly, existing virtual teachers cannot comprehend the dynamic interactions within groups, thus failing to provide guiding questions, supplementary materials, or interim summaries at crucial moments, significantly diminishing the efficiency and depth of collaborative learning. Summary of the Invention

[0006] This invention provides a method, device, system, electronic device, and storage medium for dynamic scheduling of virtual dual-teacher teaching, in order to overcome the deficiencies existing in related technologies.

[0007] This invention provides a method for dynamic scheduling of virtual dual-teacher teaching, comprising: Real-time acquisition of multimodal teaching data during the classroom teaching process; Based on the multimodal teaching data, identify the target teaching activity type in the current teaching scenario; Based on the aforementioned target teaching activity types, the target operation mode of the virtual teacher is determined; The system invokes the target behavior strategy library under the target operating mode, and controls the virtual teacher to execute the teaching task corresponding to the target teaching activity type based on the target behavior strategy library.

[0008] According to the present invention, a method for dynamic scheduling of virtual dual-teacher teaching, wherein identifying the target teaching activity type in the current teaching scenario based on the multimodal teaching data includes: Extract the multimodal temporal features from the multimodal teaching data, and align the multimodal temporal features in the time dimension to obtain aligned temporal features; The aligned temporal features are input into the temporal multimodal fusion model to obtain the probability that the current teaching scenario belongs to different teaching activity types, as output by the temporal multimodal fusion model. Based on the probability corresponding to each of the aforementioned teaching activity types, the target teaching activity type is determined from among the aforementioned teaching activity types.

[0009] According to the present invention, a virtual dual-teacher teaching dynamic scheduling method is provided, wherein the multimodal teaching data includes visual data streams, auditory data streams, and interactive data streams generated by different presentation terminals in the classroom during the teacher's classroom teaching process; The multimodal temporal features include visual feature sequences, auditory feature sequences, and interaction feature sequences; the visual feature sequences include teacher pose sequences, student pose sequences, and student facial expression sequences in the visual data stream. The auditory feature sequence includes the keyword sequence, speaker role, speech rate, emotional feature sequence, and dialogue behavior classification sequence of the transcribed text corresponding to the auditory data stream; The interaction feature sequence includes the statistical feature sequence in the interaction data stream.

[0010] According to the present invention, a method for dynamic scheduling of virtual dual-teacher teaching, wherein determining the target operating mode of the virtual teacher based on the target teaching activity type includes: Based on the strategy mapping table, the target operation mode corresponding to the target teaching activity type is determined; wherein, the strategy mapping table includes a predefined correspondence between different teaching activity types and different virtual teacher operation modes.

[0011] According to the present invention, a virtual dual-teacher teaching dynamic scheduling method is provided, wherein controlling the virtual teacher to execute teaching tasks corresponding to the target teaching activity type based on the target behavior strategy library includes: Based on the target behavior strategy library, the virtual teacher is controlled to perform different teaching tasks corresponding to the target teaching activity type on different presentation terminals.

[0012] According to the present invention, a virtual dual-teacher teaching dynamic scheduling method is provided, wherein the presentation terminal includes a teacher terminal and a student terminal; The step of controlling the virtual teacher to perform different teaching tasks corresponding to the target teaching activity type on different presentation terminals based on the target behavior strategy library includes: Based on the target behavior strategy library, a first decision instruction is determined on the teacher's end, and the virtual teacher is controlled to execute the first decision instruction on the teacher's end; the first decision instruction is used to characterize the teacher's end teaching task corresponding to the target teaching activity type; Based on the target behavior strategy library, a second decision instruction is determined on the student's end, and the virtual teacher is controlled to execute the second decision instruction on the student's end; the second decision instruction is used to characterize the student's end teaching task corresponding to the target teaching activity type.

[0013] The virtual dual-teacher teaching dynamic scheduling method provided by the present invention further includes: Receive the scheduling instructions from the teacher for the virtual teacher; The virtual teacher is controlled to execute different teaching tasks corresponding to the scheduling instructions on different presentation terminals.

[0014] The present invention also provides a virtual dual-teacher teaching dynamic scheduling system, comprising: a multimodal sensing device array and a computing and storage server, wherein the multimodal sensing device array is connected to the computing and storage server; The multimodal sensing device array is used to collect multimodal teaching data during the teacher's classroom teaching process and transmit the multimodal teaching data to the computing and storage server; The computing and storage server is used to execute the aforementioned virtual dual-teacher teaching dynamic scheduling method.

[0015] According to the virtual dual-teacher teaching dynamic scheduling system provided by the present invention, the multimodal sensing device array includes: visual sensors, auditory sensors, and different presentation terminals deployed in different locations in the classroom; The visual sensor is used to collect the visual data stream in the multimodal teaching data; The auditory sensor is used to collect auditory data streams from the multimodal teaching data; The different presentation terminals are used to collect the interactive data streams they generate.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the virtual dual-teacher teaching dynamic scheduling method as described above.

[0017] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the virtual dual-teacher teaching dynamic scheduling method as described above.

[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the virtual dual-teacher teaching dynamic scheduling method as described above.

[0019] The present invention provides a virtual dual-teacher teaching dynamic scheduling method, device, system, electronic device, and storage medium. First, it acquires multimodal teaching data from the teacher's classroom teaching process in real time. Then, it uses this multimodal teaching data to identify the target teaching activity type in the current teaching scenario. Next, it uses the target teaching activity type to determine the target operation mode of the virtual teacher. Finally, by calling the target behavior strategy library under the target operation mode, it can control the virtual teacher to execute the teaching tasks corresponding to the target teaching activity type. This method endows the virtual teacher with the ability to perceive and understand the classroom situation, enabling it to break free from dependence on teacher instructions. Based on the natural flow of teaching activities, it makes autonomous decisions and actively intervenes, forming an intelligent and dynamic collaborative relationship with the real teacher. This ensures that every student in the classroom receives attention, and timely intervention and guidance can be provided to students who encounter difficulties and stagnate or deviate from their learning tasks due to inattention, truly realizing the teaching concept of "virtual dual-teacher."

[0020] Moreover, the autonomous operation of virtual teachers minimizes the need for real teachers to operate and intervene in the classroom, allowing teachers to focus more on higher-level teaching activities such as instructional design, classroom guidance, and emotional communication. This reduces the additional technical burden on teachers, achieving seamless technological assistance. Specifically, in the dominant mode, virtual teachers can take over individual exercises and group discussions, providing continuous and personalized monitoring and guidance to each student or group. This effectively solves the problem of real teachers being overwhelmed, enabling large-scale personalized teaching and improving its precision and breadth. This method achieves a complete closed loop from environmental perception and intelligent decision-making to virtual teacher control, allowing virtual teachers to dynamically adjust their roles and behaviors based on real teaching activities, forming an efficient and intelligent dual-teacher collaboration with real teachers. This significantly enhances the depth and value of artificial intelligence applications in education. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating the virtual dual-teacher teaching dynamic scheduling method provided by the present invention.

[0023] Figure 2 This is a schematic diagram of the structure of the virtual dual-teacher teaching dynamic scheduling device provided by the present invention.

[0024] Figure 3 This is a schematic diagram of the structure of the virtual dual-teacher teaching dynamic scheduling system provided by the present invention.

[0025] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] Because existing virtual teachers are functionally closer to a responsive, single-function "tool," they fall far short of the standard of a "partner teacher" or "co-teacher" capable of proactively perceiving, intelligently making decisions, and dynamically collaborating with real teachers. Current virtual teachers cannot truly integrate into the complex and ever-changing teaching process to realize the educational value of "virtual dual-teacher" teaching. Therefore, this invention provides a method for dynamically scheduling virtual dual-teacher teaching that enables virtual teachers to proactively perceive the teaching context, intelligently decide their own role, and dynamically collaborate with real teachers.

[0028] Figure 1 This is a flowchart illustrating a virtual dual-teacher teaching dynamic scheduling method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: S1, real-time acquisition of multimodal teaching data during the classroom teaching process; S2, Based on the multimodal teaching data, identify the target teaching activity type in the current teaching scenario; S3, Based on the target teaching activity type, determine the target operation mode of the virtual teacher; S4, invoke the target behavior strategy library under the target operation mode, and based on the target behavior strategy library, control the virtual teacher to execute the teaching task corresponding to the target teaching activity type.

[0029] Specifically, the virtual dual-teacher teaching dynamic scheduling method provided in this embodiment of the invention is executed by a virtual dual-teacher teaching dynamic scheduling device. This device can be configured in a computing and storage server, which can be a locally deployed high-performance graphics processing unit (GPU) server or a cloud computing cluster, responsible for carrying out the training and real-time inference tasks of complex algorithm models. The computing and storage server can be equipped with a storage system for persistently saving the original multimodal teaching data and the results obtained from each step.

[0030] First, step S1 is executed to acquire multimodal teaching data in real time during the teacher's classroom teaching process. Here, the teacher is a real teacher; in this embodiment of the invention, the virtual dual-teacher refers to a combination of a real teacher and a virtual teacher. The multimodal teaching data can be collected by a multimodal sensing device array and transmitted to a computing and storage server, and may include visual data streams, auditory data streams, and interactive data streams generated by different presentation terminals during the teacher's classroom teaching process.

[0031] The multimodal sensing device array may include: visual sensors, auditory sensors, and different presentation terminals deployed at different locations within the classroom. The visual sensors, used to acquire visual data streams from the multimodal teaching data, may include at least one wide-angle camera deployed at the back of the classroom and one or more close-up cameras aimed at the student area.

[0032] The wide-angle camera captures panoramic images and videos of the classroom, showing the entire classroom, the teacher's movement trajectory, and teaching posture. This posture can be represented by the teacher's handwriting on the blackboard or their actions while operating the large screen. The close-up camera captures images and videos of students for facial expression and behavioral recognition. Recognized facial expressions can include confusion, focus, and excitement, while recognized behavioral actions can include raising a hand, looking down, and sitting posture. Accordingly, the visual data stream can include panoramic images and videos of the classroom and images and videos of the students.

[0033] Auditory sensors are used to acquire auditory data streams in multimodal instructional data and may include directional microphone arrays deployed on the classroom ceiling or podium. The directional microphone arrays utilize beamforming technology to localize sound sources, i.e., to determine whether the speaker is a teacher or a student in a specific area, and to acquire high-quality speech signals. Accordingly, the auditory data stream may include speech signals.

[0034] Different presentation terminals can each function as interactive data acquisition devices, collecting the interactive data streams they generate. Presentation terminals can include teacher and student terminals; the teacher terminal can be a smart screen, and the student terminal can be a tablet or personal computer, etc. Both teacher and student terminals can record user operation logs, answer data, browsing behavior, etc.

[0035] Then, step S2 is executed to identify the target teaching activity type in the current teaching scenario using multimodal teaching data. Here, identifying the target teaching activity type means determining "what is happening in the classroom?". This can be achieved by first extracting features from the multimodal teaching data using various intelligent algorithms, and then determining the target teaching activity type in the current teaching scenario through the mapping relationship between the extracted features and the teaching activity type.

[0036] In this embodiment of the invention, the types of teaching activities involved may include at least one of teacher-led instruction, Q&A session, group collaboration, and individual practice. The target teaching activity type is one of these teaching activity types.

[0037] The combination of features that are related to teacher lecturing may include teacher's voice as the dominant source, steady voice, being located in the podium area, having actions such as writing on the blackboard or operating the large screen, most students facing the podium, and no interactive answering on the student's end.

[0038] Feature combinations that map to teacher-student question-and-answer sessions can include a teacher asking a question with keywords, followed by a student raising their hand, then activation of the student's area sound source, and alternating between the teacher's and student's voices. The teacher's question keywords can include phrases like "Who will answer?"

[0039] Features that are associated with group collaboration may include the simultaneous activation of multiple dispersed sound sources in the classroom, students sitting in a circular posture, high volume of discussion, and overall vocal energy higher than that of a single teacher's lecture.

[0040] Features that are mapped to individual practice may include low overall classroom volume, most students looking down, high frequency of student tablets generating response or writing interaction data, and teachers circulating in the classroom.

[0041] Next, step S3 is executed to determine the target operational mode of the virtual teacher using the target teaching activity type. The target operational mode refers to the operational mode corresponding to the target teaching activity type. In other words, determining the target operational mode means determining "what the virtual teacher should do now?"

[0042] For each type of teaching activity, there is a pre-defined operating mode for the virtual teacher. Here, the operating mode of the virtual teacher refers to the role played by the virtual teacher relative to the real teacher, which may include Assist Mode and Lead Mode.

[0043] In the assisted mode, the real teacher is the leader of the teaching, while the virtual teacher plays a low-interference, high-efficiency assistant role. Its behavior is aimed at enhancing the teaching content and optimizing the teaching process, and it is usually executed silently in the background or presented in a non-intrusive manner.

[0044] In the dominant mode, the virtual teacher plays a leading role in specific teaching segments, responsible for issuing tasks, organizing activities, and providing guidance and feedback. In this mode, the human teacher transforms into an observer and supporter, providing human intervention in exceptional cases.

[0045] In this embodiment of the invention, the target operating mode is either auxiliary mode or active mode.

[0046] Finally, step S4 is executed. Each operating mode of the virtual teacher corresponds to a behavior strategy library, which includes the virtual teacher's behavior strategies for each type of teaching activity.

[0047] For the behavioral strategy library in the assisted mode, the behavioral strategies can include content enhancement. That is, when a teacher teaches a certain knowledge point, the virtual teacher automatically pushes a 3D model of the knowledge point, related experimental videos, or a collection of historical wrong questions to the sidebar of the student's tablet for the student to access independently.

[0048] Behavioral strategies can also include process management, such as when a teacher calls out a student's name during a Q&A session, the virtual teacher automatically highlights the student's seat number on the smart screen and selectively turns on their microphone while displaying the question text on the screen.

[0049] Behavioral strategies can also include learning record keeping, which involves silently recording key classroom events, such as a student asking a high-quality question or the general answers to a particular knowledge point, and automatically generating a summary of the post-class learning analysis report.

[0050] Behavioral strategies can also include proactive intervention. When the system detects a student exhibiting "confusion" (such as answering two questions incorrectly in a row), the virtual teacher will send a simple hint to the student via a private chat bubble, avoiding public interruption of the class. The hint could be something like, "Would you like to see a hint on how to solve this problem?"

[0051] For the behavioral strategy library in the dominant mode, behavioral strategies can include task distribution and guidance. That is, at the beginning of the group collaboration session, the virtual teacher appears on the students' end in each group in full-screen or half-screen form, clearly displaying the discussion topic, rules and time limits, and starting a countdown.

[0052] Behavioral strategies can also include personalized tutoring. In individual practice sessions, a virtual teacher can be assigned to each student to monitor their progress and accuracy in real time. For students who answer quickly and accurately, extended questions are automatically provided; for students who are slow or have a high error rate, voice or text conversations are initiated to provide step-by-step guidance or review links for relevant basic knowledge points.

[0053] Behavioral strategies can also include discipline and focus management. For example, during self-study periods, if the screen content detection or behavior analysis on the student's end reveals that the student has been browsing websites or applications unrelated to learning for an extended period of time, the virtual teacher will issue a gentle voice or pop-up reminder, such as "We are conducting self-study, please stay focused."

[0054] Behavioral strategies can also include activity summaries and reports. After group discussions or individual practice, the virtual teacher can automatically summarize the discussion results of each group, such as keyword clouds, or the practice situation of the whole class, such as the overall accuracy rate and frequently missed questions, and present them visually on the smart screen for real teachers to comment on and explain.

[0055] After determining the target operating mode, the target behavior strategy library under that mode can be invoked. The virtual teacher's behavior strategies for the target teaching activity type can then be used to control the virtual teacher in executing the corresponding teaching tasks. Here, the virtual teacher's behavior strategies for the target teaching activity type can include decision instructions. These instructions guide the virtual teacher's behavior to complete the teaching tasks.

[0056] The decision instruction can be the first decision instruction on the teacher's end. At this time, the teaching task can be the teaching task on the teacher's end. Through the first decision instruction, the behavior of the virtual teacher on the teacher's end can be guided to complete the corresponding teaching task on the teacher's end.

[0057] The decision instruction can also be a second decision instruction on the student's end. In this case, the teaching task can be a teaching task on the student's end. Through the second decision instruction, the behavior of the virtual teacher on the student's end can be guided to complete the corresponding teaching task on the student's end.

[0058] The virtual dual-teacher teaching dynamic scheduling method provided in this embodiment of the invention first acquires multimodal teaching data during the teacher's classroom teaching process; then, using the multimodal teaching data, it identifies the target teaching activity type of the current teaching scenario; subsequently, it uses the target teaching activity type to determine the target operation mode of the virtual teacher; finally, by calling the target behavior strategy library under the target operation mode, it can control the virtual teacher to execute the teaching tasks corresponding to the target teaching activity type. This method endows the virtual teacher with the ability to perceive and understand the classroom situation, enabling it to break free from dependence on teacher instructions, make autonomous decisions and actively intervene according to the natural flow of teaching activities, and form an intelligent and dynamic collaborative relationship with the real teacher. This ensures that every student in the classroom receives attention, and timely intervention and guidance can be provided for students who encounter difficulties and stagnate or deviate from their learning tasks due to lack of concentration, truly realizing the teaching concept of "virtual dual-teacher".

[0059] Moreover, the autonomous operation of virtual teachers minimizes the need for real teachers to operate and intervene in the classroom, allowing teachers to focus more on higher-level teaching activities such as instructional design, classroom guidance, and emotional communication. This reduces the additional technical burden on teachers, achieving seamless technological assistance. Specifically, in the dominant mode, virtual teachers can take over individual exercises and group discussions, providing continuous and personalized monitoring and guidance to each student or group. This effectively solves the problem of real teachers being overwhelmed, enabling large-scale personalized teaching and improving its precision and breadth. This method achieves a complete closed loop from environmental perception and intelligent decision-making to virtual teacher control, allowing virtual teachers to dynamically adjust their roles and behaviors based on real teaching activities, forming an efficient and intelligent dual-teacher collaboration with real teachers. This significantly enhances the depth and value of artificial intelligence applications in education.

[0060] Based on the above embodiments, controlling the virtual teacher to execute the teaching task corresponding to the target teaching activity type based on the target behavior strategy library includes: Based on the target behavior strategy library, the virtual teacher is controlled to perform different teaching tasks corresponding to the target teaching activity type on different presentation terminals.

[0061] Specifically, the presentation terminals can include teacher and student terminals. Different content needs to be displayed on different terminals, and the virtual teacher has different teaching tasks. Therefore, the decision-making instructions for the virtual teacher's behavioral strategies for the target teaching activity type can simultaneously include a first decision instruction from the teacher terminal and a second decision instruction from the student terminal. In this case, the teaching task can simultaneously include teaching tasks from both the teacher and student terminals. Through the first and second decision instructions, the virtual teacher's behavior can be simultaneously guided on both the teacher and student terminals to complete the different teaching tasks corresponding to the teacher and student terminals.

[0062] For example, on the teacher's side, the virtual teacher typically exists in the form of a "data panel" or "teaching dashboard," primarily aimed at real-world teachers. The virtual teacher can display real-time charts showing the entire class's attention curves, average accuracy rates on practice questions, and a list of help signals—information invisible to students. This provides real-world teachers with a comprehensive view of the situation. On the student's side, the virtual teacher appears as a specific "learning partner" or "mentor," and its presentation can be dynamically adjusted according to the target operating mode. In auxiliary mode, the virtual teacher may only exist as a small icon in the corner of the screen or a sidebar that can be collapsed at any time, avoiding distractions; in dominant mode, the virtual teacher can appear in half-screen or even full-screen mode, engaging in deeper interaction with students.

[0063] Based on the above embodiments, identifying the target teaching activity type in the current teaching scenario based on the multimodal teaching data includes: Extract the multimodal temporal features from the multimodal teaching data, and align the multimodal temporal features in the time dimension to obtain aligned temporal features; The aligned temporal features are input into the temporal multimodal fusion model to obtain the probability that the current teaching scenario belongs to different teaching activity types, as output by the temporal multimodal fusion model. Based on the probability corresponding to each of the aforementioned teaching activity types, the target teaching activity type is determined from among the aforementioned teaching activity types.

[0064] Specifically, when identifying the type of target teaching activity, multiple intelligent algorithms can be used to extract multimodal temporal features from multimodal teaching data and align these features in the time dimension to obtain aligned temporal features.

[0065] Subsequently, a pre-trained temporal multimodal fusion model can be introduced, which can be a fusion model based on the Transformer architecture.

[0066] The alignment time-series features are input into the time-series multimodal fusion model. The time-series multimodal fusion model performs fusion analysis on the alignment time-series features and identifies in real time the probability that the current teaching scenario belongs to different teaching activity types.

[0067] Then, the teaching activity type with the highest probability corresponding to each teaching activity type can be selected as the target teaching activity type.

[0068] In this embodiment of the invention, the accuracy and reliability of the identification can be guaranteed by using a temporal multimodal fusion model to identify the type of target teaching activity.

[0069] Based on the above embodiments, the multimodal temporal features include visual feature sequences, auditory feature sequences, and interaction feature sequences; The visual feature sequence includes the teacher pose sequence, student pose sequence, and student expression sequence in the visual data stream; The auditory feature sequence includes the keyword sequence, speaker role, speech rate, emotional feature sequence, and dialogue behavior classification sequence of the transcribed text corresponding to the auditory data stream; The interaction feature sequence includes the statistical feature sequence in the interaction data stream.

[0070] Specifically, the visual feature sequence is obtained by feature extraction from the visual data stream, and can include teacher pose sequence, student pose sequence, and student expression sequence from the visual data stream. Among these, object detection algorithms such as the YOLO series can identify faces and bodies, while pose estimation algorithms such as OpenPose can extract the skeletal key points of teachers and students, forming teacher pose sequence and student pose sequence. Simultaneously, expression recognition models such as FERNet can analyze student facial regions to obtain emotion classifications such as focus, confusion, and happiness.

[0071] Auditory feature sequences are obtained by extracting features from the auditory data stream. These sequences may include keyword sequences, speaker roles, speech rates, emotional feature sequences, and dialogue behavior classification sequences in the transcribed text corresponding to the auditory data stream. Specifically, a sound source localization algorithm is used to determine the speaker's location, and facial region and / or voiceprint recognition technology is used to determine whether the speaker is a teacher or a student. Each speech segment from the auditory data stream is input into an Automatic Speech Recognition (ASR) engine to obtain the transcribed text for each speech segment. Natural Language Processing (NLP) techniques are used to extract keywords from the transcribed text, obtaining keyword sequences such as "Did you understand?" and "Who knows this question?". Dialogue behavior classification is performed on the transcribed text, yielding classification results such as questioning, answering, and discussion. Finally, sentiment analysis is performed on the transcribed text to obtain emotional features.

[0072] Interactive feature sequences can be obtained directly by statistically analyzing the interactive data stream of the presentation terminal. These sequences can include statistical feature sequences in the interactive data stream, such as the submission time, accuracy rate, and number of modifications of practice questions; and the page turning frequency and key annotations of courseware collected from the teacher's large screen.

[0073] In this embodiment of the invention, by providing the content of multimodal temporal features, the corresponding target teaching activity type can be obtained quickly and accurately based on the multimodal temporal features, which facilitates the implementation of subsequent solutions.

[0074] Based on the above embodiments, determining the target operation mode of the virtual teacher based on the target teaching activity type includes: Based on the strategy mapping table, the target operation mode corresponding to the target teaching activity type is determined; wherein, the strategy mapping table includes a predefined correspondence between different teaching activity types and different virtual teacher operation modes.

[0075] Specifically, when determining the target operating mode, a policy mapping table can be introduced, as shown in Table 1. In Table 1, the triggering conditions are only a combination of some features of the corresponding teaching activity type.

[0076] Table 1 Strategy Mapping Table

[0077] The strategy mapping table can predefine mapping rules between different types of teaching activities and different virtual teacher operation modes. By selecting the target teaching activity type, the corresponding virtual teacher operation mode can be found in the strategy mapping table as the target operation mode.

[0078] In this embodiment of the invention, a predefined strategy mapping table can improve the efficiency of determining the target operating mode, thereby improving the efficiency of the virtual teacher in executing teaching tasks.

[0079] Based on the above embodiments, the presentation terminal includes a teacher's terminal and a student's terminal; The step of controlling the virtual teacher to perform different teaching tasks corresponding to the target teaching activity type on different presentation terminals based on the target behavior strategy library includes: Based on the target behavior strategy library, a first decision instruction is determined on the teacher's end, and the virtual teacher is controlled to execute the first decision instruction on the teacher's end; the first decision instruction is used to characterize the teacher's end teaching task corresponding to the target teaching activity type; Based on the target behavior strategy library, a second decision instruction is determined on the student's end, and the virtual teacher is controlled to execute the second decision instruction on the student's end; the second decision instruction is used to characterize the student's end teaching task corresponding to the target teaching activity type.

[0080] Specifically, when controlling virtual teachers to perform different teaching tasks, the target behavior strategy library can be used to determine the first decision instruction of the teaching task on the teacher's end that represents the type of target teaching activity, and control the virtual teacher to execute the first decision instruction on the teacher's end, so as to transform the first decision instruction into the specific and appropriate behavior of the virtual teacher on the teacher's end.

[0081] Furthermore, the target behavior strategy library can be used to determine the second decision instruction for the student-side teaching task corresponding to the target teaching activity type, and control the virtual teacher to execute the second decision instruction on the student-side, so as to transform the second decision instruction into the specific and appropriate behavior of the virtual teacher on the student-side.

[0082] The virtual teacher's interaction mode on different presentation terminals can be intelligently selected based on the corresponding decision-making instructions and the degree of interference in the current teaching context. For example, regarding the interaction mode of the virtual teacher on the teacher's end, when the teacher is teaching key concepts, prompts for students should prioritize silent text or icons; while during individual practice, voice dialogue can be used to provide more efficient guidance.

[0083] The image and performance of a virtual teacher on different presentation devices, such as the virtual teacher's 3D model, voice, facial expressions, and body movements, are all driven by corresponding decision-making instructions. For example, when assigning tasks in the dominant mode, the virtual teacher's tone of voice will be clearer and more authoritative, and their posture more formal; while when providing personalized prompts in the auxiliary mode, their expression will be more friendly and their voice softer, thus making their behavior highly consistent with the teaching role they are currently playing.

[0084] Based on the above embodiments, it also includes: Receive the scheduling instructions from the teacher for the virtual teacher; The virtual teacher is controlled to execute different teaching tasks corresponding to the scheduling instructions on different presentation terminals.

[0085] Specifically, in this embodiment of the invention, in addition to the automated control of the virtual teacher based on multimodal teaching data, the virtual teacher can also be manually controlled by the teacher.

[0086] When teachers need to personalize the virtual teacher, they can input scheduling instructions into the virtual dual-teacher teaching dynamic scheduling device. These instructions can be in text or voice format, without any specific limitations.

[0087] After receiving a scheduling instruction, the dual-teacher teaching dynamic scheduling device can control the virtual teacher to execute different teaching tasks corresponding to the scheduling instruction on different presentation terminals.

[0088] In this embodiment of the invention, a scheme that supports teachers manually controlling virtual teachers can be provided, which makes it convenient for teachers to use virtual teachers to assist teaching according to the actual situation.

[0089] like Figure 2 As shown, based on the above embodiments, this embodiment of the invention provides a virtual dual-teacher teaching dynamic scheduling device, comprising: Data acquisition module 21 is used to acquire multimodal teaching data during the teacher's classroom teaching process; The teaching activity type identification module 22 is used to identify the target teaching activity type in the current teaching scenario based on the multimodal teaching data. The motion mode determination module 23 is used to determine the target operation mode of the virtual teacher based on the target teaching activity type; The virtual teacher control module 24 is used to call the target behavior strategy library under the target operation mode, and based on the target behavior strategy library, control the virtual teacher to perform different teaching tasks corresponding to the target teaching activity type on different presentation terminals.

[0090] Based on the above embodiments, the virtual dual-teacher teaching dynamic scheduling device provided in this embodiment of the invention, wherein the teaching activity type identification module is specifically used for: Extract the multimodal temporal features from the multimodal teaching data, and align the multimodal temporal features in the time dimension to obtain aligned temporal features; The aligned temporal features are input into the temporal multimodal fusion model to obtain the probability that the current teaching scenario belongs to different teaching activity types, as output by the temporal multimodal fusion model. Based on the probability corresponding to each of the aforementioned teaching activity types, the target teaching activity type is determined from among the aforementioned teaching activity types.

[0091] Based on the above embodiments, the virtual dual-teacher teaching dynamic scheduling device provided in this embodiment of the invention includes multimodal teaching data such as visual data stream, auditory data stream and interactive data stream generated by different presentation terminals in the classroom during the teacher's classroom teaching process; The multimodal temporal features include visual feature sequences, auditory feature sequences, and interaction feature sequences; the visual feature sequences include teacher pose sequences, student pose sequences, and student facial expression sequences in the visual data stream. The auditory feature sequence includes the keyword sequence, speaker role, speech rate, emotional feature sequence, and dialogue behavior classification sequence of the transcribed text corresponding to the auditory data stream; The interaction feature sequence includes the statistical feature sequence in the interaction data stream.

[0092] Based on the above embodiments, the virtual dual-teacher teaching dynamic scheduling device provided in this embodiment of the invention, wherein the motion mode determination module is specifically used for: Based on the strategy mapping table, the target operation mode corresponding to the target teaching activity type is determined; wherein, the strategy mapping table includes a predefined correspondence between different teaching activity types and different virtual teacher operation modes.

[0093] Based on the above embodiments, the virtual dual-teacher teaching dynamic scheduling device provided in this embodiment of the invention includes a teacher terminal and a student terminal in the presentation terminal; The virtual teacher control module is specifically used for: Based on the target behavior strategy library, a first decision instruction is determined on the teacher's end, and the virtual teacher is controlled to execute the first decision instruction on the teacher's end; the first decision instruction is used to characterize the teacher's end teaching task corresponding to the target teaching activity type; Based on the target behavior strategy library, a second decision instruction is determined on the student's end, and the virtual teacher is controlled to execute the second decision instruction on the student's end; the second decision instruction is used to characterize the student's end teaching task corresponding to the target teaching activity type.

[0094] Based on the above embodiments, the virtual dual-teacher teaching dynamic scheduling device provided in this embodiment of the invention further includes a virtual teacher control module specifically used for: Receive the scheduling instructions from the teacher for the virtual teacher; The virtual teacher is controlled to execute different teaching tasks corresponding to the scheduling instructions on different presentation terminals.

[0095] Specifically, the functions of each module in the virtual dual-teacher teaching dynamic scheduling device provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above method-like embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0096] like Figure 3 As shown, based on the above embodiments, this embodiment of the invention provides a virtual dual-teacher teaching dynamic scheduling system, including: a multimodal sensing device array and a computing and storage server, wherein the multimodal sensing device array and the computing and storage server can be connected through an application programming interface (API); The multimodal sensing device array is used to collect multimodal teaching data during the teacher's classroom teaching process and transmit the multimodal teaching data to the computing and storage server; The computing and storage server is used to execute the virtual dual-teacher teaching dynamic scheduling method provided in the above embodiments.

[0097] Specifically, in this embodiment of the invention, the virtual dual-teacher teaching dynamic scheduling system may include a data perception layer, a data presentation layer, and an application presentation layer. The multimodal perception device array is located in the data perception layer, the computing and storage server is located in the data presentation layer, and different presentation terminals such as the teacher's end and the student's end are located in the application presentation layer.

[0098] In this embodiment of the invention, a virtual dual-teacher teaching dynamic scheduling system is constructed through a multi-layered structure. This system enables virtual teachers to execute different teaching tasks corresponding to different types of target teaching activities on different presentation terminals. It empowers virtual teachers with the ability to perceive and understand classroom situations, allowing them to break free from dependence on teacher instructions. Based on the natural flow of teaching activities, they can make autonomous decisions and actively intervene, forming an intelligent and dynamic collaborative relationship with real teachers. This ensures that every student in the classroom receives attention, truly realizing the teaching concept of "virtual dual-teacher".

[0099] Based on the above embodiments, the virtual dual-teacher teaching dynamic scheduling system provided in this embodiment of the invention includes a multimodal sensing device array comprising: visual sensors, auditory sensors, and different presentation terminals deployed in different locations within the classroom; The visual sensor is used to collect the visual data stream in the multimodal teaching data; The auditory sensor is used to collect auditory data streams from the multimodal teaching data; The different presentation terminals are used to collect the interactive data streams they generate.

[0100] Specifically, in this embodiment of the invention, the visual sensors, auditory sensors, and different presentation terminals in the multimodal sensing device array can all be connected to the computing and storage server via API.

[0101] The visual sensors may include at least one wide-angle camera deployed at the back of the classroom and one or more close-up cameras aimed at the student area, for extracting visual data streams at different levels. The auditory sensors may include a directional microphone array deployed on the classroom ceiling or podium, for extracting auditory data streams at different levels.

[0102] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the virtual dual-teacher teaching dynamic scheduling method provided in the above embodiments.

[0103] Electronic devices can include Spark Smart Classroom, Spark Smart Blackboard, and Smart Window All-in-One Machine, etc. Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the virtual dual-teacher teaching dynamic scheduling method provided in the above embodiments.

[0105] In another aspect, the present invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the virtual dual-teacher teaching dynamic scheduling method provided in the above embodiments. This computer-readable storage medium can be either a non-transitory computer-readable storage medium or a transient computer-readable storage medium, and is not specifically limited herein.

[0106] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0107] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for dynamic scheduling of virtual dual-teacher teaching, characterized in that, include: Real-time acquisition of multimodal teaching data during the classroom teaching process; Based on the multimodal teaching data, identify the target teaching activity type in the current teaching scenario; Based on the aforementioned target teaching activity types, the target operation mode of the virtual teacher is determined; The system invokes the target behavior strategy library under the target operating mode, and controls the virtual teacher to execute the teaching task corresponding to the target teaching activity type based on the target behavior strategy library.

2. The virtual dual-teacher teaching dynamic scheduling method according to claim 1, characterized in that, The process of identifying the target teaching activity type in the current teaching scenario based on the multimodal teaching data includes: Extract the multimodal temporal features from the multimodal teaching data, and align the multimodal temporal features in the time dimension to obtain aligned temporal features; The aligned temporal features are input into the temporal multimodal fusion model to obtain the probability that the current teaching scenario belongs to different teaching activity types, as output by the temporal multimodal fusion model. Based on the probability corresponding to each of the aforementioned teaching activity types, the target teaching activity type is determined from among the aforementioned teaching activity types.

3. The virtual dual-teacher teaching dynamic scheduling method according to claim 2, characterized in that, The multimodal teaching data includes visual data streams, auditory data streams, and interactive data streams generated by different presentation terminals during the teacher's classroom teaching process. The multimodal temporal features include visual feature sequences, auditory feature sequences, and interaction feature sequences; the visual feature sequences include teacher pose sequences, student pose sequences, and student facial expression sequences in the visual data stream. The auditory feature sequence includes the keyword sequence, speaker role, speech rate, emotional feature sequence, and dialogue behavior classification sequence of the transcribed text corresponding to the auditory data stream; The interaction feature sequence includes the statistical feature sequence in the interaction data stream.

4. The virtual dual-teacher teaching dynamic scheduling method according to any one of claims 1-3, characterized in that, The determination of the target operating mode of the virtual teacher based on the target teaching activity type includes: Based on the strategy mapping table, the target operation mode corresponding to the target teaching activity type is determined; wherein, the strategy mapping table includes a predefined correspondence between different teaching activity types and different virtual teacher operation modes.

5. The virtual dual-teacher teaching dynamic scheduling method according to any one of claims 1-3, characterized in that, The step of controlling the virtual teacher to execute teaching tasks corresponding to the target teaching activity type based on the target behavior strategy library includes: Based on the target behavior strategy library, the virtual teacher is controlled to perform different teaching tasks corresponding to the target teaching activity type on different presentation terminals.

6. The virtual dual-teacher teaching dynamic scheduling method according to claim 5, characterized in that, The presentation terminal includes a teacher's terminal and a student's terminal; The step of controlling the virtual teacher to perform different teaching tasks corresponding to the target teaching activity type on different presentation terminals based on the target behavior strategy library includes: Based on the target behavior strategy library, a first decision instruction is determined on the teacher's end, and the virtual teacher is controlled to execute the first decision instruction on the teacher's end; the first decision instruction is used to characterize the teacher's end teaching task corresponding to the target teaching activity type; Based on the target behavior strategy library, a second decision instruction is determined on the student's end, and the virtual teacher is controlled to execute the second decision instruction on the student's end; the second decision instruction is used to characterize the student's end teaching task corresponding to the target teaching activity type.

7. The virtual dual-teacher teaching dynamic scheduling method according to any one of claims 1-3, characterized in that, Also includes: Receive the scheduling instructions from the teacher for the virtual teacher; The virtual teacher is controlled to execute different teaching tasks corresponding to the scheduling instructions on different presentation terminals.

8. A virtual dual-teacher teaching dynamic scheduling system, characterized in that, include: A multimodal sensing device array and a computing and storage server, wherein the multimodal sensing device array is connected to the computing and storage server; The multimodal sensing device array is used to collect multimodal teaching data during the teacher's classroom teaching process and transmit the multimodal teaching data to the computing and storage server; The computing and storage server is used to execute the virtual dual-teacher teaching dynamic scheduling method as described in any one of claims 1-7.

9. The virtual dual-teacher teaching dynamic scheduling system according to claim 8, characterized in that, The multimodal sensing device array includes: visual sensors, auditory sensors, and different presentation terminals deployed in different locations within the classroom; The visual sensor is used to collect the visual data stream in the multimodal teaching data; The auditory sensor is used to collect auditory data streams from the multimodal teaching data; The different presentation terminals are used to collect the interactive data streams they generate.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the virtual dual-teacher teaching dynamic scheduling method as described in any one of claims 1-7.