Intelligent teacher system and method for integrated teaching of virtual humans and real teachers
Through live broadcast of MR equipment, voice interaction and spatial coordinate adjustment, the delay and synchronization problems in the integration of virtual people and real teachers are solved, and efficient remote interaction and teaching quality optimization are achieved.
Patent Information
- Application Number
- CN202510474480.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When existing virtual people and real teachers integrate teaching, there are problems such as large live broadcast delay, insufficient interaction, poor synchronization and difficulty in real-time correction of spatial coordinates, resulting in poor teaching experience for students.
The MR access synchronization classroom module, virtual human interaction module, spatial coordinate real-time correction module and pixel streaming service module are adopted, and combined with the smart teacher module, the integrated teaching of virtual people and real teachers is realized, including live video streaming through MR equipment, voice and virtual human interaction, spatial coordinate dynamic adjustment and pixel streaming services, integrating pre-class preview, classroom assisted teaching and after-class tutoring for teaching activities.
It realizes that the live broadcast delay is small, and students can observe the teacher demonstration through the smart blackboard screen, support voice and virtual people interaction, synchronize virtual content with the real environment, provide a smooth and efficient remote interaction experience, ensure seamless connection and intelligent assistance of teaching activities, and improve teaching quality.
Smart Images

Figure CN119996722B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of smart education and virtual reality technology, and more specifically, to a smart teacher system and method for integrated teaching between virtual humans and real teachers. Background Art
[0002] With the rapid development of information technology, the field of education is undergoing unprecedented changes. Traditional teaching methods can no longer meet the current diversified and personalized teaching needs. With the continuous innovation of virtual reality (VR) and augmented reality (AR) technologies, the field of smart education is gradually integrating virtual and real teaching scenarios, forming a collaborative teaching model between virtual and real teachers. As an important tool to enhance interactive experience and improve learning outcomes, virtual human technology has been widely used in distance learning, online education, and virtual classrooms.
[0003] However, when existing virtual humans and real teachers are integrated into teaching, there are often problems such as large live broadcast delays, insufficient interactivity, and poor synchronization, which leads to a poor teaching experience for students; and it is also inconvenient to correct spatial coordinates in real time.
[0004] In view of this, the present invention proposes an intelligent teacher system and method for integrated teaching of virtual humans and real teachers to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned shortcomings of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solutions: a smart teacher system and method for integrated teaching of virtual humans and real teachers, 1. The smart teacher system for integrated teaching of virtual humans and real teachers is characterized by comprising:
[0006] MR access synchronous classroom module, which is used to connect the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast;
[0007] A virtual human interaction module, which is used to interact with a virtual human through voice. The virtual human interaction module includes a voice control SDK and client voice interaction logic;
[0008] A spatial coordinate real-time correction module, which is used to dynamically adjust the spatial coordinates;
[0009] Pixel Streaming service module, which is used to establish the Pixel Streaming service on the Ubuntu system;
[0010] The smart teacher module is used to integrate pre-class preparation, classroom auxiliary teaching and after-class tutoring in teaching activities.
[0011] Furthermore, the method of connecting the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast includes:
[0012] The instructor conducts a live demonstration through the MR device, which starts the video stream and projects it onto the laptop. The laptop then starts the synchronous classroom client and synchronizes the local video stream to the server. Subsequently, the synchronous classroom client pulls the live video stream to observe the live broadcast, and the students watch the instructor's live demonstration through the synchronous classroom client.
[0013] Furthermore, the voice control SDK includes three working states: STATE_IDLE, STATE_READY and STATE_WORKING.
[0014] Furthermore, the client voice interaction logic includes:
[0015] The client needs to integrate the wake-up SDK and voice control SDK;
[0016] The client establishes a websocket connection with the speech analysis service and performs login authentication.
[0017] Furthermore, the method of interacting with the virtual person through voice includes:
[0018] When the user says the wake-up word, the wake-up SDK throws a wake-up event, the client generates a wake-up ID, and then sends the wake-up event to the speech analysis service. The client then calls the pixel push service to play the virtual human's response.
[0019] After the HoloLens finishes playing the wake-up response, it sends a voice stream to the voice analysis service. The voice stream ID must be included when sending the voice stream until the backend responds to the command to stop the voice stream.
[0020] After receiving the voice stream from the HoloLens, the speech parsing service will eventually respond with a set of instructions to the HoloLens.
[0021] Furthermore, the method of dynamically adjusting the spatial coordinates includes:
[0022] Place a holographic object in HoloLens and fix its position and rotation by adding spatial anchor points.
[0023] Use HoloLens' built-in sensors and cameras to capture and analyze environmental information, and create a spatial map by scanning the surrounding environment and identifying feature points;
[0024] Add a spatial anchor component to the root GameObject of the holographic object, and attach a spatial anchor component with a relative position offset to its child GameObject;
[0025] Use HoloLens to capture and identify QR codes in the scene, and use the camera and image processing system to complete real-time detection and decoding of the information in the QR code. After the QR code is recognized, the location information can be obtained;
[0026] Subsequently, the position and orientation of the virtual content in the mixed reality environment are adjusted according to the location information of the QR code.
[0027] Furthermore, the method of establishing the pixel streaming service on the Ubuntu system includes:
[0028] Build a server environment that supports pixel streaming and deploy applications and content that need to be streamed on the server;
[0029] Then, after deploying the corresponding version of Unreal Engine on the Ubuntu system, it is transmitted to the terminal device in the form of a video stream through the pixel streaming plug-in for display and operation.
[0030] Furthermore, the pre-class preparation includes a learning situation analysis unit, a teaching activity classification & intelligent test composition & auxiliary review unit, a knowledge base and intelligent question and answer unit, a resource library and resource retrieval unit, and a word cloud analysis unit;
[0031] The classroom auxiliary teaching includes a virtual-reality integrated scene generation unit, an MR & flat-panel simulation training unit, a large-screen queuing and calling unit, an MR mixed reality training unit, and a large-screen step statistics unit;
[0032] The after-class tutoring includes a class content summary unit, a personalized resource recommendation unit, a class quality analysis unit, an electronic teaching plan unit, and a statistical analysis unit.
[0033] The intelligent teacher method for integrated teaching of virtual humans and real teachers includes the following steps:
[0034] S1. Connect the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast;
[0035] S2. Interacting with a virtual human through voice. The virtual human interaction module includes a voice control SDK and client voice interaction logic.
[0036] S3. Dynamically adjust the spatial coordinates;
[0037] S4. Establish pixel streaming service on Ubuntu system;
[0038] S5. Integrate pre-class preparation, classroom auxiliary teaching and after-class tutoring in teaching activities.
[0039] The technical effects and advantages of the intelligent teacher system and method for integrated teaching of virtual humans and real teachers of the present invention are as follows:
[0040] The present invention can reduce the live broadcast delay, making it easier for students to observe the instructor's live demonstration through the smart blackboard screen; it can control the audio and video equipment and interact with the back-end voice analysis service, thereby enabling users to interact with virtual people through voice; it can automatically adjust the position of virtual content when the user moves or the environment changes to ensure its synchronization with the real world, and can transmit the virtual human application screen on the server to the remote client device in real time, and provide users with a smooth and efficient remote interactive experience, which can achieve seamless connection and intelligent assistance of teaching activities, and ensure the continuous optimization and improvement of teaching quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A schematic diagram of the structure of the intelligent teacher system for integrated teaching of virtual humans and real teachers according to the present invention;
[0042] Figure 2 A flowchart of the intelligent teacher method for integrated teaching of virtual humans and real teachers according to the present invention;
[0043] Figure 3 This is a schematic diagram of the bridging solution for MR access to the synchronous classroom module of the present invention;
[0044] Figure 4 This is a physical topology diagram of the bridging solution for MR access to the synchronous classroom module of the present invention;
[0045] Figure 5 This is a schematic diagram of the state transition relationship of the voice control SDK of the present invention;
[0046] Figure 6 This is a schematic diagram of the client voice interaction logic of the present invention;
[0047] Figure 7 A schematic diagram of the voice interaction logic of the present invention;
[0048] Figure 8 Schematic diagram of the smart teacher module of the present invention. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0050] Example 1
[0051] See also Figure 1 、 Figures 3 to 8As shown, the smart teacher system for integrated teaching of virtual humans and real teachers described in this embodiment includes:
[0052] MR access synchronous classroom module, which is used to connect the first-person video stream in the MR device to the synchronous classroom for live broadcast. Among them, MR can only provide video streams in webrtc format, and the synchronous classroom can only receive rtsp video streams;
[0053] A virtual human interaction module, which is used to interact with a virtual human through voice. The virtual human interaction module includes a voice control SDK and client voice interaction logic;
[0054] A spatial coordinate real-time correction module, which is used to dynamically adjust the spatial coordinates;
[0055] Pixel Streaming service module, which is used to establish the Pixel Streaming service on the Ubuntu system;
[0056] The smart teacher (teacher) module is used to integrate pre-class preparation, classroom auxiliary teaching and after-class tutoring in teaching activities.
[0057] Furthermore, if Figure 3 As shown, the method of connecting the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast includes:
[0058] The instructor conducts a live demonstration using an MR device. The MR device starts the video stream and projects it onto a laptop. The laptop then starts the synchronous classroom client and synchronizes the local video stream to the server. The synchronous classroom client then pulls the live video stream to watch the live broadcast. Students then watch the instructor's live demonstration using the synchronous classroom client.
[0059] Specifically, this live broadcast solution has a relatively low latency and no noticeable image lag, but the process logic is relatively complex and the link is relatively long, so this solution was selected.
[0060] Among them, the screen projection solution is:
[0061] Confirm that the device has a built-in unlimited screen mirror function, which requires Internet access to download (or install third-party software);
[0062] Screen projection depends on the network card, and the device network card must support the Miracast protocol;
[0063] The physical topology of live streaming is:
[0064] The physical topology of the entire live broadcast is as follows Figure 4 As shown, the field equipment is connected to the teaching building network through optical fiber, and the field equipment is connected to the local area network;
[0065] The equipment in the preparation area includes an 86-inch all-in-one computer, a podium host, and an audio host. The outdoor equipment includes MR equipment, a live broadcast laptop, and Bluetooth headphones.
[0066] The instructor broadcasts live through the MR device outside the experiment field, and the MR device synchronizes the video stream to the laptop. The AI classroom client on the laptop will take the video stream on the local computer and upload it to the server. The AI classroom client on the smart blackboard in the preparation area outside the experiment field will pull the live stream from the server, and the students will observe the instructor's live demonstration through the smart blackboard screen.
[0067] Furthermore, the Voice Control SDK includes three working states: STATE_IDLE (service not started), STATE_READY (awaiting state), and STATE_WORKING (working state), which are described in the following table. Specifically, during the SDK operation, it is in different states at different stages, and different states can handle different operations:
[0068] Status Name illustrate STATE_IDLE If the service is not enabled, you can only perform the start (enable service) operation. STATE_READY In the ready-to-wake-up state, you can wake up the service by voice (wake-up word) or by sending a CMD_WAKEUP message directly to the service. After calling AIUIAgent.createAgent to create the object, the service is in the ready state. STATE_WORKING In working state, you can input voice and text to interact with the AIUI background.
[0069] The client can control the running state of the SDK through SDKMessage. When SDKMessage is just created, the SDK is in the service-inactive state. After sending the CMD_START message, the SDK will switch to the wake-up state. When the SDK receives the wake-up command, the SDK will switch to the working state. At this time, it can interact with the backend service through voice or text. The specific conversion relationship is as follows: Figure 5 As shown in the table below, the states of the state transition diagram are as follows:
[0070] Operation Name illustrate start After startup, the system will be in the default state or send a CMD_START message to the SDK. stop Send a CMD_STOP message to the SDK. wakeup Say the custom wake-up word ("ding dong ding dong" by default), or send a CMD_WAKEUP message to the SDK. reset_wakeup Send a CMD_RESET_WAKEUP message to the SDK. sleep Sleep, when no valid interaction (semantics) occurs for a period of time. re_wakeup In the STATE_WORKING state, say the wake-up word again or send a CMD_WAKEUP message to the SDK.
[0071] Specifically, the status query can be sent to the SDK by constructing a CMD_GET_STATE query message. The SDK will return the current status through the EVENT_STATE event. The arg1 parameter of the EVENT_STATE event indicates the status value, which can be:
[0072] 1=>STATE_IDLE (idle state);
[0073] 2=>STATE_READY (ready state, waiting to be woken up);
[0074] 3=>STATE_WORKING (working state, awakened).
[0075] Open and close the SDK by using CMD_START and CMD_STOP to control the start and stop of the SDK;
[0076] The SDK does not perform any operations in the stopped state, and the power consumption is also the lowest at this time. It cannot be awakened in the stopped state and needs to enter the ready state through CMD_START to wake up;
[0077] CMD_RESET is used to reset the service in case of a fatal error that cannot be recovered or to reread the configuration file.
[0078] SDK wake-up and sleep are as follows:
[0079] When the SDK is in sleep mode, you can wake it up by saying a custom wake-up word (such as "Hello, Qihang") or sending a CMD_WAKEUP message to the SDK.
[0080] After entering the working state, you can interact through voice or text, but if there is no effective interaction for a period of time (configurable in the configuration file interact_timeout), it will enter the ready state. You can also enter the dormant state by manually sending CMD_RESET_WAKEUP;
[0081] Both of the above sleep modes will throw the EVENT_SLEEP event to indicate that the SDK has entered the sleep state. The arg1 field indicates the mode of entry into sleep mode.
[0082] 0=>TYPE_AUTO (automatic sleep, i.e. interaction timeout),
[0083] 1=>TYPE_COMPEL (external forced sleep, i.e. sending CMD_RESET_WAKEUP).
[0084] The wake-up result is as follows: after the SDK enters the wake-up state, the corresponding wake-up event is thrown through the EVENT_WAKEUP type message.
[0085] Furthermore, the client voice interaction logic includes:
[0086] The client needs to integrate the wake-up SDK and voice control SDK;
[0087] The client establishes a websocket connection with the voice analysis service and performs login authentication. The authentication information includes device ID, login name, login password, user information, etc.
[0088] Specifically, the client carries the image of a virtual person, can control the audio and video equipment and interact with the back-end voice analysis service, so as to realize the function of users interacting with the virtual person through voice. Figure 6 shown.
[0089] Furthermore, if Figure 7As shown, the ways to interact with virtual humans through voice include:
[0090] Wake-up: When the user says the wake-up word, the wake-up SDK throws a wake-up event, and the client generates a unique wake-up ID (wakeid). The wake-up ID must be unique and non-duplicate. The client then sends the wake-up event to the speech analysis service, which must include the device ID, login information, user information, etc. The client then calls the pixel push service to play the virtual human response (for example: Hello, I'm here).
[0091] Send the voice stream. After the HoloLens plays the wake-up response, it sends the voice stream to the voice analysis service. The voice stream ID (ASRI) must be included when sending the voice stream. The ID must be unique and non-duplicate. The voice stream is continuously sent until the backend responds to the command to stop the voice stream.
[0092] In response to the instruction set, the speech analysis service will perform speech recognition, semantic understanding, and dialogue processing after receiving the voice stream from the HoloLens. Finally, it will respond to the instruction set to the HoloLens. The instruction set includes displaying the recognition results, stopping the voice stream, playing the speech, sending the voice stream, opening the browser, opening the file, opening the audio and video, opening a link, etc. The HoloLens executes the instruction set in sequence.
[0093] Specifically, the specific explanation of each instruction is as follows: Display recognition results: display the voice recognition results at the appropriate position of HoloLens; Stop voice stream: HoloLens stops inputting voice stream to the voice analysis service; Play speech: HoloLens calls the pixel stream service collection voice synthesis engine to play the virtual human audio and video stream speech; Send voice stream: The voice analysis service returns an instruction set, which will contain multiple instructions. If it is a multi-round conversation, the returned instruction set will first stop the voice stream, play the speech, and then continue to send the voice stream to the voice analysis service, and then interact with the voice analysis service for multiple rounds. When executing this instruction, a new voice stream ID needs to be regenerated; Open browser: open the HoloLens local browser; Open file: open a specific file; Open audio and video stream: open an audio and video file; Open a link: open a connection address, etc.
[0094] Furthermore, the method of dynamically adjusting the spatial coordinates includes:
[0095] By placing a holographic object in HoloLens and fixing its position and rotation by adding spatial anchors (so that the holographic object maintains its relative position even when the user moves their head or changes their viewing angle);
[0096] Use the HoloLens' built-in sensors and cameras to capture and analyze environmental information, and create a spatial map by scanning the surrounding environment and identifying feature points (these feature points can be walls, furniture, or other fixed objects in the room; once the spatial map is created, the precise location of holographic objects can be determined based on the feature points in the map);
[0097] Add a spatial anchor component to the root GameObject of the holographic object, and attach a spatial anchor component with a relative position offset to its child GameObjects (in HoloLens development, spatial anchors are implemented through scripting. When a holographic object is placed in a scene, its position and rotation are anchored to a specific point in space). When the user leaves the scene or closes the app, the hologram in the scene is saved at the same location (thus, when the user re-enters the scene or reopens the app, the holographic content in the previous scene can be accurately restored);
[0098] Use HoloLens to capture and recognize QR codes in the scene, and use the camera and image processing system to complete real-time detection and decoding of the information in the QR code. After recognizing the QR code, the location information can be obtained (HoloLens corrects the spatial coordinates through spatial calculation and coordinate conversion based on the extracted location information);
[0099] Then, the position and orientation of the virtual content (model and interface) in the mixed reality environment are adjusted according to the position information of the QR code to make it consistent with the real environment;
[0100] Specifically, HoloLens uses its built-in multiple cameras and sensors to capture and analyze information about the surrounding environment. These cameras can capture depth information, RGB images, and infrared data, while sensors are used to detect movements, rotations, and tilts of the device. Based on the collected environmental data, HoloLens creates a virtual spatial map, which not only contains the location of objects, but also their shape, size, and positional relationship relative to the user. Once the spatial map is created, HoloLens begins to track the user's head and hand movements in real time, which is usually done through sensors such as accelerometers, gyroscopes, and magnetometers. Through these data, HoloLens can Updates the user's position and orientation in the virtual space. When the user moves or changes their head posture, HoloLens dynamically calibrates the spatial map based on its real-time tracking data. This means that virtual content adjusts in real time to the user's actual position and head orientation, ensuring that it always aligns with the user's perspective. HoloLens also includes a user feedback mechanism that allows users to fine-tune spatial correction through gestures or voice commands. This interactive feedback loop helps improve the accuracy of spatial correction and the user experience. Furthermore, by using real-time QR code scanning to achieve real-time correction of spatial coordinates, the synchronization between virtual content and the real environment within the project is improved, providing users with a more natural and realistic experience.
[0101] It should be noted that the above content is continuously detected and recognized by the HoloLens device during operation, and the spatial coordinates are dynamically adjusted according to the position information of the QR code, thereby ensuring the accurate alignment of virtual content with the real environment. It can automatically adjust the position of virtual content when the user moves or the environment changes to ensure its synchronization with the real world; and can maximize the accuracy of the system's virtual and real superposition when it comes to process operations and interaction with the real cockpit and helicopter.
[0102] Furthermore, the method for setting up the Pixel Streaming service on an Ubuntu system includes:
[0103] Build a server environment that supports Pixel Streaming (the server environment includes installing and configuring the necessary software and hardware resources), and deploy the applications and operational content that need to be streamed on the server;
[0104] Then, after deploying the corresponding version of Unreal Engine on the Ubuntu system, the video is transmitted to the terminal device (client) in the form of a video stream through the pixel streaming plug-in. The client device can be any device that supports a web browser, such as a computer, tablet or mobile phone. The client accesses the web page on the web server through the browser and receives the video stream from the server. The client currently used is HoloLens. On the client device, you can use HTML5 <video>Tags or other related technologies to display the received video stream) for display and operation (the Pixel Streaming plugin is part of the Unreal Engine and is responsible for capturing the application's rendering output and encoding it into a video stream. Furthermore, the Pixel Streaming plugin runs on the server side and can send the application's real-time image to the client with minimal latency through efficient encoding and decoding algorithms). This also includes a signaling server deployed on the Ubuntu system, which uses the WebSocket communication protocol;
[0105] Specifically, through the collaborative work of the above components, the pixel streaming technology architecture based on the Ubuntu system can transmit the virtual human application screen on the server to the remote client device in real time, providing users with a smooth and efficient remote interaction experience.
[0106] Furthermore, if Figure 8 As shown, the pre-class preparation includes a learning situation analysis unit, a teaching activity classification & intelligent test composition & auxiliary review unit, a knowledge base and intelligent question and answer unit, a resource library and resource retrieval unit, and a word cloud analysis unit:
[0107] The learning situation analysis unit is used to collect students' test data, evaluate their knowledge mastery, and track their weaknesses. Specifically, by collecting and analyzing students' test data in real time, it saves teachers' time. The assessment can accurately capture students' real-time knowledge mastery. The weak point tracking can summarize and update the weak knowledge points of each class, supporting effect verification and improvement.
[0108] The teaching activity classification, intelligent test-taking, and assisted grading unit distinguishes learning activity types and monitors learning progress and duration in real time. Objective questions are used to assess students' mastery of knowledge points, and automated test-taking is used to intelligently select questions, alleviating the burden on instructors and ensuring comprehensive coverage of knowledge points. Intelligent assisted grading compares students' key answers with standard answers and provides AI-generated scores.
[0109] The knowledge base and intelligent question-and-answer unit allows administrators to upload teaching materials, and the intelligent teacher module automatically learns and generates a structured knowledge base. When students ask questions via voice, the intelligent teacher provides answers based on the learned knowledge base content. This flexible and convenient learning method can enhance students' question-asking and problem-solving experience.
[0110] In the resource library and resource search unit, teachers and students can quickly find specific resources through voice commands; it also includes resource classification management of the resource library, thereby enhancing the search function and improving the utilization rate of teaching resources;
[0111] In the word cloud analysis unit, students share their questions and ideas in real time through word clouds; this makes it easier for students to express their doubts, and teachers can understand students' thoughts and needs before, during, and after class in real time. It can also help optimize teaching content and interaction methods, and improve teaching effectiveness.
[0112] Classroom auxiliary teaching includes virtual and real scene generation unit, MR & tablet simulation training unit, large screen queuing unit, MR mixed reality training unit and large screen step counting unit;
[0113] In the virtual-reality integrated scene generation unit, the instructor's perspective is projected onto the screen in real time, allowing students to participate in discussions immersively, deepening their memory and sense of interaction. MR enhances the presentation of teaching content, promoting students' in-depth understanding and application, and sparking profound discussions. This transcends traditional spatial limitations, allowing students to learn in different locations and enrich their teaching experience and scenario applications.
[0114] In the MR & flat-panel simulation training unit, students use simulation technology to simulate various flight environments and unexpected situations, thereby enhancing their emergency response capabilities. This allows students to learn in a safe environment, reduce operational risks, and minimize personal injury and equipment damage. By using customized software and general hardware, they can reduce dependence on complex equipment and simplify system updates and maintenance.
[0115] In the large-screen queuing and calling unit, students are automatically assigned the best time and equipment through the intelligent queuing system. Automatic calling is combined with instructor assignments, thereby avoiding resource waste and time conflicts, eliminating the inefficiency and confusion of traditional scheduling, ensuring the effective use of scarce resources and classroom time, and optimizing training efficiency.
[0116] In the MR mixed reality training unit, trainees can use MR glasses combined with digital human guidance to reduce equipment operation errors and improve training effectiveness and efficiency. By combining real equipment with virtual elements, training guidance can be provided to deeply understand equipment operation and enhance the trainees' training experience. Through trainees' exposure to and use of MR technology, they can cultivate innovative thinking and the ability to adapt to new technologies. Furthermore, MR technology can help trainees adapt to high-tech warfare in advance and enhance their ability to respond to future battlefields.
[0117] In the large-screen step statistics unit, students' immediate operation results are given real-time feedback, data visualization analysis, and automatic recording and archiving. As a result, students can obtain intuitive and immediate operation results, promote error learning and skill improvement, and instructors optimize training through data analysis to improve teaching effectiveness and student skills. All operation records are automatically saved and available for review and analysis at any time, supporting personalized training and skill development.
[0118] After-class tutoring includes a unit on class content summary, personalized resource recommendation, class quality analysis, electronic lesson plan, and statistical analysis.
[0119] In the course content summary unit, by quickly summarizing the key points of the course, you can save time and energy, avoid the tediousness of reviewing the entire video, and deepen your understanding and memory through the knowledge structure, helping you grasp the core concepts.
[0120] In the personalized resource recommendation unit, relevant resources are recommended based on the learning situation, thus realizing personalized resource recommendations. Through intelligent matching of video clips and exercises, precise review of weak points can be carried out, so as to accurately fill in knowledge gaps and improve learning effects. By optimizing the learning path, it can reduce ineffective learning time, build a complete knowledge system, and improve students' confidence and academic performance.
[0121] The classroom quality analysis unit provides detailed reports that accurately record and analyze teaching details, support precise adjustment of teaching strategies, monitor listening status in real time, identify attention troughs, and guide instructors to optimize classroom content and methods. This data-based teaching improves student engagement and learning outcomes, ensuring that teaching objectives are achieved.
[0122] In the electronic teaching plan unit, the teaching process is recorded and transcribed through multiple cameras for automatic recording and transcription, which can save time in editing and ensure the accuracy of the teaching plan. Through rapid positioning and adjustment, knowledge points are automatically switched and positioned, so that the teaching can be optimized in combination with classroom analysis to improve teaching effectiveness. Through one-click download and update, the teaching plan update process can be simplified, accurately meeting the needs of teachers and improving teaching efficiency.
[0123] In the statistical analysis unit, by analyzing course resources, student performance, and classroom quality, teachers can gain a deeper understanding of learning progress and teaching effectiveness. By identifying and intervening in abnormal situations in learning and teaching, they can improve teaching quality and efficiency, provide accurate feedback and adjustment suggestions, and ensure that teaching reaches the optimal level by the end of the semester.
[0124] Specifically, the Smart Instructor module is designed around theory, labs, and equipment practice courses. Small-scale pilot teaching has been conducted on helicopter aeronautical electrical engineering, helicopter fire control equipment, and helicopter equipment support practices. By integrating pre-class preparation, classroom training, and post-class review, it achieves seamless and intelligent teaching support. During the pre-class preparation phase, instructors use learning scores to assess students' mastery of historical knowledge points, allowing them to make targeted decisions about whether to revisit weak points in the next class. Through categorized teaching activities, instructors can issue pre-study tasks and monitor students' progress, ensuring efficient reading and learning. Intelligent test publishing and performance analysis further help instructors understand students' knowledge mastery after pre-study review, allowing them to adjust their teaching plans. Students can use voice search to identify weak points and utilize high-quality courseware, videos, and other materials for independent reinforcement learning. When students encounter difficult questions, they can also obtain timely and accurate answers through intelligent Q&A. During classroom teaching and training, seamless attendance in theory classes eliminates the need for instructors to call roll. It also automatically records and links attendance to regular grades, improving classroom efficiency. The borderless classroom enables cross-domain live streaming, effectively integrating theory and practice, both on and off campus, and promoting resource sharing and academic exchange. During lectures, instructors can collaborate with the intelligent instructor to access case studies, courseware, and conduct quick in-class quizzes in real time, making classes more flexible and engaging and improving student attention. Intelligent supervision, such as behavioral analysis and random roll call, helps instructors maintain classroom order and encourages students to pay attention through real-time behavioral analysis. Advanced simulation and mixed reality technologies ensure a fully realistic and safe immersive training environment for students during lab and equipment practice sessions. Intelligent queuing and calling ensures efficient scheduling of equipment in preparation and training areas. Data analysis and visualization technologies provide students with intuitive, real-time feedback on operational results, and data is archived and recorded in real time, helping them learn from mistakes and improve. During post-class review, students can quickly review the key content of each lesson with the Smart Instructor to deepen their memory. Based on individual review tests and learning analysis, they can receive personalized recommendations to strengthen weak points. Instructors can use classroom quality analysis tools to review and summarize their own and their students' performance in each class period and optimize the teaching rhythm. The electronic lesson plan function provides instructors with a convenient way to update lesson plans, ensuring the continuous evolution of teaching content. Furthermore, statistical analysis allows instructors to timely review teaching activities, knowledge points, and changing trends in classroom quality over a period of time, identifying issues and implementing interventions, ensuring the continuous optimization and improvement of teaching quality.
[0125] In this embodiment, through the MR access synchronous classroom module, the instructor connects the first-person video stream from the MR device to the synchronous classroom for live broadcast, which reduces the time delay during the live broadcast and facilitates students to observe the instructor's live demonstration through the smart blackboard screen. The client carries the image of the virtual human, which can control the audio and video equipment and interact with the back-end voice analysis service, thereby enabling users to interact with the virtual human through voice. The real-time spatial coordinate correction module can automatically adjust the position of the virtual content when the user moves or the environment changes, ensuring its synchronization with the real world. Through the collaborative work of the Unreal Engine and the application, pixel streaming plug-in, signaling server, web server, network transport layer, and client device components, the Ubuntu system-based pixel streaming can transmit the virtual human application screen on the server to the remote client device in real time, providing users with a smooth and efficient remote interaction experience. Through the combination of pre-class preparation, classroom auxiliary teaching, and after-class tutoring in the smart instructor module, seamless connection and intelligent assistance of teaching activities are achieved, and problems can be discovered and intervened in a timely manner, ensuring the continuous optimization and improvement of teaching quality.
[0126] Example 2
[0127] See also Figure 2 As shown, the smart teacher method for integrated teaching of virtual humans and real teachers described in this embodiment includes the following steps:
[0128] S1. Connect the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast;
[0129] S2. Interacting with a virtual human through voice. The virtual human interaction module includes a voice control SDK and client voice interaction logic.
[0130] S3. Dynamically adjust the spatial coordinates;
[0131] S4. Establish pixel streaming service on Ubuntu system;
[0132] S5. Integrate pre-class preparation, classroom auxiliary teaching and after-class tutoring in teaching activities.
[0133] In this embodiment, the live broadcast can have a small delay, which makes it convenient for students to observe the instructor's live demonstration through the smart blackboard screen; it can control the audio and video equipment and interact with the back-end voice analysis service, so as to realize the function of users interacting with virtual people through voice; it can automatically adjust the position of virtual content when the user moves or the environment changes to ensure its synchronization with the real world, and can transmit the virtual human application screen on the server to the remote client device in real time, and provide users with a smooth and efficient remote interactive experience, and can also achieve seamless connection and intelligent assistance of teaching activities, timely discover problems and implement intervention to ensure the continuous optimization and improvement of teaching quality.
[0134] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0135] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only one type. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0136] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
[0137] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.< / video>
Claims
1. The intelligent teacher system for integrated teaching of virtual humans and real teachers is characterized by: include: MR access synchronous classroom module, which is used to connect the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast; A virtual human interaction module, which is used to interact with a virtual human through voice. The virtual human interaction module includes a voice control SDK and client voice interaction logic; A spatial coordinate real-time correction module, which is used to dynamically adjust the spatial coordinates; Pixel Streaming service module, which is used to establish the Pixel Streaming service on the Ubuntu system; The smart teacher module is used to integrate pre-class preparation, classroom auxiliary teaching and after-class tutoring in teaching activities; The method of dynamically adjusting the spatial coordinates includes: Place a holographic object in HoloLens and fix its position and rotation by adding spatial anchor points. Use HoloLens' built-in sensors and cameras to capture and analyze environmental information, and create a spatial map by scanning the surrounding environment and identifying feature points; Add a spatial anchor component to the root GameObject of the holographic object, and attach a spatial anchor component with a relative position offset to its child GameObject; Use HoloLens to capture and identify QR codes in the scene, and use the camera and image processing system to complete real-time detection and decoding of the information in the QR code. After the QR code is recognized, the location information can be obtained; Subsequently, the position and orientation of the virtual content in the mixed reality environment are adjusted according to the location information of the QR code.
2. The intelligent teacher system for integrated teaching of virtual humans and real teachers according to claim 1 is characterized in that: The method of connecting the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast includes: The instructor conducts a live demonstration through the MR device, which starts the video stream and projects it onto the laptop. The laptop then starts the synchronous classroom client and synchronizes the local video stream to the server. Subsequently, the synchronous classroom client pulls the live video stream to observe the live broadcast, and the students watch the instructor's live demonstration through the synchronous classroom client.
3. The intelligent teacher system for integrated teaching of virtual humans and real teachers according to claim 1 is characterized in that: The voice control SDK includes three working states: STATE_IDLE, STATE_READY and STATE_WORKING.
4. The intelligent teacher system for integrated teaching of virtual humans and real teachers according to claim 1 is characterized in that: The client voice interaction logic includes: The client needs to integrate the wake-up SDK and voice control SDK; The client establishes a websocket connection with the speech analysis service and performs login authentication.
5. The intelligent teacher system for integrated teaching of virtual humans and real teachers according to claim 1 is characterized in that: The method of interacting with the virtual person through voice includes: When the user says the wake-up word, the wake-up SDK throws a wake-up event, the client generates a wake-up ID, and then sends the wake-up event to the speech analysis service. The client then calls the pixel push service to play the virtual human's response. After the HoloLens finishes playing the wake-up response, it sends a voice stream to the voice analysis service. The voice stream ID must be included when sending the voice stream until the backend responds to the command to stop the voice stream. After receiving the voice stream from the HoloLens, the speech parsing service will eventually respond with a set of instructions to the HoloLens.
6. The intelligent teacher system for integrated teaching of virtual humans and real teachers according to claim 1 is characterized in that: The method for establishing the Pixel Streaming service on a Ubuntu system includes: Build a server environment that supports pixel streaming and deploy applications and content that need to be streamed on the server; Then, after deploying the corresponding version of Unreal Engine on the Ubuntu system, it is transmitted to the terminal device in the form of a video stream through the pixel streaming plug-in for display and operation.
7. The intelligent teacher system for integrated teaching of virtual humans and real teachers according to claim 1 is characterized in that: The pre-class preparation includes a learning situation analysis unit, a teaching activity classification & intelligent test composition & auxiliary review unit, a knowledge base and intelligent question and answer unit, a resource library and resource retrieval unit, and a word cloud analysis unit; The classroom auxiliary teaching includes a virtual-reality integrated scene generation unit, an MR & flat-panel simulation training unit, a large-screen queuing and calling unit, an MR mixed reality training unit, and a large-screen step statistics unit; The after-class tutoring includes a class content summary unit, a personalized resource recommendation unit, a class quality analysis unit, an electronic teaching plan unit, and a statistical analysis unit.
8. A smart teacher method for integrated teaching of virtual humans and real teachers, applied to a smart teacher system for integrated teaching of virtual humans and real teachers as described in any one of claims 1 to 7, characterized in that: The following steps are involved: S1. Connect the first-person perspective video stream in the MR device to the synchronous classroom for live broadcast; S2. Interacting with a virtual human through voice. The virtual human interaction module includes a voice control SDK and client voice interaction logic. S3. Dynamically adjust the spatial coordinates; S4. Establish pixel streaming service on Ubuntu system; S5. Integrate pre-class preparation, classroom supplementary teaching and after-class tutoring in teaching activities; The method of dynamically adjusting the spatial coordinates includes: Place a holographic object in HoloLens and fix its position and rotation by adding spatial anchor points. Use HoloLens' built-in sensors and cameras to capture and analyze environmental information, and create a spatial map by scanning the surrounding environment and identifying feature points; Add a spatial anchor component to the root GameObject of the holographic object, and attach a spatial anchor component with a relative position offset to its child GameObject; Use HoloLens to capture and identify QR codes in the scene, and use the camera and image processing system to complete real-time detection and decoding of the information in the QR code. After the QR code is recognized, the location information can be obtained; Subsequently, the position and orientation of the virtual content in the mixed reality environment are adjusted according to the location information of the QR code.
Citation Information
Patent Citations
Multifunctional flight teaching platform
CN105976672A
Intelligent driving training system and method based on augment virtual reality man-machine interaction
CN106710360A
Cited By
An adaptive seamless video interaction method and system based on state perception and timing scheduling
CN122765242A