A cognitive driving-based real-time intelligent teaching system and method for college courses

CN122885104APending Publication Date: 2026-10-09WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610927148.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0005]为解决现有课程智能辅助教学中知识回答缺少课程资料约束、复杂问答链路难以动态组织、交互方式单一以及学习状态无法持续更新的问题,本发明提供了一种融合RAG知识检索、Agent工作流与数字人交互的认知事件驱动课程智能辅助教学方法,该方法通过以RAG课程知识检索增强为知识基础、以Agent式工作流编排为任务组织方式、以数字人语音视频交互为反馈形式、以认知事件回流为学习状态建模依据的技术手段,达到了使课程问答受知识库约束、复杂问答链路可动态编排、数字人讲解与问答深度融合、学习状态可持续建模与更新的目的

Benefits of technology

1. 通过RAG课程知识检索增强约束回答来源。系统在大模型生成前先检索课程知识库,将FAISS语义检索、BM25关键词检索和多轮上下文共同作为回答依据,使课程问答更贴合教材、课件和课程资料,减少通用大模型脱离课程内容生成答案的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122885104A_ABST
    Figure CN122885104A_ABST
Patent Text Reader

Abstract

The application discloses a college course real-time intelligent teaching system and method based on cognitive driving. The system comprises a course knowledge base construction module, an Agent workflow arrangement module, a hybrid retrieval module, a large model generation module, a digital human interaction module and a cognitive event backflow module. The course knowledge base construction module constructs vector index and keyword index; the Agent workflow arrangement module schedules multiple configurable nodes; the hybrid retrieval module performs semantic retrieval and keyword retrieval and fuses and rearranges the results; the large model generation module generates a text answer; the digital human interaction module generates a mouth shape synchronous video through voice synthesis and lip shape synchronization; and the cognitive event backflow module encapsulates cognitive events and triggers a timing graph atlas refresh. Through the fusion of RAG retrieval enhancement, Agent workflow arrangement, digital human interaction and cognitive event backflow, the application realizes intelligent question answering and personalized teaching feedback under the course knowledge constraint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence-based smart education, retrieval enhancement generation, semantic retrieval of course knowledge bases, agent-based workflow orchestration, multimodal learning interaction, digital human speech synthesis and lip-syncing, learning state tracking and personalized learning recommendation, and particularly to a cognitive event-driven intelligent auxiliary teaching method that integrates RAG knowledge retrieval, agent workflow and digital human interaction. Background Technology

[0002] With the iterative upgrades of large-scale artificial intelligence models and the advancement of digital transformation of campus education, intelligent teaching aids for university courses have gradually evolved from simple courseware and video resource display carriers into comprehensive teaching service platforms that integrate intelligent Q&A, online in-class quizzes, learning behavior statistical analysis, personal learning profile construction, and customized content delivery.

[0003] University professional courses are characterized by abstract and obscure theoretical concepts, closely interconnected knowledge points, strong reliance on prerequisite foundations, complex knowledge structures, and a wide range of exercise types. In their daily learning, students not only need to resolve theoretical doubts and break down core knowledge points in real time, but also need to combine online video self-study, targeted practice exercises, reviewing and summarizing incorrect answers, and periodically identifying and addressing knowledge gaps in order to fully build a systematic algorithm knowledge framework.

[0004] Currently, mainstream intelligent teaching tools for courses suffer from core problems such as fragmented functional modules and disconnected data. First, general-purpose large-scale models for Q&A are prone to detachment from course materials. While these models possess strong natural language generation capabilities, their answers typically rely on internal model parameters and lack strong constraints related to the course textbook, courseware, video materials, and teacher database. When students ask questions about course details, algorithm steps, complexity analysis, or code screenshots, the lack of course knowledge base retrieval results can lead to answers that are overly generalized, inconsistent with course chapters, and lack traceability. Second, the interaction format of simple text-based Q&A is limited. Existing intelligent Q&A tools primarily rely on text input and output, which is insufficient to match the interactive habits of "explanation, demonstration, follow-up questioning, and repetition" in classroom teaching. In video learning scenarios, students often want to ask questions directly at a specific point in the video and receive feedback from a digital human similar to a teacher's explanation. If the system only returns text answers, the interactive immersion is insufficient, and it fails to meet teaching needs such as pausing the course video for questions or receiving real-time digital human explanations. Third, existing course question-and-answer systems mostly employ fixed question-and-answer chains, lacking agent-based workflow orchestration capabilities. When faced with diverse inputs such as text questions, image questions, video-based questions, and quiz feedback, the systems struggle to automatically organize multiple steps, including course relevance assessment, query rewriting, mixed retrieval, image understanding, answer generation, digital human output, and event feedback, resulting in insufficient intelligent processing capabilities in complex teaching scenarios. Fourth, RAG question-and-answer, digital human feedback, and learning state modeling are typically independent. Even when some systems introduce RAG knowledge retrieval or digital human explanations, they often treat them merely as tools for single question-and-answer sessions or single video generation, failing to integrate data such as retrieval questions, answer content, video time points, digital human generation results, and quiz scores into a unified, calculable learning event. Therefore, the systems struggle to continuously update students' knowledge mastery, forgetting risks, and learning paths, failing to develop long-term intelligent assistance capabilities tailored to individual students. Summary of the Invention

[0005] To address the problems of existing intelligent assisted teaching methods in courses, such as the lack of course material constraints in knowledge responses, the difficulty in dynamically organizing complex question-and-answer chains, the limited interaction methods, and the inability to continuously update learning states, this invention provides a cognitive event-driven intelligent assisted teaching method that integrates RAG knowledge retrieval, agent-based workflow, and digital human interaction. This method achieves the following: RAG course knowledge retrieval is used as the knowledge foundation; agent-based workflow orchestration is used as the task organization method; digital human voice and video interaction is used as the feedback form; and cognitive event feedback is used as the basis for learning state modeling. This results in knowledge base-constrained question-and-answer sessions, dynamically orchestratable complex question-and-answer chains, deep integration of digital human explanation and question-and-answer, and continuous modeling and updating of learning states.

[0006] According to one aspect of this invention, a cognitive-driven real-time intelligent teaching system for university courses is provided, deployed on a server. The system includes: a course knowledge base construction module, used to parse course materials into text fragments and construct vector indexes and keyword indexes respectively; an Agent workflow orchestration module, used to respond to course questions submitted by students and schedule multiple configurable nodes based on a workflow orchestration framework, including at least course relevance judgment, query rewriting, hybrid retrieval, retrieval result fusion, answer generation, knowledge point extraction, and event feedback; and a hybrid retrieval module, used to perform semantic retrieval based on vector indexes and keyword indexes respectively in the hybrid retrieval nodes. The system includes a keyword index for keyword retrieval, which merges, deduplicates, and rearranges the two types of search results to obtain knowledge fragments related to the course content; a large model generation module for inputting the knowledge fragments, course questions, and contextual messages into a large language model to generate text answers; a digital human interaction module for converting the text answers into speech data through a speech synthesis model, and generating lip-synced videos based on the speech data using a lip-sync model, which then returns the lip-synced videos to the student's end; and a cognitive event feedback module for encapsulating the data generated during the question-and-answer process into cognitive event objects, writing them into the event database, and triggering a refresh of the temporal cognitive graph.

[0007] As a further technical solution, the course knowledge base construction module is also used to: generate vector representations of the text fragments using text-embedding-v4 and construct a FAISS vector index, while constructing a BM25 keyword index.

[0008] As a further technical solution, the Agent workflow orchestration module is built on LangGraph; the Agent workflow orchestration module is also used to select the execution path according to the modal type or function switch of the course question, so that text questions, image enhancement questions and video positioning questions can reuse a unified orchestration framework.

[0009] As a further technical solution, the digital human interaction module is also used to: call the GPT-SoVITS model to synthesize the text answer into speech data, call the Wav2Lip model to drive the digital human avatar to generate lip-sync video based on the speech data; and, when the streaming output conditions are met, return the streaming video address and frame rate information, and when streaming generation fails or the conditions are not met, roll back to generate a complete audio and video file.

[0010] As a further technical solution, a multimodal question-answering module is also included. This module is used to: respond to image question-answering requests submitted by students, call qwen-vl-max to identify the image type and extract content, concatenate the image analysis results with the student's question to form an enhanced question, and then enter the Agent workflow orchestration module; respond to video-based point-to-point question-answering requests submitted by students during course video playback, obtain the video identifier, playback time point, and question text, read the most recent multiple rounds of question-answering from the same video as context, enter the Agent workflow orchestration module, and save the question-answering related information of the video-based point-to-point question-answering to a VideoQARecord object.

[0011] As a further technical solution, a learning state update module is also included. This module is used to: based on the cognitive event object, employ a simplified Bayesian knowledge tracing model, combining initial mastery probability, learning transfer probability, guessing probability, error probability, and forgetting decay factor, to dynamically correct the student's mastery level of each knowledge point and update the KnowledgeMastery object; aggregate cognitive events within a specified time window to generate a PersonaTKGDailySnapshot, which includes a graph structure, a learning profile summary, and feature data.

[0012] As a further technical solution, a personalized recommendation module is also included. The personalized recommendation module is used to: generate a review plan based on the risk of forgetting; generate a learning path recommendation based on the knowledge mastery and the prerequisite relationship of knowledge points; wherein, the learning path recommendation prioritizes knowledge points whose prerequisite conditions have been met but whose mastery is insufficient, and includes the recommendation reasons, difficulty, and suggested resources.

[0013] As a further technical solution, a test management module is also included. The test management module is used to: respond to the test submission request from the student end, score multiple-choice questions by comparing them with the standard answer, and score open-ended questions by calling a large language model based on the reference answer and scoring criteria; write the question-level results and conversation-level results into a cognitive event object, and update the knowledge mastery level based on the test event.

[0014] As a further technical solution, the cognitive event feedback module is also used to: store cognitive event objects according to source type, event type, knowledge point, score, correctness, and associated question identifier; add the corresponding date range to PersonaTKGDirtyQueue to trigger asynchronous refresh of the time-series cognitive graph.

[0015] According to one aspect of the present invention, a cognitive-driven real-time intelligent teaching method for university courses is provided, applied to a server. The method includes: parsing course materials into text fragments and constructing vector indexes and keyword indexes respectively; responding to course questions submitted by students, scheduling multiple configurable nodes based on a workflow orchestration framework, including at least course relevance judgment, query rewriting, hybrid retrieval, retrieval result fusion, answer generation, knowledge point extraction, and event feedback; in the hybrid retrieval node, performing semantic retrieval based on vector indexes and keyword retrieval based on keyword indexes respectively, fusing, deduplicating, and rearranging the two types of retrieval results to obtain knowledge fragments related to the course content; inputting the knowledge fragments, course questions, and contextual messages into a large language model to generate text answers; converting the text answers into speech data through a speech synthesis model, and generating lip-sync video based on a digital human avatar driven by the speech data through a lip-sync model, returning it to the student; encapsulating the data generated during this question-and-answer process into cognitive event objects, writing them into an event database, and triggering a refresh of the temporal cognitive graph.

[0016] Compared with traditional teaching aids, the advantages of this invention are as follows: 1. Enhance the constraint of answer sources through RAG course knowledge retrieval. Before generating the large model, the system searches the course knowledge base, using FAISS semantic retrieval, BM25 keyword retrieval, and multi-round context as the basis for the answer. This makes the course questions and answers more closely aligned with the textbook, courseware, and course materials, reducing the problem of general large models generating answers that are detached from the course content.

[0017] 2. Enhance the ability to handle complex teaching tasks through agent-based workflow orchestration. The system organizes nodes such as course relevance judgment, query rewriting, mixed retrieval, answer generation, knowledge point extraction, mastery update, and event feedback based on LangGraph, enabling text-based question answering, image-based question answering, and video-based point-of-sight question answering to automatically enter the appropriate processing chain according to the input type.

[0018] 3. The integration of vector retrieval and BM25 addresses both semantic understanding and terminology recall. Vector retrieval is suitable for handling synonyms and semantically related content in students' natural language questions, while BM25 keyword retrieval is suitable for recalling algorithm names, complexity symbols, technical terms, and code concepts. The integration of the two improves the knowledge localization capabilities in course-based question answering, image-based question answering, and video-based question answering.

[0019] 4. Transform text-based answers into human-like explanations through digital human interaction. The system uses GPT-SoVITS to generate course response audio and Wav2Lip to generate lip-synced video, supporting streaming output and full audio / video rewind, enabling students to receive interactive feedback closer to teacher explanations in video learning and question-and-answer scenarios.

[0020] 5. The system records RAG question-and-answer sessions, Agent workflow node results, video question-and-answer sessions, and digital human feedback results through cognitive events. All interaction results are uniformly aggregated into CognitiveEvents and further used for mastery updates, temporal cognitive graphs, review reminders, and learning path recommendations. This ensures that RAG retrieval and digital human interaction not only serve single question-and-answer sessions but also continuously support personalized teaching. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 An overall architecture diagram of a cognitive-driven real-time intelligent teaching system for university courses is provided in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the RAG, Agent workflow, and digital human-driven student interaction closed-loop process provided in this embodiment of the invention. Figure 3 A flowchart of RAG course knowledge retrieval enhancement and agent orchestration provided for embodiments of the present invention; Figure 4 A flowchart of intelligent testing and cognitive event feedback provided in an embodiment of the present invention; Figure 5 This is a flowchart of the cognitive event-driven temporal cognitive graph construction provided in an embodiment of the present invention; Figure 6 A flowchart illustrating the digital human interaction feedback process provided in an embodiment of the present invention. Detailed Implementation

[0023] The terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this invention are intended to cover a non-exclusive inclusion, such as a process, method, system, product, or apparatus that includes a series of steps or units, not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form new technical solutions. Such combinations are not bound by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0025] This invention addresses the need for intelligent teaching assistance in university courses. It uses RAG course knowledge retrieval enhancement as the foundation for knowledge answering, LangGraph-based agent-style workflow as the task orchestration method, digital human voice and video interaction as the form of teaching feedback, and cognitive event feedback as the basis for updating learning status. This enables technical synergy between course Q&A, video questioning, digital human explanation, intelligent testing, and personalized recommendations.

[0026] This invention provides a cognitively driven real-time intelligent teaching system for university courses, deployed on a server side, such as... Figure 1 As shown, it includes a course knowledge base construction module, an agent workflow orchestration module, a hybrid retrieval module, a large model generation module, a digital human interaction module, a cognitive event feedback module, as well as a multimodal question answering module, a learning status update module, a personalized recommendation module, and a test management module.

[0027] 1. Course Knowledge Base Construction Module

[0028] The course knowledge base construction module is used to parse course materials into text fragments and build vector indexes and keyword indexes respectively.

[0029] In the course material preprocessing stage, the system parses the course PDF, teaching materials, knowledge point text, and video-related text into several text fragments, retaining metadata such as titles, page numbers, chapters, or knowledge points. Each text fragment is generated into a vector representation using `text-embedding-v4` and written into the FAISS vector library to build a vector index; simultaneously, a BM25 keyword index is built based on the word segmentation results for precise retrieval of technical terms and algorithm names.

[0030] 2. Agent Workflow Orchestration Module

[0031] like Figure 2 and Figure 3 As shown, the Agent workflow orchestration module is used to respond to course questions submitted by students. The Agent-based workflow framework built on LangGraph includes at least several configurable nodes, such as course relevance judgment, query rewriting, hybrid retrieval, retrieval result fusion, answer generation, knowledge point extraction, and knowledge tracking.

[0032] In the workflow node scheduling phase, the system constructs an Agent-based RAG workflow based on LangGraph, organizing steps such as course relevance assessment, query rewriting, hybrid retrieval, search result fusion, answer generation, knowledge point extraction, and knowledge tracking into configurable nodes. This workflow selects the execution path based on the question type and function switches, enabling the reuse of a unified orchestration framework for ordinary text questions, image enhancement questions, and video positioning questions.

[0033] During the student question comprehension phase, the system addresses POST / ask, POST / ask_stream, POST / ask_with_image, and POST / api / videos / <video_id> Questions submitted via APIs such as / ask are processed uniformly. The system reads the most recent rounds of question-and-answer sessions in the current session or video, forms a contextual message sequence, and determines course relevance. For image-based question-and-answer, the system first calls qwen-vl-max to perform type identification and content extraction on the image, and then merges the image analysis results with the student's question to form an enhanced question.

[0034] 3. Hybrid Search Module

[0035] The hybrid retrieval module is used in the hybrid retrieval node to perform semantic retrieval based on vector index and keyword retrieval based on keyword index, respectively, and to merge, deduplicate and rearrange the two types of retrieval results to obtain knowledge fragments related to the course content.

[0036] In the hybrid retrieval and fusion stage, the system performs FAISS vector retrieval and BM25 keyword retrieval on student questions and augmented questions, respectively. FAISS retrieval is used to recall course texts with similar semantics, while BM25 retrieval is used to recall texts containing algorithmic terms, data structure names, complexity expressions, or key code concepts. The system then merges, deduplicates, and performs necessary rearrangements on the two types of retrieval results to obtain the final course context input to the large model.

[0037] 4. Large Model Generation Module

[0038] The large model generation module is used to input the knowledge fragments, course questions, and contextual messages into the large language model to generate text answers.

[0039] In the large-scale model answer generation stage, the system inputs retrieved course segments, student questions, multi-turn context, and necessary user memories into qwen-plus or qwen-max to generate course question-and-answer results. The generated results can be returned all at once or via a streaming interface according to preset chunk lengths. The system writes the question, answer, course relevance, response latency, and session information into a Question object and records RAG question-and-answer cognitive events.

[0040] 5. Multimodal Question Answering Module

[0041] The multimodal question answering module supports image-based question answering and video-based point-of-sight question answering. This module supports different learning entry points based on a unified RAG pipeline.

[0042] For plain text question-and-answer sessions, students submit questions via POST / ask, and the system generates a complete answer after RAG retrieval enhancement. For real-time streaming question-and-answer sessions, students submit questions via POST / ask_stream, and the system generates a complete answer and outputs it in segments, allowing students to view the answer content step by step.

[0043] For image-based question-and-answer sessions, students submit images and optional questions via POST / ask_with_image. The system uses qwen-vl-max to identify whether the image is an algorithm flowchart, pseudocode, data structure diagram, complexity analysis diagram, or code screenshot, and extracts the algorithm name, key concepts, code content, problem points, and learning suggestions. If the student also submits a question, the system combines the image analysis results with the question and sends it to the RAG (Research, Analysis, and Question) process; if the student does not submit a question, the system directly returns the image analysis results. This process is also recorded in the Question object and cognitive event.

[0044] For video-based targeted Q&A, students can use the POST request to access / api / videos / during the course video viewing process.<video_id\> The ` / ask` command submits the video identifier, playback time, and question text. The system reads the most recent rounds of Q&A for that video as context, generates an answer via the RAG (Rapid Access Query) framework, and saves information such as the video identifier, video time, question text, answer text, and audio or video path to a `VideoQARecord` object. This design enables the system to retrieve and explain course knowledge based on the current content being explained in the video.

[0045] 6. Digital Human Interaction Module

[0046] like Figure 6 As shown, the digital human interaction module is used to convert the text answers generated by RAG into audio and video explanations. Its core includes TTS configuration parsing, speech synthesis, lip-sync driven video generation, streaming output, and rollback processing.

[0047] During the TTS configuration parsing phase, the system selects available GPT-SoVITS models, reference audio, and reference text based on the digital human voice materials and TTS model configuration configured by the administrator. If a default configuration exists, the system will use it first; if a voice material is specified, the TTS configuration bound to that material can be used.

[0048] In the speech synthesis stage, the system inputs the response text generated by RAG into GPT-SoVITS to generate a lecture audio file. This audio file can be used as independent audio feedback or as input for subsequent lip-sync driven video generation.

[0049] In the lip-syncing stage, the system inputs digital human head images and synthesized speech into Wav2Lip to generate digital lip-syncing video synchronized with the response speech. If real-time digital human streaming output is enabled, the system prioritizes the streaming generation capability and returns information such as stream_url, stream_fps, and output mode; if streaming generation fails, the service is not ready, or the current configuration does not support streaming output, the system retains the generated audio and reverts to the complete video generation process.

[0050] In the scenario of compositing instructional video layers, the system can further combine Celery, Redis, MuseTalk, ffmpeg, and optional RobustVideoMatting to perform asynchronous processing, integrating the digital human layer with course video resources. Both successful and unsuccessful digital human generation results can be recorded as cognitive events from the digital_human source for subsequent learning state analysis.

[0051] 7. Cognitive event feedback module, test control module, and learning status update module

[0052] like Figure 4 and Figure 5 As shown, the cognitive event feedback module is used to encapsulate the data generated during this question-and-answer process into cognitive event objects, write them into the event database, and trigger the refresh of the temporal cognitive graph.

[0053] The test management module is used to respond to test submission requests from students. For multiple-choice questions, it uses standard answers to compare and score them. For open-ended questions, it calls a large language model to score them based on the reference answers and scoring criteria. It writes the question-level results and conversation-level results into a cognitive event object and updates the knowledge mastery level based on the test events.

[0054] The learning status update module is used to update the student's knowledge mastery based on the cognitive event object and generate a snapshot of the temporal cognitive graph.

[0055] The system uniformly writes text-based RAG Q&A, streaming RAG Q&A, image-based RAG Q&A, video-based RAG Q&A, digital human generation, question scoring, and quiz submission into a `CognitiveEvent` object. Each event includes information such as source type, event type, knowledge point, score, correctness, associated question identifier, associated video Q&A identifier, or associated quiz identifier. After an event is written, the system adds the corresponding date range to a `PersonaTKGDirtyQueue` to trigger a temporal cognitive graph refresh.

[0056] For test results, the system creates a test session via POST / api / quiz / attempts, and sends the test results via POST / api / quiz / attempts / <attempt_id> The ` / submit` command submits answers and performs scoring. Multiple-choice questions are compared against a standard answer; open-ended questions are scored using a large model based on the reference answer and scoring criteria. After scoring, the system updates `QuizAttempt` and `QuizAttemptItem`, and writes the question-level and session-level results to the cognitive event.

[0057] The system updates KnowledgeMastery objects based on test events. Mastery updates employ a simplified Bayesian knowledge tracing model, combining initial mastery probability, learning transfer probability, guessing probability, error probability, and forgetting decay factor to dynamically adjust students' mastery of each knowledge point. Subsequently, the system can aggregate cognitive events within a specified time window to generate a PersonaTKGDailySnapshot, which includes the graph structure, learning profile summary, and feature data.

[0058] 8. Personalized Recommendation Module

[0059] The personalized recommendation module generates review plans based on the risk of forgetting and recommends learning paths based on knowledge mastery and prior knowledge relationships. This module utilizes cognitive events and mastery results to provide feedback for subsequent teaching.

[0060] Students can view their knowledge mastery, weaknesses, and strengths via GET / analyze / knowledge_mastery; view a review plan generated based on forgetting patterns via GET / analyze / review_schedule; view a snapshot of their cognitive map for a given day via GET / analyze / persona_tkg_snapshot; and obtain learning path recommendations that combine prior knowledge relationships via GET / analyze / learning_path.

[0061] In the learning path recommendation, the system prioritizes recommending knowledge points whose prerequisites are met but whose mastery is insufficient, based on the prerequisite relationships of algorithm course knowledge points and the student's current level of mastery. It also provides reasons for the recommendations, their difficulty, and suggested resources. This recommendation is not generated in isolation but is based on continuous learning data derived from RAG question answering, video RAG question answering, digital human interaction, and feedback from quiz scoring.

[0062] The administrator can manage students, question banks, teaching videos, digital human materials, and TTS configurations. The question bank object QuizBankItem supports fields such as knowledge category, knowledge point, question type, difficulty, reference answer, and scoring criteria; the quiz objects QuizAttempt and QuizAttemptItem are used to save quiz sessions and question snapshots; StudentKnowledgePointCache is used to cache student knowledge point statistics results, improving the response efficiency of the statistical analysis interface.

[0063] In summary, this invention enhances the basis for course answers by strengthening knowledge retrieval in RAG courses, organizes complex question-and-answer tasks through agent-style workflow orchestration, improves teaching interaction through digital human voice and lip-reading, and connects question-and-answer, video explanations, quizzes, and recommendations into a continuous learning data link through cognitive event feedback. This results in an intelligent auxiliary teaching method that combines course knowledge constraints, intelligent process orchestration, human-like interaction, and personalized feedback capabilities.

[0064] This invention also provides a cognitively driven real-time intelligent teaching method for university courses. Addressing the intelligent auxiliary teaching needs of university courses, it uses RAG course knowledge retrieval enhancement as the foundation for knowledge answering, a LangGraph-based agent-style workflow as the task orchestration method, digital human voice and video interaction as the form of teaching feedback, and cognitive event feedback as the basis for updating learning status. This achieves technical synergy between course question-and-answer, video questioning, digital human explanation, intelligent quizzes, and personalized recommendations. The specific steps are as follows.

[0065] Step S1: Building the Course Knowledge Base

[0066] The course PDF, teaching materials, knowledge point text, and video-related text are parsed into several text fragments, retaining metadata such as titles, page numbers, chapters, or knowledge points. Each text fragment is used to generate a vector representation using text-embedding-v4 and a FAISS vector index is constructed. Simultaneously, a BM25 keyword index is built based on the word segmentation results.

[0067] Step S2: Agent workflow node scheduling and hybrid retrieval

[0068] In response to course questions submitted by students, the agent-based RAG workflow built on LangGraph schedules configurable nodes for course relevance judgment, query rewriting, hybrid retrieval, retrieval result fusion, answer generation, knowledge point extraction, and knowledge tracking, and selects the execution path according to the question type and function switches.

[0069] In the hybrid retrieval node, FAISS semantic retrieval and BM25 keyword retrieval are performed respectively. The two types of retrieval results are then merged, deduplicated, and rearranged to obtain knowledge fragments related to the course content.

[0070] Step S3: Generate the answer for the large model

[0071] Input knowledge fragments, course questions, and multi-turn context into qwen-plus or qwen-max to generate text answers; these answers can be returned all at once or in a stream of fragments. The system writes the questions, answers, course relevance, response latency, and session information into a Question object and records RAG question-and-answer cognitive events.

[0072] Step S4: Multimodal Question Answering

[0073] For image-based question answering, qwen-vl-max is called to identify the image type (algorithm flowchart, pseudocode, data structure diagram, complexity analysis diagram, or code screenshot), extract the algorithm name, key concepts, code content, problem points, and learning suggestions, and then concatenate the image analysis results with the student's question to form an enhanced question before entering the RAG process; if the student does not ask a question, the image analysis results are returned directly.

[0074] For video-based point-to-point question and answer, the video identifier, playback time point, and question text are obtained. The most recent rounds of question and answer in the same video are read as context. The answer is generated through the RAG link, and information such as video identifier, video time point, question text, answer text, audio path or video path is saved to the VideoQARecord object.

[0075] Step S5: Digital Human Interaction Feedback

[0076] Based on the administrator's configuration, select the available GPT-SoVITS model, reference audio, and reference text. Input the response text into GPT-SoVITS to generate course explanation audio. Then, input the digital human head image material and synthesized audio into Wav2Lip to generate lip-synced video. If real-time digital human streaming output is enabled, prioritize returning the streaming video address and frame rate information; if streaming generation fails or conditions are not met, revert to generating a complete audio and video file.

[0077] Step S6: Cognitive Event Regression and Learning Status Update

[0078] Text-based RAG Q&A, streaming RAG Q&A, image-based RAG Q&A, video-based RAG Q&A, digital human generation, question scoring, and quiz submission are all written into a CognitiveEvent object. Each event includes fields such as source type, event type, knowledge point, score, correctness, and associated question identifier. The corresponding date range is added to PersonaTKGDirtyQueue to trigger a time-series cognitive graph refresh.

[0079] A simplified Bayesian knowledge tracing model is employed, combining initial mastery probability, learning transfer probability, guessing probability, error probability, and forgetting decay factor to dynamically adjust students' mastery of each knowledge point and update the KnowledgeMastery object. Cognitive events within a specified time window are aggregated to generate a PersonaTKGDailySnapshot, which includes the graph structure, learning profile summary, and feature data.

[0080] Step S7: Test Control

[0081] In response to the test submission request, multiple-choice questions are scored by comparing with the standard answer, while open-ended questions are scored by calling the large language model based on the reference answer and scoring criteria. The question-level results and conversation-level results are written into the cognitive event object, and the knowledge mastery is updated based on the test event.

[0082] Step S8: Personalized Recommendations

[0083] A review plan is generated based on the risk of forgetting; a learning path recommendation is generated based on the level of knowledge mastery and the prerequisite relationships of knowledge points, prioritizing knowledge points whose prerequisite conditions have been met but whose mastery is insufficient, and including the reasons for the recommendation, the difficulty level, and suggested resources.

[0084] In summary, the intelligent auxiliary teaching method for university courses proposed in this invention is deployed in a teaching platform composed of a browser and a server. The server includes a Flask main service, a course knowledge base retrieval module, a LangGraph-based agent-based RAG workflow orchestration module, a large model generation module, a multimodal question answering module, a digital human interaction module, and a database and cognitive event analysis module. This method does not simply stack functions according to the teaching process, but rather focuses on four core technologies: enhanced RAG course knowledge retrieval, agent-based task orchestration, digital human interactive feedback, and cognitive event feedback. This allows students to receive teaching feedback constrained by the course knowledge base, presented in a human-like manner, and usable for subsequent learning analysis in text-based question answering, image-based question answering, video-based targeted question answering, and testing scenarios.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A cognitively driven real-time intelligent teaching system for university courses, characterized in that, Deployed on the server side, the system includes: The course knowledge base construction module is used to parse course materials into text fragments and build vector indexes and keyword indexes respectively; The Agent workflow orchestration module is used to respond to course questions submitted by students. Based on the workflow orchestration framework, it schedules multiple configurable nodes, including at least course relevance judgment, query rewriting, hybrid retrieval, retrieval result fusion, answer generation, knowledge point extraction, and event feedback. The hybrid retrieval module is used to perform semantic retrieval based on vector index and keyword retrieval based on keyword index in the hybrid retrieval node, and to merge, deduplicate and rearrange the two types of retrieval results to obtain knowledge fragments related to the course content. The large model generation module is used to input the knowledge fragments, course questions, and contextual messages into the large language model to generate text answers. The digital human interaction module is used to convert the text answer into speech data through a speech synthesis model, and to drive the digital human avatar to generate a lip-sync video based on the speech data through a lip-sync model, and then return the lip-sync video to the student's end. The cognitive event feedback module is used to encapsulate the data generated during this question-and-answer process into cognitive event objects, write them into the event database, and trigger the refresh of the temporal cognitive graph.

2. The cognitive-driven real-time intelligent teaching system for university courses according to claim 1, characterized in that, The course knowledge base construction module is also used to: generate vector representations of the text fragments using text-embedding-v4 and construct a FAISS vector index, while constructing a BM25 keyword index.

3. The real-time intelligent teaching system for university courses based on cognitive drive according to claim 1, characterized in that, The Agent workflow orchestration module is built on LangGraph; the Agent workflow orchestration module is also used to select the execution path according to the modal type or function switch of the course question, so that text questions, image enhancement questions and video positioning questions can reuse a unified orchestration framework.

4. The real-time intelligent teaching system for university courses based on cognitive drive according to claim 1, characterized in that, The digital human interaction module is also used to: call the GPT-SoVITS model to synthesize the text answer into speech data, call the Wav2Lip model to drive the digital human avatar to generate lip-sync video based on the speech data; and, when the streaming output conditions are met, return the streaming video address and frame rate information, and when streaming generation fails or the conditions are not met, roll back to generate a complete audio and video file.

5. The real-time intelligent teaching system for university courses based on cognitive drive according to claim 1, characterized in that, It also includes a multimodal question answering module, which is used to: respond to image question answering requests submitted by students, call qwen-vl-max to identify the image type and extract the content, and then combine the image analysis results with the student's question to form an enhanced question before entering the Agent workflow orchestration module; In response to a video-based question-and-answer request submitted by a student during the course video playback, the system obtains the video identifier, playback time point, and question text, reads the most recent multiple rounds of questions and answers for the same video as context, enters the Agent workflow orchestration module, and saves the question-and-answer related information of the video-based question-and-answer request to a VideoQARecord object.

6. The real-time intelligent teaching system for university courses based on cognitive drive according to claim 1, characterized in that, It also includes a learning status update module, which is used to: based on the cognitive event object, adopt a simplified Bayesian knowledge tracing model, and combine the initial mastery probability, learning transfer probability, guessing probability, error probability and forgetting decay factor to dynamically correct the student's mastery of each knowledge point and update the KnowledgeMastery object; aggregate cognitive events within a specified time window to generate PersonaTKGDailySnapshot, which includes a graph structure, learning profile summary and feature data.

7. The real-time intelligent teaching system for university courses based on cognitive drive according to claim 6, characterized in that, It also includes a personalized recommendation module, which is used to generate a review plan based on the risk of forgetting. Learning path recommendations are generated based on knowledge mastery and prerequisite relationships of knowledge points; wherein, the learning path recommendations prioritize knowledge points whose prerequisite conditions have been met but whose mastery is insufficient, and include the reasons for the recommendations, difficulty, and suggested resources.

8. The real-time intelligent teaching system for university courses based on cognitive drive according to claim 1, characterized in that, It also includes a test management module, which is used to: respond to test submission requests from students, score multiple-choice questions by comparing them with standard answers, score open-ended questions by calling a large language model based on reference answers and scoring criteria; write question-level results and conversation-level results into a cognitive event object, and update knowledge mastery based on test events.

9. The real-time intelligent teaching system for university courses based on cognitive drive according to claim 1, characterized in that, The cognitive event feedback module is also used to: store cognitive event objects according to source type, event type, knowledge point, score, correctness, and associated question identifier; add the corresponding date range to PersonaTKGDirtyQueue to trigger asynchronous refresh of the time-series cognitive graph.

10. A cognitively driven real-time intelligent teaching method for university courses, characterized in that, Applied to the server side, the method includes: The course materials are parsed into text fragments, and vector indexes and keyword indexes are constructed respectively; In response to course-related questions submitted by students, the workflow orchestration framework schedules at least several configurable nodes, including course relevance assessment, query rewriting, hybrid retrieval, retrieval result fusion, answer generation, knowledge point extraction, and event reflow. In the hybrid retrieval node, semantic retrieval based on vector index and keyword retrieval based on keyword index are performed respectively. The two types of retrieval results are then merged, deduplicated, and rearranged to obtain knowledge fragments related to the course content. The knowledge fragments, course questions, and contextual messages are input into a large language model to generate text answers. The text answer is converted into speech data using a speech synthesis model, and a lip-sync model is used to generate a lip-synced video based on the speech data using a digital human avatar, which is then returned to the student's device. The data generated during this Q&A process is encapsulated into cognitive event objects, written into the event database, and the temporal cognitive graph is refreshed.