Lightweight digital human lesson preparation system based on intelligent agent

The lightweight digital human lesson preparation system built through intelligent body technology solves the problems of low efficiency of traditional teaching lesson preparation and difficult resource integration, and realizes efficient and personalized teaching content generation and intelligent teaching plan assistance, improving teaching quality and flexibility.

CN120543330APending Publication Date: 2025-08-26TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 19 Cited by

Patent Information

Application Number
CN202510435093.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Traditional teaching lesson preparation relies on manual integration of multiple modal resources, with large workload and difficult to fully explore related information between resources. Digital life generation requires high hardware, and insufficient teaching flexibility and personalized expression.

Method used

The lightweight digital human lesson preparation system based on the agent is adopted, and through the knowledge graph construction and reasoning module, the multi-modal cognitive intelligent agent module, the lightweight digital human generation module and the intelligent teaching plan auxiliary module, the full process intelligent support from courseware analysis to classroom explanation is realized.

Benefits of technology

Significantly improve lesson preparation efficiency, generate personalized and vivid explanation content, support diversified teaching scenarios, provide intelligent teaching plans and resource recommendations, reduce repetitive labor, and improve teaching quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543330A_ABST
    Figure CN120543330A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight digital human lesson preparation system based on an intelligent agent, and belongs to the field of intelligent teaching. Through collaborative operation of four core modules of knowledge graph construction and reasoning, multi-modal cognitive agent, lightweight digital human generation and intelligent teaching plan assistance, the problems of low efficiency of resource integration, teaching content homogenization, insufficient digital human interaction experience and the like in traditional lesson preparation are solved. The knowledge graph construction and reasoning module is used for constructing a structured knowledge graph and realizing knowledge point association mining and teaching logic reasoning; the multi-modal cognitive agent module is used for generating personalized explanation content according with a teaching target by fusing multi-modal courseware analysis, semantic understanding and lecture style dynamic adaptation functions; the lightweight digital human generation module is combined with model pruning and emotion modeling technologies to synchronously output natural voice and a high-simulation digital human image; the intelligent teaching plan auxiliary module helps the teacher to intelligently generate a teaching plan and a teaching outline according to the courseware content based on the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent teaching technology, and in particular to a lightweight digital human lesson preparation system based on an intelligent agent. Background Art

[0002] Traditional lesson preparation relies primarily on teachers manually integrating multiple modal resources, including documents, images, videos, exercises, and syllabi. This approach is not only labor-intensive, but also hinders the formation of a systematic and logically coherent teaching content due to the difficulty in fully exploring the connections between these resources. Furthermore, while traditional digital human generation technology can simulate human figures and movements, it typically requires high hardware resources, making it difficult to run smoothly on standard devices. Furthermore, it lacks adaptability to diverse teaching scenarios, limiting the flexibility and personalized expression of teaching.

[0003] In recent years, multimodal fusion agent technology has made significant progress. Traditional intelligent systems typically rely on a single modality (such as text or images) for information processing. Multimodal technology, by simultaneously integrating multiple data sources such as text, images, video, and audio, enables richer semantic expression and information interaction. With the rapid development of large language models, multimodal fusion, graph neural networks, deep reinforcement learning, and speech synthesis and digital human generation technologies, the education sector is experiencing a technological revolution. These technologies not only demonstrate significant advantages in automatic summarization, text generation, and style transfer, but also enable deep semantic analysis of multimodal teaching resources, thereby constructing structured knowledge networks and enabling logical reasoning and dynamic evolution between knowledge points. This makes it possible to build a comprehensive intelligent lesson preparation system, from courseware analysis and knowledge graph construction to personalized lecture script generation, natural speech output, digital human image presentation, and intelligent teaching plan development. Summary of the Invention

[0004] In order to solve the problems of cumbersome lesson preparation for teachers, difficulty in changing teaching styles, inflexible classroom presentation formats, and high resource requirements for digital human generation, this application proposes a lightweight digital human lesson preparation system based on intelligent agents. It uses artificial intelligence technology to improve teachers' lesson preparation efficiency and teaching quality, and realizes intelligent support for the entire process from courseware analysis to classroom explanation.

[0005] The technical solution adopted in this application is: a lightweight digital human lesson preparation system based on an intelligent agent, including: Knowledge graph construction and reasoning module: used to conduct in-depth semantic analysis of multimodal teaching resources, build structured knowledge graphs, and realize knowledge point association mining and teaching logic reasoning; Multimodal cognitive agent module: By integrating multimodal courseware analysis, semantic understanding, and dynamic adaptation of lecture styles, it generates personalized teaching content that meets teaching objectives, including: Intelligent parsing unit: used to deploy a multimodal hierarchical parsing framework; Style decision unit: used to construct a teaching style decision tree and implement dynamic style selection using a deep reinforcement learning framework; Content generation unit: adopts a multi-agent collaborative architecture and a distributed consensus algorithm to achieve collaborative decision-making under multi-dimensional constraints; Adaptive optimization unit: used to build a closed-loop system for teaching effect feedback; Lightweight Digital Human Generation Module: This module converts lecture notes into natural and fluent speech output, supporting different intonations and emotional expressions to suit different teaching scenarios. It generates a synchronized digital human image based on the speech output, including facial expressions, lip sync, and movements, to achieve digital human presentation. Intelligent teaching plan assistance module: Based on the knowledge graph, it helps teachers intelligently generate teaching plans and syllabuses based on the courseware content, and provides targeted teaching suggestions to optimize the teaching structure.

[0006] Furthermore, the knowledge graph construction and reasoning module includes: Cross-modal semantic fusion and feature alignment unit: This unit uses a vision-language pre-training model and a cross-modal attention mechanism to jointly encode semantic entities in text, images, and videos, and establish cross-modal feature mapping relationships. It also uses a contrastive learning algorithm to align the semantic spaces of different modalities, extract multi-dimensional feature vectors of conceptual entities in teaching resources, and construct a triple knowledge representation system of entity-relationship-attribute. Dynamic Knowledge Graph Evolution Engine: This engine builds a scalable knowledge topology network based on graph neural networks. It uses graph convolutional networks and graph attention networks to dynamically model the association strength and weight distribution between knowledge points, and uses time-series graph embedding technology to capture the dynamic evolution of the knowledge system. When new teaching resources are added, the system automatically triggers the optimization process of the graph topology structure, adaptively adjusts the association weights of knowledge points, and verifies the integrity of the teaching logic chain, ensuring that the knowledge system is always updated in sync with teaching needs. Visual Interaction and Knowledge Reconstruction Unit: This unit provides a scalable knowledge network visualization interface, enabling teachers to interactively edit knowledge graphs. It also features a built-in conflict detection algorithm that automatically identifies logical breaks, hierarchical contradictions, or redundant relationships in real time when teachers manually adjust knowledge point associations, and generates repair suggestions based on a rule engine. Teaching Path Reasoning and Recommendation Unit: Combining probabilistic graphical models with path planning algorithms, the unit searches for the optimal teaching sequence in the knowledge graph based on the student cognitive level profile and curriculum standard requirements, and outputs personalized path plans that include core knowledge point coverage, teaching time allocation, and advanced route branches, providing a scientific basis for subsequent explanation content generation and teaching plan formulation.

[0007] Furthermore, the intelligent parsing unit includes: Structural analysis layer: Using a deep learning-based layout analysis model, through the joint encoding of visual elements and text features, it identifies the title hierarchy, paragraph segmentation, and mixed text and image structure in the courseware, and dynamically generates the document's logical skeleton; Semantic parsing layer: Through the joint embedding space alignment technology of the multimodal large model, cross-modal semantic fusion of text, formulas, charts and video key frames is performed to extract semantic units and generate a list of explanation elements with time sequence tags; Teaching logic layer: Build a teaching framework based on a dynamic knowledge graph, automatically mark the association between knowledge points and the distribution of teaching focus through graph structure reasoning, and support dynamic planning of teaching paths.

[0008] Furthermore, the state space in the style decision unit contains multi-dimensional feature encoding that integrates the complexity of teaching content, student cognitive level assessment values, and historical teaching effect data; the action space covers a variety of basic styles and combination modes; the reward function adopts a multi-objective reinforcement learning framework, and improves the model adaptability by combining a composite reward function that integrates classroom interaction rate prediction, knowledge point mastery estimation, and style consistency scoring, combined with an offline strategy optimization algorithm.

[0009] Furthermore, the content generation unit includes: Knowledge Distillation Agent: Based on the logical reasoning chain of the knowledge graph and the implicit knowledge verification of the pre-trained language model, it realizes dual-channel review of the consistency of concept representation. It also builds a three-dimensional knowledge verification system, specifically the logical reasoning chain verification of the knowledge graph, the implicit association verification of the pre-trained language model, and the boundary verification of domain expert feedback. It also introduces a dynamic concept alignment mechanism to achieve cross-modal knowledge representation consistency through comparative learning. It also develops a multi-source heterogeneous knowledge fusion module to support the knowledge distillation and reorganization of text / formulas / charts. Style rendering agent: inserts stylized elements based on the selected style template through context-aware text generation technology; constructs a style feature transfer matrix containing a multi-dimensional style quantification indicator system; builds a case interpolation algorithm based on attention gating and reinforcement learning-driven question strategy selector that combines context-aware generation technology; develops a multimodal style transfer engine that supports real-time switching between multiple teaching styles; Teaching compliance intelligent body: Build a rule constraint library based on curriculum standards to detect out-of-syllabus and knowledge point coverage of generated content; establish a four-layer constraint system: curriculum standards, cognitive development model, regional teaching outline, and ethical review standards; develop an intelligent calibration system that can realize knowledge point coverage analysis based on concept topology map + out-of-syllabus detection based on semantic dependency tree; build a dynamic adjustment mechanism to update the regional teaching policy knowledge base in real time through online learning.

[0010] Furthermore, the adaptive optimization unit collects classroom interaction data in real time through an online learning mechanism, and dynamically adjusts the weight parameters of the style decision model through real-time teaching data stream analysis to adapt to changes in teaching scenarios; through incremental fine-tuning strategies, the adapter components of the multimodal large model are updated based on newly added teaching cases to achieve low-cost knowledge transfer; a generator-discriminator adversarial framework is constructed, and the teachability of the generated content is evaluated through the teaching applicability discriminant network to drive model optimization.

[0011] Furthermore, the coordinated control architecture of the lightweight digital human generation module adopts a lightweight design, and several digital human images are pre-installed in the module. It also supports users to upload photos to generate digital human images that are highly matched with them; the module uses pruning, quantization and model compression technologies to run efficiently on devices with limited resources.

[0012] Furthermore, the lightweight digital human generation module dynamically adjusts the speech speed, tone and emotional expression based on lightweight speech synthesis technology, so that the speech is both in line with the teaching content and natural and realistic; and the module is integrated with optimized speech generation algorithms and intelligent scheduling strategies to minimize the delay of speech generation while ensuring output clarity and coherence; the module's miniaturized multilingual model is both efficient and accurate when dealing with complex cross-language teaching environments; the module uses a lightweight action generation algorithm to achieve real-time synchronization of speech and the lip shape, expression and action of the digital human's explanation, ensuring the smoothness and naturalness of the output process.

[0013] Furthermore, the intelligent teaching plan assistance module first extracts core knowledge points from teaching resources through multimodal semantic understanding technology, and constructs a knowledge point topological network that includes logical relationships, teaching dependencies, difficulty levels and importance weights; on this basis, combined with the preset teaching objectives and students' cognitive level characteristics, a constraint satisfaction algorithm is used to generate a teaching plan draft that includes the order of knowledge point explanation, classroom activity design, time allocation and evaluation plan. Based on the teaching plan draft, it automatically matches and recommends relevant teaching resources. By analyzing the knowledge density and logical structure in the lecture content and combining historical teaching data, it predicts the possible understanding obstacles of students, provides personalized recommendations and targeted teaching suggestions based on teachers' usage preferences, and supports rapid preview and import of resources.

[0014] Furthermore, the intelligent teaching plan assistance module also has a built-in interdisciplinary knowledge mapping engine, which can automatically identify cross-domain connection points such as the derivation relationship between mathematical formulas and physical experiments, and the temporal connection between historical events and literary works, supporting teachers to quickly build interdisciplinary integrated courses; and the module also has a built-in resource recommendation subsystem, which intelligently recommends auxiliary materials from local libraries or cloud-based educational resource platforms based on collaborative filtering and content matching algorithms, and provides resource preview, one-click import and copyright compliance check functions.

[0015] The beneficial effects of this application compared to the prior art are: 1. Efficient lesson preparation and resource integration: Utilizing multimodal courseware analysis and knowledge graph construction technology, the system can automatically extract key information from various teaching resources and build a structured knowledge network, effectively integrating text, images, videos and other information, greatly reducing the workload of teachers in manually organizing materials and significantly improving lesson preparation efficiency.

[0016] 2. Personalized teaching and dynamic adaptation: The system uses a multimodal cognitive agent module, combined with students' cognitive level, teaching objectives and historical feedback, to dynamically generate explanation texts with diverse styles and precise content to meet the personalized needs of different classroom scenarios and student groups, and achieve flexible adjustment of teaching content.

[0017] 3. Vivid and lifelike digital human explanations: Through a lightweight digital human generation module, natural and emotionally rich speech is generated. Adjustable intonation, speaking speed, and emotion are supported, ensuring that the speech closely matches the lecture style, enhancing the listening experience and increasing classroom interactivity. The system generates a synchronized digital human image based on the speech content. Combining lip movements, facial expressions, and movements makes the explanation process more vivid and visually appealing, increasing student participation and understanding.

[0018] 4. Intelligent teaching plan assistance and decision support: Based on dynamic knowledge graphs and teaching path reasoning, the system can automatically generate scientific and reasonable teaching plans and outlines, and combine interdisciplinary mapping, resource recommendations and teaching feedback to provide teachers with optimization suggestions and decision support to ensure the rigor of the course structure and the efficient achievement of teaching objectives.

[0019] 5. Full-process automation and intelligent closed loop: This application integrates full-process automation technology from courseware analysis, lecture generation to digital human presentation. It not only reduces repetitive work, but also continuously improves the quality of teaching content through real-time data feedback and adaptive optimization mechanisms, freeing up more time for teachers to focus on teaching innovation and personalized guidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present application will be further described below with reference to the accompanying drawings: Figure 1 A system block diagram provided for an embodiment of the present application; Figure 2 A flowchart of the system provided in the embodiment of the present application; Figure 3 This is a structural block diagram of the multimodal cognitive agent module provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] like Figures 1 to 3As shown, the present application provides a lightweight digital human lesson preparation system based on intelligent agents. The system integrates a knowledge graph construction and reasoning module, a multimodal cognitive intelligent agent module, a lightweight digital human generation module, and an intelligent teaching plan auxiliary module to realize the full process of intelligent operation from courseware to classroom explanation. The system uses a visual-language pre-training model and a graph neural network to perform cross-modal semantic fusion of teaching resources such as text, images, and videos, and constructs a dynamic knowledge graph to realize knowledge point association mining and teaching logic reasoning; relying on a hierarchical parsing framework and deep reinforcement learning technology to generate personalized lecture notes in 9 adjustable styles such as academic and interactive, and ensure content accuracy and teaching adaptability through a multi-agent collaborative mechanism; combining model pruning and emotion modeling technology to synchronously output natural speech and highly realistic digital human images, supporting speech speed, intonation, dynamic emotional adjustment, and precise matching of lip shape, expression, and action, and adapting to smooth operation of low-power devices; based on the knowledge point topology network and path planning algorithm, it automatically generates a teaching plan with class time allocation and cross-disciplinary association, providing difficulty prediction and intelligent resource recommendation. The system significantly improves lesson preparation efficiency, increases classroom interaction rates through emotional digital human interaction, and supports interdisciplinary course design and adaptive difficulty adjustment, providing an efficient and lightweight one-stop solution for the digital transformation of education.

[0022] Among them, the knowledge graph construction and reasoning module is used to conduct in-depth semantic analysis of multimodal teaching resources, build a structured knowledge network, and realize knowledge point association mining and teaching logical reasoning; the multimodal cognitive intelligent agent module generates personalized explanation content that meets the teaching objectives by integrating multimodal courseware analysis, semantic understanding and dynamic adaptation of lecture style; the lightweight digital human generation module is used to convert lectures into natural and fluent voice output, and supports different tones and emotional expressions to adapt to different teaching scenarios. According to the voice output, a synchronized digital human image is generated, including facial expressions, lip synchronization and movements, to realize the presentation of the digital human; the intelligent teaching plan auxiliary module is based on the knowledge graph to help teachers intelligently generate teaching plans and teaching outlines according to the courseware content, and provide targeted teaching suggestions to optimize the teaching structure.

[0023] The knowledge graph construction and reasoning module uses multimodal semantic fusion and dynamic graph evolution technology to achieve deep structural processing of teaching resources and teaching logic reasoning, including: 1) Cross-modal semantic fusion and feature alignment unit: Using a vision-language pre-training model and a cross-modal attention mechanism, semantic entities in text, images, and videos are jointly encoded, establishing cross-modal feature mappings. A contrastive learning algorithm aligns the semantic spaces of different modalities, extracting multidimensional feature vectors of conceptual entities (such as formulas, diagram elements, and experimental procedures) in teaching resources. This constructs a knowledge representation system of entity-relationship-attribute triples, effectively addressing the semantic fragmentation and lack of relevance of multimodal resources in traditional lesson preparation.

[0024] 2) Dynamic knowledge graph evolution engine: Based on graph neural networks (GNNs), the system builds a scalable knowledge topology network. It employs graph convolutional networks (GCNs) and graph attention networks (GATs) to dynamically model the strength and weight distribution of associations between knowledge points. Furthermore, it uses temporal graph embedding technology to capture the dynamic evolution of the knowledge system. When new teaching resources are added, the system automatically triggers an optimization process for the graph topology, adaptively adjusting the weights associated with knowledge points while verifying the integrity of the teaching logic chain, ensuring that the knowledge system is always updated in sync with teaching needs.

[0025] 3) Visualization Interaction and Knowledge Reconstruction Unit: This platform provides a scalable knowledge network visualization interface, allowing teachers to interactively edit knowledge graphs through dragging and dropping nodes, boundary selection, and semantic search. A built-in conflict detection algorithm automatically identifies logical breaks, hierarchical contradictions, or redundant relationships when teachers manually adjust knowledge point relationships, and generates repair suggestions based on a rule engine. The edited knowledge graph can be exported as an OWL ontology file or JSON-LD structured data, facilitating integration with other educational systems.

[0026] 4) Teaching Path Reasoning and Recommendation Unit: Combining the probabilistic graphical model and path planning algorithm, the optimal teaching sequence is searched in the knowledge graph according to the student's cognitive level portrait and curriculum standard requirements, and a personalized path plan is output that includes the coverage of core knowledge points, teaching time allocation, and advanced route branches, providing a scientific basis for subsequent explanation content generation and teaching plan formulation.

[0027] The multimodal cognitive agent module includes: 1) Intelligent parsing unit: used to deploy a multimodal hierarchical parsing framework, including: a. Structural analysis layer: This layer uses a deep learning-based layout analysis model to identify the title hierarchy, paragraph segmentation, and mixed text and image structure in the courseware through the combined encoding of visual elements and text features, and dynamically generates the document's logical skeleton. b. Semantic parsing layer: This layer uses the joint embedding space alignment technology of a multimodal large model to perform cross-modal semantic fusion of text, formulas, charts, and video keyframes, extract semantic units, and generate a time-stamped list of explanation elements. c. Teaching logic layer: Build a teaching framework based on a dynamic knowledge graph, automatically annotate the associations between knowledge points and the distribution of teaching priorities through graph structure reasoning, and support dynamic planning of teaching paths.

[0028] 2) Style Decision Unit: This unit is used to construct a teaching style decision tree and implements dynamic style selection using a deep reinforcement learning framework. The state space contains a multi-dimensional feature encoding that integrates the complexity of teaching content (calculated based on the weighted cognitive difficulty of knowledge points), student cognitive level assessment values ​​(modeled through historical learning behavior), and historical teaching effectiveness data. The action space covers nine basic styles and combination modes, including academic rigor, case-driven, and Socratic question-and-answer. The reward function uses a multi-objective reinforcement learning framework, integrating a composite reward function that integrates classroom interaction rate prediction, knowledge point mastery estimation, and style consistency scoring, combined with an offline policy optimization algorithm (PPO) to improve model adaptability.

[0029] 3) Content generation unit: adopts a multi-agent collaborative architecture, including: a. Knowledge Distillation Agent: Based on the logical reasoning chain of the knowledge graph and the implicit knowledge verification of the pre-trained language model, a dual-channel audit is implemented to ensure the consistency of concept representation. A three-dimensional knowledge verification system is constructed, which includes the logical reasoning chain verification of the knowledge graph (explicit knowledge), the implicit association verification of the pre-trained language model (latent knowledge), and the boundary verification of domain expert feedback (experiential knowledge). A dynamic concept alignment mechanism is introduced to achieve cross-modal knowledge representation consistency through comparative learning. A multi-source heterogeneous knowledge fusion module is developed to support the knowledge distillation and reorganization of text, formulas, and diagrams. b. Style Rendering Agent: Based on the selected style template, it inserts stylized elements such as interactive questions and case analogies through context-aware text generation technology. It constructs a style feature transfer matrix that includes a style quantification indicator system with 12 dimensions, such as interaction intensity, cognitive load, and emotional temperature. It also builds a question strategy selector that combines context-aware generation technology, an attention-gated case interpolation algorithm, and reinforcement learning. It also develops a multimodal style transfer engine that supports real-time switching between seven teaching styles, including narrative, debate, and inquiry. c. Teaching compliance intelligent body: Build a rule constraint library based on curriculum standards to detect out-of-syllabus content and analyze knowledge point coverage of generated content; establish a four-layer constraint system: curriculum standards (basic layer) + cognitive development model (psychology layer) + regional teaching outline (institutional layer) + ethical review standards (value layer); develop an intelligent calibration system that can realize knowledge point coverage analysis based on concept topology map + out-of-syllabus detection based on semantic dependency tree; build a dynamic adjustment mechanism to update the regional teaching policy knowledge base in real time through online learning.

[0030] Each intelligent agent achieves collaborative decision-making under multi-dimensional constraints through a distributed consensus algorithm.

[0031] 4) Adaptive Optimization Unit: Construct a closed-loop system for teaching effect feedback by: a. Online learning mechanism: Real-time collection of classroom interaction data (such as student question frequency and attention heatmaps), analysis of real-time teaching data streams, and dynamic adjustment of the weight parameters of the style decision model to adapt to changes in teaching scenarios; b. Incremental fine-tuning strategy: Using efficient parameter fine-tuning technology, the adapter components of the multimodal large model are updated based on newly added teaching cases to achieve low-cost knowledge transfer; c. Adversarial training module: Build a generator-discriminator adversarial framework, evaluate the teachability of generated content through a teaching applicability discriminant network, and drive model optimization.

[0032] The lightweight digital human generation module efficiently converts text generated by the multimodal cognitive agent module into natural and fluent speech output. The module comes with several pre-installed digital human avatars and also supports users uploading photos to generate highly compatible digital human avatars. By employing pruning, quantization, and model compression techniques, the module significantly reduces computational complexity and hardware requirements, enabling efficient operation on resource-constrained devices and ensuring high-quality speech and digital human generation in a low-power environment. The module supports customized generation of multiple voice styles and multi-level emotional expression. Based on lightweight speech synthesis technology, it dynamically adjusts speech rate, intonation, and emotional expression, ensuring that the speech is both relevant to the teaching content and natural and realistic. Through optimized speech generation algorithms and intelligent scheduling strategies, the module minimizes speech generation latency while ensuring output clarity and coherence. The module supports multilingual speech synthesis and utilizes a compact multilingual model, ensuring both efficiency and accuracy in complex cross-language teaching environments. Furthermore, the module employs a low-latency synchronization mechanism, using a lightweight motion generation algorithm to synchronize speech with the digital human's lip movements, facial expressions, and movements in real time, ensuring smooth and natural output. The module's coordinated control architecture is designed to be lightweight. By loading functions on demand and dynamically allocating resources, it minimizes unnecessary calculations and memory usage, significantly improving system efficiency. Combining teacher needs, courseware content, and teaching objectives, the module can automatically optimize the generation process with optimal resource usage. The module also supports highly customized digital human generation. Teachers can choose or design digital human images that suit the course requirements while ensuring smooth and consistent appearance, movements, and voice performance. Through the comprehensive application of lightweight technology, the module not only achieves efficient processing of complex functions, but also ensures flexible adaptability and economy in a variety of teaching scenarios, providing a solution for environments with limited educational resources.

[0033] The intelligent teaching plan assistance module, based on knowledge graphs and teaching path inference results, provides teachers with comprehensive intelligent support from course design to resource integration. The module first uses multimodal semantic understanding technology to extract core knowledge points from teaching resources and construct a topological network of knowledge points, including logical relationships, instructional dependencies, difficulty levels, and importance weights. Based on this, it employs a constraint satisfaction algorithm, combining pre-set teaching objectives with student cognitive characteristics, to generate a draft teaching plan that includes the order of knowledge point presentation, classroom activity design (such as group discussions and quizzes), time allocation, and an assessment plan. Based on the draft teaching plan, it automatically recommends relevant teaching resources (such as courseware, videos, and exercises), providing personalized recommendations based on teacher preferences and supporting fast resource preview and import. The module also features a built-in interdisciplinary knowledge mapping engine that automatically identifies cross-disciplinary connections, such as the derivational relationships between mathematical formulas and physical experiments, and the temporal connections between historical events and literary works, enabling teachers to quickly build interdisciplinary courses. The module dynamically generates customized syllabi for different grade levels by adjusting the depth of knowledge point explanation and the complexity of case studies (for example, transitioning from flowcharts to code implementation in a programming course). The module also includes a built-in resource recommendation subsystem. Based on collaborative filtering and content matching algorithms, this subsystem intelligently recommends supplementary materials such as courseware, exercises, and popular science videos from local libraries or cloud-based educational resource platforms. It also provides resource previews, one-click importing, and copyright compliance checks. Furthermore, by analyzing the knowledge density and logical structure of lecture notes and combining historical teaching data with predictions of potential student comprehension barriers (such as abstract concepts and multi-step reasoning), the module generates targeted teaching suggestions for teachers. For example, when explaining circuit principles, it can suggest adding a physical wiring demonstration or introducing a teaching metaphor similar to water flow. All generated teaching plans and suggestions can be edited and adjusted using a visual timeline and exported as standardized lesson plan templates, significantly reducing the workload for teachers in course design.

[0034] This application utilizes a vision-language pre-training model and a cross-modal attention mechanism, enabling the system to jointly encode data from different modalities, thereby establishing a unified semantic space. Using a contrastive learning algorithm, it aligns features of semantic entities in text, images, and videos, ensuring high consistency and accuracy of information extracted from multimodal resources.

[0035] Based on in-depth analysis of multimodal data, the intelligent agent can automatically extract key concepts, formulas, diagrams, and experimental procedures from the courseware. It then uses graph neural networks (GNNs), graph convolutional networks (GCNs), and graph attention networks (GATs) to construct a dynamically evolving knowledge graph. This approach not only captures the logical relationships between knowledge points but also optimizes the graph structure as teaching resources are continuously updated.

[0036] The system utilizes a multi-agent collaborative architecture, with each specialized agent (e.g., knowledge distillation, style rendering, or teaching compliance) handling a specific task and reaching a final decision through a distributed consensus algorithm. Combined with deep reinforcement learning and online feedback mechanisms, the agents can adjust their strategies in real time to adapt to different teaching scenarios and needs.

[0037] The digital human in this application uses an optimized speech synthesis engine to convert the processed lecture text into smooth and natural speech output. Simultaneously, with the help of a speech emotion and intonation adjustment algorithm, the speech speed, emotion, and intonation are dynamically adjusted according to the teaching content and preset style, ensuring that the explanation is both clear and engaging.

[0038] The digital human generation module not only handles speech output but also simultaneously generates a matching virtual avatar. By simulating lip movements, facial expressions, and movements, the system enables highly realistic and natural interaction during the digital human presentation. The lightweight design, model pruning, and quantization techniques ensure efficient operation even on resource-constrained devices.

[0039] Digital human technology supports diverse teaching scenarios, such as classroom lectures, online course recording, and the production of teaching materials. Teachers can customize their output styles (e.g., formal, relaxed, or interactive) based on specific needs, enabling personalized interactions with different student groups and enhancing classroom engagement and teaching effectiveness.

[0040] Taking the course "Big Data Technology" as an example, the teacher uses the following steps: 1. Courseware upload and preprocessing: Teachers upload the lesson plan package containing the following contents on the courseware upload interface: Technical principles: HDFS architecture diagram, MapReduce workflow diagram, YARN resource scheduling PPT; Programming Practice: Word Experiment Manual (HBase Shell Operation Guide), Java Code File (SparkWordCount Example); Supplementary materials: Zookeeper election process animation; The system automatically identifies the technical level and marks "Storm Topology Design" as advanced content, suggesting that you set it as an elective module.

[0041] 2. Multimodal parsing and structured lecture generation: The system activates a multimodal parsing engine: extracting titles, paragraphs, and formulas, identifying knowledge hierarchies (e.g., core concepts and further reading), creating an architecture diagram (noting the HDFS NameNode / DataNode collaboration relationship), capturing key video frames (decomposing the three phases of the Zookeeper election process), and parsing code (e.g., annotating Mapper / Reducer methods in MapReduce). The parsing results are integrated into a large model to generate a structured lecture transcript containing chapter divisions, a knowledge point dependency diagram, annotated lecture text, and automatically inserted question points (e.g., "What is the core difference between HBase and relational databases?").

[0042] Example: When analyzing the "How MapReduce Works" PPT, automatically generate a diagram of the data flow in the Shuffle phase and annotate core concepts that undergraduates need to master, such as "ring buffer."

[0043] 3. Customization and optimization of speech style: Choose from nine basic modes, including "Academic Rigor," "Case-Driven," and "Socratic Q&A." Use the slider to adjust "Interaction Density" (the frequency of questions).

[0044] After the teacher chooses the "case-driven" style, the system automatically inserts the "Real-time Monitoring of Urban Traffic Flow" example into the "Real-time Data Processing Technology" explanation, simulating how to dynamically analyze road sensor data through Apache Kafka and Spark Streaming, and generating data flow diagrams and traffic heat map diagrams; when explaining "Data Cleaning", it associates the "E-commerce User Behavior Log Cleaning" case, demonstrates the operational procedures for missing value filling and outlier filtering, and simultaneously generates a data quality comparison radar chart.

[0045] 4. Generate and preview the digital human explanation: In the lightweight digital human generation module, teachers configure the relevant parameters for the explanation voice and flexibly adjust the speech speed, intonation, and emotion to ensure that the generated voice aligns with the teaching content and lecture style. The system defaults to providing a natural and smooth speech synthesis output, which teachers can manually adjust based on their preferences. After configuring the voice, users can select a pre-set digital human image or upload a photo to have a digital human image generated synchronously based on the speech content. Teachers can preview the digital human's performance and make fine-tuning adjustments as needed.

[0046] 5. Intelligent teaching plan assistance: Based on the generated lecture content, the system automatically generates a syllabus and teaching plan, including course objectives, key points, and class schedules. Teachers can manually adjust the generated plan and further optimize the teaching design based on the course content. This includes a matching question bank and extended reading.

[0047] The system generates a 16-week teaching plan: Weeks 1-4: Basics Overview of the Hadoop ecosystem; HDFS architecture and CLI operations; MapReduce programming model; YARN resource scheduling experiment; Weeks 5-8 Core Spark RDD programming (compared to MapReduce); HBase data model and Shell operations; Zookeeper distributed collaboration; Weeks 9-12 Advanced Basics of Storm stream processing; Cluster environment construction (including Hadoop+Spark integration); Comprehensive experiment: e-commerce user behavior analysis; Week 13-16 Project Practice Log analysis system development (HDFS+MapReduce); Real-time hot word statistics (Storm+Redis); Guidance on student course design; Automatically associate resources suitable for undergraduates, such as MIT's "Introduction to Hadoop" open course videos and Cloudera experimental environments.

[0048] 6. Export and share: Experimental package: includes VirtualBox image and experimental report template; Mobile Review Pack: Transforms complex concepts like HBase storage architecture into 3D interactive models; Automatically generate a wrong question book: organize common misunderstandings based on classroom question and answer records.

[0049] The teacher will export the generated digital human explanation video and upload it to the online learning platform, where students can watch and review the classroom content after class.

[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A lightweight digital human lesson preparation system based on an intelligent agent, characterized by: include: Knowledge graph construction and reasoning module: used to conduct in-depth semantic analysis of multimodal teaching resources, build structured knowledge graphs, and realize knowledge point association mining and teaching logic reasoning; Multimodal cognitive agent module: By integrating multimodal courseware analysis, semantic understanding, and dynamic adaptation of lecture styles, it generates personalized teaching content that meets teaching objectives, including: Intelligent parsing unit: used to deploy a multimodal hierarchical parsing framework; Style decision unit: used to construct a teaching style decision tree and implement dynamic style selection using a deep reinforcement learning framework; Content generation unit: adopts a multi-agent collaborative architecture and a distributed consensus algorithm to achieve collaborative decision-making under multi-dimensional constraints; Adaptive optimization unit: used to build a closed-loop system for teaching effect feedback; Lightweight Digital Human Generation Module: This module converts lecture notes into natural and fluent speech output, supporting different intonations and emotional expressions to suit different teaching scenarios. It generates a synchronized digital human image based on the speech output, including facial expressions, lip sync, and movements, to achieve digital human presentation. Intelligent teaching plan assistance module: Based on the knowledge graph, it helps teachers intelligently generate teaching plans and syllabuses based on the courseware content, and provides targeted teaching suggestions to optimize the teaching structure.

2. The lightweight digital human lesson preparation system based on an agent according to claim 1, characterized in that: The knowledge graph construction and reasoning modules include: Cross-modal semantic fusion and feature alignment unit: This unit uses a vision-language pre-training model and a cross-modal attention mechanism to jointly encode semantic entities in text, images, and videos, and establish cross-modal feature mapping relationships. It also uses a contrastive learning algorithm to align the semantic spaces of different modalities, extract multi-dimensional feature vectors of conceptual entities in teaching resources, and construct a triple knowledge representation system of entity-relationship-attribute. Dynamic Knowledge Graph Evolution Engine: This engine builds a scalable knowledge topology network based on graph neural networks. It uses graph convolutional networks and graph attention networks to dynamically model the association strength and weight distribution between knowledge points, and uses time-series graph embedding technology to capture the dynamic evolution of the knowledge system. When new teaching resources are added, the system automatically triggers the optimization process of the graph topology structure, adaptively adjusts the association weights of knowledge points, and verifies the integrity of the teaching logic chain, ensuring that the knowledge system is always updated in sync with teaching needs. Visual Interaction and Knowledge Reconstruction Unit: This unit provides a scalable knowledge network visualization interface, enabling teachers to interactively edit knowledge graphs. It also features a built-in conflict detection algorithm that automatically identifies logical breaks, hierarchical contradictions, or redundant relationships in real time when teachers manually adjust knowledge point associations, and generates repair suggestions based on a rule engine. Teaching Path Reasoning and Recommendation Unit: Combining probabilistic graphical models with path planning algorithms, the unit searches for the optimal teaching sequence in the knowledge graph based on the student cognitive level profile and curriculum standard requirements, and outputs personalized path plans that include core knowledge point coverage, teaching time allocation, and advanced route branches, providing a scientific basis for subsequent explanation content generation and teaching plan formulation.

3. The lightweight digital human lesson preparation system based on an agent according to claim 1, characterized in that: The intelligent parsing unit includes: Structural analysis layer: Using a deep learning-based layout analysis model, through the joint encoding of visual elements and text features, it identifies the title hierarchy, paragraph segmentation, and mixed text and image structure in the courseware, and dynamically generates the document's logical skeleton; Semantic parsing layer: Through the joint embedding space alignment technology of the multimodal large model, cross-modal semantic fusion of text, formulas, charts and video key frames is performed to extract semantic units and generate a list of explanation elements with time sequence tags; Teaching logic layer: Build a teaching framework based on a dynamic knowledge graph, automatically mark the association between knowledge points and the distribution of teaching focus through graph structure reasoning, and support dynamic planning of teaching paths.

4. The lightweight digital human lesson preparation system based on an agent according to claim 1, characterized in that: The state space in the style decision unit contains multi-dimensional feature encoding that integrates the complexity of teaching content, students' cognitive level evaluation values, and historical teaching effect data; The action space covers a variety of basic styles and combination modes; the reward function adopts a multi-objective reinforcement learning framework, which improves the model adaptability by integrating a composite reward function that integrates classroom interaction rate prediction, knowledge point mastery estimation, and style consistency scoring, combined with an offline strategy optimization algorithm.

5. The lightweight digital human lesson preparation system based on an agent according to claim 1, characterized in that: The content generation unit includes: Knowledge Distillation Agent: Based on the logical reasoning chain of the knowledge graph and the implicit knowledge verification of the pre-trained language model, it realizes dual-channel review of the consistency of concept representation. It also builds a three-dimensional knowledge verification system, specifically the logical reasoning chain verification of the knowledge graph, the implicit association verification of the pre-trained language model, and the boundary verification of domain expert feedback. It also introduces a dynamic concept alignment mechanism to achieve cross-modal knowledge representation consistency through comparative learning. It also develops a multi-source heterogeneous knowledge fusion module to support the knowledge distillation and reorganization of text / formulas / charts. Style rendering agent: inserts stylized elements based on the selected style template through context-aware text generation technology; constructs a style feature transfer matrix containing a multi-dimensional style quantification indicator system; builds a case interpolation algorithm based on attention gating and reinforcement learning-driven question strategy selector that combines context-aware generation technology; develops a multimodal style transfer engine that supports real-time switching between multiple teaching styles; Teaching compliance intelligent body: Build a rule constraint library based on curriculum standards to detect out-of-syllabus and knowledge point coverage of generated content; establish a four-layer constraint system: curriculum standards, cognitive development model, regional teaching outline, and ethical review standards; develop an intelligent calibration system that can realize knowledge point coverage analysis based on concept topology map + out-of-syllabus detection based on semantic dependency tree; build a dynamic adjustment mechanism to update the regional teaching policy knowledge base in real time through online learning.

6. The lightweight digital human lesson preparation system based on an agent according to claim 1, characterized in that: The adaptive optimization unit collects classroom interaction data in real time through an online learning mechanism. By analyzing the real-time teaching data stream, it dynamically adjusts the weight parameters of the style decision model to adapt to changes in the teaching scenario. Through an incremental fine-tuning strategy, it updates the adapter components of the multimodal large model based on newly added teaching cases, achieving low-cost knowledge transfer. Construct a generator-discriminator adversarial framework, evaluate the teachability of generated content through a teaching applicability discriminant network, and drive model optimization.

7. The lightweight digital human lesson preparation system based on an agent according to claim 1, characterized in that: The coordinated control architecture of the lightweight digital human generation module adopts a lightweight design, and several digital human images are pre-installed in the module. It also supports users to upload photos and generate digital human images that highly match them. The module uses pruning, quantization and model compression technologies to run efficiently on devices with limited resources.

8. The lightweight digital human lesson preparation system based on an agent according to claim 6, characterized in that: The lightweight digital human generation module dynamically adjusts speech speed, intonation and emotional expression based on lightweight speech synthesis technology, making the speech both consistent with the teaching content and natural and realistic; and the module integrates optimized speech generation algorithms and intelligent scheduling strategies to minimize the delay of speech generation while ensuring output clarity and coherence; the module's miniaturized multilingual model is both efficient and accurate when dealing with complex cross-language teaching environments; the module uses a lightweight action generation algorithm to achieve real-time synchronization of speech and the digital human's lip shape, expression and action during explanation, ensuring the smoothness and naturalness of the output process.

9. The lightweight digital human lesson preparation system based on an agent according to claim 1, characterized in that: The intelligent teaching plan assistance module first extracts core knowledge points from teaching resources through multimodal semantic understanding technology, and constructs a knowledge point topological network that includes logical relationships, teaching dependencies, difficulty levels and importance weights; on this basis, combined with preset teaching objectives and students' cognitive level characteristics, a constraint satisfaction algorithm is used to generate a teaching plan draft that includes the order of knowledge point explanation, classroom activity design, time allocation and evaluation plan. Based on the teaching plan draft, it automatically matches and recommends relevant teaching resources. By analyzing the knowledge density and logical structure in the lecture content and combining historical teaching data, it predicts the possible understanding obstacles of students, provides personalized recommendations and targeted teaching suggestions based on teachers' usage preferences, and supports rapid preview and import of resources.

10. The lightweight digital human lesson preparation system based on an intelligent agent according to claim 9, characterized in that: The intelligent teaching plan assistance module also has a built-in interdisciplinary knowledge mapping engine, which can automatically identify cross-domain connection points such as the derivation relationship between mathematical formulas and physical experiments, and the temporal connection between historical events and literary works, supporting teachers to quickly build interdisciplinary integrated courses; and the module also has a built-in resource recommendation subsystem, which intelligently recommends auxiliary materials from local libraries or cloud-based educational resource platforms based on collaborative filtering and content matching algorithms, and provides resource preview, one-click import and copyright compliance check functions.

Citation Information

Cited By

  • Knowledge graph-driven textbook automatic generation method and system

    CN120832870A

  • A knowledge graph driven textbook automatic generation method and system

    CN120832870B

  • AI digital human interactive response method based on large language model

    CN121144484A

  • An AI Digital Human Interactive Response Method Based on a Large Language Model

    CN121144484B

  • Cooperative management system for higher vocational education based on multi-source knowledge graph and intelligent agent

    CN121169032A