A 5e paradigm textbook construction system based on large model and retrieval enhancement generation

CN122820409APending Publication Date: 2026-09-25NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611272651.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]具体而言,现有技术存在以下问题:(1)现有各类教材或教学内容生成方案均缺失5E教学范式的结构化约束,其设计内核仅围绕传统知识点的匹配与个性化推送,生成内容各模块相互独立,缺乏递进式教学逻辑,无法形成闭环的问题探究链路,产出的教材不能直接应用于5E教学场景

Benefits of technology

1.本发明针对Excitation、Exploration、Enhancement、Execution、Evaluation以及Application各阶段分别设计专属提示词模板,并采用链式Prompt调用机制,在后续阶段生成时自动引入前序阶段生成结果作为上下文依据,使教材内容严格遵循“提出问题-探索问题-学习知识-解决问题-评价反思-学以致用”的递进式教学路径。本发明能够保证教材内容各环节之间具备明确的逻辑依赖关系和知识传递关系,实现符合POT-OBE理念的问题驱动型教材自动生成,使生成教材能够直接应用于5E教学场景,提高教材的教学适配性与教育价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820409A_ABST
    Figure CN122820409A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence and knowledge engineering, and particularly discloses a 5E paradigm textbook construction system based on a large model and retrieval enhancement generation, which comprises a course definition and user interaction module, a course knowledge point subset generation module, an outline planning module and a 5E stage content generation module; the course definition and user interaction module is used for acquiring teaching reference materials; the course knowledge point subset generation module is used for generating course knowledge point subsets; the outline planning module is used for generating a textbook outline by using a large language model according to the course knowledge point subsets; the 5E stage content generation module is used for generating textbook contents for each chapter of the textbook outline in the order of the stages of arousing interest and putting forward a question stage, a stage of exploring the essence of a problem by using a first principle, a stage of learning necessary knowledge and ability for solving a problem, a stage of practically solving a problem, an evaluation and reflection stage and a stage of learning for practical use, and the large language model is used for generating the textbook contents by using a semantic block unit one by one; and the application constructs an end-to-end full-automatic textbook generation process and can generate a structured textbook conforming to a 5E new teaching paradigm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and knowledge engineering technology, and in particular to a 5E paradigm textbook construction system based on large model and retrieval enhancement generation. Background Technology

[0002] The rapid development of artificial intelligence technology is profoundly transforming the theoretical and practical models of education. To address the new demands on talent cultivation in the AI ​​era, and to implement the POT-OBE (Problem-Oriented Outcome-Based Education) concept, the "5E Teaching Paradigm," oriented towards problem-based inquiry, has emerged. This paradigm restructures the entire teaching process into five progressive stages: Excitation (arousing interest and posing questions), Exploration (exploring the essence of the problem using first principles), Enhancement (learning the necessary knowledge and skills to solve the problem), Execution (hands-on problem-solving), and Evaluation (evaluation and reflection). In this paradigm, teachers act as curriculum writers, designing the entire problem-based inquiry script and guiding the learning process in the classroom; students move beyond passive listening to become active learners; and artificial intelligence transcends its role as a mere question-and-answer tool, acting as a collaborative learning partner throughout the process.

[0003] However, rewriting traditional textbooks to fit the new 5E teaching paradigm is an extremely complex and systematic project. If it relies entirely on manual writing, not only is the compilation process lengthy, but the quality of the content is also heavily constrained by the individual teaching and research capabilities of teachers, making it difficult for this paradigm to be quickly and widely implemented.

[0004] Meanwhile, while large language models demonstrate powerful text generation capabilities, existing technologies are significantly ill-suited for the automatic generation of textbooks. Simply relying on native large language models to directly generate textbook content easily leads to problems such as fabricated knowledge (knowledge illusion), confused professional logic, and misaligned structural elements, failing to guarantee the scientific rigor and accuracy of the textbooks. Retrieval-enhanced generation techniques introduced to suppress these illusions are currently mostly applied to shallow interactive scenarios such as intelligent Q&A and fragmented courseware generation, lacking a systematic solution for compiling complete textbooks.

[0005] Specifically, the existing technologies have the following problems: (1) The existing textbooks or teaching content generation schemes lack the structured constraints of the 5E teaching paradigm. Their core design revolves only around the matching and personalized push of traditional knowledge points. The generated content modules are independent of each other, lack progressive teaching logic, and cannot form a closed-loop problem exploration link. The produced textbooks cannot be directly applied to the 5E teaching scenario. (2) The underlying retrieval enhancement generation technology does not have the ability to perceive the teaching stage. Its retrieval strategy takes answering the current question as the only goal. It cannot provide targeted knowledge supply according to the differentiated needs of reference material types in different stages of 5E (such as the explanation stage and the practice stage). This reduces the professionalism of the content and makes it difficult to avoid the problems of knowledge point illusion and non-standard citation from the root. (3) Textbooks are ultra-long texts. Existing large language models are prone to content truncation and context breakage when generating long texts. Moreover, they lack a corresponding automated quality assessment and iterative correction closed loop, resulting in high costs for manual review and proofreading in the later stage.

[0006] In summary, there is an urgent need for a systematic solution that can deeply integrate retrieval-enhanced generation technology, large language models, and the closed-loop structural logic of the 5E teaching paradigm to solve the current dilemma of large workload and long cycle of purely manual writing, while existing large model technology cannot adapt to the needs of generating specialized teaching materials. Summary of the Invention

[0007] This invention aims to solve the aforementioned problems. To this end, it provides a 5E paradigm textbook construction system based on large-scale modeling and retrieval enhancement. Starting with user-inputted textbook descriptions and teaching references, and aiming to output structured textbooks conforming to the 5E new teaching paradigm, it constructs an end-to-end fully automated textbook generation process. Customized prompt word templates are provided for each stage of 5E teaching, employing a Prompt chaining method to constrain the sequence of teaching stages and link content. In subsequent stages, prompt words incorporate previously generated content as contextual basis, ensuring the textbook content strictly follows the progressive teaching logic of the 5E paradigm. Simultaneously, a stage-aware multi-round progressive RAG mechanism is introduced, dynamically constructing retrieval semantics based on different teaching stages (Enhancement, Execution) to achieve accurate matching of chapter knowledge points and references. Furthermore, the system integrates a dual-stage content truncation detection and recovery mechanism and an AI scoring iterative optimization mechanism, effectively addressing the pain points of incomplete long textbooks and difficulty in controlling content quality.

[0008] This invention provides a 5E paradigm textbook construction system based on large model and retrieval enhancement, and the technical solution adopted is as follows: including: The course definition and user interaction module is used to obtain textbook descriptions and / or teaching reference materials; and to generate a subset of course knowledge points based on the textbook descriptions and the knowledge point pool. The reference material preprocessing and knowledge base construction module is used to semantically segment teaching reference materials according to title level to obtain semantic segmentation units; and to extract knowledge points from semantic segmentation units to build a knowledge point pool. The outline planning module is used to generate a textbook outline based on a subset of course knowledge points using a large language model. Each chapter of the textbook outline includes a title, a description of the main content, a list of knowledge points, and prerequisite knowledge points. The 5E content generation module is used to generate textbook content for each chapter of the textbook outline in the following order: the stage of arousing interest and raising questions, the stage of exploring the essence of the problem using first principles, the stage of learning the knowledge and skills necessary to solve the problem, the stage of solving the problem by doing practical work, the stage of evaluation and reflection, and the stage of applying what has been learned. It calls the large language model to generate textbook content using semantic block units. The quality assurance module group is used to retrieve semantic block units based on the content of each chapter, use the retrieval results to constrain the generation of textbook content, evaluate the content quality of the textbook content, and detect whether there are any interruptions or semantic omissions in the textbook content. When the content quality assessment is unqualified or there are interruptions or semantic omissions, the content generation modules of each stage of 5E will regenerate the textbook content of that stage. The quality assurance module group includes a multi-round progressive RAG retrieval submodule, a truncation detection and repair submodule, and an evaluation and quality iteration submodule; among them, The multi-round progressive RAG retrieval submodule is used to retrieve semantic block units based on the content of each chapter in the stages of learning the knowledge and skills necessary to solve problems and in the actual hands-on problem-solving stage, and to use the retrieval results to constrain the generation of textbook content. The truncation detection and repair submodule is used to evaluate the mid-term truncation and semantic incompleteness of the teaching material in the stage of learning the knowledge and skills necessary to solve the problem and the stage of actual hands-on problem solving. When the teaching material has mid-term truncation and semantic incompleteness, the teaching material content of each stage is regenerated by the 5E stage content generation module. The evaluation and quality iteration submodule is used to evaluate the content quality of the teaching materials in the stages of stimulating interest and raising questions, exploring the essence of the problem using first principles, learning the knowledge and skills necessary to solve the problem, solving the problem by doing, and evaluation and reflection. When the content quality of the teaching materials is not up to standard, the content generation module of each stage of 5E will regenerate the teaching materials for that stage. The textbook output module is used to output textbook content.

[0009] Furthermore, the course definition and user interaction module should obtain at least one of the following: textbook description and teaching reference materials; The textbook description includes the knowledge points, teaching objectives, and teaching requirements.

[0010] Furthermore, the teaching reference materials are semantically segmented according to the heading level, and the specific process for obtaining semantic segmentation units is as follows: The teaching reference materials are parsed to obtain Markdown text content and JSON structured data with heading hierarchy information; Traverse the JSON structured data in document node order to extract the original units; or split the Markdown text content according to heading regular expressions and blank line paragraph rules to obtain the original units. Based on the number of tokens, the original units are merged and split to obtain semantic block units.

[0011] Furthermore, the specific process for extracting knowledge points from semantic chunks and constructing a knowledge point pool is as follows: Semantic blocks are grouped according to structured rules to form chapter units with semantic integrity; Candidate knowledge points are extracted from chapter units; The candidate knowledge points are semantically merged and deduplicated to obtain a knowledge point pool.

[0012] Furthermore, when generating teaching material content in the downstream stage, all content already generated in the upstream stage will be used as context input; By using pre-set textbook prompts, the generated content includes multiple rounds of dialogue and interaction between learners and learning partners in the stages of exploring the essence of a problem using first principles and learning the knowledge and skills necessary to solve the problem.

[0013] Furthermore, the workflow of the multi-round progressive RAG retrieval submodule is as follows: Using the main content descriptions and knowledge points of the chapters as search text, dense vector search and BM25 sparse keyword search were performed on the semantic block units respectively to obtain two preliminary search results. The RRF reciprocal ranking fusion algorithm is used to fuse the two preliminary retrieval results, and the fusion score of the candidate fragments in the preliminary retrieval results is calculated. Candidate segments are sorted based on their fusion scores. The top-scoring candidate segments are then injected into prompts in the content generation modules of each stage of 5E to constrain them to regenerate textbook content.

[0014] Furthermore, the multi-round progressive RAG retrieval submodule adopts a sharding multi-round retrieval strategy, specifically: splitting the main content descriptions and knowledge points of chapters according to fixed sharding standards; For each subquery text obtained from the splitting, add targeted keywords based on the current stage, and then enable dense vector retrieval and BM25 sparse keyword retrieval in parallel; All candidate fragments obtained from multiple rounds of retrieval are summarized, and duplicate and redundant content is removed to obtain two preliminary retrieval results.

[0015] Furthermore, the workflow of the truncation detection and repair submodule is as follows: Conduct heuristic rapid testing of teaching materials to identify truncation defects; When there is a truncation defect in the teaching material, the corresponding truncation prompt is generated and injected into the prompt words of the content generation module of each stage of 5E, so that the teaching material content can be regenerated. When there are no truncation defects in the teaching materials, perform LLM refined integrity detection on the teaching materials to identify truncation defects and semantic incompleteness; When the teaching material has truncation defects and semantic incompleteness, corresponding truncation and incompleteness prompts are generated and injected into the prompt words of the content generation modules of each stage of 5E, which then complete the regeneration of the teaching material content.

[0016] Furthermore, the assessment and quality iteration submodule evaluates the content quality of the teaching materials from the dimensions of teaching logic, accuracy of knowledge points, learner experience, and standardization of text language.

[0017] The above-described one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: 1. This invention designs dedicated prompt templates for each stage—Excitation, Exploration, Enhancement, Execution, Evaluation, and Application—and employs a chained Prompt call mechanism. When generating subsequent stages, the results generated in previous stages are automatically incorporated as contextual information, ensuring that the teaching material strictly follows a progressive teaching path of "posing questions - exploring questions - learning knowledge - solving problems - evaluating and reflecting - applying knowledge." This invention guarantees clear logical dependencies and knowledge transfer relationships between each stage of the teaching material, enabling the automatic generation of problem-driven teaching materials that conform to the POT-OBE (Plan-Do-Check-Act) principle. This allows the generated teaching materials to be directly applied to 5E (Excitation, Exploration, Enhancement, Execution, Evaluation, and Application) teaching scenarios, improving the teaching adaptability and educational value of the materials.

[0018] 2. This invention introduces a dialogue mechanism between the learner (I) and the learning partner (AI assistant) in the Exploration and Enhancement phases. Through the Prompt project, it constructs human-computer collaborative interactive content that conforms to the characteristics of Socratic heuristic teaching, integrating the exploration and knowledge construction processes into the textbook text in a natural dialogue format. This invention transforms traditional static textbooks into intelligent textbooks with interactive features, allowing learners to experience the complete thought process of problem discovery, analysis, and solution while reading the textbook, improving learning participation, exploration depth, and knowledge internalization, and better meeting the needs of human-computer collaborative teaching in the era of artificial intelligence.

[0019] 3. This invention innovatively introduces a stage-aware, multi-round progressive RAG mechanism to dynamically construct retrieval semantics based on different teaching stages. Specifically, the Enhancement stage focuses on retrieving information related to concepts, definitions, principles, and explanations, while the Execution stage focuses on retrieving information related to cases, examples, steps, and practical methods. It also combines dense vector retrieval, BM25 keyword retrieval, RRF fusion ranking, and a re-ranking model for multi-level filtering. This invention can automatically match the most suitable knowledge sources according to the teaching stage, improving the utilization rate of reference materials and the accuracy of knowledge supply, effectively reducing problems such as knowledge illusion in large models, content distortion, and non-standard knowledge citations, and enhancing the professional reliability of textbook content.

[0020] 4. This invention constructs a two-layer quality assurance system consisting of a truncation detection and recovery mechanism and an AI evaluation and optimization mechanism. The truncation detection and recovery mechanism employs a dual verification method combining heuristic rapid detection and LLM refined integrity detection to automatically check and recover generated content. The AI ​​evaluation and optimization mechanism automatically scores content from multiple dimensions, including teaching logic, knowledge point accuracy, learner experience, and text language standardization, and performs iterative optimization based on the scoring results. This invention can automatically detect and repair generated defects with little or no human intervention, significantly reducing the cost of manual review during textbook generation and improving the integrity, consistency, and stability of long textbook content.

[0021] 5. This invention eliminates the need for pre-constructing complex teaching knowledge graphs or fixed template libraries. It automatically completes the entire process of knowledge extraction, syllabus planning, 5E textbook content generation, quality optimization, and textbook export solely based on textbook descriptions, teaching reference materials, and a large language model. This invention significantly shortens the textbook development cycle, improves textbook production efficiency, and enables rapid migration and application across different educational stages, courses, and language environments, providing technical support for large-scale intelligent textbook construction.

[0022] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0024] Figure 1 This is a system structure block diagram provided by the present invention.

[0025] Figure 2 This is a flowchart of the generation process of semantic block units provided by the present invention.

[0026] Figure label: 1. Course definition and user interaction module; 2. Reference material preprocessing and knowledge base construction module; 3. Syllabus planning module; 4. Content generation module for each stage of 5E; 5. Quality assurance module group; 6. Textbook output module. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. The following embodiments are used to illustrate this invention but should not be used to limit the scope of this invention.

[0028] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0029] The following is combined Figure 1 and Figure 2 The present invention will be further described in detail below, describing a 5E paradigm textbook construction system based on large model and retrieval enhancement: In this embodiment, as Figure 1 As shown, a 5E paradigm textbook construction system based on large model and retrieval enhancement is provided, including: 1. Course definition and user interaction module; 2. Reference material preprocessing and knowledge base construction module; 3. Syllabus planning module; 4. Content generation module for each stage of 5E; 5. Quality assurance module group; and 6. Textbook output module.

[0030] The course definition and user interaction module is used to obtain textbook descriptions and / or teaching reference materials; and to generate a subset of course knowledge points based on the textbook descriptions and the knowledge point pool.

[0031] Users create courses through the front-end interface or API in the course definition and user interaction module, inputting a course requirement description and optionally uploading teaching reference materials (PDF / DOCX / PPTX, etc.). The course requirement description includes the textbook name, textbook description, applicable grade level, language type, and expected number of chapters. The course definition and user interaction module records the binding relationship between files and courses and tracks file processing status (pending / processing / success / failure). Users must provide at least one of the following in the course definition and user interaction module: a textbook description and teaching reference materials.

[0032] When a user creates a course, the fields required to describe the course requirements are shown in Table 1.

[0033] Table 1

[0034] Users can independently input textbook descriptions, briefly outlining the core knowledge points, teaching objectives, and overall teaching requirements. This module supports AI-powered professional polishing and optimization of textbook descriptions. The polishing process begins with content compliance verification, eliminating redundant information irrelevant to the course theme while retaining effective content such as knowledge points, learning objectives, skills development, and application value. Furthermore, the polishing and optimization process deeply integrates the POT-OBE outcome-oriented education philosophy and the core ideas of the 5E new teaching paradigm, adhering to a problem-oriented and outcome-based design approach. It emphasizes inquiry-based learning driven by real-world problems, highlighting the cultivation of learners' problem-solving logical cognitive patterns, higher-order cognition, and innovative abilities.

[0035] When users upload teaching reference materials, they can select all or part of the knowledge points from the knowledge point pool built based on these materials to obtain a subset of course knowledge points. This subset of course knowledge points serves as the input source for generating the textbook syllabus. The syllabus planning module strictly adheres to the user's selection range, neither introducing unselected knowledge points nor adding, rewriting, or omitting selected knowledge points.

[0036] If the user does not upload teaching reference materials, the description of the teaching materials entered by the user will be polished; the polished description of the teaching materials will be parsed, the course knowledge points will be extracted, and a subset of course knowledge points will be obtained.

[0037] Finally, the textbook name and a subset of course knowledge points are simultaneously transmitted to the syllabus planning module, providing preliminary constraints and design basis for the intelligent generation of subsequent structured course syllabi.

[0038] The reference material preprocessing and knowledge base construction module is used to semantically segment teaching reference materials according to the title level to obtain semantic segmentation units; knowledge points are extracted from the semantic segmentation units to build a knowledge point pool.

[0039] The reference material preprocessing and knowledge base construction module performs structured processing on user-uploaded teaching reference materials, converting them from binary document format into a structured set of chunks suitable for subsequent LLM (Large Language Model) semantic understanding and vector retrieval. The specific operations are as follows: S2.1: Parse the teaching reference materials and semantically segment them according to heading levels (e.g., chapters, sections, subsections) to obtain semantic block units. For example... Figure 2 As shown.

[0040] This module first performs standardized document parsing on various teaching reference materials uploaded by users. This embodiment uses IBM open-source Docling as the core parsing engine, compatible with common formats such as PDF, DOCX, PPTX, XLSX / XLS / CSV, TXT / MD, and HTML. Dedicated main parsers and fallback parsing schemes are configured for different file types. PDF format has built-in OCR capabilities, and various formats automatically switch to the corresponding backup parsing library when the main parser fails, ensuring the compatibility and stability of document parsing.

[0041] During parsing, a document parsing record is first created and marked as "parsing in progress." The system then routes to the corresponding parsing branch based on the file type. Doling is used to perform the document structure transformation, outputting both Markdown text content and JSON structured data with heading hierarchy information. The Markdown text content is used for subsequent large model prompt injection, while the JSON structure retains key information such as heading hierarchy, page numbers, and content block types, providing a structural basis for semantic segmentation. To address memory anomalies during parsing, the system has a built-in degradation retry strategy, as shown in Table 2, which automatically reduces image resolution and retryes parsing to improve parsing fault tolerance.

[0042] Table 2

[0043] After document parsing is complete, semantic chunking is performed using rules-driven methods without the involvement of a large model. Chunking primarily relies on the JSON structured data output by Doling.

[0044] If JSON structured data exists, extract it from the JSON structured data, traverse it in the document node order and maintain the title hierarchy stack (chapter, section, subsection), extract the original units such as text content, title path, page number and block type, and obtain the original units.

[0045] If the JSON structure is missing, it will automatically degrade to extracting from the Markdown text content, splitting the content according to heading regular expressions and blank line paragraph rules to obtain the original unit.

[0046] This module performs fine-tuning of the original units based on the token threshold, merges adjacent units with too few tokens, directly retains units with medium tokens, and performs sentence-level fine-tuning of extremely long units according to a fixed budget, ultimately generating standardized, independent, and searchable semantic chunks. The original unit tuning rules are shown in Table 3.

[0047] Table 3

[0048] Each semantic chunk unit has fixed structured fields such as index, title path, page range, content, character count, and token count. A group key (group_key) is generated for each semantic chunk unit. The group key is generated first based on the title path, and if there is no title path, the slide number and starting page number are used as fallbacks to ensure that the relevant semantic chunk units can be efficiently aggregated based on the group key in the subsequent knowledge point extraction stage. The final result is a standardized and searchable set of semantic chunks, laying a standardized data foundation for subsequent knowledge point extraction, vector embedding, and RAG knowledge base index construction.

[0049] S2.2: Extract knowledge points from semantic block units and construct a knowledge point pool.

[0050] The knowledge point pool construction is a core step in the reference preprocessing workflow, relying on the Large Language Model (LLM). It aims to further refine standardized semantic chunks into structured, reusable teaching knowledge points, providing precise knowledge anchors for textbook content generation. This step completes the construction of a course-level knowledge point pool from semantic chunks through section grouping, candidate point extraction, and cross-section merging.

[0051] S2.21: Chapter Grouping. Semantic chunks are grouped according to structured rules based on the Chunk list of the semantic chunk set, forming chapter units with semantic integrity. During grouping according to structured rules, auxiliary semantic chunks such as tables and image descriptions are temporarily stored and added to the main semantic chunks (text or lists) when they are matched. The grouping of main semantic chunks strictly follows three rules: isomorphism, adjacency, and budget control. Only semantic chunks with consistent title paths, grouping keys, or page number structures, consecutive Chunk indices, and a merged token count not exceeding 1200 are merged. If any condition is not met, a new chapter unit is immediately created to ensure that each chapter unit has semantically independent content and a moderate size.

[0052] S2.22: Candidate Knowledge Point Extraction. For each chapter unit, the large language model, which is responsible for knowledge point extraction, is invoked to extract structured candidate knowledge points based on a preset knowledge point extraction prompt template. To ensure stable and reliable output, the large language model employs low-randomness parameter configuration and sets reasonable retries and timeouts. It also includes a built-in fallback mechanism to automatically correct abnormal outputs such as empty summaries or exceeding the knowledge point limit, improving the overall robustness of the process. In this embodiment, illustratively, the knowledge point extraction prompt can be: "Taking the combined text as input, the model output strictly follows the specified JSON format specification, including structured content such as section overview, candidate knowledge point list, core keywords, and chapter type." S2.23: Cross-Chapter Merging. To eliminate redundancy in knowledge points across different chapter units, candidate knowledge points across the entire course are uniformly summarized and merged. All candidate knowledge points formed in step S2.22 are collected to form a global candidate pool. A large language model is used to semantically merge and deduplicate the global candidate pool, generating final course-level knowledge points and identifying their associated candidate knowledge points. Based on this, for each final knowledge point, the system generates a standardized summary based on its associated original text evidence (semantic chunk units), rigorously verifying the legality of the evidence sources. Finally, a version-based management mechanism is used to obtain the knowledge point pool. The knowledge point pool is open to users for display and selection; users can select the required knowledge points as needed, providing precise reference for the intelligent planning of subsequent course outlines.

[0053] S2.3: Construct an index for the semantic block unit to obtain the semantic block unit embedding vector index.

[0054] Vector index construction is the core step in converting standardized semantic blocks into dense vector representations that can be retrieved by machines. This process is automatically triggered after semantic blocks are completed. Through a complete process of embedding services, batch encoding, and transactional database storage, an efficient and stable vector retrieval foundation is built, providing accurate semantic matching support for subsequent hybrid retrieval.

[0055] This embodiment uses the singleton pattern to design the vector embedding service (EmbeddingService), and the relevant configurations can be flexibly adjusted, as shown in Table 4.

[0056] Table 4

[0057] The index building process is as follows: S2.31: Index Status Initialization: Get / create course-level RAG index status records, marked as "Processing", to track the index building progress; S2.32: Filtering of semantic chunks to be processed: Supports two modes: incremental construction (processing only semantic chunks that have not been indexed) and full reconstruction, adapting to different update scenarios; S2.33: Batch Encoding and Database Storage: Concatenate the title path (separated by ">") and content of semantic block units to form index text; then call the vector embedding service to batch encode the index text, generating a 1024-dimensional dense vector; finally, use update_or_create (update if it exists, create if it doesn't exist) to create / update semantic block unit embedding records for each semantic block unit in the database transaction; the failure of encoding a single semantic block unit does not affect other semantic block units in the batch; S2.34: Progress Update and Exception Handling: After each batch is completed, update the number of indexed items and the number of failed items in the RAG index status record. Mark it as "Completed" when all items are completed, and mark it as "Failed" when there is an exception.

[0058] Finally, a semantic block unit embedding vector index is generated, which, combined with subsequent sparse retrieval (BM25), forms a hybrid retrieval foundation of dense and sparse, providing high-precision semantic matching capabilities for RAG enhancement generation.

[0059] The syllabus planning module is used to generate a textbook syllabus based on a subset of course knowledge points using a large language model. The textbook syllabus adopts a structured JSON Schema format, and each chapter of the textbook syllabus includes a title, a description of the main content, a list of knowledge points, and prerequisite knowledge points.

[0060] The module inputs a subset of course knowledge points into the large language model of the syllabus generation role. Through strong constraints imposed by syllabus generation prompts, it automatically generates a structured textbook syllabus that conforms to the OBE (Outcome-Based Education) and 5E (5E) new teaching paradigms. This module can automatically match Chinese and English prompt templates based on the user-defined language type. Under the unified constraints of the syllabus generation prompts, the large language model scientifically organizes and allocates chapters for selected knowledge points according to the principles of progressive learning paths, problem-driven learning, and highly cohesive knowledge structures.

[0061] This embodiment achieves multi-dimensional and robust control through built-in rules for outline generation prompts. For example, outline generation prompts could be: "The number of knowledge points per chapter is limited to 2-5 to avoid scattered accumulation; each chapter revolves around a clear core issue, with knowledge points forming a progressive and supporting logical relationship; the learning path for each chapter progresses step-by-step, with subsequent chapters inheriting prior knowledge, avoiding jumps or advanced content. The design of prior knowledge points follows the principle of precision and concreteness, prioritizing already learned content and prohibiting the use of vague, general, or unlearned knowledge as a foundation." The main content and knowledge points of each chapter in the textbook syllabus are strictly aligned, and each item in each chapter corresponds to at least one knowledge point, ensuring that the content is highly focused and free of redundant information.

[0062] The syllabus planning module forces the output of a textbook syllabus in a structured JSON Schema format. The textbook syllabus includes fields such as course name, language, and chapter. The chapter includes a title, a description of the main content, a list of knowledge points, and prerequisite knowledge points.

[0063] The 5E content generation module is used to generate textbook content for each chapter of the textbook outline in the following order: stimulating interest and raising questions, exploring the essence of problems using first principles, learning the necessary knowledge and skills to solve problems, solving problems hands-on, evaluation and reflection, and applying what has been learned. It calls upon the large language model to generate textbook content using semantic chunk units.

[0064] The 5E content generation module, with strong constraints from prompt words at its core, automatically generates teaching materials that conform to the 5E new teaching paradigm and problem-based cognitive model according to five teaching stages: Excitation, Exploration, Enhancement, Execution, and Evaluation, combined with Application (learning by doing). The system supports bilingual generation in Chinese and English, dynamically loading corresponding prompt words according to stage and language through a unified interface, achieving standardized, structured, and reusable content output.

[0065] This module, based on the RTGO (Role-Task-Goal-Objective) prompt structure, configures independent and professional textbook prompt templates for each stage. It clearly defines role positioning, input variables, output format, narrative perspective, and strong constraints. The entire learning process is presented in an immersive first-person "I" perspective, prohibiting the use of "you / we" perspectives to ensure coherent writing and logical closure. The textbook prompts include placeholders for variables such as titles, main content descriptions, knowledge points, prerequisite knowledge points, and semantic blocks. When generating textbook content, these placeholders are automatically injected into the course and chapter context, ensuring that the content is strongly aligned with the chapter's knowledge points, does not advance ahead, does not diverge, and does not deviate from the teaching objectives.

[0066] In this embodiment, the illustrative teaching material prompts for the six stages are shown in Table 5.

[0067] Table 5

[0068] There is a strong dependency between the stages. When the downstream stage generates teaching materials, it will use all the main text content generated in the upstream stage as context input.

[0069] Furthermore, through textbook prompts, the generated content in the Exploration and Enhancement stages includes multi-turn dialogue interactions between learners and their learning partners. The learner is represented by "I" (the learner's first-person perspective), while the learning partner is represented by an "AI assistant." This multi-turn dialogue interaction between "I" and the "AI assistant" is embedded immersively into the textbook text rather than as a separate dialogue system, allowing students to experience a human-computer collaborative inquiry learning process while reading the textbook. This interaction aims to align with the new teaching role positioning in the AI ​​era, where students are the main participants and AI is the learning partner. Following a problem-based cognitive model, it recreates the inquiry-based learning process through a human-computer question-and-answer format, transforming the textbook from a static knowledge presentation into an immersive collaborative inquiry scenario, thus meeting the inherent teaching requirements of the 5E paradigm's exploration and enhancement stages.

[0070] The quality assurance module group is used to retrieve semantic chunks based on the content of each chapter, use the retrieval results to constrain the generation of textbook content, evaluate the content quality of the textbook content, and detect whether there are any interruptions or semantic omissions in the textbook content. When the content quality assessment is unsatisfactory or there are interruptions or semantic omissions, the content generation modules at each stage of the 5E process regenerate the textbook content for that stage. The intervention of the quality assurance module group at each stage of the 5E content generation modules can improve the quality of the generated textbook content.

[0071] The quality assurance module group includes a multi-round progressive RAG retrieval submodule, a truncation detection and repair submodule, and an evaluation and quality iteration submodule.

[0072] The multi-round progressive RAG retrieval submodule is activated in two stages: the knowledge and skills stage necessary for solving problems in the textbook generation process, and the practical problem-solving stage. It retrieves semantic blocks based on the content of each chapter, using the retrieval results to constrain the generation of textbook content. These two stages focus on knowledge point instruction and problem-solving, respectively, with significantly different needs for reference material retrieval, making them the core stages of the system's knowledge retrieval.

[0073] The multi-round progressive RAG retrieval submodule adopts a hybrid retrieval architecture of dense vectors and sparse keywords. The vector embedding uses the text-embedding-v4 model, and the relevance reordering relies on the qwen3-rerank model to achieve fine-grained filtering. The whole is divided into two major execution logics: general RAG basic process and multi-round shard retrieval, and is equipped with an automatic degradation mechanism for empty results.

[0074] The execution flow of the multi-round progressive RAG retrieval submodule includes: When content generation is initiated during the Enhancement and Execution phases, the main content descriptions and knowledge points of each chapter are used as the search text to start the search task. Vector indexes are embedded in semantic chunks, and dense vector search and BM25 sparse keyword search are performed on the semantic chunks to obtain two preliminary search results. The Reciprocal Rank Fusion algorithm is used to fuse the two preliminary search results and calculate the fusion score of the candidate segments in the preliminary search results.

[0075] The calculation formula for the RRF reciprocal ranking fusion algorithm is as follows: in, The inverse ranking of candidate segments is merged into a score. As a smoothing constant, let it be 60. is the ranking number of the candidate fragment in the i-th preliminary search result.

[0076] The RRF reciprocal ranking fusion algorithm calculates scores for each candidate segment in each preliminary search result and then sums them up. The higher the ranking, the higher the score of a single segment. This achieves a weighted fusion of vector semantic retrieval and keyword retrieval results, taking into account both semantic similarity and keyword hit rate.

[0077] Then, the qwen3-rerank reordering model is called to perform a refined relevance ranking of candidate segments based on the fusion score, and the best candidates (highly relevant segments) with the highest fusion scores are selected. The highly relevant segments are then injected into the textbook prompt words, and the textbook content is regenerated based on the content generation modules of each stage of 5E constrained by the real original text.

[0078] To avoid token overruns and decreased search accuracy caused by excessive content in a single search, the multi-round progressive RAG search submodule adopts a sharded multi-round search strategy: (1) Search round splitting: Set the number of items per round (items_per_round) to 2, split the main content description and knowledge points of the chapter according to the fixed splitting standard, obtain multiple sub-query texts, and automatically calculate the search rounds that need to be split; (2) Stage-aware query construction: For each sub-query text obtained by splitting, add targeted keywords in combination with the current teaching stage to achieve search intent adaptation: for the Enhancement stage, add search keywords such as concept, principle, definition, explanation, understanding, interpretation, and reason; for the Execution stage, add keywords such as exercise, example, application, problem-solving, steps, methods, and case studies to accurately match the information search preferences of different stages; (3) Round-by-round hybrid retrieval: The subquery text generated in each round is traversed in sequence. The native RAG hybrid retrieval mechanism is reused in a single round: text-embedding-v4 dense vector retrieval (top_k=20) and BM25 sparse keyword retrieval (top_k=20) are enabled in parallel. After fine sorting by the qwen3-rerank model, the top_K candidate fragments of this round are fixedly output. The iteration runs until all fragment retrieval tasks are completed. (4) Post-processing of full results: Summarize all candidate fragments obtained from multiple rounds of retrieval, remove duplicate and redundant content, and obtain two preliminary retrieval results; call qwen3-rerank again to complete the secondary re-sorting, filter the top_M high-quality and highly relevant fragments with higher priority, splice and integrate them into the final reference context, inject prompt words, and use them for content generation.

[0079] The multi-round progressive RAG retrieval submodule has built-in fault-tolerant fallback logic. When the semantic block unit embedding vector index is empty or there are no valid fragments in the full round of retrieval, it automatically skips the RAG context injection process and switches to the pure LLM self-generation mode to avoid interruption of textbook generation due to missing data sources.

[0080] The truncation detection and repair submodule is also activated in the two long text generation stages of textbook generation: the knowledge and skills stage necessary for solving problems and the hands-on problem-solving stage. It is used to evaluate mid-course truncation and semantic incompleteness of textbook content. When mid-course truncation or semantic incompleteness exists, the content generation modules of each stage in 5E regenerate the textbook content for that stage; this solves the problem of mid-course truncation and semantic incompleteness caused by the output length limitation of large models. This module adopts a heuristic coarse screening + LLM precise verification two-layer detection architecture, coupled with a maximum of 3 rounds of automatic retry recovery mechanism, ensuring a complete closed loop of chapter knowledge point explanations and problem-solving steps from the source. This module supports both Chinese and English detection prompts and can automatically switch verification rules according to the course language.

[0081] The truncation detection and repair submodule employs a two-level verification logic. First, it performs a low-cost, heuristic-based rapid detection on the textbook content to quickly screen and identify truncation defects. If truncation defects are found, corresponding truncation warnings are generated and injected into the textbook, and the 5E content generation modules at each stage regenerate the textbook content. If the textbook content has no truncation defects, it uses a large model to perform LLM-based refined integrity detection, identifying truncation defects and semantic incompleteness. If truncation defects or semantic incompleteness are found, corresponding truncation / incompleteness warnings are generated and injected into the textbook, and the 5E content generation modules at each stage regenerate the textbook content.

[0082] Heuristic rapid detection (pre-screening): Quickly identifies explicit truncation defects based on fixed text rules, without calling LLM, with low runtime overhead; a content truncation defect is determined to exist if any one of the rules is met. Rule A: Markdown code block markers "```" must be in odd pairs (the code is not closed); Rule B: The last line of the content contains only a series of # heading symbols and no subsequent body text; Rule C: The sentence ends abruptly without any standard Chinese or English punctuation (such as periods, exclamation marks, etc.); Rule D: Ends with incomplete markers such as "...", "not finished", or "to be continued".

[0083] If no anomalies are found during heuristic verification, the system automatically proceeds to the refined integrity check process of LLM.

[0084] LLM Refined Completeness Detection (Post-Processed Fine-tuning): After the heuristic quick detection passes, the corresponding semantic incompleteness detection prompt words are loaded, and the textbook content and the corresponding textbook outline chapter content and semantic block units are passed in. The LLM, with its professional text analysis role, judges the completeness from two dimensions: (1) Content dimension: The Enhancement stage verifies that all knowledge points in the chapter are complete and the explanation is complete; the Execution stage verifies that the entire set of problem-solving steps is complete; (2) Logic dimension: Check whether the end of the text is natural and there is no problem of interruption in the middle. The model strictly uses a fixed JSON format to truncate the incomplete prompt content (is_truncated boolean value + truncation reason).

[0085] Automatic retry and recovery process of the truncation detection and repair submodule: After the detected content is truncated, the automatic retry logic is enabled, and the global maximum number of retries is set to 3. The overall execution steps are as follows: (1) Initial generation: The content generation modules of each stage of 5E complete the generation of textbook content for the first time and obtain the textbook content; (2) Two-layer verification: Heuristic rapid detection and LLM refined integrity detection are performed on the textbook content in sequence. If a problem is found in either step, the process will be retried. (3) Optimization and retry of prompt words: Before the next round of generation, add truncated prompt content / truncated incomplete prompt content to the beginning of the prompt words in the textbook. The prompt content includes the length constraint description of language adaptation. The model is required to prioritize the integrity of core knowledge points, simplify redundant expressions, control the total number of characters within the limit, and ensure natural ending of the text. (4) Result fallback: After a total of 3 retries, even if a problem is still detected, no further retries will be made, and the content of the textbook generated in the last round will be returned directly.

[0086] This module significantly reduces the incompleteness rate of long text generation through a fault-tolerant design of graded detection and limited retries, avoiding the loss of knowledge points and the break in problem-solving steps due to exceeding output limits, and ensuring the content quality of knowledge teaching and practical operation in the 5E system.

[0087] The evaluation and quality iteration submodule is used to evaluate the content quality of the teaching materials generated by each stage of the 5E content generation module; when the content quality of the teaching materials is unqualified, the 5E content generation module will regenerate the teaching materials for that stage.

[0088] This sub-module establishes a closed-loop quality control mechanism encompassing content generation, AI-automated evaluation, and iterative optimization of low-scoring content. The control scope covers all stages of 5E teaching (stimulating interest and raising questions, exploring the essence of problems using first principles, learning the necessary knowledge and skills to solve problems, practical problem-solving, and evaluation and reflection), excluding the application of knowledge stage.

[0089] Once the teaching materials for each stage are generated, the quality verification logic is automatically triggered, invoking the evaluation and quality iteration submodule to assess the content of a specific chapter generated at that stage. The evaluation and quality iteration submodule automatically calls the large language model to conduct quantitative scoring across four dimensions: teaching logic, knowledge point accuracy, learner experience, and text language standardization. The scores for each dimension are weighted according to preset weights to calculate a total score, with a maximum score of 10 points. Optimization suggestions are also provided. The large language model returns all evaluation data in a standardized JSON format, with returned fields including the scores for each dimension, the weighted overall score, and optimization suggestions (content optimization and rectification opinions, and text logic breakdown information).

[0090] When the weighted total score is not lower than the quality qualification threshold (8.0 points), the textbook content is qualified, and the 5E content generation module continues to generate the textbook content for the next stage.

[0091] When the weighted total score is lower than the quality pass threshold (8.0 points), the teaching material content is considered unqualified. Optimization suggestions are then incorporated into the corresponding stage's prompts, and the 5E stage content generation modules are invoked to regenerate the teaching material content. The newly generated teaching material content is then evaluated again using the evaluation and quality iteration submodule. If the cumulative iterations of unqualified teaching material content reach the upper limit of evaluation rounds (3 rounds), the evaluation and optimization process is terminated, and the version with the best overall quality among the multiple rounds of generation results (the one with the highest weighted total score across three evaluations) is archived.

[0092] When the cumulative number of unsatisfactory textbook content iterations reaches the upper limit of the evaluation rounds, manual intervention and optimization can be carried out: users can independently input customized modification requests. The customized modification requests and optimization suggestions are assembled into the corresponding textbook prompts for each stage, and the 5E content generation modules for each stage are invoked to regenerate the textbook content.

[0093] The textbook output module is used to output textbook content.

[0094] The textbook output module can export textbook content into standardized, printable documents, including both synchronous DOCX export and asynchronous PPT export options. This facilitates offline review, secondary editing, and classroom teaching, enhancing the practicality and applicability of the textbook output. The entire export process completes content integrity verification, format preprocessing, document conversion and style standardization, and file storage.

[0095] 1. DOCX export is implemented by deriving DOCXExporter from the base class BaseExporter. The base class integrates general capabilities such as chapter integrity verification and document data retrieval. Before exporting, it automatically traverses the target chapters and verifies whether all six 5E stages of content in a single chapter have been generated. After successful verification, the course and chapter information is uniformly and structurally summarized. The overall process is divided into the following four steps: (1) Markdown text splicing: Remove the cover and table of contents, sort the entire text by chapter, and splice it into a complete Markdown text as the original data source for format conversion.

[0096] (2) Markdown content preprocessing: Three fixes were made to address the shortcomings of Pandoc parsing: automatic completion of missing line breaks before and after lists and tables; batch replacement of image URLs in documents, mapping the / media / path to the real file path on the local server to avoid the problem of remote image retrieval failure.

[0097] (3) Pandoc format conversion: Call the pypandoc tool with customized parameters to convert the processed Markdown into a temporary DOCX file; parameter configuration enables the conversion of LaTeX formulas into editable Word OMML formulas, the blocking of mis-parsing of YAML content, and the syntax highlighting of code blocks.

[0098] (4) Refined post-formatting of docx: Unify the global format on the basis of the original converted file: the main text uses SimSun font with Times New Roman by default, and the code characters use Consolas font; the tables are uniformly 100% page width and black single solid line border; the first line of the main text paragraph is indented by 2 characters, and the cover, table of contents, header and footer layout styles are customized separately, and the font disorder problem of code segment is specifically fixed.

[0099] The final DOCX file is returned synchronously as an attachment file stream.

[0100] 2. The PPT export adopts an asynchronous generation architecture, which generates documents offline in the background based on the Django-Q task queue. Each chapter generates a task independently, and the front end obtains the finished file by polling the task status.

[0101] (1) Content integrity verification: After the task is started, the first step is to verify whether the content of Chapter 6, Item 5E is complete. If the content is missing, the task will be terminated and rejected.

[0102] (2) AI structured PPT outline generation: summarize the text of all stages of the chapters, call the exclusive PPT prompt words, and generate the paginated outline according to the fixed JSON specification; the fixed structure of the whole PPT is cover → table of contents → stage transition page → main text of each stage → closing page, strictly constraining the number of content items, text format and content scope of each stage on a single page.

[0103] (3) PPT file filling generation: The system has a pre-set PPT template with 5 types of layouts. Relying on python-pptx, the title and placeholder content are filled in sequentially according to the layout index of cover, table of contents, transition, body and end, and the generated file is temporarily stored in the server's temporary directory.

[0104] (4) File persistence management: Read temporary PPT files, create archive records in the system storage module, clear local temporary files after they are stored in the cloud, and provide download services based on the archive number.

[0105] The final product is made available for download via a cloud-based archive address.

[0106] The workflow of this system is as follows: Step 1: The user enters a textbook description and / or uploads teaching reference materials in the course definition and user interaction module. If teaching reference materials exist, proceed to Step 2. If only a textbook description exists, proceed to Step 4.

[0107] Step 2: The reference material preprocessing and knowledge base construction module performs semantic segmentation of teaching reference materials according to the title level to obtain semantic segmentation units; extracts knowledge points from semantic segmentation units, constructs a knowledge point pool, and builds an index for semantic segmentation units to obtain semantic segmentation unit embedding vector index.

[0108] Step 3: The user selects the knowledge points they are interested in from the knowledge point list in the knowledge point pool to generate a subset of course knowledge points. Proceed to Step 5.

[0109] Step 4: The course definition and user interaction module uses AI to professionally polish and optimize the textbook description, then parses the polished textbook description, extracts the course knowledge points, and obtains a subset of course knowledge points. Proceed to Step 5.

[0110] Step 5: The outline planning module generates a textbook outline based on a subset of course knowledge points and a large language model of the outline generation role.

[0111] Step 6: The 5E content generation module generates content for each stage according to the chapter order in the textbook outline. It sequentially calls the large language model to generate textbook content using semantic chunking units, following the order of: sparking interest and posing questions; exploring the essence of problems using first principles; learning the necessary knowledge and skills to solve problems; hands-on problem-solving; evaluation and reflection; and applying knowledge. After completing all six stages of generation for each chapter, the textbook content for that chapter is obtained. After completing the generation tasks for all chapters, the entire course textbook content is compiled.

[0112] For the stages of stimulating interest and raising questions, exploring the essence of the problem using first principles, learning the necessary knowledge and skills to solve the problem, solving the problem by doing, and evaluation and reflection, after the textbook content is generated, it is input into the evaluation and quality iteration submodule, which evaluates the quality of the textbook content. If it is unqualified, the textbook content for that stage is regenerated; if it is qualified, the textbook content for the next stage is generated.

[0113] Furthermore, for the stages of learning the necessary knowledge and skills to solve problems and the practical problem-solving stages, before executing the textbook content generation task, a multi-round progressive RAG retrieval submodule is first invoked. The main content descriptions and knowledge points of each chapter are used as retrieval text to search semantic chunking units. The resulting candidate fragments are then injected into textbook prompts before executing the textbook content generation task. After generating the textbook content, it is first input into the truncation detection and repair submodule for verification of truncation and semantic incompleteness. If truncation or semantic incompleteness exists, the textbook content generation for that stage is restarted; if no truncation or semantic incompleteness exists, the textbook content is input into the evaluation and quality iteration submodule.

[0114] Step 7: The textbook output module exports the textbook content as a standardized document that can be printed directly, either DOCX or PPT.

[0115] This invention is the first to organically integrate the closed-loop logic of the 5E teaching paradigm, the POT-OBE educational concept, the multi-round progressive retrieval enhancement generation mechanism, and the automated quality assurance system, constructing an intelligent generation method for creating textbooks based on the 5E teaching paradigm. This method not only improves the logic, professionalism, and completeness of the generated textbook content, but also enhances the textbook's human-computer interaction capabilities and teaching adaptability, demonstrating significant technological advancement and practical application value.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A 5E paradigm textbook construction system based on large model and retrieval enhancement, characterized in that, include: The course definition and user interaction module is used to obtain textbook descriptions and / or teaching reference materials; Generate a subset of course knowledge points based on textbook descriptions and a knowledge point pool; The reference material preprocessing and knowledge base construction module is used to semantically segment teaching reference materials according to title level to obtain semantic segmentation units; and to extract knowledge points from semantic segmentation units to build a knowledge point pool. The outline planning module is used to generate a textbook outline based on a subset of course knowledge points using a large language model. Each chapter of the textbook outline includes a title, a description of the main content, a list of knowledge points, and prerequisite knowledge points. The 5E content generation module is used to generate textbook content for each chapter of the textbook outline in the following order: the stage of arousing interest and raising questions, the stage of exploring the essence of the problem using first principles, the stage of learning the knowledge and skills necessary to solve the problem, the stage of solving the problem by doing practical work, the stage of evaluation and reflection, and the stage of applying what has been learned. It calls the large language model to generate textbook content using semantic block units. The quality assurance module group is used to retrieve semantic block units based on the content of each chapter, use the retrieval results to constrain the generation of textbook content, evaluate the content quality of the textbook content, and detect whether there are any interruptions or semantic omissions in the textbook content. When the content quality assessment is unqualified or there are interruptions or semantic omissions, the content generation modules of each stage of 5E will regenerate the textbook content of that stage. The quality assurance module group includes a multi-round progressive RAG retrieval submodule, a truncation detection and repair submodule, and an evaluation and quality iteration submodule; among them, The multi-round progressive RAG retrieval submodule is used to retrieve semantic block units based on the content of each chapter in the stages of learning the knowledge and skills necessary to solve problems and in the actual hands-on problem-solving stage, and to use the retrieval results to constrain the generation of textbook content. The truncation detection and repair submodule is used to evaluate the mid-term truncation and semantic incompleteness of the teaching material in the stage of learning the knowledge and skills necessary to solve the problem and the stage of actual hands-on problem solving. When the teaching material has mid-term truncation and semantic incompleteness, the teaching material content of each stage is regenerated by the 5E stage content generation module. The evaluation and quality iteration submodule is used to evaluate the content quality of the teaching materials in the stages of stimulating interest and raising questions, exploring the essence of the problem using first principles, learning the knowledge and skills necessary to solve the problem, solving the problem by doing, and evaluation and reflection. When the content quality of the teaching materials is not up to standard, the content generation module of each stage of 5E will regenerate the teaching materials for that stage. The textbook output module is used to output textbook content.

2. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 1, characterized in that, The course definition and user interaction module must obtain at least one of the following: textbook description and teaching reference materials; The textbook description includes the knowledge points, teaching objectives, and teaching requirements.

3. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 1, characterized in that, The teaching reference materials are semantically segmented according to heading levels. The specific process for obtaining semantic segmentation units is as follows: The teaching reference materials are parsed to obtain Markdown text content and JSON structured data with heading hierarchy information; Traverse the JSON structured data in document node order to extract the original units; or split the Markdown text content according to heading regular expressions and blank line paragraph rules to obtain the original units. Based on the number of tokens, the original units are merged and split to obtain semantic block units.

4. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 1, characterized in that, The specific process for extracting knowledge points from semantic chunks and building a knowledge point pool is as follows: Semantic blocks are grouped according to structured rules to form chapter units with semantic integrity; Candidate knowledge points are extracted from chapter units; The candidate knowledge points are semantically merged and deduplicated to obtain a knowledge point pool.

5. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 1, characterized in that, When generating teaching materials in the downstream stage, all content generated in the upstream stage will be used as context input. By using pre-set textbook prompts, the generated content includes multiple rounds of dialogue and interaction between learners and learning partners in the stages of exploring the essence of a problem using first principles and learning the knowledge and skills necessary to solve the problem.

6. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 1, characterized in that, The workflow of the multi-round progressive RAG retrieval submodule is as follows: Using the main content descriptions and knowledge points of the chapters as search text, dense vector search and BM25 sparse keyword search were performed on the semantic block units respectively to obtain two preliminary search results. The RRF reciprocal ranking fusion algorithm is used to fuse the two preliminary retrieval results, and the fusion score of the candidate fragments in the preliminary retrieval results is calculated. Candidate segments are sorted based on their fusion scores. The top-scoring candidate segments are then injected into prompts in the content generation modules of each stage of 5E to constrain them to regenerate textbook content.

7. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 6, characterized in that, The multi-round progressive RAG search submodule adopts a sharding multi-round search strategy, specifically: splitting the main content descriptions and knowledge points of chapters according to fixed sharding criteria; For each subquery text obtained from the splitting, add targeted keywords based on the current stage, and then enable dense vector retrieval and BM25 sparse keyword retrieval in parallel; All candidate fragments obtained from multiple rounds of retrieval are summarized, and duplicate and redundant content is removed to obtain two preliminary retrieval results.

8. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 1, characterized in that, The workflow of the truncation detection and repair submodule is as follows: Conduct heuristic rapid testing of teaching materials to identify truncation defects; When there is a truncation defect in the teaching material, the corresponding truncation prompt is generated and injected into the prompt words of the content generation module of each stage of 5E, so that the teaching material content can be regenerated. When there are no truncation defects in the teaching materials, perform LLM refined integrity detection on the teaching materials to identify truncation defects and semantic incompleteness; When the teaching material has truncation defects and semantic incompleteness, corresponding truncation and incompleteness prompts are generated and injected into the prompt words of the content generation modules of each stage of 5E, which then complete the regeneration of the teaching material content.

9. The 5E paradigm textbook construction system based on large model and retrieval enhancement as described in claim 1, characterized in that, The assessment and quality iteration submodule evaluates the content quality of the teaching materials from the dimensions of teaching logic, accuracy of knowledge points, learner experience, and standardization of text language.