Teaching content generation method and system based on large model instruction and context knowledge dual drive
By using a teaching content generation method driven by both large model instructions and contextual knowledge, the problem of insufficient integration of professional knowledge in the field of education and poor multi-dimensional adaptability in existing technologies has been solved. This method enables the precise and personalized generation of teaching content, thereby improving teaching quality and efficiency.
Patent Information
- Application Number
- CN202511463586.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-13
AI Technical Summary
Existing teaching content generation technologies lack in-depth integration of professional knowledge in the field of education, making it difficult to accurately adapt to multi-dimensional teaching scenarios. The generated content is not scientific or practical enough, resulting in a heavy workload for teachers.
We adopt a teaching content generation method driven by both large model instructions and contextual knowledge. By preprocessing, structuring, and vectorizing knowledge in the education field, and combining multi-dimensional conditional instructions and contextual knowledge, we generate teaching content that conforms to educational principles and subject requirements, and then perform post-processing optimization.
This approach enables the precise and personalized generation of teaching content, improves its scientific rigor and educational applicability, reduces teachers' lesson preparation burden, and enhances the practicality and efficiency of teaching content.
Smart Images

Figure CN121328684A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of educational artificial intelligence and educational big data processing technology, and in particular to a method and system for generating teaching content based on a dual-drive approach of large model instructions and contextual knowledge. Background Technology
[0002] With the rapid development of artificial intelligence technology, the education sector is actively exploring intelligent teaching tools to improve teaching efficiency and quality. In the fields of educational AI and educational big data processing, content generation technology has become a key area of research and application. Currently, most mainstream content generation technologies rely on a single large language model, using simple instructions to drive the model and generate teaching content. However, this existing technology has many drawbacks and limitations.
[0003] On the one hand, existing technologies lack in-depth integration and application of professional knowledge in the field of education. They fail to fully integrate core educational resources such as curriculum standards and teaching syllabi, making it easy for large language models to become disconnected from the subject knowledge system when generating teaching content. This makes it difficult to accurately grasp teaching objectives and meet the cognitive development patterns and learning needs of students at different educational stages. For example, when generating explanations of knowledge points in a specific subject, a lack of accurate understanding of subject-specific terminology and knowledge structures may lead to imprecise and inaccurate content, reducing the scientific rigor and educational applicability of the teaching content.
[0004] On the other hand, existing methods for generating teaching content are inadequate in addressing the complexity and diversity of teaching scenarios. Their constraints on content generation are relatively simplistic, making it difficult to accurately and flexibly generate content tailored to specific teaching needs based on multiple dimensions such as subject, grade level, grade, lesson type, teaching theme, teaching model, and teaching stage. For example, when designing interactive classroom questions for different grade levels, existing technologies cannot effectively combine the knowledge level and thinking characteristics of students at each level. The generated questions may exceed or fall below students' comprehension abilities, failing to achieve the expected interactive teaching effects.
[0005] Furthermore, existing technologies have shortcomings in the post-generation processing of teaching content. The lack of a robust post-processing mechanism makes it impossible to standardize the format of the generated content, correct issues such as vague descriptions, logical inconsistencies, or redundant repetitions, and optimize the information presentation structure according to the type of teaching content. This results in the generated teaching content being difficult to apply directly to actual teaching scenarios, requiring teachers to perform extensive secondary processing and modifications, increasing their workload and reducing the efficiency and practicality of teaching content generation. Summary of the Invention
[0006] The purpose of this invention is to provide a teaching content generation method and system based on a dual-driven approach of large model instructions and contextual knowledge. This aims to address issues in existing technologies such as insufficient integration of teaching content generation with professional knowledge in the education field, poor adaptability to multi-dimensional teaching scenarios, and weak practicality of generated results. By combining large model instruction-driven and contextual knowledge-driven approaches, the invention achieves precise and personalized generation of teaching content, ensuring that the generated content conforms to educational principles, subject curriculum standards, and students' cognitive characteristics. This effectively improves the scientific rigor, accuracy, and educational applicability of teaching content, while reducing teachers' lesson preparation burden. It also provides flexible and adaptable content solutions for different teaching scenarios, promoting the efficient implementation of smart teaching.
[0007] The specific technical solution for achieving the objective of this invention is as follows:
[0008] A method for generating teaching content based on a dual-drive approach of large model instructions and contextual knowledge, the method comprising the following specific steps:
[0009] Step 1: Knowledge Preprocessing in the Educational Field
[0010] A1: Knowledge Document Analysis
[0011] This method involves structuring raw knowledge documents in the education field, including curriculum standards, teaching outlines, teaching plans, and teaching handouts. First, the document format type is identified. Then, key elements are extracted, and information units are analyzed and extracted. Finally, noise filtering is applied to these information units to remove redundant formatting tags, irrelevant comments, and erroneous coding content, resulting in a structured set of knowledge documents. The format types include PDF documents, DOC documents, DOCX documents, PPT presentations, PPTX presentations, XLS worksheets, XLSX worksheets, text documents, and Markdown documents. The key elements include titles, paragraphs, tables, figures, and formulas. The information units include knowledge points, teaching content, teaching objectives, teaching requirements, key teaching points and difficulties, and subject literacy requirements.
[0012] A2: Document content slicing
[0013] The content in the knowledge document collection is segmented. By analyzing the logical relationships and thematic consistency between sentences, continuous text is divided into several semantically complete knowledge units. For long paragraphs, further subdivision is performed based on the knowledge granularity in the education field to obtain knowledge document slice units. Each slice unit contains independent and complete knowledge information while maintaining semantic relevance to the context. The logical relationships include causal, progressive, and parallel relationships. The knowledge granularity includes concept definitions, principle explanations, case analyses, and exercise solutions.
[0014] A3: Content Text Vectorization
[0015] Each knowledge document slice unit is encoded using a pre-trained language model, transforming textual knowledge into a low-dimensional, dense vector representation. During the encoding process, the model is fine-tuned by incorporating a professional vocabulary in the education field to enhance the vector's ability to represent subject-specific terms and semantics. Finally, each knowledge document slice unit is converted into a corresponding feature vector, forming a feature library containing all knowledge document slice units, providing a foundation for subsequent contextual knowledge retrieval and matching.
[0016] Step 2: Define the prompt template for generating teaching content
[0017] B1: Define the large model instructions and their variables. The large model instructions include meta-instructions, task instructions, conditional instructions, output format instructions, and extended instructions. Among them, task instructions, conditional instructions, and output format instructions contain several fixed variables, forming a teaching content generation prompt template containing the large language model instructions and their variables; specifically:
[0018] B11: Meta-instruction, defines the core principles for generating teaching content from large models, clarifies that teaching content must conform to educational laws, subject curriculum standards, and the laws of students' cognitive development, and emphasizes the scientific nature, accuracy, and educational applicability of the content;
[0019] B12: Task instructions, which include variables for generating task types from teaching content, specifying the task types for generating specific teaching content, including course introduction design, knowledge point explanation writing, classroom interaction question design, teacher lesson plan generation, after-class exercise creation, and teaching evaluation language generation;
[0020] B13: Conditional instructions, which include variables such as subject, grade level, grade, lesson type, teaching theme, teaching mode, and teaching stage. They can specify the constraints for generating teaching content through combinations of multiple dimensions of conditions, thereby enhancing the accuracy of teaching content generation.
[0021] B14: Output format instructions, including text structure, length range, special elements, and format specification variables, specifying the presentation format of the generated teaching content; the text structure includes point-by-point discussion, paragraph format, and question-and-answer format; the length range includes word count range and number of paragraphs; the special elements include keyword annotation, prompts for inserting charts and graphs, and interactive element markers; the format specifications include heading levels, symbol usage, and requirements for the expression of professional terminology;
[0022] B15: Extended instructions provide supplementary guidance to meet the personalized and innovative needs of teaching content, including requirements for interdisciplinary integration and the embedding of contemporary elements; the requirements for interdisciplinary integration include incorporating engineering practice cases into mathematics teaching and connecting historical background analysis into Chinese language teaching; the embedding of contemporary elements includes explaining knowledge points by combining the latest scientific and technological achievements and social hot topics.
[0023] B2: Define a contextual knowledge content block, which includes knowledge points, teaching content, teaching objectives, teaching requirements, key teaching points and difficulties, and subject literacy requirements. This block, combined with the aforementioned large model instructions and their variables, forms a complete teaching content generation prompt template; specifically:
[0024] B21: Knowledge points, extracting core concepts, formulas, theorems, and basic units from the curriculum standards and teaching materials, and marking the logical hierarchy;
[0025] B22: Teaching content, integrating the knowledge modules in the syllabus with the case materials in the lecture notes to form structured teaching materials;
[0026] B23: Teaching objectives, which align with the competency indicators in the curriculum standards and are further refined into quantifiable behavioral verb expressions;
[0027] B24: Teaching requirements, and the establishment of differentiated implementation guidelines based on student learning analysis and teaching models;
[0028] B25: Teaching focus and difficulties, and defining breakthrough paths based on the difficulty of knowledge cognition and students' common mistakes;
[0029] B26: Subject literacy requirements, transforming the core literacy framework into key points for integration into specific teaching activities;
[0030] Step 3: Workflow Arrangement for Generating Teaching Content
[0031] C1: Instruction-driven logic for building large models
[0032] Based on the complete teaching content generation prompt template described in step 2, a generation logic chain centered on instructions is constructed. By parsing meta-instructions, the fundamental principles of content generation are established. Specific generation goals are clarified according to task instructions. Multi-dimensional variables in conditional instructions are used to precisely constrain the contextual knowledge retrieval and teaching content generation direction. At the same time, the presentation format of the content is standardized according to output format instructions. Personalized and innovative requirements are injected through extended instructions, forming a complete instruction transmission path from core principles to specific execution details. The multi-dimensional variables include subject, grade level, and teaching theme.
[0033] The precise constraint is implemented by scalar filtering, that is, based on the set multi-dimensional variables, the first set of candidate knowledge document slice units that meet the conditions is selected to reduce the retrieval space of subsequent context knowledge and improve retrieval efficiency.
[0034] C2: Constructing contextual knowledge-driven logic
[0035] A knowledge retrieval mechanism is established based on the feature library containing all knowledge document slice units obtained in step 1 through preprocessing. By analyzing the correlation between the teaching content generation task and the contextual knowledge content block, the knowledge document slice unit with the highest matching degree is retrieved from the feature library and integrated as background knowledge into the generation process. This achieves the invocation of knowledge points, alignment with teaching objectives, grasp of teaching key points and difficulties, and reflection of subject literacy requirements. At the same time, the semantic relationship between knowledge units is maintained, so that the generated teaching content is not only closely aligned with professional knowledge in the field of education, but also deeply integrated with specific teaching scenarios, enhancing the scientificity and applicability of the content. The contextual knowledge content block includes knowledge points, teaching objectives, and teaching key points and difficulties.
[0036] The step of retrieving the knowledge document slice unit with the highest matching degree from the feature library specifically includes:
[0037] C21: Semantic retrieval based on dense vectors. The cosine similarity is calculated between the feature vector of the teaching topic and the feature vector of the knowledge document slice unit in the knowledge base. The knowledge document slice unit with the highest similarity is obtained and used as the second set of candidate knowledge document slice units with the highest matching degree with the current teaching content generation task.
[0038] C22: Keyword retrieval based on sparse vectors. The BM25 algorithm is used to calculate the knowledge document slice unit with the highest similarity to the feature vector of the teaching topic, which is used as the third set of candidate knowledge document slice units with the highest matching degree with the current teaching content generation task.
[0039] C23: Hybrid retrieval based on reordering. The RRF (Reciprocal Rank Fusion) ranking fusion algorithm is used to reorder the second set of candidate knowledge document slice units and the third set of candidate knowledge document slice units to obtain a set of target knowledge document slice units that takes into account both semantic relevance and keyword accuracy. Finally, it is filled into the context knowledge content block.
[0040] Step 4: Implementation of Teaching Content Generation and Post-processing of Results
[0041] D1: Dual-Driven Teaching Content Generation and Implementation
[0042] After the aforementioned large-scale model instruction-driven logic and the aforementioned context knowledge-driven logic are completed, the results of the two are merged and sent to the downstream large language model instance. After analysis and processing, the large language model generates initial teaching content. During the generation process, the large language model simultaneously receives the constraints and context knowledge content in the instruction chain. Through a deep understanding of the teaching objectives, subject requirements, and knowledge connections, it integrates the standardized instructions with educational knowledge to form an initial teaching content output that meets the needs of the teaching scenario. This ensures that the generation direction aligns with the preset task requirements and that the accuracy of knowledge references and the coherence of the teaching logic are both guaranteed.
[0043] D2: Post-processing of teaching content generation results
[0044] Post-processing of the initial teaching content includes format revision, content polishing, and layout adjustment, specifically:
[0045] D21: Formatting revision mainly involves addressing formatting flaws in the original teaching content, standardizing the format of text content and repairing its structure. For Markdown formatted content, it involves checking and fixing heading levels, list indentation, code block identifiers, and link reference formatting. For JSON formatted content, it involves checking and fixing missing quotes, mismatched brackets, and key-value pair formatting errors.
[0046] D22: Content polishing, mainly correcting vague expressions, logical gaps, or redundant repetitions in the original teaching content, and supplementing the standardized expressions of subject terminology, so that the text not only conforms to the professional context of the education field, but also matches the cognitive and comprehension level of students in the target grade level.
[0047] D23: Layout adjustment. Optimize the information presentation structure according to the type of teaching content. Enhance the organization and ease of use of teaching content by dividing it into sections, setting up labels, and adjusting the content density.
[0048] A teaching content generation system based on a dual-drive approach of large model instructions and contextual knowledge, implementing the above method, includes: an education domain knowledge preprocessing module, a teaching content generation prompt template definition module, a teaching content generation workflow orchestration module, and a teaching content generation implementation and result post-processing module. Users perform knowledge document parsing, document content slicing, and content text vectorization operations in the education domain knowledge preprocessing module through human-computer collaborative interaction; define large model instructions and their variables and define contextual knowledge content blocks in the teaching content generation prompt template definition module; construct large model instruction-driven logic and construct contextual knowledge-driven logic in the teaching content generation workflow orchestration module; and perform dual-drive teaching content generation implementation and teaching content generation result post-processing operations in the teaching content generation implementation and result post-processing module.
[0049] This invention achieves its goals by constructing a dual-driven mechanism of large-scale model instructions and contextual knowledge, combining four core steps: knowledge preprocessing in the education domain, definition of prompt templates for generating teaching content, workflow arrangement for teaching content generation, and post-processing of results. First, the original knowledge documents in the education domain are parsed, sliced, and vectorized to form a structured knowledge document set and feature library. Second, prompt templates containing various types of instructions and variables, such as meta-instructions and task instructions, are defined, and contextual knowledge content blocks such as knowledge points and teaching objectives are clearly defined. Next, instruction-driven and knowledge-driven logics are arranged, and the retrieval space is constrained by scalar filtering. Dense vector semantic retrieval, sparse vector keyword retrieval, and hybrid reordering algorithms are used to obtain matching knowledge units. Finally, the results of the dual-driven logic are integrated into a large language model to generate initial content, which is then post-processed after format revision, content polishing, and layout adjustment. The corresponding system includes modules for knowledge preprocessing in the education domain, prompt template definition, workflow arrangement, generation implementation, and post-processing. It supports human-computer collaborative interaction and achieves precise teaching content generation that conforms to educational principles and teaching needs.
[0050] In this invention, the large language model serves as the core generation engine, responsible for receiving compound prompt words composed of "heterogeneous contextual knowledge" and "structured generation instructions," and generating high-quality lesson plans and teaching designs that conform to specific teaching models, curriculum standards, and subject literacy requirements.
[0051] This invention constructs a dual-drive prompting framework of "heterogeneous contextual knowledge + structured generation instructions". This framework combines unstructured educational professional knowledge with highly structured generation logic, ensuring that the large language model can generate educational professional texts that are goal-oriented, logically rigorous, content-accurate, and conform to specific teaching paradigms.
[0052] In this invention, knowledge base construction and retrieval technologies are used to achieve the dynamic assembly of "heterogeneous contextual knowledge". When a user selects a specific textbook topic, the system uses retrieval technology to accurately extract heterogeneous information fragments related to that topic from the knowledge base, such as core concepts, knowledge points, teaching objectives, and competency requirements, and organizes them into logically coherent contextual modules.
[0053] The method of this invention guides a large language model to generate high-quality lesson plans and instructional designs that meet specific teaching models, curriculum standards and subject literacy requirements through the deep integration of "heterogeneous contextual knowledge" and "structured generation instructions".
[0054] This invention innovatively constructs a dual-drive prompting engineering system of "heterogeneous contextual knowledge + structured generation instructions," achieving deep integration of diverse educational elements such as teaching models, textbook themes, curriculum standards, and subject literacy requirements with the generated content. This significantly enhances the professionalism, structuring, and intelligence of the large language model in educational scenarios. Through the structured processing and dynamic organization of heterogeneous contextual knowledge, the system can accurately match teaching needs and generate lesson plans and instructional designs that conform to teaching logic and subject characteristics. The structured generation instruction framework possesses high reusability and scalability, flexibly adapting to different teaching models and subject scenarios.
[0055] The teaching content generation method and system proposed in this invention, driven by both large model instructions and contextual knowledge, effectively overcomes the shortcomings and limitations of existing technologies through in-depth preprocessing of knowledge in the education field, scientific definition of teaching content generation prompt templates, rigorous arrangement of teaching content generation workflow, and comprehensive implementation and post-processing of teaching content generation results. It significantly improves the accuracy, scientificity, and practicality of teaching content generation, providing a higher-quality and more efficient intelligent solution for education and teaching. Attached Figure Description
[0056] Figure 1 This is a flowchart of the method of the present invention;
[0057] Figure 2 This is a system structure diagram of the present invention. Detailed Implementation
[0058] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0059] like Figure 1 As shown, the present invention provides a teaching content generation method based on a dual-drive approach of large model instructions and contextual knowledge, comprising the following specific steps:
[0060] Step 1: Knowledge preprocessing in the education field;
[0061] Step 2: Define the prompt template for generating teaching content;
[0062] Step 3: Workflow arrangement for generating teaching content;
[0063] Step 4: Generation and implementation of teaching content and post-processing of results.
[0064] like Figure 2 As shown, the core of the teaching content generation system based on the dual-drive of large model instructions and contextual knowledge, which implements the above method, lies in the implementation of the "dual-drive prompting engineering" method. The specific implementation methods of its three key points will be elaborated below:
[0065] A. Methods for constructing and organizing "heterogeneous contextual knowledge"
[0066] 1) Structured Preprocessing of Educational Knowledge: Transforming unstructured educational knowledge into structured data that is machine-understandable and efficiently retrieved. Using natural language processing technology, key information units such as core concepts, knowledge points, teaching objectives, key difficulties, and competency requirements are automatically identified and extracted from massive amounts of text. These extracted key information units are then converted into vectors using an embedding model and stored in a vector database, constructing a structured educational knowledge base. This lays a solid data foundation for subsequent accurate knowledge retrieval and context construction.
[0067] 2) Semantic retrieval and knowledge extraction based on user topics. When a user selects a specific "textbook topic" through the teaching and research intelligent agent, the system initiates the following process:
[0068] Topic Analysis: The "textbook topic" selected by the user is analyzed to generate its core word vectors, and the search terms are expanded using a thesaurus.
[0069] Textbook Knowledge Base Retrieval: The system employs a hybrid retrieval strategy within the textbook knowledge base to accurately locate core concepts, definitions, and other knowledge points directly related to the topic. This strategy executes two retrieval paths in parallel: first, semantic retrieval (dense vector retrieval), which calculates cosine similarity between the parsed topic word vectors and the dense vectors of all knowledge fragments in the knowledge base to capture deep semantic connections; second, keyword retrieval (sparse vector retrieval), which uses the BM25 algorithm to precisely match keywords within the topic, efficiently filtering out knowledge fragments containing core terms. Finally, the system uses RRF (Reciprocal Rank Fusion) ranking fusion technology to intelligently reorder the preliminary results of the two retrieval methods, generating a comprehensive result list that balances semantic relevance and keyword accuracy.
[0070] Curriculum Standards Knowledge Base Retrieval: For the curriculum standards knowledge base, the system first performs scalar filtering, quickly selecting a range of curriculum standard documents that meet the user's preset metadata ("subject", "grade"), thus narrowing down the search space. Within this range, the system then employs a hybrid retrieval mechanism, combining semantic retrieval (dense vectors) and keyword retrieval (sparse vectors / BM25) to accurately extract key information such as teaching objectives and content requirements related to the selected curriculum standards.
[0071] Subject literacy knowledge base retrieval: This step also employs a hybrid retrieval mode. The system applies the queries after topic parsing to both the dense and sparse vector indexes of the subject literacy knowledge base. Semantic retrieval uncovers implicit subject literacy requirements related to topic-based teaching activities; keyword retrieval directly targets standard terminology within literacy entries. The two retrieval results are then reordered using RRF (Reference Randomization) to ensure that the final returned subject literacy requirements are both comprehensive and highly relevant.
[0072] 3) Logical Assembly of Heterogeneous Knowledge: Heterogeneous knowledge fragments related to the selected topic retrieved from the structured knowledge base are structurally integrated through variable filling. Based on the semantic relevance, importance, and teaching objectives of the knowledge fragments, the system matches and fills the retrieved heterogeneous knowledge fragments into the corresponding variable positions. Simultaneously, a preset rule engine performs logical connection and sequence optimization on the filled knowledge nodes, ultimately forming a complete and rich "heterogeneous contextual knowledge package." This package serves as the core input, providing a comprehensive, accurate, and pedagogically intelligent contextual background for subsequent large language model prompting.
[0073] B. Instruction framework design and implementation of "structured instruction generation"
[0074] 1) The instruction framework adopts a three-layer structure of "meta-instruction - task instruction - format instruction" to ensure the explicitness, reusability, and scalability of instructions. The three-layer structure is as follows:
[0075] Meta-instructions: Located at the very beginning of the prompt words, they are used to set the role, task tone, and core constraints of the large language model.
[0076] Task instructions: Clearly specify the specific type and core requirements of the task to be generated.
[0077] Formatting instructions: Strictly define the format, style, and length of the output content to ensure the standardization and consistency of the generated results.
[0078] 2) Dynamic assembly of instructions: Based on the user's selection in the interactive interface (subject, grade, course type, teaching mode, teaching stage, etc.), the system automatically selects the corresponding instruction modules from the instruction library and dynamically splices them to form the final "structured generated instructions".
[0079] C. Integration Mechanism of the "Dual-Driven Hint Engineering" Method
[0080] 1) Constructing a composite prompt word: This involves organically integrating a "heterogeneous contextual knowledge package" with "structured generation instructions" to construct a composite prompt word that guides the large language model in generating precise content. This prompt word accurately guides the large language model to generate high-quality, highly relevant lesson plan content based on the provided professional knowledge base, while adhering to specific teaching logic and format specifications.
[0081] 2) Large Language Model-Driven Generation: The system takes the constructed compound prompt words as input and sends them to a pre-selected or fine-tuned large language model. Upon receiving the prompt words, the large language model's internal network architecture simultaneously focuses on both "structured generation instructions" and "heterogeneous contextual knowledge." The instruction part provides strong logical constraints and directional guidance for the model's generation process, ensuring that its output does not deviate from professional teaching standards and specific model requirements. The contextual knowledge part provides the model with a solid factual foundation and content materials, ensuring that each part of its generation is based on evidence and rich in content.
[0082] 3) Post-processing and User Presentation: The raw text generated by the model may contain minor formatting flaws or redundant information. The system standardizes the format and repairs the structure of the text content. For Markdown output, the system automatically checks and corrects formatting issues such as heading levels, list indentation, code block identifiers, and link references to ensure readability and consistency. For JSON content containing structured data, the system automatically repairs it by calling the jsonrepair tool, handling common syntax problems such as missing quotes, mismatched brackets, and incorrect key-value pair formats. Through the system's post-processing, the integrity and usability of the data are guaranteed, and the final generated lesson plan content is presented to the user through the user interface.
[0083] 1. The features of this invention are: a structured preprocessing method for knowledge in the field of education: including parsing original documents in the field of education (curriculum standards, teaching outlines, etc.) to extract key information units, slicing documents based on semantic logic and knowledge granularity, and combining a domain-specific glossary to fine-tune the pre-trained model to achieve content text vectorization, forming a complete technical solution for an efficient retrieval knowledge feature database.
[0084] 2. Construction method of dual-drive prompt template: It consists of a large model instruction system that includes meta-instructions, task instructions (including variable for generating task type), conditional instructions (including multi-dimensional variables such as subject / grade level), output format instructions (including variables such as structure / length), and extension instructions. This system is combined with contextual knowledge content blocks such as knowledge points, teaching objectives, and key points and difficulties of teaching to form a specific structure of prompt template that is both restrictive and knowledge-supporting.
[0085] 3. A dual-driven workflow for generating teaching content based on instructions and knowledge: This workflow achieves precise constraints driven by instructions through scalar filtering, and precise matching driven by knowledge through a hybrid retrieval process combining semantic retrieval, keyword retrieval, and RRF re-ranking. The content is then integrated with these two methods and input into a large model. Following post-processing steps such as format revision, content polishing, and layout adjustment, a complete generation chain is formed. This invention, through the collaborative design of dual-driven instructions and contextual knowledge, solves the core problems of ambiguity in instructions leading to directional deviations and knowledge gaps leading to unprofessional content in traditional teaching content generation. This invention constructs a professional knowledge feature database through knowledge preprocessing in the education field, proposes a hybrid retrieval mechanism to ensure that the generated content accurately references authoritative knowledge such as curriculum standards and teaching outlines, and strictly aligns with teaching objectives, key points, and subject literacy requirements. The multi-dimensional conditional variables and scalar filtering mechanism in the prompt template can precisely constrain the generation direction, meeting the personalized needs of different teaching scenarios. Simultaneously, the extended instructions support interdisciplinary integration and the embedding of contemporary elements, enhancing content innovation. The instruction-driven logical chain clarifies the generation criteria and execution details, while the knowledge-driven retrieval mechanism narrows the matching range; the dual-drive synergy reduces invalid generation from large models. Post-processing, including format revision, content polishing, and layout adjustments, further ensures the professionalism, standardization, and usability of the output content.
Claims
1. A method for generating teaching content based on a dual-driven approach of large-scale model instructions and contextual knowledge, characterized in that, The method includes the following specific steps: Step 1: Knowledge Preprocessing in the Educational Field A1: Knowledge Document Analysis This method involves structuring raw knowledge documents in the education field, including curriculum standards, teaching outlines, teaching plans, and teaching handouts. First, the document format type is identified. Then, key elements are extracted, and information units are analyzed and extracted. Finally, noise filtering is applied to these information units to remove redundant formatting tags, irrelevant comments, and erroneous coding content, resulting in a structured set of knowledge documents. The format types include PDF documents, DOC documents, DOCX documents, PPT presentations, PPTX presentations, XLS worksheets, XLSX worksheets, text documents, and Markdown documents. The key elements include titles, paragraphs, tables, figures, and formulas. The information units include knowledge points, teaching content, teaching objectives, teaching requirements, key teaching points and difficulties, and subject literacy requirements. A2: Document content slicing The content in the knowledge document collection is segmented. By analyzing the logical relationships and thematic consistency between sentences, continuous text is divided into several semantically complete knowledge units. For long paragraphs, further subdivision is performed based on the knowledge granularity in the education field to obtain knowledge document slice units. Each slice unit contains independent and complete knowledge information while maintaining semantic relevance to the context. The logical relationships include causal, progressive, and parallel relationships. The knowledge granularity includes concept definitions, principle explanations, case analyses, and exercise solutions. A3: Content Text Vectorization Each knowledge document slice unit is encoded using a pre-trained language model, transforming textual knowledge into a low-dimensional, dense vector representation. During the encoding process, the model is fine-tuned by incorporating a professional vocabulary list in the field of education to enhance the vector's ability to represent subject terms and semantics. Finally, each knowledge document slice unit is converted into a corresponding feature vector, forming a feature library containing all knowledge document slice units, which provides a foundation for subsequent contextual knowledge retrieval and matching. Step 2: Define the prompt template for generating teaching content B1: Define the large model instructions and their variables. The large model instructions include meta-instructions, task instructions, conditional instructions, output format instructions, and extended instructions. Among them, task instructions, conditional instructions, and output format instructions contain several fixed variables, forming a teaching content generation prompt template containing the large language model instructions and their variables; specifically: B11: Meta-instruction, defines the core principles for generating teaching content from large models, clarifies that teaching content must conform to educational laws, subject curriculum standards, and the laws of students' cognitive development, and emphasizes the scientific nature, accuracy, and educational applicability of the content; B12: Task instructions, which include variables for generating task types from teaching content, specifying the task types for generating specific teaching content, including course introduction design, knowledge point explanation writing, classroom interaction question design, teacher lesson plan generation, after-class exercise creation, and teaching evaluation language generation; B13: Conditional instructions, which include variables such as subject, grade level, grade, lesson type, teaching theme, teaching mode, and teaching stage. They can specify the constraints for generating teaching content through combinations of multiple dimensions of conditions, thereby enhancing the accuracy of teaching content generation. B14: Output format instructions, including text structure, length range, special elements, and format specification variables, specifying the presentation format of the generated teaching content; the text structure includes point-by-point discussion, paragraph format, and question-and-answer format; the length range includes word count range and number of paragraphs; the special elements include keyword annotation, prompts for inserting charts and graphs, and interactive element markers; the format specifications include heading levels, symbol usage, and requirements for the expression of professional terminology; B15: Extended instructions provide supplementary guidance to meet the personalized and innovative needs of teaching content, including requirements for interdisciplinary integration and the embedding of contemporary elements; the requirements for interdisciplinary integration include incorporating engineering practice cases into mathematics teaching and connecting historical background analysis into Chinese language teaching; the embedding of contemporary elements includes explaining knowledge points by combining the latest scientific and technological achievements and social hot topics. B2: Define a contextual knowledge content block, which includes knowledge points, teaching content, teaching objectives, teaching requirements, key teaching points and difficulties, and subject literacy requirements. This block, combined with the aforementioned large model instructions and their variables, forms a complete teaching content generation prompt template; specifically: B21: Knowledge points, extracting core concepts, formulas, theorems, and basic units from the curriculum standards and teaching materials, and marking the logical hierarchy; B22: Teaching content, integrating the knowledge modules in the syllabus with the case materials in the lecture notes to form structured teaching materials; B23: Teaching objectives, which align with the competency indicators in the curriculum standards and are further refined into quantifiable behavioral verb expressions; B24: Teaching requirements, and the establishment of differentiated implementation guidelines based on student learning analysis and teaching models; B25: Teaching focus and difficulties, and defining breakthrough paths based on the difficulty of knowledge cognition and students' common mistakes; B26: Subject literacy requirements, transforming the core literacy framework into key points for integration into specific teaching activities; Step 3: Workflow Arrangement for Generating Teaching Content C1: Instruction-driven logic for building large models Based on the complete teaching content generation prompt template described in step 2, a generation logic chain centered on instructions is constructed. By parsing meta-instructions, the fundamental principles of content generation are established. Specific generation goals are clarified according to task instructions. Multi-dimensional variables in conditional instructions are used to precisely constrain the contextual knowledge retrieval and teaching content generation direction. At the same time, the presentation format of the content is standardized according to output format instructions. Personalized and innovative requirements are injected through extended instructions, forming a complete instruction transmission path from core principles to specific execution details. The multi-dimensional variables include subject, grade level, and teaching theme. The precise constraint is implemented by scalar filtering, that is, based on the set multi-dimensional variables, the first set of candidate knowledge document slice units that meet the conditions is selected to reduce the retrieval space of subsequent context knowledge and improve retrieval efficiency. C2: Constructing contextual knowledge-driven logic A knowledge retrieval mechanism is established based on the feature library containing all knowledge document slice units obtained in step 1 through preprocessing. By analyzing the correlation between the teaching content generation task and the contextual knowledge content block, the knowledge document slice unit with the highest matching degree is retrieved from the feature library and integrated as background knowledge into the generation process. This achieves the invocation of knowledge points, alignment with teaching objectives, grasp of teaching key points and difficulties, and reflection of subject literacy requirements. At the same time, the semantic relationship between knowledge units is maintained, so that the generated teaching content is not only closely aligned with professional knowledge in the field of education, but also deeply integrated with specific teaching scenarios, enhancing the scientificity and applicability of the content. The contextual knowledge content block includes knowledge points, teaching objectives, and teaching key points and difficulties. The step of retrieving the knowledge document slice unit with the highest matching degree from the feature library specifically includes: C21: Semantic retrieval based on dense vectors. The cosine similarity is calculated between the feature vector of the teaching topic and the feature vector of the knowledge document slice unit in the knowledge base. The knowledge document slice unit with the highest similarity is obtained and used as the second set of candidate knowledge document slice units with the highest matching degree with the current teaching content generation task. C22: Keyword retrieval based on sparse vectors. The BM25 algorithm is used to calculate the knowledge document slice unit with the highest similarity to the feature vector of the teaching topic, which is used as the third set of candidate knowledge document slice units with the highest matching degree with the current teaching content generation task. C23: Hybrid retrieval based on reordering. The RRF ranking fusion algorithm is used to reorder the second set and the third set of candidate knowledge document slice units to obtain a target knowledge document slice unit set that takes into account both semantic relevance and keyword accuracy. Finally, it is filled into the context knowledge content block. Step 4: Implementation of Teaching Content Generation and Post-processing of Results D1: Dual-Driven Teaching Content Generation and Implementation After the aforementioned large-scale model instruction-driven logic and the aforementioned context knowledge-driven logic are completed, the results of the two are merged and sent to the downstream large language model instance. After analysis and processing, the large language model generates initial teaching content. During the generation process, the large language model simultaneously receives the constraints and context knowledge content in the instruction chain. Through a deep understanding of the teaching objectives, subject requirements, and knowledge connections, it integrates the standardized instructions with educational knowledge to form an initial teaching content output that meets the needs of the teaching scenario. This ensures that the generation direction aligns with the preset task requirements and that the accuracy of knowledge references and the coherence of the teaching logic are both guaranteed. D2: Post-processing of teaching content generation results Post-processing of the initial teaching content includes format revision, content polishing, and layout adjustment, specifically: D21: Formatting Revision addresses formatting flaws in the original teaching content, standardizes and repairs the text content's format and structure. For Markdown formatted content, it checks and corrects heading levels, list indentation, code block identifiers, and link reference formatting. For JSON formatted content, it checks and corrects missing quotes, mismatched brackets, and key-value pair formatting errors. D22: Content polishing, correcting vague expressions, logical gaps or redundant repetitions in the original teaching content, and supplementing the standardized expressions of subject terminology to make the text both in line with the professional context of the education field and in line with the cognitive and comprehension level of students in the target grade level. D23: Layout adjustment. Optimize the information presentation structure according to the type of teaching content. Enhance the organization and ease of use of teaching content by dividing it into sections, setting up labels, and adjusting the content density.
2. A teaching content generation system based on a dual-drive approach of large model instructions and contextual knowledge, implementing the method of claim 1, characterized in that, The system includes: an education domain knowledge preprocessing module, a teaching content generation prompt template definition module, a teaching content generation workflow orchestration module, and a teaching content generation implementation and result post-processing module. Users interact with the system via human-computer collaboration. In the education domain knowledge preprocessing module, they perform operations such as knowledge document parsing, document content slicing, and content text vectorization. In the teaching content generation prompt template definition module, they define large model instructions and their variables, and define contextual knowledge content blocks. In the teaching content generation workflow orchestration module, they construct large model instruction-driven logic and contextual knowledge-driven logic. In the teaching content generation implementation and result post-processing module, they perform dual-drive teaching content generation implementation and teaching content generation result post-processing.
Citation Information
Cited By
Pumping well production condition analysis method based on cue word self-adaptive generation
CN121882220A