Education courseware generation method based on deep learning
The structured PPT is generated through sparse gated MoE network and pigeon flock optimization algorithm based on deep learning, which solves the problem of fragmentation of slide content and incoordination of visual layout in teaching scenarios, and realizes logical unified and visually consistent teaching PPT generation.
Patent Information
- Application Number
- CN202510624868.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to generate structured PPTs with clear logic and unified visuals in teaching scenarios, and lacks the ability to multi-objective global optimization, resulting in fragmented slide content and incoordinated visual layout.
A sparse gated MoE network and pigeon flock optimization algorithm based on deep learning are used to construct teaching structures through teaching semantic graphs, combined with sparse expert activation strategies and layout parameter optimization, a structured PPT draft is generated, and the teaching structure integrity and visual hierarchy consistency are checked.
It realizes the logical unity of teaching content and the consistency of visual style, reduces the teacher's adjustment workload, and improves the generation efficiency and the teachingability of results.
Smart Images

Figure CN120541246A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of educational courseware, and in particular to a method for generating educational courseware based on deep learning. Background Art
[0002] With the continuous integration and development of educational informatization and artificial intelligence technologies, intelligent courseware generation has become a key area of support for teaching content design and personalized classroom interaction. In university classrooms, corporate training, and online course scenarios, teachers or content producers need to quickly convert large amounts of teaching materials into logically clear, visually appealing, and structured PPTs to accommodate diverse teaching paces and cognitive levels. However, current automated PPT generation tools focus on shallow modules such as natural language summarization, template filling, or image-text pairing, and are unable to meet the unique structural expression requirements of teaching documents. When faced with heterogeneous multi-source syllabi, textbook texts, handouts, and case studies, existing methods struggle to maintain consistency in teaching logic and a progressive slide hierarchy.
[0003] Among existing technologies, although template-driven PPT generation methods have achieved automatic splicing of content and layout, they are highly dependent on manually designed templates, have poor template generalization capabilities, and are difficult to flexibly adapt to various teaching structures; generation methods based on large models, such as the Transformer series, also face the problems of high inference costs and insufficient structural modeling capabilities for long documents. In teaching scenarios, it is necessary to generate teaching content that fully covers the "teaching objectives-knowledge points-cases-exercises" structural blocks and match them with slide pages with clear levels and visual unity. However, traditional methods can often only generate content summaries and it is difficult to construct PPT results that coordinate semantics and layout. In addition, existing methods lack multi-objective global optimization capabilities, resulting in fragmented content and inconsistent visual layout of the generated slides, and an inability to make global path adjustments based on the semantic structure of teaching.
[0004] To sum up, there is an urgent need for a technical approach that takes into account the semantic understanding of teaching, the structured expression of generated content, and the ability to optimize and schedule multiple objectives. Summary of the Invention
[0005] One purpose of the present invention is to propose a method for generating educational courseware based on deep learning. The present invention ensures that the final output structured PPT is unified and teachable in terms of expression logic, page layout, and visual style, greatly reducing the workload of teachers' manual adjustments.
[0006] A method for generating educational courseware based on deep learning according to an embodiment of the present invention includes the following steps:
[0007] S1. Collect heterogeneous teaching documents from multiple sources, including teaching syllabi, textbook texts, handouts, and case studies, to form a raw set of teaching documents. This set of teaching documents is then processed using unified character encoding, paragraph boundary standardization, redundant symbol removal, and missing placeholder completion to generate a standardized set of teaching documents.
[0008] S2. Input the standardized teaching document set into the improved sparse gated MoE network. The improved sparse gated MoE network uses the expert sparse activation mechanism to identify and annotate teaching objectives, knowledge points, case information, and exercises to obtain a set of teaching elements.
[0009] S3. Construct a teaching semantic graph based on the teaching element set. The graph uses nodes to represent teaching elements and directed edges to represent the teaching logical sequence. This generates teaching semantic structure data. The improved sparse gated MoE network is then invoked to assign tasks to experts based on the teaching semantic structure data, performing content reorganization, key point summarization, chart matching, and layout draft generation. The resulting output is a draft slide sequence, a draft sparse expert activation strategy, and a draft layout parameter vector.
[0010] S4. The draft of the sparse expert activation strategy, the draft of the layout parameter vector, and the draft of the slide sequence are jointly encoded into optimization individuals. A pigeon flock optimization algorithm is introduced to perform a global-local coupled search under a multi-objective fitness function that comprehensively considers the teaching structure integrity goal, the PPT visual hierarchy clarity goal, and the computing resource constraint goal. The optimized sparse expert activation strategy, optimized layout parameter vector, and optimized slide sequence are obtained.
[0011] S5. Feedback the optimized sparse expert activation strategy to the improved sparse gated MoE network, driving the improved sparse gated MoE network to re-infer based on the same teaching semantic structure data to generate an optimized content set. Apply the optimized layout parameter vector to typeset and render the optimized content set to obtain a structured PPT draft.
[0012] S6. Perform teaching structure integrity check and visual level consistency check on the structured PPT draft. If the check fails, the check result is converted into a fitness penalty signal and fed back to the pigeon swarm optimization algorithm and returns to step S4; if the check passes, the final structured PPT draft is output.
[0013] Optionally, step S1 includes the following contents:
[0014] S11. Collect multi-source heterogeneous teaching documents with structural expression value in teaching scenarios raw :
[0015]
[0016] Among them, d irepresents the i-th teaching document fragment, N is the total number of teaching document fragments, t i The timestamp information of the teaching document segment indicates the position of the teaching document in the teaching sequence. i Represents the teaching structure of the teaching document fragment, including the syllabus, textbook text, handout content and case materials, m i Indicates the original text content contained in the teaching document fragment, c i The course identifier of the teaching document fragment is used to uniquely correspond to the course scenario;
[0017] S12. Performing unified character encoding processing on the original text content field of each teaching document fragment in the multi-source heterogeneous teaching documents. The unified character encoding processing is used to standardize all original text content into a unified character representation format, and the character encoding format used is UTF-8 format;
[0018] S13. Perform paragraph boundary normalization on each teaching document segment in the multi-source heterogeneous teaching document after character encoding standardization. The paragraph boundary normalization process is used to divide the original text content into multiple paragraph units with teaching structural significance. The original text of each teaching document segment is divided into a group of paragraph units during the paragraph boundary normalization process. The paragraph unit division rules are set based on the teaching document language characteristics and format logic, resulting in a standardized paragraph set;
[0019] S14. Perform redundant symbol removal on each paragraph unit in the standardized paragraph set to remove repeated characters, abnormal punctuation, and redundant spaces that do not have any teaching significance in the paragraph unit. The redundant symbol removal process retains the core content of the teaching semantics according to the teaching expression standard to obtain a redundant paragraph set;
[0020] S15. Perform missing placeholder completion processing on each de-redundant paragraph unit in the de-redundant paragraph set to identify paragraph units in the teaching structure that have placeholders but lack corresponding filling content. Placeholders include expression markers with structural guidance significance for teaching objectives, knowledge points, case content and exercise prompts. For the identified placeholder paragraph units, the missing teaching information is automatically filled in by analyzing the context content of the original teaching document fragment where the paragraph is located and its adjacent paragraph units to form a placeholder completion paragraph unit. All placeholder completion paragraph units are recombined into a standardized teaching document set.
[0021] Optionally, step S2 includes the following contents:
[0022] S21. Input the standardized teaching document collection into the improved sparse gated MoE network model, perform deep teaching semantic feature encoding at the paragraph unit granularity, and obtain the paragraph unit teaching embedding vector in, Represents the teaching goal orientation and knowledge structure semantic information implied by the j-th paragraph unit in the i-th teaching document fragment, and d is the embedding vector dimension;
[0023] S22. Improve the sparse gating mechanism for the structured PPT generation scenario of teaching documents and design a dynamic adaptive gating function guided by teaching semantics Dynamically adaptive gating function to teach embedding vectors in paragraph units As input, adaptively predict the degree of match between each expert sub-network and the current paragraph unit teaching objectives and knowledge structure, and obtain the expert activation score vector Among them, the vector dimension K is the total number of experts in the expert sub-network set;
[0024] S23. Introducing the expert load balancing factor ρ in the scenario of generating structured PPTs for teaching documents balance , used to dynamically adjust the expert activation score vector Forming a sparse expert activation strategy for optimizing the teaching content structure, the expert load balancing factor ρ balance By real-time counting the number of times each expert sub-network is activated in the process of generating structured PPTs, online balancing of the expert sub-network load and optimal allocation of computing resources are performed;
[0025] S24. Based on the sparse expert activation strategy, the optimized sparse activation operation TopK is used opt (·) activates the scoring vector for the expert Screen and determine the expert sub-network set that best matches the current paragraph unit teaching structure The number of experts activated in the expert sub-network set is limited by the sparse activation number k;
[0026] S25. Activated expert sub-network set Teaching embedding vectors to input paragraph units Perform parallel multi-teaching task recognition, which includes deep semantic recognition of teaching objectives, mapping of knowledge points and course chapters, matching of teaching case content, and annotation of the relevance between exercises and knowledge points. Output of structured teaching element recognition results through multi-expert parallel reasoning
[0027]
[0028] in, They represent the deep semantic recognition results of teaching objectives, the mapping results of knowledge points and course chapters, the matching results of teaching case content, and the annotation results of the association between exercises and knowledge points.
[0029] S26. Structural integration of all structured teaching element recognition results to form a teaching element set Y for the entire standardized teaching document set teach :
[0030]
[0031] Among them, M i is the number of segmented paragraphs of the i-th segment.
[0032] Optionally, the sparse expert activation strategy includes the following:
[0033] S241. Setting the expert activation scoring vector in the teaching scenario in, Denotes the expert subnetwork E k The teaching semantic matching score of the jth paragraph unit in the i-th teaching document fragment is given. The vector dimension K represents the total number of expert subnetworks in the improved sparse gated MoE network.
[0034] S242. Define the expert subnetwork load vector as u global ={u (1) ,u (2) ,...,u (K)}, where u (k) Denotes the expert subnetwork E k The normalized value of the historical activation frequency during the structured PPT generation process is used to measure the computational load ratio of each expert sub-network in the entire training iteration;
[0035] S243. Activate the expert scoring vector Perform load regularization adjustment to construct the balanced expert activation score vector Load regularization adjustment uses the load suppression factor β to penalize the expert sub-network with high historical activation frequency:
[0036]
[0037] S244. Constructing a mask vector for the correlation between teaching structures in Denotes the expert subnetwork E k Whether the history task has a high matching task label under the teaching structure of the current paragraph unit. If so, the value is 1, otherwise 0;
[0038] S245. According to the teaching structure matching relationship, the balanced expert activation score vector is structurally biased to form the final expert activation score vector guided by the teaching structure. The structural bias factor is δ, which represents the boost weight that the structural association expert sub-network should obtain:
[0039]
[0040] S246. Activate the final expert scoring vector Perform a sorting operation and select the top k expert sub-networks with the highest scores to form the expert activation set of the current paragraph unit:
[0041]
[0042] in, represents the sparse activation expert sub-network set corresponding to the j-th paragraph unit in the i-th teaching document fragment. The size of the expert sub-network set is controlled by the sparse activation number k and satisfies
[0043] Optionally, step S3 includes the following contents:
[0044] S31. Extract each structured teaching element recognition result from the structured teaching element set as a node element of the teaching semantic graph to form a node set. Each structured teaching element recognition result includes a teaching objective, a knowledge point, a case content, or an exercise prompt. The node set is used to construct the structural skeleton of the teaching semantic graph.
[0045] S32. Based on the timestamp value of the original teaching document fragment to which each structured teaching element belongs and its sequence label in the four-element teaching structure of teaching objectives-knowledge points-case studies-exercises, establish directed connections between the nodes. The rule for establishing directed connections is that if a node precedes another node in terms of timestamp value and its structure label sequence is equal to or earlier than that of the node, then a directed edge is established between the two nodes, forming a directed edge set.
[0046] S33. Combining the node set and the directed edge set to construct a complete teaching semantic graph, and representing the teaching semantic graph as a structured teaching semantic structure data, the teaching semantic structure data includes a node set and a directed adjacency matrix, the adjacency matrix is used to indicate whether there is a teaching logical connection relationship between the nodes;
[0047] S34. Input the teaching semantic structure data into the improved sparse gated MoE network and divide the tasks according to the preset expert division of labor, including content reconstruction experts, key point summary experts, chart matching experts, and layout draft generation experts. Based on the structured teaching element attributes of each node in the teaching semantic structure data, an expert routing matrix is constructed to mark whether each structured teaching element is processed by a certain expert sub-network. The expert routing matrix is used to indicate the expert processing path of each structured teaching element and is an intermediate expression of the sparse expert activation strategy.
[0048] S35. The expert subnetworks process their respective tasks in parallel, performing content reorganization, key point summary extraction, diagram matching generation, and layout draft prediction on the structured teaching element nodes. For each node, a structural representation of a slide segment is generated. This structural representation is composed of content reorganization, key point summary extraction, diagram matching generation, and layout draft prediction. All slide segments are then sequentially combined according to the topological order of the teaching semantic graph to form a draft slide sequence.
[0049] S36. Record the expert activation status and layout parameter output corresponding to each structured teaching element node during the generation process, forming a preliminary draft of the sparse expert activation strategy and a preliminary draft of the layout parameter vector. The preliminary draft of the sparse expert activation strategy is used to indicate which expert sub-networks participate in the processing of each structured teaching element. The preliminary draft of the layout parameter vector is used to describe the visual presentation parameters of each structural element in the slide, including the proportion of the title, the contrast ratio of the graphic content density and the color scheme.
[0050] Optionally, step S4 includes the following contents:
[0051] S41. The draft of the sparse expert activation strategy, the draft of the layout parameter vector, and the draft of the slide sequence are represented as the expert activation code, the layout parameter sequence, and the content structure sequence, respectively, and jointly constitute the original individual population of the pigeon flock optimization algorithm.
[0052] S42. Define the teaching structure coverage index ψ (t) , which is used to evaluate whether each slide in the current structured slide sequence in the tth iteration completely covers the coverage identification value of the teaching structure of teaching objectives - knowledge points - cases - exercises. The coverage identification value is used to measure whether the slide on this page contains complete teaching structure information. The coverage identification value is judged by detecting whether the slide on this page contains the teaching objective elements, knowledge point elements, case content elements and exercise prompt elements at the same time. If all are included, it is recorded as 1, and if any structural element is missing, it is recorded as 0. The coverage identification values of all slides are summed up and divided by the total number of slides to obtain the teaching structure coverage index ψ (t) , teaching structure coverage index ψ (t) Used to measure the completeness of the current round of slide structure in terms of overall teaching elements:
[0053]
[0054] in, Indicates whether the slide k contains the complete teaching structure, 1 means it contains, and 0 means it is missing;
[0055] S43. Define the visual consistency deviation function φ vis, used to measure whether the generated slide layout parameters are consistent with teaching visual standards in terms of image and text density, white space ratio, and color palette contrast:
[0056]
[0057] in, Indicates the image and text density ratio of page k, is the target image density threshold, Indicates the page blank ratio, θ ref Leave blank reference values for expectations;
[0058] S44. Introducing the expert routing entropy regularization term H route , used to evaluate whether the activation distribution of the expert sub-network is balanced:
[0059]
[0060] Among them, π l Denotes the expert subnetwork E l The probability distribution of being activated in the current individual, K is the total number of expert sub-networks;
[0061] S45. Constructing the joint fitness function F of comprehensive teaching structure coverage index, visual consistency deviation function and expert routing entropy regularization term task , as the main objective function of the pigeon flock optimization algorithm:
[0062] F task =λ1·ψ (t) -λ2·φ vis -λ3·H route ;
[0063] Among them, λ1, λ2, and λ3 are the adjustment weights of each sub-goal respectively;
[0064] S46. Construct a navigation matrix to guide the individual structure of the pigeon flock to adjust its direction. According to the teaching semantic graph adjacency matrix A teach The constructed guidance factor matrix, navigation matrix The role of the navigation matrix is to introduce the topological relationship between the nodes of each structured teaching element in the teaching semantic graph into the optimization process, which is used to constrain the direction of individual position update in the pigeon flock optimization algorithm and ensure that the optimized sparse expert activation strategy is consistent with the teaching structure. By subtracting the teaching semantic graph adjacency matrix A from the identity matrix I teachThe weighted form is obtained, where the weighting factor is the graph topology influence factor γ, which is a positive real number between zero and one. It is used to control the degree of guidance of the teaching semantic graph on the individual update path of the pigeon group. The larger the graph topology influence factor γ, the stronger the constraint of the navigation matrix on the optimization direction. The smaller the graph topology influence factor γ, the weaker the guiding effect of the navigation matrix. The adjacency matrix A of the teaching semantic graph is teach Derived from the teaching semantic structure data, the element at each position in the teaching semantic graph adjacency matrix indicates whether there is a teaching logical connection relationship between two structured teaching element nodes in the teaching semantic graph. If there is a connection relationship, the corresponding position element is one, and if there is no connection relationship, the corresponding position element is zero. The navigation matrix As a flight direction correction factor, it is embedded in the individual pigeon update formula, guiding the position update of each optimized individual in the pigeon optimization algorithm. This ensures that the structure evolution path generated by the slides follows the logical order of the teaching structure, effectively preventing structural dislocation and confusion in teaching semantics during the optimization process.
[0065] S47. Evolutionary adjustment of the structure perception of the individual population of the pigeon swarm optimization algorithm based on the position update formula modified by the navigation matrix:
[0066]
[0067] Among them, η is the flight step length control factor, represents the individual structure with the best fitness in round t, represents the current position vector of the k-th pigeon group optimization individual at the t-th iteration, represents the position vector of the k-th optimized individual in the pigeon group at the t+1-th iteration;
[0068] S48. When the value of the joint fitness function of the pigeon population does not improve in multiple consecutive rounds, or the number of iterations reaches the maximum number of rounds, the optimization process is terminated and the optimal fitness individual is output. Its corresponding sparse expert activation strategy G opt , layout parameter vector L opt and slide structure content sequence P opt .
[0069] Optionally, step S5 includes the following contents:
[0070] S51. Feedback the optimized sparse expert activation strategy to the improved sparse gated MoE network, replacing the original sparse expert activation strategy draft, as the basis for expert activation control in the re-inference phase, and used to guide the sparse gated MoE network to activate only the optimized expert sub-network set when regenerating structured content. Each of the optimized sparse expert activation strategies The expert routing state corresponding to the content structure of the slide on page k represents which expert sub-networks should participate in generating the content of the page structure;
[0071] S52. The optimized sparse expert activation strategy and the teaching semantic structure data are jointly input into the improved sparse gated MoE network model to maintain the teaching semantic graph structure G teach Unchanged, the expert sub-network division of labor structure is used to re-reason the teaching semantic structure data, generate an optimized content set, and optimize each This includes optimized content reorganization information, key point summary information, chart suggestion information, and content components that are semantically consistent with their teaching nodes;
[0072] S53. Apply the optimized layout parameter vector to each item in the optimized content set, and The title ratio, image and text density, white space ratio and color matching style are used to visually arrange the components of the content module, adjust the margins and bind the colors;
[0073] S54. Structural splicing of all optimized slide pages and output of structured PPT draft Represents the structured output content of slide k.
[0074] Optionally, the structured output content consists of the following structure:
[0075] If the slide contains the teaching objective elements, knowledge point elements, case content elements, and exercise prompt elements, and the layout structure meets the requirements of the title ratio ≥ 15%, the image and text density is in the range of [30%, 70%], and the white space ratio is not less than 10%, then the page is output as a structurally consistent page;
[0076] If the slide on this page lacks any structured teaching elements, or the image and text density exceeds 80% and the title ratio is less than 10%, the page will be output as a content-heavy page;
[0077] If the slide on this page contains all the teaching elements but the color palette contrast is lower than the set threshold, the color matching style does not conform to the high contrast principle, or the white space rate is less than 5%, the page will be output as a visually unbalanced page.
[0078] The beneficial effects of the present invention are:
[0079] (1) The present invention constructs a dynamic adaptive gating function based on paragraph embedding vectors to predict the matching degree between the semantic content of each teaching paragraph and each expert sub-network, and realizes on-demand sparse activation in combination with the expert load balancing factor. It not only realizes the structured extraction of multiple teaching elements such as teaching objective recognition, knowledge point chapter mapping, case content extraction and exercise structure association, but also reduces the number of expert activations during reasoning by introducing sparse activation constraints, improves the scalability and computational efficiency of long document processing, and breaks through the problem of high video memory overhead of traditional Transformer when processing teaching documents with tens of thousands of words.
[0080] (2) The present invention designs a joint fitness function that integrates the teaching structure coverage index, the visual consistency deviation function and the expert routing entropy regularization term, and introduces a navigation matrix constructed based on the adjacency matrix of the teaching semantic graph as a correction term for the flight direction of the pigeon flock. This significantly improves the collaborative consistency between the slide structure generation and the semantic topology, and can guide the structured PPT to form a logical progressive relationship between teaching nodes, avoiding structural dislocation and content jumping problems. It has obvious advantages in organizing cross-chapter and cross-case teaching content.
[0081] (3) The present invention not only outputs sparse expert activation strategy and layout parameter vector for controlling MoE network and rendering engine in the generation stage, but also performs double verification on teaching structure integrity and visual level consistency after the structured PPT draft is generated. If the verification fails, the verification result is converted into a fitness penalty item and fed back to the pigeon swarm optimization algorithm and re-enters the optimization cycle, thereby forming a closed-loop feedback chain of content-structure-rendering, effectively avoiding the problem of static generation results or style fragmentation, ensuring that the final output structured PPT is unified and teachable in expression logic, page layout, and visual style, and greatly reducing the workload of teachers' manual adjustments. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0083] Figure 1 This is a flowchart of a method for generating educational courseware based on deep learning proposed by the present invention. DETAILED DESCRIPTION
[0084] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0085] refer to Figure 1 , a method for generating educational courseware based on deep learning, comprising the following steps:
[0086] S1. Collect heterogeneous teaching documents from multiple sources, including teaching syllabi, textbook texts, handouts, and case studies, to form a raw set of teaching documents. This set of teaching documents is then processed using unified character encoding, paragraph boundary standardization, redundant symbol removal, and missing placeholder completion to generate a standardized set of teaching documents.
[0087] S2. Input the standardized teaching document set into the improved sparse gated MoE network. The improved sparse gated MoE network uses the expert sparse activation mechanism to identify and annotate teaching objectives, knowledge points, case information, and exercises to obtain a set of teaching elements.
[0088] S3. Construct a teaching semantic graph based on the teaching element set. The graph uses nodes to represent teaching elements and directed edges to represent the teaching logical sequence. This generates teaching semantic structure data. The improved sparse gated MoE network is then invoked to assign tasks to experts based on the teaching semantic structure data, performing content reorganization, key point summarization, chart matching, and layout draft generation. The resulting output is a draft slide sequence, a draft sparse expert activation strategy, and a draft layout parameter vector.
[0089] S4. The draft of the sparse expert activation strategy, the draft of the layout parameter vector, and the draft of the slide sequence are jointly encoded into optimization individuals. A pigeon flock optimization algorithm is introduced to perform a global-local coupled search under a multi-objective fitness function that comprehensively considers the teaching structure integrity goal, the PPT visual hierarchy clarity goal, and the computing resource constraint goal. The optimized sparse expert activation strategy, optimized layout parameter vector, and optimized slide sequence are obtained.
[0090] S5. Feedback the optimized sparse expert activation strategy to the improved sparse gated MoE network, driving the improved sparse gated MoE network to re-infer based on the same teaching semantic structure data to generate an optimized content set. Apply the optimized layout parameter vector to typeset and render the optimized content set to obtain a structured PPT draft.
[0091] S6. Perform teaching structure integrity check and visual level consistency check on the structured PPT draft. If the check fails, the check result is converted into a fitness penalty signal and fed back to the pigeon swarm optimization algorithm and returns to step S4; if the check passes, the final structured PPT draft is output.
[0092] In this embodiment, step S1 includes the following contents:
[0093] S11. Collect multi-source heterogeneous teaching documents with structural expression value in teaching scenarios raw :
[0094]
[0095] Among them, di represents the i-th teaching document fragment, N is the total number of teaching document fragments, t i The timestamp information of the teaching document segment indicates the position of the teaching document in the teaching sequence. i Represents the teaching structure of the teaching document fragment, including the syllabus, textbook text, handout content and case materials, m i Indicates the original text content contained in the teaching document fragment, c i The course identifier of the teaching document fragment is used to uniquely correspond to the course scenario;
[0096] S12. Performing unified character encoding processing on the original text content field of each teaching document fragment in the multi-source heterogeneous teaching documents. The unified character encoding processing is used to standardize all original text content into a unified character representation format, and the character encoding format used is UTF-8 format;
[0097] S13. Perform paragraph boundary normalization on each teaching document segment in the multi-source heterogeneous teaching document after character encoding standardization. The paragraph boundary normalization process is used to divide the original text content into multiple paragraph units with teaching structural significance. The original text of each teaching document segment is divided into a group of paragraph units during the paragraph boundary normalization process. The paragraph unit division rules are set based on the teaching document language characteristics and format logic, resulting in a standardized paragraph set;
[0098] S14. Perform redundant symbol removal on each paragraph unit in the standardized paragraph set to remove repeated characters, abnormal punctuation, and redundant spaces that do not have any teaching significance in the paragraph unit. The redundant symbol removal process retains the core content of the teaching semantics according to the teaching expression standard to obtain a redundant paragraph set;
[0099] S15. Perform missing placeholder completion processing on each de-redundant paragraph unit in the de-redundant paragraph set to identify paragraph units in the teaching structure that have placeholders but lack corresponding filling content. Placeholders include expression markers with structural guidance significance for teaching objectives, knowledge points, case content and exercise prompts. For the identified placeholder paragraph units, the missing teaching information is automatically filled in by analyzing the context content of the original teaching document fragment where the paragraph is located and its adjacent paragraph units to form a placeholder completion paragraph unit. All placeholder completion paragraph units are recombined into a standardized teaching document set.
[0100] In this embodiment, step S2 includes the following contents:
[0101] S21. Input the standardized teaching document collection into the improved sparse gated MoE network model, perform deep teaching semantic feature encoding at the paragraph unit granularity, and obtain the paragraph unit teaching embedding vector in, Represents the teaching goal orientation and knowledge structure semantic information implied by the j-th paragraph unit in the i-th teaching document fragment, and d is the embedding vector dimension;
[0102] S22. Improve the sparse gating mechanism for the structured PPT generation scenario of teaching documents and design a dynamic adaptive gating function guided by teaching semantics Dynamically adaptive gating function to teach embedding vectors in paragraph units As input, adaptively predict the degree of match between each expert sub-network and the current paragraph unit teaching objectives and knowledge structure, and obtain the expert activation score vector Among them, the vector dimension K is the total number of experts in the expert sub-network set;
[0103] S23. Introducing the expert load balancing factor ρ in the scenario of generating structured PPTs for teaching documents balance , used to dynamically adjust the expert activation score vector Forming a sparse expert activation strategy for optimizing the teaching content structure, the expert load balancing factor ρ balance By real-time counting the number of times each expert sub-network is activated in the process of generating structured PPTs, online balancing of the expert sub-network load and optimal allocation of computing resources are performed;
[0104] S24. Based on the sparse expert activation strategy, the optimized sparse activation operation TopK is used opt (·) activates the scoring vector for the expert Screen and determine the expert sub-network set that best matches the current paragraph unit teaching structure The number of experts activated in the expert sub-network set is limited by the sparse activation number k;
[0105] S25. Activated expert sub-network set Teaching embedding vectors to input paragraph units Perform parallel multi-teaching task recognition, which includes deep semantic recognition of teaching objectives, mapping of knowledge points and course chapters, matching of teaching case content, and annotation of the relevance between exercises and knowledge points. Output of structured teaching element recognition results through multi-expert parallel reasoning
[0106]
[0107] in, They represent the deep semantic recognition results of teaching objectives, the mapping results of knowledge points and course chapters, the matching results of teaching case content, and the annotation results of the association between exercises and knowledge points.
[0108] The goal of deep semantic recognition of teaching objectives is to determine whether the current paragraph contains teaching objective information and extract its corresponding teaching objective description and cognitive level attributes. The input is the embedded representation of the current paragraph unit and the corresponding semantic context. The recognition process is structurally guided by the preset teaching intention template vocabulary library. Through semantic pattern matching and contextual semantic weight aggregation, it is determined whether there are teaching objective expressions representing "mastering...", "understanding...", "being able to...do...", and combining Bloom's cognitive taxonomy to further divide the objectives into levels such as "memory", "understanding", and "application". The final output is the teaching objective existence mark, the target text description, and the cognitive level category, which are used to clarify the expression content and difficulty level of the teaching objective node in the subsequent teaching structure composition.
[0109] The goal of knowledge point and course chapter mapping recognition is to extract specific teaching knowledge point names from paragraphs and determine their attribution within the course chapter system. The input is the paragraph unit embedding vector and the course syllabus chapter index set. The recognition process calculates the semantic relevance between the paragraph semantics and the semantic relevance of each chapter title, and uses a soft matching strategy to select the closest chapter as the structural mapping location of the knowledge point. This recognition process comprehensively considers the idiomatic expressions of course terminology and the semantic center of the chapter theme to ensure that the chapter to which the output knowledge point belongs is pedagogically consistent. The final output is the knowledge point name of the current paragraph and its mapping to the chapter number in the course syllabus, which is used to bind the knowledge structure path in the teaching semantic graph.
[0110] The goal of teaching case content matching and recognition is to determine whether a paragraph is a teaching case and identify scenario keywords in the case description and their corresponding knowledge points. The input is the current paragraph's embedded representation and its associated knowledge point semantic labels. The recognition process matches the case expression structure of "in practice...", "for example...", "a company...", and "the application scenario is...", and combines the presence of real-world situations, people, organizations, and problem vocabulary entities in the context to determine whether the paragraph is a teaching case description. It then further extracts the knowledge application points referred to by the case. The final output includes a case presence identifier, case scenario keywords, and associated knowledge points, which are used to form a teaching scenario presentation unit in the PPT.
[0111] The goal of the exercise and knowledge point correlation annotation recognition is to determine whether the paragraph is a teaching exercise content, and to mark the correlation between its question type, core examination points and target knowledge points. The input is the embedded representation of the paragraph unit and the set of knowledge point labels that have been identified in the early stage. The recognition process performs structural classification judgment on whether the "multiple choice question", "fill in the blank question" and "true or false question" question type keywords appear in the paragraph, and combines the proposition language, question stem structure and knowledge point keywords in the paragraph content for semantic matching to identify the main knowledge point range tested by the question and its corresponding course structure position. The final output is the question type category, the main semantic keywords of the exercise question stem and the most relevant knowledge point number, which is used to support the construction of the exercise structure module in PPT generation.
[0112] The recognition tasks for these four types of structured instructional elements are performed in parallel by four expert sub-networks within the sparsely gated MoE network. Each expert is activated by a gating function to participate in the recognition process based on the paragraph's instructional semantics and task preferences. The outputs of each recognition task collectively constitute the structured instructional element annotation results for the paragraph unit.
[0113] S26. Structural integration of all structured teaching element recognition results to form a teaching element set Y for the entire standardized teaching document set teach :
[0114]
[0115] Among them, M i is the number of segmented paragraphs of the i-th segment.
[0116] In this embodiment, the sparse expert activation strategy includes the following:
[0117] S241. Setting the expert activation scoring vector in the teaching scenario in, Denotes the expert subnetwork E k The teaching semantic matching score of the jth paragraph unit in the i-th teaching document fragment is given. The vector dimension K represents the total number of expert subnetworks in the improved sparse gated MoE network.
[0118] S242. Define the expert subnetwork load vector as u global ={u (1) ,u (2) ,...,u (K)}, where u (k) Denotes the expert subnetwork E k The normalized value of the historical activation frequency during the structured PPT generation process is used to measure the computational load ratio of each expert sub-network in the entire training iteration;
[0119] S243. Activate the expert scoring vector Perform load regularization adjustment to construct the balanced expert activation score vector Load regularization adjustment uses the load suppression factor β to penalize the expert sub-network with high historical activation frequency:
[0120]
[0121] S244. Constructing a mask vector for the correlation between teaching structures in Denotes the expert subnetwork E k Whether the history task has a high matching task label under the teaching structure of the current paragraph unit. If so, the value is 1, otherwise 0;
[0122] S245. According to the teaching structure matching relationship, the balanced expert activation score vector is structurally biased to form the final expert activation score vector guided by the teaching structure. The structural bias factor is δ, which represents the boost weight that the structural association expert sub-network should obtain:
[0123]
[0124] S246. Activate the final expert scoring vector Perform a sorting operation and select the top k expert sub-networks with the highest scores to form the expert activation set of the current paragraph unit:
[0125]
[0126] in, represents the sparse activation expert sub-network set corresponding to the j-th paragraph unit in the i-th teaching document fragment. The size of the expert sub-network set is controlled by the sparse activation number k and satisfies
[0127] In this embodiment, step S3 includes the following contents:
[0128] S31. Extract each structured teaching element recognition result from the structured teaching element set as a node element of the teaching semantic graph to form a node set. Each structured teaching element recognition result includes a teaching objective, a knowledge point, a case content, or an exercise prompt. The node set is used to construct the structural skeleton of the teaching semantic graph.
[0129] S32. Based on the timestamp value of the original teaching document fragment to which each structured teaching element belongs and its sequence label in the four-element teaching structure of teaching objectives-knowledge points-case studies-exercises, establish directed connections between the nodes. The rule for establishing directed connections is that if a node precedes another node in terms of timestamp value and its structure label sequence is equal to or earlier than that of the node, then a directed edge is established between the two nodes, forming a directed edge set.
[0130] S33. Combining the node set and the directed edge set to construct a complete teaching semantic graph, and representing the teaching semantic graph as a structured teaching semantic structure data, the teaching semantic structure data includes a node set and a directed adjacency matrix, the adjacency matrix is used to indicate whether there is a teaching logical connection relationship between the nodes;
[0131] S34. Input the teaching semantic structure data into the improved sparse gated MoE network and divide the tasks according to the preset expert division of labor, including content reconstruction experts, key point summary experts, chart matching experts, and layout draft generation experts. Based on the structured teaching element attributes of each node in the teaching semantic structure data, an expert routing matrix is constructed to mark whether each structured teaching element is processed by a certain expert sub-network. The expert routing matrix is used to indicate the expert processing path of each structured teaching element and is an intermediate expression of the sparse expert activation strategy.
[0132] S35. The expert subnetworks process their respective tasks in parallel, performing content reorganization, key point summary extraction, diagram matching generation, and layout draft prediction on the structured teaching element nodes. For each node, a structural representation of a slide segment is generated. This structural representation is composed of content reorganization, key point summary extraction, diagram matching generation, and layout draft prediction. All slide segments are then sequentially combined according to the topological order of the teaching semantic graph to form a draft slide sequence.
[0133] S36. Record the expert activation status and layout parameter output corresponding to each structured teaching element node during the generation process, forming a preliminary draft of the sparse expert activation strategy and a preliminary draft of the layout parameter vector. The preliminary draft of the sparse expert activation strategy is used to indicate which expert sub-networks participate in the processing of each structured teaching element. The preliminary draft of the layout parameter vector is used to describe the visual presentation parameters of each structural element in the slide, including the proportion of the title, the contrast ratio of the graphic content density and the color scheme.
[0134] In this embodiment, step S4 includes the following contents:
[0135] S41. The draft of the sparse expert activation strategy, the draft of the layout parameter vector, and the draft of the slide sequence are represented as the expert activation code, the layout parameter sequence, and the content structure sequence, respectively, and jointly constitute the original individual population of the pigeon flock optimization algorithm.
[0136] S42. Define the teaching structure coverage index ψ (t) , which is used to evaluate whether each slide in the current structured slide sequence in the tth iteration completely covers the coverage identification value of the teaching structure of teaching objectives - knowledge points - cases - exercises. The coverage identification value is used to measure whether the slide on this page contains complete teaching structure information. The coverage identification value is judged by detecting whether the slide on this page contains the teaching objective elements, knowledge point elements, case content elements and exercise prompt elements at the same time. If all are included, it is recorded as 1, and if any structural element is missing, it is recorded as 0. The coverage identification values of all slides are summed up and divided by the total number of slides to obtain the teaching structure coverage index ψ (t) , teaching structure coverage index ψ (t) Used to measure the completeness of the current round of slide structure in terms of overall teaching elements:
[0137]
[0138] in, Indicates whether the slide k contains the complete teaching structure, 1 means it contains, and 0 means it is missing;
[0139] S43. Define the visual consistency deviation function φ vis , used to measure whether the generated slide layout parameters are consistent with teaching visual standards in terms of image and text density, white space ratio, and color palette contrast:
[0140]
[0141] in, Indicates the image and text density ratio of page k, is the target image density threshold, Indicates the page blank ratio, θ ref Leave blank reference values for expectations;
[0142] S44. Introducing the expert routing entropy regularization term H route , used to evaluate whether the activation distribution of the expert sub-network is balanced:
[0143]
[0144] Among them, π l Denotes the expert subnetwork E l The probability distribution of being activated in the current individual, K is the total number of expert sub-networks;
[0145] S45. Constructing the joint fitness function F of comprehensive teaching structure coverage index, visual consistency deviation function and expert routing entropy regularization term task , as the main objective function of the pigeon flock optimization algorithm:
[0146] F task =λ1·ψ (t) -λ2·φ vis -λ3·H route ;
[0147] Among them, λ1, λ2, and λ3 are the adjustment weights of each sub-goal respectively;
[0148] S46. Construct a navigation matrix to guide the individual structure of the pigeon flock to adjust its direction. According to the teaching semantic graph adjacency matrix A teach The constructed guidance factor matrix, navigation matrix The role of the navigation matrix is to introduce the topological relationship between the nodes of each structured teaching element in the teaching semantic graph into the optimization process, which is used to constrain the direction of individual position update in the pigeon flock optimization algorithm and ensure that the optimized sparse expert activation strategy is consistent with the teaching structure. By subtracting the teaching semantic graph adjacency matrix A from the identity matrix I teach The weighted form is obtained, where the weighting factor is the graph topology influence factor γ, which is a positive real number between zero and one. It is used to control the degree of guidance of the teaching semantic graph on the individual update path of the pigeon group. The larger the graph topology influence factor γ, the stronger the constraint of the navigation matrix on the optimization direction. The smaller the graph topology influence factor γ, the weaker the guiding effect of the navigation matrix. The adjacency matrix A of the teaching semantic graph is teach Derived from the teaching semantic structure data, the element at each position in the teaching semantic graph adjacency matrix indicates whether there is a teaching logical connection relationship between two structured teaching element nodes in the teaching semantic graph. If there is a connection relationship, the corresponding position element is one, and if there is no connection relationship, the corresponding position element is zero. The navigation matrix As a flight direction correction factor, it is embedded in the individual pigeon update formula, guiding the position update of each optimized individual in the pigeon optimization algorithm. This ensures that the structure evolution path generated by the slides follows the logical order of the teaching structure, effectively preventing structural dislocation and confusion in teaching semantics during the optimization process.
[0149] S47. Evolutionary adjustment of the structure perception of the individual population of the pigeon swarm optimization algorithm based on the position update formula modified by the navigation matrix:
[0150]
[0151] Among them, η is the flight step length control factor, represents the individual structure with the best fitness in round t, represents the current position vector of the k-th pigeon group optimization individual at the t-th iteration, represents the position vector of the k-th optimized individual in the pigeon group at the t+1-th iteration;
[0152] S48. When the value of the joint fitness function of the pigeon population does not improve in multiple consecutive rounds, or the number of iterations reaches the maximum number of rounds, the optimization process is terminated and the optimal fitness individual is output. Its corresponding sparse expert activation strategy G opt , layout parameter vector L opt and slide structure content sequence P opt .
[0153] In this embodiment, step S5 includes the following contents:
[0154] S51. Feedback the optimized sparse expert activation strategy to the improved sparse gated MoE network, replacing the original sparse expert activation strategy draft, as the basis for expert activation control in the re-inference phase, and used to guide the sparse gated MoE network to activate only the optimized expert sub-network set when regenerating structured content. Each of the optimized sparse expert activation strategies The expert routing state corresponding to the content structure of the slide on page k represents which expert sub-networks should participate in generating the content of the page structure;
[0155] S52. The optimized sparse expert activation strategy and the teaching semantic structure data are jointly input into the improved sparse gated MoE network model to maintain the teaching semantic graph structure G teach Unchanged, the expert sub-network division of labor structure is used to re-reason the teaching semantic structure data, generate an optimized content set, and optimize each This includes optimized content reorganization information, key point summary information, chart suggestion information, and content components that are semantically consistent with their teaching nodes;
[0156] S53. Apply the optimized layout parameter vector to each item in the optimized content set, and The title ratio, image and text density, white space ratio and color matching style are used to visually arrange the components of the content module, adjust the margins and bind the colors;
[0157] S54. Structural splicing of all optimized slide pages and output of structured PPT draft Represents the structured output content of slide k.
[0158] In this embodiment, the structured output content consists of the following structure:
[0159] If the slide contains the teaching objective elements, knowledge point elements, case content elements, and exercise prompt elements, and the layout structure meets the requirements of the title ratio ≥ 15%, the image and text density is in the range of [30%, 70%], and the white space ratio is not less than 10%, then the page is output as a structurally consistent page;
[0160] If the slide on this page lacks any structured teaching elements, or the image and text density exceeds 80% and the title ratio is less than 10%, the page will be output as a content-heavy page;
[0161] If the slide on this page contains all the teaching elements but the color palette contrast is lower than the set threshold, the color matching style does not conform to the high contrast principle, or the white space rate is less than 5%, the page will be output as a visually unbalanced page.
[0162] Example 1:
[0163] At 9:00 AM on September 12, 2024, Mr. Zhang, a course development specialist at a university's School of Information Science and Technology, launched a system for automatically generating structured PPTs for teaching documents. This system was used to generate initial teaching materials for the upcoming course "Introduction to Intelligent Recommender Systems." Mr. Zhang's initial input included the following four types of teaching documents, which were uploaded to the system's task pool:
[0164] File 1: Course Syllabus Document (File ID: D20240912_A001, Source: Exported from the Undergraduate Course Center)
[0165] File 2: Textbook Chapters 1 to 4 (File ID: D20240912_A002, Source: PDF provided by the publisher, OCR pre-processed)
[0166] File 3: Previous lecture notes and supplementary notes (File ID: D20240912_A003, 52 pages in total after combining scans)
[0167] File 4: Teaching Cases and Exercises (File ID: D20240912_A004, structured JSON, 138 items)
[0168] After the upload was complete, the system automatically applied a unified character encoding (UTF-8 format), standardized paragraph structure, and removed redundant symbols. During the cleaning process, 282 illegal characters were identified, 1,289 punctuation marks were corrected, and 16 placeholders for teaching elements were filled. After the cleaning was complete, the system constructed a standardized teaching document collection with a total of 2,884 paragraph units and 57,312 words.
[0169] Subsequently, the data is input into the sparse gated MoE network of the present invention for teaching semantic recognition. The system starts model reasoning at 9:17:04 on September 12, 2024, and ends at 9:17:45, which takes 41 seconds. Each paragraph unit is converted into an embedding vector, activating the Top-2 expert subnetwork under the sparse gating mechanism. Taking paragraph 387 (content is "Learning the sparsity feature modeling of the user-item rating matrix through the basic principles of collaborative filtering algorithm") as an example, the system activates the "knowledge point recognition expert E2" and the "teaching goal extraction expert E1", and outputs the following annotation results:
[0170] Teaching objective identification: Master the basic formula modeling method of collaborative filtering, cognitive level: application;
[0171] Knowledge point mapping: Chapter number CH2.3, keywords "scoring matrix" and "sparse modeling";
[0172] Case matching: not triggered (threshold similarity 0.38 is lower than the set 0.45);
[0173] Exercise association: Bound to question 42 in the question bank, "Item recommendation process design question based on rating matrix."
[0174] During the entire semantic recognition phase, the system output 1,022 teaching objectives, 874 knowledge point labels, 212 case matches, and 145 exercise bindings. After completion, the system began constructing a teaching semantic graph, which contained 1,925 teaching element nodes and 2,286 directed edges. The graph structure contained 38 teaching quadruple paths, each of which consisted of the structure "teaching objective → knowledge point → case → exercise."
[0175] The MoE network performs preliminary generation based on the semantic graph, outputting a 74-page slide structure sequence. Each page contains the following structure draft: title block, summary of key points, recommended charts, candidate color schemes, and page number. For page 11, the system output is as follows:
[0176] Slide title: Collaborative filtering and rating matrix modeling;
[0177] Abstract: The rating matrix reflects the sparsity of user interests. Collaborative filtering achieves recommendations through similarity calculation. The key formula is modeling the user-item bias term.
[0178] Image recommendation: Chart library image number #T0432, content is the rating matrix heat map;
[0179] Recommended layout parameters: title accounted for 17.4%, image-text ratio was 63:37, white space ratio was 11%, and the color palette was a contrasting gray-blue.
[0180] The system encapsulated the three data points generated above into optimized individuals and introduced a pigeon flock optimization algorithm. The navigation matrix was constructed based on the semantic graph adjacency matrix, and the system initiated 32 rounds of iterative search. The optimal fitness score reached 89.5% in the 19th round, a 19.4% improvement over the first round. Slide 43 was corrected from a structurally incomplete page (missing teaching objectives) to a structurally consistent page in the 7th round. The system completed the optimized output at 10:18. The slide sequence structure was adjusted to 72 pages, and the structural completeness coverage increased to 98.9%.
[0181] The system automatically returns the optimized expert activation strategy and layout parameter vector to the MoE network and re-executes the inference. Taking the optimized page 18 as an example, its final structure is:
[0182] Title: "Content-based Recommender Systems and Their Limitations";
[0183] Teaching objectives: "Understand the content-based feature extraction process and identify its impact on the cold start problem";
[0184] Image and text density: images account for 45%, text accounts for 50%, and blank space accounts for 5%;
[0185] Structure tags: teaching objectives√, knowledge points√, cases√, exercises√;
[0186] Visual component number: Icon#T0510, color template#CLR_GRAY3;
[0187] The structured PPT draft is generated. The system initiates the structural consistency check and visual style assessment module, reporting two inconsistent slides (pages 32 and 58). The system issues a fitness penalty signal, which is automatically passed back to the pigeon flock optimization module, triggering a second round of fine-tuning and optimization until the structural consistency reaches 100%, generating the final PPT draft.
[0188] To verify the effectiveness, the team compared the method of the present invention with the traditional manual method and the commercial API method (using a well-known large model + Canva automatic typesetting template). The selected evaluation indicators and results are as follows:
[0189] Table 1 Comparative data of the method of the present invention, the traditional manual method and the commercial API method
[0190]
[0191] In addition, in a real teaching scenario, comparing the PPT presentation effect of the "Collaborative Filtering" chapter, the proportion of students who chose "the PPT structure of this class is clear and the rhythm is appropriate" in the questionnaire feedback on the day of teaching was 94.5%, which was significantly higher than the course using the template generation method (81.2%).
[0192] In summary, Example 1 demonstrates the deployment path, data processing flow, generation results and performance comparison of the method of the present invention in actual teaching tasks, and comprehensively verifies the feasibility and advantages of the "structure preservation + semantic guidance + visual optimization" three-in-one mechanism proposed by the present invention in real scenarios.
[0193] The present invention constructs a dynamic adaptive gating function based on paragraph embedding vectors to predict the matching degree between the semantic content of each teaching paragraph and each expert sub-network, and combines the expert load balancing factor to realize on-demand sparse activation. It not only realizes the structured extraction of multiple teaching elements such as teaching objective identification, knowledge point chapter mapping, case content extraction and exercise structure association, but also reduces the number of expert activations during reasoning by introducing sparse activation constraints, improves the scalability and computational efficiency of long document processing, and overcomes the problem of excessive memory overhead of traditional Transformer when processing teaching documents with tens of thousands of words.
[0194] The present invention designs a joint fitness function that integrates the teaching structure coverage index, the visual consistency deviation function and the expert routing entropy regularization term, and introduces a navigation matrix constructed based on the adjacency matrix of the teaching semantic graph as a correction term for the flight direction of the pigeon flock. This significantly improves the collaborative consistency between the slide structure generation and the semantic topology, and can guide the structured PPT to form a logical progressive relationship between teaching nodes, avoiding structural dislocation and content jumping problems. It has obvious advantages in organizing teaching content across chapters and cases.
[0195] The present invention not only outputs sparse expert activation strategies and layout parameter vectors for controlling the MoE network and rendering engine during the generation phase, but also performs double verification on the teaching structure integrity and visual hierarchy consistency after the structured PPT draft is generated. If the verification fails, the verification result is converted into a fitness penalty item and fed back to the pigeon swarm optimization algorithm and re-enters the optimization loop, thereby forming a closed-loop feedback chain of content-structure-rendering, effectively avoiding the problems of static generation results or style fragmentation, ensuring that the final output structured PPT is unified and teachable in expression logic, page layout, and visual style, and greatly reducing the workload of teachers' manual adjustments.
[0196] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for generating educational courseware based on deep learning, characterized in that: The steps include: S1. Collect and preprocess the original set of teaching documents to generate a standardized set of teaching documents; S2. Input the standardized teaching document set into the improved sparse gated MoE network to obtain the teaching element set; S3. Construct a teaching semantic graph based on the teaching element set, generate teaching semantic structure data, call the improved sparse gated MoE network, and output the first draft of the slide sequence, the first draft of the sparse expert activation strategy, and the first draft of the layout parameter vector; S4. The draft of the sparse expert activation strategy, the draft of the layout parameter vector, and the draft of the slide sequence are jointly encoded into an optimization individual. The pigeon flock optimization algorithm is introduced to perform a global-local coupled search to obtain the optimized sparse expert activation strategy, layout parameter vector, and slide sequence. S5. Feedback the optimized sparse expert activation strategy to the improved sparse gated MoE network, driving the improved sparse gated MoE network to re-infer based on the same teaching semantic structure data to generate an optimized content set. Apply the optimized layout parameter vector to typeset and render the optimized content set to obtain a structured PPT draft. S6. Perform teaching structure integrity check and visual level consistency check on the structured PPT draft. If the check fails, the check result is converted into a fitness penalty signal and fed back to the pigeon swarm optimization algorithm and returns to step S4; if the check passes, the final structured PPT draft is output.
2. The method for generating educational courseware based on deep learning according to claim 1, characterized in that: The step S1 includes the following contents: S11. Collect multi-source heterogeneous teaching documents with structural expression value in teaching scenarios raw : Among them, d i represents the i-th teaching document fragment, N is the total number of teaching document fragments, t i The timestamp information of the teaching document segment indicates the position of the teaching document in the teaching sequence. i Represents the teaching structure of the teaching document fragment, including the syllabus, textbook text, handout content and case materials, m i Indicates the original text content contained in the teaching document fragment, c i The course identifier of the teaching document fragment, which is used to uniquely correspond to the course scenario; S12. Performing unified character encoding processing on the original text content field of each teaching document fragment in the multi-source heterogeneous teaching documents. The unified character encoding processing is used to standardize all original text content into a unified character representation format, and the character encoding format used is UTF-8 format; S13. After character encoding standardization, paragraph boundary standardization is performed on each teaching document fragment in the multi-source heterogeneous teaching document to obtain a standardized paragraph set; S14 performs redundant symbol removal processing on each paragraph unit in the standardized paragraph set to obtain a redundant paragraph set; S15. Perform missing placeholder completion processing on each de-redundant paragraph unit in the de-redundant paragraph set to form a placeholder completion paragraph unit, and all placeholder completion paragraph units are recombined into a standardized teaching document set.
3. The method for generating educational courseware based on deep learning according to claim 2, characterized in that: The step S2 includes the following contents: S21. Input the standardized teaching document collection into the improved sparse gated MoE network model, perform deep teaching semantic feature encoding at the paragraph unit granularity, and obtain the paragraph unit teaching embedding vector in, Represents the teaching goal orientation and knowledge structure semantic information implied by the j-th paragraph unit in the i-th teaching document fragment, and d is the embedding vector dimension; S22. Improve the sparse gating mechanism for the structured PPT generation scenario of teaching documents and design a dynamic adaptive gating function guided by teaching semantics Dynamically adaptive gating function to teach embedding vectors in paragraph units As input, adaptively predict the degree of match between each expert sub-network and the current paragraph unit teaching objectives and knowledge structure, and obtain the expert activation score vector Among them, the vector dimension K is the total number of experts in the expert sub-network set; S23. Introducing the expert load balancing factor ρ in the scenario of generating structured PPTs for teaching documents balance , used to dynamically adjust the expert activation score vector Forming a sparse expert activation strategy for optimizing the teaching content structure, the expert load balancing factor ρ balance By real-time counting the number of times each expert sub-network is activated in the process of generating structured PPTs, online balancing of the expert sub-network load and optimal allocation of computing resources are performed; S24. Based on the sparse expert activation strategy, the optimized sparse activation operation TopK is used opt (·) activates the scoring vector for the expert Screen and determine the expert sub-network set that best matches the current paragraph unit teaching structure S25. Activated expert sub-network set Teaching embedding vectors to input paragraph units Perform parallel multi-teaching task recognition, which includes deep semantic recognition of teaching objectives, mapping of knowledge points and course chapters, matching of teaching case content, and annotation of the relevance between exercises and knowledge points. Output of structured teaching element recognition results through multi-expert parallel reasoning in, They represent the deep semantic recognition results of teaching objectives, the mapping results of knowledge points and course chapters, the matching results of teaching case content, and the annotation results of the relevance between exercises and knowledge points. S26. Structural integration of all structured teaching element recognition results to form a teaching element set Y for the entire standardized teaching document set teach .
4. The method for generating educational courseware based on deep learning according to claim 3, characterized in that: The sparse expert activation strategy includes the following: S241. Setting the expert activation scoring vector in the teaching scenario in, Denotes the expert subnetwork E k The teaching semantic matching score of the jth paragraph unit in the i-th teaching document fragment is given. The vector dimension K represents the total number of expert subnetworks in the improved sparse gated MoE network. S242. Define the expert subnetwork load vector as u global ={u (1) ,u (2) ,...,u (K) }, where u (k) Denotes the expert subnetwork E k The normalized value of the historical activation frequency during the structured PPT generation process is used to measure the computational load ratio of each expert sub-network in the entire training iteration; S243. Activate the expert scoring vector Perform load regularization adjustment to construct the balanced expert activation score vector Load regularization adjustment uses a load suppression factor β to penalize expert sub-networks with high historical activation frequencies; S244. Constructing a mask vector for the correlation between teaching structures in Denotes the expert subnetwork E k Whether the history task has a high matching task label under the teaching structure of the current paragraph unit. If so, the value is 1, otherwise 0; S245. According to the teaching structure matching relationship, the balanced expert activation score vector is structurally biased to form the final expert activation score vector guided by the teaching structure. The structural bias factor is δ, which represents the improvement weight that the structural association expert sub-network should obtain; S246. Activate the final expert scoring vector Perform a sorting operation and select the top k expert sub-networks with the highest scores to form the expert activation set of the current paragraph unit: in, represents the set of sparsely activated expert sub-networks corresponding to the j-th paragraph unit in the i-th teaching document fragment.
5. The method for generating educational courseware based on deep learning according to claim 3, characterized in that: The step S3 includes the following contents: S31. Extract each structured teaching element recognition result from the structured teaching element set as a node element of the teaching semantic graph to form a node set. Each structured teaching element recognition result includes a teaching objective, a knowledge point, a case content, or an exercise prompt. The node set is used to construct the structural skeleton of the teaching semantic graph. S32. Based on the timestamp value of the original teaching document fragment to which each structured teaching element belongs and its sequence label in the four-element teaching structure of teaching objectives-knowledge points-case studies-exercises, establish directed connections between the nodes. The rule for establishing directed connections is that if a node precedes another node in terms of timestamp value and its structure label sequence is equal to or earlier than that of the node, then a directed edge is established between the two nodes, forming a directed edge set. S33. Combining the node set and the directed edge set to construct a complete teaching semantic graph, and representing the teaching semantic graph as a structured teaching semantic structure data, the teaching semantic structure data includes a node set and a directed adjacency matrix, the adjacency matrix is used to indicate whether there is a teaching logical connection relationship between the nodes; S34. Input the teaching semantic structure data into the improved sparse gated MoE network, divide the tasks according to the preset expert division of labor, including content reconstruction experts, key point summary experts, chart matching experts, and layout draft generation experts. Based on the structured teaching element attributes of each node in the teaching semantic structure data, an expert routing matrix is constructed to mark whether each structured teaching element is processed by a specific expert sub-network; S35. The expert subnetworks process their respective tasks in parallel, performing content reorganization, key point summary extraction, diagram matching generation, and layout draft prediction on the structured teaching element nodes. For each node, a structural representation of a slide segment is generated. This structural representation is composed of content reorganization, key point summary extraction, diagram matching generation, and layout draft prediction. All slide segments are then sequentially combined according to the topological order of the teaching semantic graph to form a draft slide sequence. S36. Record the expert activation status and layout parameter output corresponding to each structured teaching element node during the generation process, forming a preliminary draft of the sparse expert activation strategy and a preliminary draft of the layout parameter vector. The preliminary draft of the sparse expert activation strategy is used to indicate which expert sub-networks participate in the processing of each structured teaching element. The preliminary draft of the layout parameter vector is used to describe the visual presentation parameters of each structural element in the slide, including the proportion of the title, the contrast ratio of the graphic content density and the color scheme.
6. The method for generating educational courseware based on deep learning according to claim 1, characterized in that: The step S4 includes the following contents: S41. The draft of the sparse expert activation strategy, the draft of the layout parameter vector, and the draft of the slide sequence are represented as the expert activation code, the layout parameter sequence, and the content structure sequence, respectively, and jointly constitute the original individual population of the pigeon flock optimization algorithm. S42. Define the teaching structure coverage index ψ (t) , used to evaluate the coverage identification value of each slide in the current structured slide sequence in the tth iteration to see whether it completely covers the teaching objective-knowledge point-case-exercise four-element teaching structure. The coverage identification value is determined by detecting whether the slide contains the teaching objective element, knowledge point element, case content element, and exercise prompt element at the same time. If all of them are included, it is recorded as 1, and if any of the structural elements are missing, it is recorded as 0. The coverage identification values of all slides are summed up and divided by the total number of slides to obtain the teaching structure coverage index; S43. Define the visual consistency deviation function φ vis , used to measure whether the generated slide layout parameters are consistent with teaching visual standards in terms of image and text density, white space ratio, and color palette contrast: in, Indicates the image and text density ratio of page k, is the target image density threshold, Indicates the page blank ratio, θ ref Leave blank reference values for expectations; S44. Introducing the expert routing entropy regularization term H route , used to evaluate whether the activation distribution of the expert sub-network is balanced: Among them, π l Denotes the expert subnetwork E l The probability distribution of being activated in the current individual, K is the total number of expert sub-networks; S45. Constructing the joint fitness function F of comprehensive teaching structure coverage index, visual consistency deviation function and expert routing entropy regularization term task , as the main objective function of the pigeon flock optimization algorithm: F task =λ1·ψ (t) -λ2·φ vis -λ3·H route ; Among them, λ1, λ2, and λ3 are the adjustment weights of each sub-goal respectively; S46. Construct a navigation matrix to guide the individual structure of the pigeon flock to adjust its direction. According to the teaching semantic graph adjacency matrix A teach The constructed guidance factor matrix, navigation matrix By subtracting the teaching semantic graph adjacency matrix A from the identity matrix I teach The weighted form of , where the weighting factor is the graph topology influence factor γ, which is a positive real number between zero and one; S47. Evolutionary adjustment of the structure perception of the individual population of the pigeon swarm optimization algorithm based on the position update formula modified by the navigation matrix: Among them, η is the flight step length control factor, represents the individual structure with the best fitness in round t, represents the current position vector of the k-th pigeon group optimization individual at the t-th iteration, represents the position vector of the k-th optimized individual in the pigeon group at the t+1-th iteration; S48. When the value of the joint fitness function of the pigeon population does not improve in multiple consecutive rounds, or the number of iterations reaches the maximum number of rounds, the optimization process is terminated and the optimal fitness individual is output. Its corresponding sparse expert activation strategy G opt , layout parameter vector L opt and slide structure content sequence P opt .
7. The method for generating educational courseware based on deep learning according to claim 1, characterized in that: The step S5 includes the following contents: S51. Feedback the optimized sparse expert activation strategy to the improved sparse gated MoE network, replacing the original sparse expert activation strategy draft, as the basis for expert activation control in the re-inference phase, and used to guide the sparse gated MoE network to activate only the optimized expert sub-network set when regenerating structured content. Each of the optimized sparse expert activation strategies The expert routing state corresponding to the content structure of the slide on page k represents which expert sub-networks should participate in generating the content of the page structure; S52. The optimized sparse expert activation strategy and the teaching semantic structure data are jointly input into the improved sparse gated MoE network model to maintain the teaching semantic graph structure G teach Unchanged, the expert sub-network division of labor structure is used to re-reason the teaching semantic structure data, generate an optimized content set, and optimize each This includes optimized content reorganization information, key point summary information, chart suggestion information, and content components that are semantically consistent with their teaching nodes; S53. Apply the optimized layout parameter vector to each item in the optimized content set, and The title ratio, image and text density, white space ratio and color matching style are used to visually arrange the components of the content module, adjust the margins and bind the colors; S54. Structural splicing of all optimized slide pages and output of structured PPT draft Represents the structured output content of slide k.
8. The method for generating educational courseware based on deep learning according to claim 1, characterized in that: The structured output content consists of the following structure: If the slide contains the teaching objective elements, knowledge point elements, case content elements, and exercise prompt elements, and the layout structure meets the requirements of the title ratio ≥ 15%, the image and text density is in the range of [30%, 70%], and the white space ratio is not less than 10%, then the page is output as a structurally consistent page; If the slide on this page lacks any structured teaching elements, or the image and text density exceeds 80% and the title ratio is less than 10%, the page will be output as a content-heavy page; If the slide on this page contains all the teaching elements but the color palette contrast is lower than the set threshold, the color matching style does not conform to the high contrast principle, or the white space rate is less than 5%, the page will be output as a visually unbalanced page.
Citation Information
Cited By
Intelligent document generation system and control method thereof
CN120781990A
PPT style conversion method and system
CN121788662A