Digital teaching video courseware automatic generation method

By constructing a multi-level conceptual network and dynamic content topology, the problem of insufficient expression of complex knowledge systems in existing technologies is solved, enabling high-quality generation of digital teaching video courseware and ensuring the logical coherence and semantic accuracy of the content.

CN122053931APending Publication Date: 2026-05-15BEIJING ZHENSHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHENSHI TECHNOLOGY CO LTD
Filing Date
2025-12-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods for generating digital teaching video courseware cannot effectively express the cross-references and multi-dimensional relationships in complex knowledge systems. As a result, the generated courseware content is superficial and lacks logical coherence. Furthermore, the keyword matching mechanism has a shallow understanding of semantics and cannot distinguish the subtle differences of the same word in different teaching contexts, which affects the accuracy of the content and the effectiveness of teaching.

Method used

By analyzing the internal logical structure of the original teaching materials, a multi-level conceptual network is constructed. Key nodes are located in the semantic space and semantic vectors are labeled. Multimedia content fragments are matched from the material library to construct a dynamic content topology, plan the generation path, and monitor the assembly process in real time to optimize the trajectory, ensuring the semantic fluency and reasonable rhythm of the content assembly.

Benefits of technology

It achieves in-depth analysis and mathematical representation of knowledge structure, accurately matches multimedia materials, reduces the risk of irrelevant or ambiguous content, ensures high quality and teaching relevance of courseware, and avoids contextual disconnect or logical jumps caused by fixed processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053931A_ABST
    Figure CN122053931A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital education resource automatic generation, and discloses a digital teaching video courseware automatic generation method. The method comprises the steps that teaching material logic is analyzed to construct a concept network, semantic vectors are marked for key nodes, and according to the semantic vectors, multimedia materials are accurately matched to form standardized content primitives; and constructing a dynamic content topology based on a network structure and primitive semantics, planning a content generation path and calculating a node expected state. In the process of driving the primitives to be assembled in a serialized mode along the path, state deviation is monitored in real time, a track is dynamically optimized and generated, and finally, courseware is output after assembly products are subjected to coherence stitching and rendering. According to the method, the logic fluency of content evolution is ensured through dynamic path planning and real-time optimization, and meanwhile, the matching precision of materials and teaching concepts is improved by utilizing semantic vector matching, so that the self-adaption and intelligence of the courseware generation process are realized, and the quality of the generated courseware and the teaching effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic generation technology of digital educational resources, specifically a method for automatically generating digital teaching video courseware. Background Technology

[0002] Currently, the automated generation of digital teaching video courseware mainly relies on preset templates and rule bases. The production process typically involves cutting teaching materials into linear sequences, retrieving corresponding multimedia elements from a resource library through keyword matching, and finally splicing them together sequentially. This method treats teaching content as a static, flat sequence, lacking a deep understanding of the inherent logical hierarchy and conceptual connections of knowledge, as well as the ability to dynamically organize it.

[0003] Existing technical solutions have shortcomings. Assembly models based on linear or simple tree structures struggle to express the cross-referencing and multidimensional relationships within complex knowledge systems, resulting in superficial courseware content with abrupt transitions between knowledge points and a lack of logical coherence. Furthermore, simple keyword matching mechanisms offer only a superficial understanding of semantics, failing to distinguish the subtle differences in the same word across different teaching contexts. This easily introduces irrelevant or semantically conflicting materials, impacting content accuracy and teaching effectiveness. The connection between teaching content and multimedia materials remains at a shallow level, unable to be dynamically adjusted and optimized based on the actual state of the content assembly process. Summary of the Invention

[0004] The purpose of this invention is to provide a method for automatically generating digital teaching video courseware to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides a method for automatically generating digital teaching video courseware, the method comprising: Analyze the internal logical structure of the original teaching materials, and construct a multi-level concept network based on the internal logical structure; Locate the key nodes of the multi-level concept network in the semantic space, and label the key nodes with semantic vectors; Based on the semantic vector, the corresponding multimedia content fragments are matched and extracted from the original material library to form an original content primitive set; Semantic integrity verification and conflict detection are performed on the original content primitive set to generate verified and normalized content primitives; Based on the structure of the multi-level conceptual network and the semantic attributes of the standardized content primitives, a dynamic content topology is constructed. Plan the content evolution generation path on the dynamic content topology, and calculate the expected state of each node on the generation path; Based on the expected state, the content synthesis engine is scheduled to drive the normalized content primitives to be serialized and assembled along the generation path; The serialization assembly process is monitored in real time, the state offset is captured, and the generated path is optimized in real time based on the state offset. After the serialization assembly is completed, multi-level coherent stitching and adaptive rendering are performed on the assembled product; Output the final generated digital teaching video courseware stream.

[0006] Preferably, the step of analyzing the internal logical structure of the original teaching materials and constructing a multi-level concept network based on the internal logical structure includes: Identify heading levels, knowledge point annotations, and logical connectors in original teaching materials; Using the title level and knowledge point annotation as candidate nodes, and the logical relationships defined by the logical connectors as edges, an initial concept map is generated. Calculate the centrality and connectivity density of each node in the initial concept graph, and filter out core concept nodes and derived concept nodes based on the centrality and connectivity density. Logical hierarchy labels are attached to the core concept nodes, and the initial concept graph is expanded into a multi-layered concept network containing a core layer, an explanatory layer, and an instance layer based on the logical hierarchy labels.

[0007] Preferably, locating key nodes of the multi-level concept network in the semantic space and labeling the key nodes with semantic vectors includes: Input all node names and descriptive text from the multi-level conceptual network into a pre-trained semantic encoder; Obtain the high-dimensional vector representation of each node in the semantic space output by the semantic encoder; Calculate the cluster center of the high-dimensional vector representation in the semantic space, and determine the nodes closest to the cluster center as the key nodes of the multi-level concept network; Assign a corresponding high-dimensional vector representation to the key node, which serves as the semantic vector of the key node.

[0008] Preferably, the step of matching and extracting corresponding multimedia content fragments from the original material library based on the semantic vector to form an original content primitive set includes: The semantic vector of the key node is compared with the index vector of all materials in the original material library to calculate the similarity. Materials with a similarity exceeding a predetermined threshold are selected as candidate materials for the key node; Perform timestamp or fragment boundary analysis on the candidate materials to cut out the independent fragments that are most semantically related to the key nodes; All independent segments corresponding to key nodes are aggregated and their source key node identifiers are marked to form the original content primitive set.

[0009] Preferably, the step of performing semantic integrity verification and conflict detection on the original content primitive set to generate verified normalized content primitives includes: Analyze the internal information density of each original content primitive and compare it with the information carrying capacity of the key nodes associated with the original content primitive to detect missing or redundant information. Compare whether there are factual contradictions or logical conflicts between different original content primitives; For original content primitives with missing information, a supplementary retrieval process is initiated to obtain compensation fragments and then merge them; For original content primitives that contain redundant or conflicting information, perform content trimming or semantic rewriting operations. The content primitives that have undergone verification, supplementation, trimming, or rewriting operations are marked as normalized content primitives.

[0010] Preferably, the construction of the dynamic content topology based on the structure of the multi-level conceptual network and the semantic attributes of the normalized content primitives includes: Extract the connection relationships between all nodes in the multi-level conceptual network and map them into a topological connection skeleton; The duration, media type, and sentiment semantic attributes of the standardized content primitives are used as dynamic parameters and attached to the corresponding nodes of the topological connection skeleton. Define the transmission rules for the mutual influence between different semantic attributes, and simulate the propagation and diffusion of the dynamic parameters on the topological connection skeleton according to the transmission rules; Based on the results of propagation and diffusion, a dynamic content topology containing node parameters, edge weights, and state transition probabilities is generated.

[0011] Preferably, the step of planning the content evolution generation path on the dynamic content topology and calculating the expected state of each node on the generation path includes: The path starts at the logical starting node in the dynamic content topology and ends at the logical ending node. Based on the state transition probabilities in the dynamic content topology, a path search algorithm is used to calculate multiple candidate content evolution paths from the starting point of the path to the ending point of the path; Evaluate the total coherence cost and cognitive load cost of each candidate content evolution path; The candidate content evolution path with the optimal overall cost is selected and determined as the final content evolution generation path; Based on the node parameters in the dynamic content topology, the media state, knowledge concentration state, and rhythm state that each node should present when the content evolves along the generation path are deduced as the expected state.

[0012] Preferably, the step of scheduling the content synthesis engine according to the expected state and driving the normalized content primitives to be serialized and assembled along the generation path includes: The sequence of nodes on the generated path and the expected state of each node are converted into a sequence of synthetic instructions. The synthesis instruction sequence is matched and bound to the standardized content primitive pool to generate a primitive assembly queue with timestamps and transition requirements; The content synthesis engine sequentially reads and executes the instructions in the primitive assembly queue, calls the corresponding standardized content primitives, and applies visual transitions, audio mixing, and subtitle synchronization operations according to the instructions to achieve serialized assembly.

[0013] Preferably, the real-time monitoring of the serialization assembly process, capturing the state offset, and performing real-time trajectory optimization of the generated path based on the state offset includes: During the serialization assembly process, the actual media features and rhythmic features of the assembled portions are periodically sampled; The actual features sampled are compared with the expected state of the current node to calculate the state offset. When the state offset exceeds the tolerance threshold, the trajectory optimization mechanism is triggered; The trajectory optimization mechanism takes the current actual assembly state as a new starting point and replans the generation path of the remaining nodes in the dynamic content topology. The redesigned generation path is updated in subsequent synthesis instruction sequences.

[0014] Preferably, the step of performing multi-level coherent stitching and adaptive rendering on the assembled product after the serialization assembly is completed includes: Detect the audio waveforms and video streams of the serialized assembly products, and identify jumps and discontinuities at the segment connections. The audio crossfade-in / fade-out algorithm and the video motion compensation algorithm are used to smoothly stitch together the jumps and discontinuities. Based on the format requirements and bitrate limitations of the target playback platform, the encoding parameters of the stitched complete stream are adaptively configured and transcoded for rendering. Global color correction and audio loudness equalization are applied to generate the final digital teaching video courseware stream.

[0015] Compared with the prior art, the beneficial effects of the present invention are: By constructing a dynamic content topology and planning generation paths on it, the system transforms the organization of teaching content from a static, predetermined process into a dynamically evolving, networked process. During the serialization and assembly process, the system monitors and captures the offset between the current assembly state and the expected state calculated based on semantic attributes in real time, and optimizes the subsequent generation paths accordingly. This enables the content assembly to be adaptive, dynamically adjusting the organization order and presentation of subsequent content based on the actual synthesis effect of preceding content. This avoids contextual disconnects or logical jumps caused by a fixed process, ensuring that the courseware maintains semantic fluency and a reasonable rhythm throughout the dynamic generation process.

[0016] By analyzing the internal logical structure of teaching materials to construct a multi-layered conceptual network and annotating key nodes with semantic vectors in the semantic space, a deep analysis and mathematical representation of the knowledge structure is achieved. This integrates the logical position and semantic information of concepts within specific teaching contexts, providing subsequent material matching with accuracy and contextual awareness far exceeding keyword matching. Based on this vector representation, multimedia content fragments are extracted from the original material library, accurately identifying materials with the highest semantic fit to specific conceptual nodes. This reduces the risk of introducing irrelevant or ambiguous content, ensuring the accuracy and relevance of basic content units, and laying a reliable content foundation for constructing high-quality courseware. Attached Figure Description

[0017] Figure 1 This is a schematic diagram illustrating the working principle of the automatic generation method for digital teaching video courseware described in this invention. Figure 2 A flowchart for constructing a multi-level conceptual network; Figure 3 A flowchart for constructing a dynamic content topology; Figure 4 Generate a path cost comparison chart for content evolution; Figure 5 A comparison chart showing the jump values ​​before and after stitching together the content of the courseware. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Please see Figure 1This invention provides a method for automatically generating digital teaching video courseware. The method includes: first, inputting original teaching materials into the system, parsing their internal logical structure, and constructing a multi-level conceptual network based on the parsed structure. Then, locating key nodes of the conceptual network in a pre-constructed semantic space and labeling these key nodes with corresponding semantic vectors. Based on these semantic vectors, the system performs matching and retrieval from a raw material library storing a large amount of video, audio, and textual data, extracting multimedia content fragments most semantically relevant to the nodes to form an original content primitive set. This set needs to undergo semantic integrity verification and logical conflict detection, supplementing, trimming, or correcting the content to generate a series of standardized and usable normalized content primitives.

[0020] By combining the structural relationships of multi-level conceptual networks with the semantic attributes of standardized content primitives, a dynamic content topology containing dynamic parameters and state transition relationships is constructed. On this topology, the system plans a content evolution generation path and calculates the expected state that each node on the path should reach during content synthesis. Based on this path and state sequence, the content synthesis engine is scheduled, driving each standardized content primitive to assemble according to strict sequence and state requirements. The assembly process is monitored in real time, and the generation path is dynamically optimized by capturing the deviation between the actual state and the expected state. After serialized assembly, the initially synthesized assembly products are smoothed at the junctions and encoded and rendered according to the target platform, ultimately outputting a high-quality digital teaching video courseware stream.

[0021] Example 1: See Figure 2 This process analyzes the internal logical structure of the original teaching materials and constructs a multi-level concept network based on this structure. It identifies heading levels, knowledge point annotations, and logical connectors within the materials. Using the identified heading levels and knowledge point annotations as candidate nodes and the logical relationships defined by the connectors as edges, an initial concept graph is generated. The centrality and connectivity density of each node in this initial concept graph are calculated, and core and derived concept nodes are selected based on these values. Logical level labels are attached to the core concept nodes, and the initial concept graph is expanded into a multi-level concept network comprising a core layer, an explanatory layer, and an example layer based on these labels.

[0022] The process of locating key nodes in a multi-level conceptual network in semantic space and labeling them with semantic vectors involves inputting all node names and descriptive text from the multi-level conceptual network into a pre-trained semantic encoder. The high-dimensional vector representation of each node in semantic space is obtained from the encoder's output. Cluster centers of these high-dimensional vector representations are calculated in semantic space, and the nodes closest to these cluster centers are identified as key nodes in the multi-level conceptual network. Finally, corresponding high-dimensional vector representations are assigned to these key nodes as their semantic vectors.

[0023] In its implementation, the automatic generation method for digital teaching video courseware involves parsing the internal logical structure of the original teaching materials and constructing a multi-level concept network, as well as locating key nodes in the semantic space and labeling semantic vectors. After the original teaching materials are input into the system, the system identifies the title levels, knowledge point annotations, and logical connectors within the materials. Using the identified title levels and knowledge point annotations as candidate nodes and the logical relationships defined by logical connectors as edges, the system generates an initial concept graph. This initial concept graph can be understood as a directed or undirected graph structure composed of nodes and edges, where nodes represent knowledge units and edges represent the logical connections between these units. The system calculates the centrality and connectivity density of each node in the initial concept graph. Centrality reflects the importance and pivotal role of a node within the entire graph, while connectivity density reflects the tightness of the connection between a node and its directly connected nodes. Based on centrality and connectivity density, core concept nodes and derived concept nodes are selected. Core concept nodes typically have high centrality, while derived concept nodes have relatively high local connectivity density. Logical level labels are attached to core concept nodes; these labels identify the abstraction level of the concept within the knowledge system. Based on the logical hierarchy labels, the initial concept map is expanded into a multi-layered concept network containing a core layer, an explanatory layer, and an instance layer. The core layer contains the most core and abstract concept nodes, the explanatory layer contains nodes that explain and expand on the core concepts, and the instance layer contains specific cases, applications, or data nodes.

[0024] After constructing a multi-level concept network, the system locates key nodes of the network in the semantic space. In practice, all node names and descriptions from the multi-level concept network are input into a pre-trained semantic encoder, which can be a Transformer-based model. The system then obtains the high-dimensional vector representation of each node in the semantic space, output by the semantic encoder. The cluster centers of these high-dimensional vector representations are calculated, representing the average or central position of the concept set within the semantic space. The cluster centers can be calculated using the mean method, with the following formula: in: Represents the cluster center vector. This represents the total number of nodes in a multi-level conceptual network. Indicates the first A high-dimensional vector representation of each node. The nodes closest to the cluster center are identified as key nodes in the multi-level conceptual network; distance calculation typically uses cosine similarity or Euclidean distance. Each key node is assigned a corresponding high-dimensional vector representation, which serves as its semantic vector. These semantic vectors are used as the basis for subsequent matching of content fragments from the source material library.

[0025] In some embodiments, the parsing process supports various formats of original teaching materials, including text documents, presentations, and structured ebooks. The system identifies heading levels and knowledge point annotations by parsing the document's style tags, paragraph structure, and specific markers. The identification of logical connectors is based on a predefined connector vocabulary, which contains words representing cause and effect, contrast, progression, and exemplification. In some embodiments, the selection of core concept nodes and derived concept nodes is based on set centrality thresholds and connection density thresholds. Optionally, the addition of logical hierarchy labels can be automatically determined based on the node's position order in the original teaching material, font format, and connection patterns with other nodes.

[0026] Example 2: This implementation method involves matching and extracting corresponding multimedia content fragments from the original material library based on semantic vectors to form an original content primitive set. The semantic vector of the key node is compared with the index vectors of all materials in the original material library to calculate similarity. Materials with similarity exceeding a predetermined threshold are selected as candidate materials for that key node. Timestamp or fragment boundary analysis is performed on the candidate materials to extract the independent fragments most semantically related to the key node. All independent fragments corresponding to the key nodes are aggregated and their source key node identifiers are marked to form the original content primitive set.

[0027] The process involves performing semantic integrity verification and conflict detection on the original content primitive set to generate verified normalized content primitives. This includes analyzing the internal information density of each original content primitive and comparing it with the information carrying capacity of the key nodes associated with that primitive to detect missing or redundant information. It also involves comparing different original content primitives to identify factual contradictions or logical conflicts. For original content primitives with missing information, a supplementary retrieval process is initiated to obtain and merge compensated fragments. For original content primitives with redundant or conflicting information, content pruning or semantic rewriting operations are performed. Content primitives that have undergone verification, supplementation, pruning, or rewriting are marked as normalized content primitives.

[0028] In its implementation, the automatic generation method for digital teaching video courseware involves matching and extracting multimedia content segments from the original material library based on the semantic vectors of key nodes, as well as performing semantic integrity verification and conflict detection. After obtaining the semantic vectors of key nodes, the system calculates the similarity between these semantic vectors and the index vectors of all materials in the original material library. The original material library is a structured database containing various types of multimedia materials, including videos, audios, images, and text-based explanations. Each material undergoes feature extraction upon entry into the library, generating an index vector representing its core semantic content. The similarity calculation aims to quantify the closeness between the semantic vectors of key nodes and the index vectors of the materials in the semantic space. The calculation formula is as follows: in: Represents the similarity value. The semantic vector representing the key node. This represents the index vector of a specific material in the original material library. Representing vectors with vector A normalized distance metric function between them. The system selects similarity values. Exceeding the predetermined threshold The system uses selected materials as candidate materials for the current key node. For each candidate material, the system performs timestamp or segment boundary analysis. Timestamp analysis is performed on video or audio materials with metadata tags, while segment boundary analysis segments independent segments by analyzing visual scene transitions, audio energy abrupt changes, or text theme shifts. The system extracts the independent segments most semantically related to the key node from the candidate materials, aggregates all independent segments from different original materials but related to the same key node, and marks the key node identifier of their source to form the original content primitive set.

[0029] After forming the initial set of content primitives, the system performs semantic integrity verification and conflict detection on the set. In practice, the internal information density of each initial content primitive is analyzed. This internal information density can be measured by calculating the frequency of key terms and the number of times concepts are mentioned per unit time. The internal information density of each initial content primitive is compared with the information carrying capacity of the key nodes associated with that primitive. The information carrying capacity of the key nodes is derived from indicators such as node centrality and connectivity calculated during the construction of the multi-level concept network. Information loss or redundancy is detected through comparison. Information loss is manifested when the internal information density of the initial content primitive is significantly lower than the information carrying capacity of the associated key nodes; information redundancy is manifested when the internal information density of the initial content primitive is significantly higher than the information carrying capacity of the associated key nodes. The system also compares whether there are factual contradictions or logical conflicts between different initial content primitives. Factual contradiction detection compares whether the specific data, definitions, and dates stated in the primitives are consistent; logical conflict detection analyzes whether the arguments and reasoning relationships between the primitives are consistent. For original content primitives with missing information, the system initiates a supplementary retrieval process to obtain compensating fragments and fuse them. The supplementary retrieval process may expand the similarity threshold range or search in an expanded material library. For original content primitives with redundant or conflicting information, the system performs content trimming or semantic rewriting operations. Content trimming directly removes redundant or contradictory parts, while semantic rewriting corrects the expression to eliminate conflicts.

[0030] In some embodiments, the normalized distance metric function used for similarity calculation It is a cosine distance function. Predetermined threshold. The value can be configured according to the knowledge depth and rigor requirements of the generated video courseware. In some embodiments, the index vector generation of the original material library uses a pre-trained semantic encoder that is the same as or semantically aligned with the semantic vectors used to generate the key node semantic vectors. Optionally, timestamp or segment boundary analysis can be combined with the transcript generated by automatic speech recognition to help determine more accurate segment cutting points through text analysis.

[0031] Example 3: See Figure 3 Based on the structure of a multi-layered conceptual network and the semantic attributes of standardized content primitives, a method for constructing dynamic content topology is proposed. This method extracts the connections between all nodes in the multi-layered conceptual network and maps them to a topological connection skeleton. The duration, media type, and sentiment semantic attributes of the standardized content primitives are used as dynamic parameters and appended to the corresponding nodes of the topological connection skeleton. Transmission rules governing the interaction between different semantic attributes are defined, and the propagation and diffusion of dynamic parameters on the topological connection skeleton are simulated according to these rules. Based on the results of propagation and diffusion, a dynamic content topology containing node parameters, edge weights, and state transition probabilities is generated.

[0032] This process involves planning the content evolution generation path on a dynamic content topology and calculating the expected state of each node along the path. The logical starting node in the dynamic content topology is used as the path's starting point, and the logical ending node as the path's ending point. Based on the state transition probabilities in the dynamic content topology, a path search algorithm is used to calculate multiple candidate content evolution paths from the starting point to the ending point. The total coherence cost and cognitive load cost of each candidate content evolution path are evaluated. The candidate content evolution path with the optimal overall cost is selected as the final content evolution generation path. Based on the node parameters in the dynamic content topology, the media state, knowledge concentration state, and rhythm state that each node should present when the content evolves along the generation path are deduced as the expected state.

[0033] In its implementation, the method for automatically generating digital teaching video courseware involves constructing a dynamic content topology based on the structure of a multi-level conceptual network and the semantic attributes of standardized content primitives, as well as planning the generation path of content evolution and calculating the expected state of nodes on the dynamic content topology. The system extracts the connection relationships between all nodes in the multi-level conceptual network, including parent-child relationships, reference relationships, and sequence relationships, and maps them to a topological connection skeleton. The topological connection skeleton is a graph structure that retains the nodes and their original connection relationships, but does not yet contain specific media content parameters. The duration, media type, and sentiment semantic attributes of the standardized content primitives are used as dynamic parameters and attached to the corresponding nodes of the topological connection skeleton. The duration parameter indicates the playback length of the standardized content primitive, the media type parameter indicates whether the standardized content primitive is video, animation, image, or audio, and the sentiment semantic parameter indicates the teaching tone carried by the standardized content primitive. Transmission rules are defined to describe how changes in the parameters of one node affect the parameters of its neighboring nodes. For example, a longer video node might increase the cognitive load weight of its successor nodes. The propagation and diffusion of dynamic parameters along the topological connection skeleton are simulated based on these transmission rules, using an iterative computation method. Based on the propagation and diffusion results, a dynamic content topology is generated, containing node parameters, edge weights, and state transition probabilities. Indicates from node Transfer to node The probability of this can be calculated based on the semantic relevance between nodes, the original weights of the connecting edges, and the nodes themselves. The cognitive load parameters are determined comprehensively, and the formula is as follows: in: Indicates from node To the node The state transition probability, Represents a node With nodes The semantic relevance strength, Representing nodes in a multi-level conceptual network With nodes The weights of the original connecting edges between them. Represents a node Cognitive load parameters, It is a fusion function used to map multiple input factors to a single probability value.

[0034] After constructing the dynamic content topology, the system plans the content evolution generation path on the dynamic content topology. In specific implementation, the logical starting node in the dynamic content topology is used as the path start point, and the logical ending node is used as the path end point. The logical starting node is usually a node representing the introduction or basic concepts, and the logical ending node is usually a node representing the summary or advanced applications. Based on the state transition probabilities in the dynamic content topology, a path search algorithm is used to calculate multiple candidate content evolution paths from the path start point to the path end point. The path search algorithm can be a heuristic graph search algorithm. The total coherence cost and cognitive load cost of each candidate content evolution path are evaluated. The total coherence cost quantifies the smoothness of topic switching between adjacent nodes in the path, and the cognitive load cost quantifies the total intellectual consumption that learners are expected to bear when learning along the path. The candidate content evolution path with the optimal comprehensive cost is selected. The comprehensive cost is the weighted sum of the total coherence cost and the cognitive load cost, and this is determined as the final content evolution generation path. Based on the node parameters in the dynamic content topology, we can deduce the media state, knowledge concentration state, and rhythm state that each node should present as the content evolves along the generation path. As expected states, the media state describes the specific visual and auditory presentation of the content at that node, the knowledge concentration state describes the density of information output per unit time, and the rhythm state describes the coordination relationship between the content at that node and the content at the preceding and following nodes in terms of playback speed.

[0035] In some embodiments, the simulation of propagation rules can be achieved by constructing an attribute propagation graph model on the topological connection skeleton. Cognitive load parameters The calculation can consider the duration of normalized content primitives, media type complexity, and the knowledge span difference with predecessor nodes. In some embodiments, the path search algorithm can employ Dijkstra's algorithm, which considers transition probabilities and local cost factors. Dijkstra's algorithm is adapted for dynamic content topologies to plan content evolution paths. The algorithm treats nodes in the dynamic content topology as vertices of a graph, and edges represent connections between nodes based on logical relationships. The weight of an edge is derived by combining state transition probabilities with local cost factors such as coherence cost and cognitive load cost. During implementation, the logical starting node is used as the source node, the distance values ​​of each node are initialized, and the edge weights are iteratively relaxed using a priority queue to dynamically update the path cost. The path cost calculation incorporates the reciprocal of the state transition probability to reflect the semantic association strength, while also incorporating coherence cost to assess the smoothness of topic switching and cognitive load cost to quantify the cumulative effect of learning difficulty. The algorithm terminates when a logical terminal node is marked as visited, and backtracks to generate the minimum comprehensive cost path, ensuring that content evolution achieves optimal balance between semantic fluency and learning load. Optionally, the determination of optimal overall cost can be achieved by setting a weighted ratio between coherence cost and cognitive load cost, which can be adjusted according to the age and knowledge background of the target learner. Optionally, the rhythm state in the expected state can be specifically defined by calculating the ratio between the duration of the current node's normalized content primitive and the duration of the primitives in the preceding and following nodes.

[0036] See Figure 4 This chart, a grouped bar chart, primarily illustrates two key costs across four paths in the courseware generation path planning: coherence cost and cognitive load cost. The chart visually presents the cost differences between different generation paths, aligning with the logic of evaluating candidate path costs and selecting the optimal overall path. It demonstrates both the impact of coherence on courseware fluency and the role of cognitive load in the learning experience, providing data support for the scientific selection of paths in courseware generation.

[0037] Example 4: An implementation method that schedules the content compositing engine based on expected states, driving the serialization and assembly of standardized content primitives along the generation path. This method transforms the sequence of nodes on the generation path and the expected state of each node into a compositing instruction sequence. The compositing instruction sequence is matched and bound to the standardized content primitive pool, generating a primitive assembly queue with timestamps and transition requirements. The content compositing engine sequentially reads and executes the instructions in the primitive assembly queue, calling the corresponding standardized content primitives, and applying visual transitions, audio mixing, and subtitle synchronization operations according to the instructions to achieve serialization and assembly.

[0038] This method involves real-time monitoring of the serialization and assembly process, capturing state offsets, and optimizing the generation path based on these offsets. During serialization and assembly, the actual media features and rhythmic features of the assembled portions are periodically sampled. The sampled features are compared with the expected state of the current node to calculate the state offset. When the state offset exceeds a tolerance threshold, a trajectory optimization mechanism is triggered. This mechanism uses the current actual assembly state as a new starting point and replans the generation path for the remaining nodes within the dynamic content topology. The replanned generation path is then updated in subsequent synthesis instruction sequences. In practical implementation, the automatic generation method for digital teaching video courseware involves scheduling the content synthesis engine for serialization and assembly based on the expected state, as well as real-time monitoring of the assembly process and trajectory optimization. The system transforms the sequence of nodes on the generation path and the expected state of each node into a synthesis instruction sequence, which is a structured, machine-readable list of commands.

[0039] In practice, the sequence of nodes on the generation path determines the playback order of normalized content primitives, and the expected state of each node specifies the specific parameter requirements that the corresponding normalized content primitive should meet during playback. Each instruction in the synthesis instruction sequence explicitly points to a unique identifier of a normalized content primitive and includes the start time, duration, target visual transition effect, audio mixing parameters, and subtitle synchronization timeline for that primitive's playback. The synthesis instruction sequence is matched and bound to the normalized content primitive pool, which is a centralized storage and index of all available normalized content primitives, generating a primitive assembly queue with timestamps and transition requirements. This primitive assembly queue can be understood as a queue of tasks to be executed, strictly ordered by timestamps. The content synthesis engine sequentially reads and executes the instructions in the primitive assembly queue, calls the corresponding normalized content primitives, and applies visual transitions, audio mixing, and subtitle synchronization operations according to the instruction requirements, achieving serialized assembly. Visual transition operations include fade-in / fade-out, sliding, and zoom effects; audio mixing operations include volume adjustment, background music addition, and sound effect insertion; and subtitle synchronization operations ensure that the timing of text information and voice playback is precisely aligned.

[0040] After the serialization assembly process begins, the system monitors the process in real time, captures state offsets, and optimizes the generated path based on these offsets. Real-time monitoring is achieved by periodically sampling the actual media and rhythmic features of the assembled segments. These samples include the average brightness of the output video stream, motion vector amplitude, loudness of the audio stream, spectral characteristics, and the ratio of the actual playback duration to the planned duration of the current segment. The sampled features are compared with the expected state of the current node to calculate the state offset. The calculation of the state offset involves the fusion of multi-dimensional feature differences. The formula for calculating the state offset E is: in: This represents the calculated normalized state offset. This represents the actual average motion vector amplitude obtained from the sampling. This represents the magnitude of the motion vector set in the expected state. This represents the actual audio loudness obtained from the sampling. This indicates the audio loudness set in the expected state. This indicates the actual playback progress percentage obtained from the sampling. This indicates the planned playback progress percentage set in the expected state. , , These are the weight coefficients corresponding to the visual, audio, and rhythm dimensions, respectively, and satisfy the following conditions: , This is a small constant added to prevent the denominator from being zero. When the calculated normalized state offset E exceeds the preset tolerance threshold, the system triggers the trajectory optimization mechanism. The trajectory optimization mechanism takes the current actual assembly state as the new starting point and replans the generation path of the remaining unassembled nodes in the dynamic content topology. The process of replanning the generation path takes into account the accumulated cognitive load caused by the currently assembled content and the actual rhythm state, and updates the state transition probability weight from the current node to the logical termination node. The replanned generation path is updated in the subsequent synthesis instruction sequence, and the content synthesis engine will continue to execute the subsequent serialization assembly operation according to the updated synthesis instruction sequence, see Table 1.

[0041] Table 1: Dimensions and Weights for State Offset Calculation In some embodiments, the synthesis instruction sequence can be described using a structured data format such as JSON or XML to facilitate parsing and execution by the content synthesis engine. In some embodiments, the frequency of periodic sampling can be adaptively adjusted according to the total duration of the video courseware; for courseware with a longer total duration, the sampling interval can be appropriately increased. Optionally, the tolerance threshold can be differentiated according to the type of different nodes on the generation path; for example, a stricter tolerance threshold can be set for core concept nodes. Optionally, the path search algorithm used when replanning the generation path can be the same as that used in the initial planning, but the initial input cost and heuristic function will be adjusted according to the current actual assembly state.

[0042] Example 5: After serialization and assembly, multi-level coherence stitching and adaptive rendering are performed on the assembled product to detect abrupt changes and discontinuities in the audio waveform and video stream at segment transitions. Audio crossfade-in / fade-out algorithms and video motion compensation algorithms are applied to smoothly stitch these abrupt changes and discontinuities. Based on the format requirements and bitrate limitations of the target playback platform, adaptive configuration of encoding parameters and transcoding rendering are performed on the stitched complete stream. Global color correction and audio loudness equalization are applied to generate the final digital teaching video courseware stream.

[0043] In practical implementation, the automatic generation method for digital teaching video courseware involves performing multi-level coherence stitching and adaptive rendering on the assembled product after serialization and assembly. The system detects abrupt changes and discontinuities in the audio waveform and video stream at segment transitions of the serialized and assembled product. The serialized and assembled product is a preliminary video file generated by the content synthesis engine after sequentially splicing standardized content primitives according to the synthesis instruction sequence. In specific implementation, detecting abrupt changes in the audio waveform at segment transitions involves analyzing abrupt changes in the amplitude and spectral energy distribution of adjacent audio segments near the transition point, and detecting discontinuities in the video stream at segment transitions involves analyzing discontinuous changes in the brightness, color histogram, and motion vector of adjacent video segments near the transition point. It can be understood that the segment transitions of the serialized and assembled product correspond to the connection points between different standardized content primitives, and these connection points may produce perceptual abruptness due to differences in the source materials. Audio crossfade-in / fade-out algorithms and video motion compensation algorithms are applied to smoothly stitch together detected jumps and discontinuities. The audio crossfade-in / fade-out algorithm creates an overlapping transition region before and after the transition point, causing the audio amplitude of the previous segment to gradually decrease while the audio amplitude of the subsequent segment gradually increases. The video motion compensation algorithm analyzes the motion trends of frames before and after the transition point and generates visually coherent transition frames through interpolation or motion vector smoothing to mask the discontinuity caused by sudden changes in scene or content in the video stream.

[0044] After smooth stitching, the system adaptively configures and transcodes the encoding parameters of the stitched complete stream according to the format requirements and bitrate limits of the target playback platform. The target playback platform can be an online education platform, a mobile application, or a local player. Format requirements include container format, video encoding format, audio encoding format, resolution, and frame rate. The bitrate limit specifies the maximum data transmission rate of the video stream. The adaptive configuration of encoding parameters is dynamically determined based on the format requirements and bitrate limits of the target playback platform, as well as the content complexity of the stitched complete stream. In practice, content complexity can be quantified by analyzing the motion intensity, texture detail, and color richness of the video scene. A method for calculating the suggested bitrate for a specific segment is also included. The formula is: in: This represents the suggested video bitrate calculated for this segment. This indicates the base bitrate determined based on the target platform and resolution. This represents the average motion intensity coefficient of the segment. This represents the complexity factor based on the texture details of the segment. The system calculates parameters using a similar formula and then performs transcoding and rendering on the stitched complete video stream. The transcoding and rendering process re-encodes the original video stream into an output stream that conforms to the target specifications.

[0045] After completing the adaptive configuration of encoding parameters and transcoding rendering, the system applies global color correction and audio loudness equalization. Global color correction aims to unify the color tone, contrast, and brightness of all segments in the entire video courseware, eliminating color differences caused by different source materials. Audio loudness equalization aims to unify the perceived volume level of all audio segments in the entire video courseware, avoiding sudden changes in volume between different segments. In practice, global color correction analyzes the color histogram of the entire video stream, calculates a global color mapping curve, and applies this curve to all frames for adjustment. Audio loudness equalization calculates the overall loudness of the entire audio stream according to standards such as ITU-RBS.1770 loudness measurement, and adjusts the gain or attenuation of local segments with excessively high or low loudness. After global color correction and audio loudness equalization, the system generates the final digital teaching video courseware stream, which is a complete video file that meets delivery requirements in terms of both technical specifications and perceived quality.

[0046] In some embodiments, the length of the transition region during smooth stitching can be adaptively adjusted based on the detected abrupt change intensity; a greater abrupt change intensity may require a longer transition region. In some embodiments, the adaptive configuration of encoding parameters also needs to consider the target playback platform's support for keyframe intervals, encoding levels, and grades. Optionally, global color correction can prioritize ensuring color accuracy and readability in face or text regions.

[0047] See Figure 5 This chart, a grouped bar chart, primarily displays the abrupt changes in the connection points of five content elements within the assembled courseware before and after the multi-layered coherence stitching operation. The chart visually verifies the technical effectiveness of the coherence stitching step: through targeted smoothing, it effectively eliminates abrupt changes and discontinuities at the joints of different content elements, making the audiovisual presentation of the courseware smoother and more natural, aligning with the core goal of improving the coherence of the assembled product.

[0048] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for automatically generating digital teaching video courseware, characterized in that, The method includes: Analyze the internal logical structure of the original teaching materials, and construct a multi-level concept network based on the internal logical structure; Locate the key nodes of the multi-level concept network in the semantic space, and label the key nodes with semantic vectors; Based on the semantic vector, the corresponding multimedia content fragments are matched and extracted from the original material library to form an original content primitive set; Semantic integrity verification and conflict detection are performed on the original content primitive set to generate verified and normalized content primitives; Based on the structure of the multi-level conceptual network and the semantic attributes of the standardized content primitives, a dynamic content topology is constructed. Plan the content evolution generation path on the dynamic content topology, and calculate the expected state of each node on the generation path; Based on the expected state, the content synthesis engine is scheduled to drive the normalized content primitives to be serialized and assembled along the generation path; The serialization assembly process is monitored in real time, the state offset is captured, and the generated path is optimized in real time based on the state offset. After the serialization assembly is completed, multi-level coherent stitching and adaptive rendering are performed on the assembled product; Output the final generated digital teaching video courseware stream.

2. The method for automatically generating digital teaching video courseware as described in claim 1, characterized in that, The process of analyzing the internal logical structure of the original teaching materials and constructing a multi-level concept network based on the internal logical structure includes: Identify heading levels, knowledge point annotations, and logical connectors in original teaching materials; Using the title level and knowledge point annotation as candidate nodes, and the logical relationships defined by the logical connectors as edges, an initial concept map is generated. Calculate the centrality and connectivity density of each node in the initial concept graph, and filter out core concept nodes and derived concept nodes based on the centrality and connectivity density. Logical hierarchy labels are attached to the core concept nodes, and the initial concept graph is expanded into a multi-layered concept network containing a core layer, an explanatory layer, and an instance layer based on the logical hierarchy labels.

3. The method for automatically generating digital teaching video courseware as described in claim 2, characterized in that, The step of locating key nodes of the multi-level concept network in the semantic space and labeling the key nodes with semantic vectors includes: Input all node names and descriptive text from the multi-level conceptual network into a pre-trained semantic encoder; Obtain the high-dimensional vector representation of each node in the semantic space output by the semantic encoder; Calculate the cluster center of the high-dimensional vector representation in the semantic space, and determine the nodes closest to the cluster center as the key nodes of the multi-level concept network; Assign a corresponding high-dimensional vector representation to the key node, which serves as the semantic vector of the key node.

4. The method for automatically generating digital teaching video courseware as described in claim 1, characterized in that, The step of matching and extracting corresponding multimedia content fragments from the original material library based on the semantic vector to form an original content primitive set includes: The semantic vector of the key node is compared with the index vector of all materials in the original material library to calculate the similarity. Materials with a similarity exceeding a predetermined threshold are selected as candidate materials for the key node; Perform timestamp or fragment boundary analysis on the candidate materials to cut out the independent fragments that are most semantically related to the key nodes; All independent segments corresponding to key nodes are aggregated and their source key node identifiers are marked to form the original content primitive set.

5. The method for automatically generating digital teaching video courseware as described in claim 4, characterized in that, The step of performing semantic integrity verification and conflict detection on the original content primitive set to generate verified normalized content primitives includes: Analyze the internal information density of each original content primitive and compare it with the information carrying capacity of the key nodes associated with the original content primitive to detect missing or redundant information. Compare whether there are factual contradictions or logical conflicts between different original content primitives; For original content primitives with missing information, a supplementary retrieval process is initiated to obtain compensation fragments and then merge them; For original content primitives that contain redundant or conflicting information, perform content trimming or semantic rewriting operations. The content primitives that have undergone verification, supplementation, trimming, or rewriting operations are marked as normalized content primitives.

6. The method for automatically generating digital teaching video courseware as described in claim 1, characterized in that, The construction of dynamic content topology based on the structure of the multi-level conceptual network and the semantic attributes of the normalized content primitives includes: Extract the connection relationships between all nodes in the multi-level conceptual network and map them into a topological connection skeleton; The duration, media type, and sentiment semantic attributes of the standardized content primitives are used as dynamic parameters and attached to the corresponding nodes of the topological connection skeleton. Define the transmission rules for the mutual influence between different semantic attributes, and simulate the propagation and diffusion of the dynamic parameters on the topological connection skeleton according to the transmission rules; Based on the results of propagation and diffusion, a dynamic content topology containing node parameters, edge weights, and state transition probabilities is generated.

7. The method for automatically generating digital teaching video courseware as described in claim 6, characterized in that, The step of planning the generation path of content evolution on the dynamic content topology and calculating the expected state of each node on the generation path includes: The path starts at the logical starting node in the dynamic content topology and ends at the logical ending node. Based on the state transition probabilities in the dynamic content topology, a path search algorithm is used to calculate multiple candidate content evolution paths from the starting point of the path to the ending point of the path; Evaluate the total coherence cost and cognitive load cost of each candidate content evolution path; The candidate content evolution path with the optimal overall cost is selected and determined as the final content evolution generation path; Based on the node parameters in the dynamic content topology, the media state, knowledge concentration state, and rhythm state that each node should present when the content evolves along the generation path are deduced as the expected state.

8. The method for automatically generating digital teaching video courseware as described in claim 7, characterized in that, The step of scheduling the content synthesis engine according to the expected state and driving the normalized content primitives to be serialized and assembled along the generation path includes: The sequence of nodes on the generated path and the expected state of each node are converted into a sequence of synthetic instructions. The synthesis instruction sequence is matched and bound to the standardized content primitive pool to generate a primitive assembly queue with timestamps and transition requirements; The content synthesis engine sequentially reads and executes the instructions in the primitive assembly queue, calls the corresponding standardized content primitives, and applies visual transitions, audio mixing, and subtitle synchronization operations according to the instructions to achieve serialized assembly.

9. The method for automatically generating digital teaching video courseware as described in claim 1, characterized in that, The real-time monitoring of the serialization assembly process, capturing state offsets, and performing real-time trajectory optimization of the generated path based on the state offsets include: During the serialization assembly process, the actual media features and rhythmic features of the assembled portions are periodically sampled; The actual features sampled are compared with the expected state of the current node to calculate the state offset. When the state offset exceeds the tolerance threshold, the trajectory optimization mechanism is triggered; The trajectory optimization mechanism takes the current actual assembly state as a new starting point and replans the generation path of the remaining nodes in the dynamic content topology. The redesigned generation path is updated in subsequent synthesis instruction sequences.

10. The method for automatically generating digital teaching video courseware as described in claim 1, characterized in that, The step of performing multi-level coherent stitching and adaptive rendering on the assembled product after the serialization assembly is completed includes: Detect the audio waveforms and video streams of the serialized assembly products, and identify jumps and discontinuities at the segment connections. The audio crossfade-in / fade-out algorithm and the video motion compensation algorithm are used to smoothly stitch together the jumps and discontinuities. Based on the format requirements and bitrate limitations of the target playback platform, the encoding parameters of the stitched complete stream are adaptively configured and transcoded for rendering. Global color correction and audio loudness equalization are applied to generate the final digital teaching video courseware stream.