Multimedia file generation method based on multi-modal large model

By using a multimodal large model generation method, the consistency and security issues of multimedia file generation in existing technologies are solved, achieving a unified structured expression of the problem-solving process and age-appropriate security control, thereby reducing generation costs and uncertainties.

CN121284362BActive Publication Date: 2026-03-31北京爱宾果科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies lack a unified, structured representation of the problem-solving process when generating multimedia files for homework assignments. This makes it difficult to guarantee consistency, accuracy, and compliance. Furthermore, the lack of fine-grained control over facts and age-appropriate safety during the generation phase makes it difficult to predict costs and timeliness.

Method used

By using a multimodal large model approach, the following steps are completed sequentially: task alignment and evidence readiness, question analysis and solution trajectory diagram construction, compilation into unified instructions for storyboard whiteboard and narration, controlled generation and co-optimization during generation, posterior judgment and traceable release. This ensures that the semantic structure is mapped to shots, whiteboard and narration. In the generation stage, fact and safety scores are calculated per shot, and local regeneration is performed to reduce rework of the entire process.

Benefits of technology

It achieves clear scope and depth of explanation, multimodal security pre-screening of the generation process, and output multimedia files that correspond one-to-one with the camera, whiteboard, and narration, reducing uncertainty and cost, and facilitating distribution within the school and verification between home and school.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121284362B_ABST
    Figure CN121284362B_ABST
Patent Text Reader

Abstract

The application discloses a multimedia file generation method based on a multimodal large model, and relates to the technical field of educational content generation.The method successively completes task alignment and evidence readiness, question analysis and problem solving trajectory graph construction, compilation into unified instructions of split-screen whiteboard and narration, controlled generation and common optimization in generation, post-determination and traceable release and learning backwriting; the scheme takes the problem solving trajectory graph as the only true source, maps the semantic structure of the problem to be solved into the shot, the board and the oral broadcast, synchronously constrains the age-appropriate target and the cognitive load; in the generation stage, the fact and the safety score are calculated according to the shot, and local regeneration is triggered, reducing the whole segment rework; before release, the source credentials are written, the viewing and evaluation backwriting learning situation are combined, and the correct, age-appropriate, traceable and self-evolving closed loop is formed. Unified numbering runs through input, processing and output, supports fixed-point backtracking and spot check, reduces uncertainty and cost, and is convenient for school distribution and school verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational content generation technology, specifically a method for generating multimedia files based on a multimodal large model. Background Technology

[0002] With the increasing prevalence of online homework grading, classroom reviews, home tutoring, and self-study reinforcement, users' demand for "automatically generated visual explanation media from questions" has increased. Existing practices either require teachers to manually create blackboard notes and voiceovers, which is time-consuming and difficult to scale; or rely on general text-to-video or document-to-presentation tools, using templates to drive generation, which fails to faithfully reflect the solution process and common mistakes of specific questions; some systems only output text or images, lacking synchronized presentation with camera, narration, and blackboard notes; overall, the correspondence between "question-steps-audiovisual presentation" is poorly controlled, and personalization and compliance are mostly handled in post-processing.

[0003] At the technical level, the task images can be used to obtain the question stem and conditions through layout positioning, optical character recognition, and formula recognition. The solution can be formed into step text using rules, general language models, or symbols. However, a verifiable mapping from "step text" to "storyboard, blackboard writing, and oral delivery sequence" is lacking. Generative image and video models are mostly driven by natural language prompts or sparse constraints, lacking strong constraints oriented towards "step-level semantics." This leads to issues such as blackboard writing position drift and asynchrony between oral delivery and formulas in different batches of the same question. Once a factual deviation occurs, the entire section often needs to be redone, which is costly and highly uncertain.

[0004] In educational practice, learners vary significantly in age, and the readability and information density of the language used in instruction need to be matched to their grade level. Existing systems often output using standardized scripts or fixed templates, lacking quantitative control over sentence length, terminology ratios, and the number of explanation levels, which can easily lead to cognitive overload or superficial explanations. Content safety requirements for minors are stringent, and post-generation overall review is commonly used, making it difficult to promptly intercept fragment-level risks. Evidence citations and source attribution often remain at the descriptive text level, failing to establish verifiable source evidence at the media level, thus affecting distribution within schools and shared use between schools and families.

[0005] In terms of large-scale production and operation, the lack of unified instructions that allow for replayability and spot checks makes it difficult to quickly pinpoint specific steps and shots to quality issues. When teaching objectives change, the entire process often needs to be redone, resulting in low reusability and uncontrollable human and computing costs. Multi-person collaboration and cross-platform deployment also lack engineering guarantees that "the same problem-solving semantics will inevitably produce the same audiovisual presentation," affecting teaching consistency and deployment efficiency. Therefore, in the automatic generation of instructional media driven by assignment images, there is a lack of a unified structured expression for the problem-solving process and a verifiable mapping mechanism between it and shots, blackboard writing, and narration. Furthermore, the lack of fine-grained control over factual accuracy and age-appropriate safety during the generation stage makes it difficult to simultaneously guarantee consistency, accuracy, and compliance, and makes it difficult to predict and converge costs and timeliness. Summary of the Invention

[0006] (a) Technical problems to be solved

[0007] To address the shortcomings of existing technologies, this invention provides a multimedia file generation method based on a multimodal large model. This method sequentially completes task alignment and evidence readiness, question analysis and solution trajectory diagram construction, compilation into unified instructions for storyboards, whiteboards, and narration, controlled generation and co-optimization during generation, posterior judgment and traceable release, and learning rewriting. The solution uses the solution trajectory diagram as the sole source of truth, mapping the semantic structure of the questions to shots, whiteboard writing, and narration, simultaneously constraining age-appropriate goals and cognitive load. During the generation stage, factual and safety scores are calculated per shot, triggering local regeneration to reduce overall rework. Source credentials are written before release, and learning progress is rewritten based on viewing and assessment, reducing uncertainty and cost, facilitating distribution within schools and verification between home and school, thus solving the technical problems described in the background section.

[0008] (II) Technical Solution

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] A multimedia file generation method based on a multimodal large model includes receiving homework images or knowledge point topics and grade information, completing course alignment and evidence set construction, setting age-appropriate target vectors and cognitive load limits, conducting security pre-screening, and outputting an initial task plan containing learning summary and risk warnings.

[0011] The task image is structured and candidate step chains are generated, which are organized into a problem-solving trajectory diagram. The nodes carry blackboard elements, oral anchor points, evidence summaries and age levels, and the edges contain sequence, dependency, substitution and emphasis, and the time sequence number and nodes that need to be reviewed are given.

[0012] Based on the problem-solving trajectory diagram, age-appropriate target vector, and cognitive load limit, the storyboard sequence, whiteboard vector operation, narration and subtitle timeline, and control layer are compiled, alignment weights are calculated, and unified into a generation control instruction set;

[0013] The system calls the generation control instruction set and evidence set to perform controlled synthesis, calculates the fact score and safety score at the shot level, and when both the fact score and safety score are lower than the preset threshold, it performs local regeneration of the corresponding shot or switches the whiteboard main perspective, and outputs the initial media.

[0014] The initial version of the media is subjected to a post-hoc judgment to obtain the overall post-hoc judgment result of the initial version of the media. If it fails, it is rolled back to a fixed point. If it passes, the source credential fingerprint is written and published. Viewing and evaluation are collected to generate learning feedback and update the learning situation for subsequent use.

[0015] Furthermore, the course mapping score is calculated by integrating three indicators: text embedding, element coverage, and structural consistency. Course items and evidence sets are determined according to thresholds, and security scores and copyright scores are used to form a gating quantity. After passing the gating, an initial task plan is generated, and the number and handover information are recorded.

[0016] Furthermore, the age-appropriate target vector is set as a four-dimensional threshold of terminology ratio, sentence length, explanation level, and information points per unit time. Based on this, a cognitive load upper limit and time budget are generated and written into the task worksheet along with the learning progress vector and course mapping score, serving as a unified input for subsequent steps.

[0017] Furthermore, the task image is processed by page layout localization, optical character recognition and formula structure tree parsing, the page semantic combination score is calculated and the effective area is selected, the question stem, conditions, variables and implicit constraints are extracted, and a compilable candidate step chain is input and archived to support subsequent generation.

[0018] Furthermore, a candidate step chain is generated by using concept matching, evidence matching, and order type constraints to solve the path. For each step, a speech anchor point and a draft of the whiteboard elements are generated, a time sequence number is given, and it is written into the task worksheet for direct reference in the diagram structure organization and subsequent compilation process.

[0019] Furthermore, the candidate step chain is organized into a problem-solving trajectory diagram, and four types of edges are established: sequential, dependent, replacement, and emphasis. The node conflict probability is calculated based on the evidence fit and symbol back-substitution results. Nodes exceeding the threshold are marked as requiring review and synchronously written back to the task worksheet.

[0020] Furthermore, based on the time sequence number, evidence fit, and node conflict probability, the shot duration and order are allocated. The whiteboard vector operation is atomized from edge relationship mapping to fade-in, erase, highlight, item-by-item replacement, and graphic annotation. Slices are added within the shot time window and cross-shot continuous markers are added.

[0021] Furthermore, the narration and subtitle timelines are generated and aligned with the whiteboard atomic operations. At the same time, a control layer is constructed, and the intersection-union ratio of the main focus bounding box, the similarity of the gradient direction histogram, and the unified reference frame indication are calculated, merged into alignment weights, and written into the generation control instruction set for execution.

[0022] Furthermore, based on the alignment weight and node conflict probability, the controlled intensity parameters are calculated, the layout, edge and posture are controlled, the whiteboard, narration and subtitles are synchronized according to the key anchor points, the shot frame sequence and key frame index are generated, and the generation log and shot list are established.

[0023] Furthermore, a shot-level fact score and a safety score are calculated. The former consists of evidence coverage, back-substitution consistency, and residuals, while the latter consists of the proportion of sensitive items, the intensity of language transgression, and alignment weight. When the scores are below the threshold, controlled enhancement, language reduction, and whiteboard main perspective regeneration are executed in sequence.

[0024] Furthermore, the fact score and safety score are aggregated by power average based on the weighted average of shot duration and node conflict probability to obtain the work-level score. A binary gating variable is used to determine whether to proceed to publication. If the submission fails, the process is reversed to the point where the trigger shot is concentrated.

[0025] Furthermore, a source credential fingerprint is generated, and the evidence set summary, the generation control instruction set summary, the shot-level scoring sequence and the timestamp are concatenated in a fixed field order and subjected to anti-collision hashing, written into the media file metadata, and the verification code and verification entry are displayed on the page.

[0026] Furthermore, a shot feedback vector is constructed, which includes viewing completion rate, dwell time, number of rewinds, and the correctness and error type of the in-class assessment. The weighted average is calculated according to the shot weight and the mastery vector is updated by the learning rate to generate a learning write-back and stored in the task worksheet.

[0027] (III) Beneficial Effects

[0028] This invention provides a method for generating multimedia files based on a multimodal large model, which has the following beneficial effects:

[0029] Beforehand, course alignment, evidence collection, and age-appropriate target vectors are completed, clarifying the scope and depth of explanation. Multimodal safety pre-screening and source hints are performed, and an initial task plan is output to unify subsequent input boundaries and constraints. The structured questions and candidate step chains are organized into a problem-solving trajectory diagram. Nodes carry whiteboard elements, oral anchors, evidence summaries, and age-appropriate levels. Edges are labeled with sequence, dependency, substitution, and emphasis, and include time sequence numbers, evidence fit, and conflict probability for easy sampling and rollback.

[0030] Based on the problem-solving trajectory diagram, the storyboard sequence, whiteboard vector operation, narration and subtitle timeline and control layer are compiled. Alignment weights are calculated and unified into a generation control instruction set, so that the semantics correspond one-to-one with the shots, whiteboard writing and voice-over. The narration and subtitles are reduced in difficulty according to the age-appropriate target vector.

[0031] Controlled synthesis is executed by invoking the generation control instruction set and evidence set. Fact scores and safety scores are calculated at the shot level. If the scores fall below a threshold, a chain of controlled intensity enhancement, language burden reduction, and local regeneration from the whiteboard main perspective is triggered, limiting errors to shot-level correction. The initial media aggregation shot-level score forms a work-level a posteriori judgment. If it fails, it reverts to step three or four. If it passes, it is published after writing the source credential fingerprint and generation chain summary, and viewing and evaluation data are collected for learning feedback and updating the learning progress. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the multimedia file generation method based on a multimodal large model according to the present invention. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] Please see Figure 1 This invention provides a method for generating multimedia files based on a multimodal large model, including:

[0035] Step 1: Before generating the lesson, determine the scope, depth, and basis of the explanation. Establish age-appropriate goals based on students' learning progress and grade level. Align the curriculum and build an evidence set. Conduct a security and copyright pre-screening of the assignment images and text. Create an initial task plan that includes a learning summary, age-appropriate thresholds, evidence list, and risk warnings. This plan serves as a hard constraint and entry point for subsequent steps. Record the source, time, and number to ensure retrieval and reuse. If the pre-screening fails, stop at this step and prompt for rectification.

[0036] If the topic is not accurately anchored to the course item, the evidence set will be biased, leading to subsequent storyboarding and blackboard writing deviating from the key points of problem-solving. If the learning situation is not quantified, the depth and speed of explanation will be difficult to match the learners' comprehension ability. Therefore, course alignment is performed first to obtain the course mapping score and target item; then, the evidence set is filtered within the semantic domain of the item, and the learners' historical answer records are extracted to generate a learning situation mastery vector and a weakness index; the results from both paths are written into the task worksheet as the input prerequisites for step two: question analysis and problem-solving trajectory diagram construction.

[0037] To avoid topic mismatches caused by crude matching based solely on text similarity, a scoring method integrating three heterogeneous indicators is adopted. First, the user's topic and candidate course entries are embedded in text; then, the coverage of question elements and the consistency of question structure are combined for a comprehensive evaluation. The comprehensive score is defined as follows:

[0038]

[0039] Where: Course mapping score : Used for sorting; weighting coefficients : and Determined by the training set; text embedding function This involves mapping Chinese text to a unit norm vector space and using a sentence vector model trained on the same corpus; the output dimension is... Fixed vectors are pre-normalized, and the similarity is calculated by taking the inner product, i.e., the cosine similarity; text The text of a candidate course entry within the curriculum framework or textbook; text The text representing the current question or topic of knowledge for the user;

[0040] cosine function Vector similarity by angle, range Here, normalization is used to map to ;Element collection These are collections of user-generated question elements and course entry elements, with the elements being standardized conceptual terms.

[0041] Question type structure consistency function :return ,in For the identification of question type labels, For course paradigm tags, return consistently Inconsistencies are attenuated based on structural similarity, as follows:

[0042] Consistency of Question Type Structure Tree edit distance (The minimum cost of node replacement, insertion, and deletion), normalized using an exponential kernel:

[0043]

[0044] A structure tree for question types / items. The scaling constant (selected offline) is used, and the structure tree is extracted by the question type parser from the textbook paradigm library and question meta-information.

[0045] In practice, the coupling of three indicators avoids course mismatch caused solely by text similarity; under the dual constraints of element coverage and question structure, the target items have an interpretable basis for placement; the resulting course mapping score... It can be directly used as a driving force for subsequent evidence screening thresholds.

[0046] At the candidate evidence item level, source-level credibility, consistency with the target item, and coverage strength with the topic elements are assessed separately, and then fused into an evidence reliability score using a generalized average, defined as:

[0047]

[0048] Where: Evidence reliability score : Used for evidence screening and ranking; weight : and It can be set offline via grid search; source-level credibility. : Calculated based on the source whitelist and publication information; item consistency. : Based on course mapping scores Derived from or calculated by the structural alignment function;

[0049] Element coverage intensity : It is obtained from the intersection-union ratio of the question elements and the evidence elements; the generalized average index. : ,when The system penalizes low-scoring items, and choosing a negative value here can reinforce the weakest link effect.

[0050] In practice, generalized average aggregation is used to prioritize high-credibility sources while imposing a strong penalty on any low-scoring factor, thus enriching the evidence set. The overall quality is dominated by the weakest element, reducing the risk of subsequent factual bias; the aggregated evidence items are written into the task worksheet in the form of numbers for direct access in step two.

[0051] Furthermore, the explanation and expression of the same knowledge point should vary with age or grade level. Without explicit goal constraints, it is easy to cause an imbalance between language burden and information density. At the same time, media targeting minors need to complete gating before the content enters the generation process.

[0052] Based on grade level or assessment results, an age-appropriate target vector is generated, from which the upper limit of cognitive load is derived; then, multimodal security and copyright pre-screening is performed on image and text materials to generate gating variables and determine the path to continue, replace, or roll back; all parameters and judgments are recorded in the task worksheet for spot checks and playback.

[0053] The age-appropriate target is abstracted as a four-dimensional threshold vector and coupled with the cognitive load ceiling, then written into the task worksheet, formally as follows:

[0054]

[0055] Where: Age-appropriate target vector Ordered quadruples are used to constrain narration and subtitles; the term "proportion threshold" is also relevant. : ;

[0056] Sentence length limit Positive integer, unit is word; upper limit of interpretation level. Positive integer; upper limit of information points per unit time. : Positive integer, unit is points / minute; cognitive load limit : A non-negative real number representing the theoretically maximum amount of information that can be processed per unit of time; Constructs a constructor Based on grade parameters age-appropriate target vector Perform mapping and provide an upper limit for cognitive load. For example, by combining piecewise linear mapping with saturation functions, it can be ensured that the expression remains monotonically constant as the grade level increases.

[0057] In use, a combination of a four-dimensional threshold vector and a cognitive load limit forms rigid boundaries for language, structure, and rhythm, allowing direct control over the subsequently compiled storyboards, whiteboards, and narration; age-appropriate target vectors. With cognitive load limit It will be used in step three for syntactic difficulty reduction and duration allocation, and in step four for information density monitoring.

[0058] Multimodal detection is performed on the uploaded image and the question text, and scores for security and copyright are calculated separately and then unified as a gating value. Step 2 is allowed only when the gating value is passed; if it fails, suggestions for occlusion or replacement are given, and rollback is performed if necessary.

[0059] The gating quantity is defined as follows:

[0060]

[0061] Where: gate quantity The value can be 0 or , Passed; Safety score : This comes from the text / image / audio detector and the source validator;

[0062] Copyright rating : From text / image / audio detectors and source validators; threshold : Preset by compliance policy; indicator function Return if condition is met Otherwise return .

[0063] In use, a binary gating system is employed to pre-emptively determine security and copyright compliance, preventing non-compliant materials from entering subsequent processes. When any score falls below a threshold, suggestions for obscuring or replacing the material are displayed on the page. If the user rejects the suggestion or the replacement fails, a rollback and work order closure are triggered. This system quantifies and sets age and workload boundaries, as well as controls the compliance of input materials, ensuring that all subsequent steps operate within verifiable upper and lower limits.

[0064] Step Two: Combine the task image with the evidence list in the initial task plan to complete the layout positioning, text and formula recognition, and variable binding. Generate a candidate step chain covering concepts, formulas, derivations, substitutions, and conclusions, and organize it into a problem-solving trajectory diagram. Nodes carry whiteboard elements, oral anchors, evidence summaries, and age-appropriate levels, with edges labeled with sequence, dependency, substitution, emphasis, and chronological numbering for direct use in subsequent compilation. Set verification marks for nodes with significant contradictions and write them back.

[0065] Without verifiable structured results and candidate step chains, subsequent storyboarding and blackboard writing will lack stable anchors, leading to a disconnect between spoken and written delivery and inconsistencies between evidence and conclusions. Therefore, it is necessary to achieve unified binding of regions, text, formulas, and variables at the question level, and generate a set of evidence accordingly. Consistent candidate step chain.

[0066] First, based on the course mapping score With evidence set The elements are used to locate the semantic domain of text and formulas in the image; then, with variable binding as the axis, a candidate step chain of concept-formula-derivation-substitution-conclusion is generated, and the evidence citation summary and oral anchor point of each step are marked; finally, the link and binding results are written into the task worksheet.

[0067] To ensure a one-to-one correspondence between the question stem, conditions, and formulas in terms of space and semantics, a combined scoring system is adopted, consisting of three quantitative factors: geometric coverage, textual consistency, and formula completeness. A geometric average is used to mitigate the risk of an excessively high score in one factor masking weaknesses.

[0068]

[0069] In the formula: layout-semantic combination score : Used to determine whether the structured result enters the candidate step chain generation; weight : non-negative and Defined by the offline annotation set; geometric coverage : The intersection-union ratio of the question stem and condition area boxes with the text-formula line blocks; text consistency. : Normalized results of matching scores based on glyph, grammar, and line order; formula completeness. : The ratio of operators and variables that can be restored after structure tree parsing; skew penalty : This originates from baseline skew and perspective distortion estimation.

[0070] Among these steps, after the page is segmented into rows and blocks and baseline corrected, the geometric coverage is first calculated. Then, text consistency is obtained by lexical-syntactic joint matching. Then, the completeness of the formula is obtained by using the node coverage of the tree structure. And assign a penalty term to the tilt metric. Combined scoring The interface rules engine determines the threshold, and areas below the threshold are marked for manual review.

[0071] In practice, by coupling geometric mean with a penalty term, three types of uncertainty—spatial, semantic, and general—are simultaneously incorporated into the score, preventing individual strengths from masking overall weaknesses; the resulting combined score... As a uniform threshold, it ensures that the materials entering the candidate step chain have usable consistency.

[0072] Candidate step chains are generated using three key elements: concept triggering, formula referencing, and evidence fit. These elements are then transformed into an energy function, and the sequence solver provides the path with the minimum energy.

[0073]

[0074] Where: Link energy : Non-negative real numbers; the smaller the number, the more the sequence conforms to the course paradigm; weight : Non-negative real numbers, defined by offline grid search; concept matching cost : ,step The cost of consistency between concepts and course entries; the complementarity of similarity between concept words and entry word vectors; the cost of evidence deviation. : ,step With evidence set Matching bias, evidence summary matching complement; sequence type number Positive integers, corresponding to the topological order of concept → formula → derivation → substitution → conclusion; number of steps. : A positive integer, representing the length of the candidate step chain.

[0075] Search the concept library for trigger words that match the entry to obtain... In the evidence set Search for the corresponding abstract in the middle to obtain The system imposes a significant adjacency difference penalty on non-compliant sequence type jumps. A sequence solver calculates the minimum energy path to obtain candidate step chains, simultaneously generating a spoken anchor point description and a draft of whiteboard elements for each step. In practice, the system unifies the three types of constraints—concept, evidence, and topological order—in an energy-minimizing manner, ensuring that the resulting candidate step chains both conform to the course paradigm and can be directly mapped to storyboards and whiteboards. The adjacency difference penalty in the energy term effectively suppresses skipping steps.

[0076] A problem-solving trajectory graph is constructed based on the candidate step chain, and node attribute filling and consistency coarse checking are completed for subsequent compilation, so that the graph structure has compilability and sampleability.

[0077] Linear step chains alone are insufficient to express error-prone points, substitution relationships, and multiple predecessor dependencies. Therefore, a graph structure is needed to carry out these relationships. At the same time, to prevent errors from propagating to the compilation stage, the topology and evidence must be roughly checked on the layer faces.

[0078] First, each step of the candidate step chain is expanded into a graph node, and edges are generated according to sequence, dependency, replacement, and emphasis. Then, each node is filled with whiteboard elements, oral anchors, evidence citation summaries, and age-appropriate readability levels. Subsequently, the topological legality and evidence fit index are calculated, and nodes that do not meet the standards are marked as needing review and the task worksheet is written back.

[0079] To ensure that the edge direction is consistent with the step sequence and to avoid loops and inversions, define Figure 1 Consistency cost and threshold determination, where:

[0080]

[0081] Where: Topology consistency cost : Non-negative real numbers; the smaller the number, the more valid the topology; weight : Non-negative real numbers, defined by offline annotation and heuristics;

[0082] edge set Node set : From the structure definition of the graph; timing number : Positive integer, corresponding to the order of the candidate step chain on the timeline; node evidence fit : From the set of evidence The local matching score is obtained.

[0083] First, the initial temporal sequence number is directly given from the candidate step chain. Then, additional edges are generated based on the replacement and emphasis relationships, and the first cost is calculated. Subsequently, based on the evidence set... The abstract similarity is assigned to each node. And calculate the second cost; apply a threshold to the topology consistency cost. The system makes a judgment, and edges and nodes exceeding the threshold are highlighted as requiring review. When in use, it encompasses two types of structural risks—time sequence inversion and evidence gaps—at a uniform cost and explicitly exposes them at the graph level, facilitating spot checks and targeted corrections. The obtained time sequence number and evidence fit can be directly used in subsequent compilation for allocating scene length and narration strength.

[0084] Furthermore, to quickly identify high-risk nodes at the graph level, a logical transformation is introduced to converge multiple factors into conflict probabilities, which then drive the marking of nodes requiring review.

[0085]

[0086] Where: Node collision probability : Used for sorting and threshold determination; Sigmoid function Defined as Mapping real numbers to Weight : Non-negative real numbers, derived from offline experience settings; node evidence fit Same as above. ; Symbol for failed back-substitution : Value can be 0 or 1; if fast back-substitution verification fails, then it is 0. ;

[0087] Semantic self-contradiction : This was obtained through a two-way inference consistency check. , This is the intermediate set obtained from the forward derivation. This is the set of prerequisites required for a reverse proof.

[0088] Perform fast symbol backsubstitution and bidirectional deduction on each node to obtain the symbol backsubstitution failure flag. Semantic self-contradiction The three factors are substituted into the above formula to calculate the node conflict probability, and the threshold of the task worksheet is used for judgment. Nodes exceeding the threshold are marked as requiring review, and their evidence citation summaries are required to be supplemented. The draft of the whiteboard elements and the voice-over anchors are set to pending confirmation in the interface. When in use, the three types of risks are aggregated with a single probability, which significantly reduces the scope of manual sampling. The marked nodes will be allocated more shot time and more explicit voice-over anchors in subsequent compilation, thereby reducing the probability of subsequent backtracking. Step 3: Based on the problem-solving trajectory diagram, the age-appropriate target vector, and the upper limit of cognitive load, the storyboard sequence is arranged and the shot time is allocated. The whiteboard vector operation is mapped according to the edge relationship to generate the narration and subtitle timeline and align it with the shot. The layout, edges, posture, and reference frames are summarized to form a control layer. The alignment weight is calculated, and the storyboard, whiteboard, narration, subtitles, and control layer are uniformly numbered to generate a control instruction set and written back. This is provided for direct reading and auditing by the controlled composition. Consistency is ensured.

[0089] The storyboard and whiteboard form the framework for subsequent scene generation and narration alignment. If the shot length allocation does not consider the relevance of key evidence... Conflict probability with nodes This can easily lead to problems such as downplaying high-risk steps and insufficient time spent on key derivations; if the whiteboard vector operations are not aligned with the edge relationships (sequence, dependency, substitution, emphasis), the whiteboard presentation will become disconnected from the reasoning rhythm. Therefore, it is necessary to establish a problem-solving trajectory diagram in both time and action dimensions. Isomorphic executable manifests.

[0090] First, based on the time sequence number Node evidence fit Conflict probability with nodes For each node, the shot duration and sequence allocation are calculated to obtain the set of shot sequences. Then, the edge relationships are translated into a sequence of whiteboard atomic actions and sorted to obtain a sequence of whiteboard vector operations. Both share a unified number so that they can be directly referenced when generating subsequent narration and subtitles.

[0091] To ensure that important and high-risk nodes have sufficient camera capacity, while controlling the total duration to not exceed the cognitive load limit. Number each shot Duration Calculate using the following formula:

[0092]

[0093] In the formula: shot duration Positive real numbers, used to determine the lens. Time budget; weighting coefficient : non-negative real number, and Offline configuration; timing mapping function : Time sequence number The mapping to a monotonic function of baseline weights can take the form of exponential decay. ; The index represents all shots, and the weighted terms of all shots are summed in the denominator;

[0094] Exponential parameters Positive real numbers, controlling the relative weights of early steps; nodal evidence fit. : interval The larger the value, the more substantial the evidence; node conflict probability : interval The larger the value, the more potential conflicts there are; cognitive load limit : A non-negative real number from step one, representing the upper limit of the total time that can be allocated per unit of content.

[0095] First, calculate the baseline weight term for each node. , and then unite and Perform weighted normalization, and finally... Scaling yields the duration of each shot. and in chronological order The sequence of shots is given in order; for example, the duration of a certain shot. If the runtime falls below the minimum executable time threshold, it is merged into an adjacent shot, and action details are preserved on the whiteboard side. In practice, this allocation integrates early definition, compensation for weak evidence, and reinforcement of high-risk actions into a single runtime budget, ensuring that key steps are visible and audible; simultaneously... The global constraints make the total duration of the sequence predictable.

[0096] Whiteboard vector operations start from the edge relationships of the graph and perform atomized mapping: first, the sequence is mapped to a time series of fade-in and erase, the dependency is mapped to a nested display of parent and child items, the replacement is mapped to item-by-item replacement, and the emphasis is mapped to highlighting and overlaying the pointer pen trajectory.

[0097] Then press the camera. Shot duration The corresponding time windows slice and sort these atomic operations, and mark the scenes that need to continue across shots with cross-shot continuity indicators to ensure that subsequent shot splicing does not disrupt the continuity of the whiteboard.

[0098] For each node First, create a draft of the elements on the blackboard. Then, generate a sequence of atomic operations based on the relationship between their incoming and outgoing edges. Finally, arrange the operations according to the shot duration. The system segments and injects fine-grained timestamps; when emphasis and substitution occur simultaneously, the three-stage process of emphasis-pause-substitution is prioritized, allowing learners to complete visual localization before symbol substitution. In practice, the blackboard presentation is strictly constrained by the semantics of the problem-solving process through the deterministic mapping of edge relationships to atomic actions; slicing and cross-camera annotations ensure the continuity of the blackboard presentation remains stable under camera splicing, avoiding formula jumps and annotation loss.

[0099] Furthermore, narration and subtitles bear the responsibility of ensuring language readability and controlling pacing; if not... and Explicit constraints can easily lead to excessive terminology stacking or information density at critical steps; without unified weights and reference frame selection rules in the control layer, positioning and pose will drift between shots. Therefore, quantitative constraints and summary scores need to be established at the language layer and alignment layer respectively.

[0100] First in each shot Generate narration text and subtitle timelines within a time window, and use... With shot duration To reduce the difficulty and control the density of hard constraints, control layers for layout, edges, and pose are generated based on the main focus and attention hotspots of the storyboard, and alignment weights are calculated. Finally, the four types of lists are uniformly numbered as a set of generation control instructions. .

[0101] Furthermore, to ensure that the proportion of terminology, sentence length, number of explanatory levels, and number of information points per unit of time do not exceed the limits, the camera... The language load is used to construct a penalty term and iteratively rewritten until the penalty term is zero:

[0102]

[0103]

[0104] Where: Lens-level language load overload strength The closer the value is to 1, the higher the language load of the shot is relative to its age / grade threshold; 0 indicates no violation of the limit, and is used as the input for the safety score in step four; total language penalty items. : Non-negative real number, zero indicates all criteria are met; smoothing constant It is a very small positive number; weight Non-negative real numbers, reflecting the relative importance of the four types of boundary crossings; terminology: proportion. : interval , indicating the lens Percentage of internal terminology;

[0105] Sentence length Positive integer, in digits, representing a shot. Average sentence length; number of explanatory levels : Positive integer, representing the lens Interpretive hierarchy; number of information points per unit time : Non-negative real numbers For information points, For shot duration; age-appropriate target vector components : from respectively The four threshold components.

[0106] Among them, after generating the initial draft of the narration, calculation Substitute this into the above formula; when the total amount of language penalties... At that time, the language penalties were progressively reduced by lowering the level of technical terms, breaking down long sentences, adjusting the level of explanation, and allocating information points to adjacent shots, until the total amount of language penalties was increased. The subtitle timeline is generated by interpolation using the timestamps of the camera edge and the whiteboard atomic operations as anchor points, ensuring that the words appear on the screen and the whiteboard writing actions are aligned.

[0107] When used, this targeting constraint ensures that the language load strictly falls within the specified range. and Within the defined readability and rhythm boundaries, avoid terminology piling up and information overload; bind the subtitle timeline with the whiteboard action to reduce misalignment of speaking without seeing words or seeing words without speaking words.

[0108] To maintain layout, edge stability, and pose at the shot level, each shot... The detection results are aligned and scored with the reference frame, and then summarized into alignment weights:

[0109]

[0110] Where: Alignment weight : interval A larger value indicates higher alignment quality; weighting coefficient : Non-negative real numbers, and their sum is ; bounding box With reference box For the lens The main focal region and the corresponding region of the reference frame; the intersection-exchange ratio function. :return This measures the degree of overlap between the two frames. ;

[0111] gradient field With reference gradient field : Histogram vector of orientation obtained from edge detection; cosine similarity function :return After linear mapping to Reference frame indication function : Value or When the camera It is 1 when a unified reference frame is used, otherwise it is 1. .

[0112] For each shot, detect the principal focus bounding box and gradient field on the keyframes, and compare them with the reference frame. Similarity calculation with cosine similarity; when alignment weights When the threshold is exceeded, the attention hotspots of the storyboard and whiteboard are traced back, the layout parameters and reference frame selection are adjusted, and the calculation is recalculated; finally, the control layers that pass the threshold and their weights are written into the list. When used, this score uniformly represents the consistency between spatial positioning and texture orientation cues and the reference frame, enabling alignment risks to be detected and corrected during compilation, thus ensuring stable image anchors in subsequent controlled generation.

[0113] Step 4: Under the drive of the generation control instruction set, perform controlled compositing, calculate factual scores based on the evidence set of each shot, and calculate safety scores for language and visuals. If any score is below the threshold, only the corresponding shot is partially regenerated. This process sequentially involves increasing the controlled intensity, reducing the language burden, and reverting the whiteboard main perspective, while retaining keyframe indexes and trigger logs. Once all shots pass, the initial media version is output for final inspection. If there are consecutive failures, the process reverts to the compilation stage for adjustments and retry.

[0114] If only free prompts are used to composite the footage, layout drift, posture shifts, and missing strokes in the whiteboard writing can easily occur, leading to desynchronization between narration, subtitles, and the visuals. Since the shot duration was already specified in the previous step... Alignment weight Vector operations with whiteboard Therefore, it is necessary to synchronize the control layer and the generator's control intensity so that lenses with high control intensity have tighter layout and pose constraints, while lenses with low control intensity allow moderate texture details to unfold freely, thereby achieving an auditable balance between consistency and image richness.

[0115] First, based on alignment weight Conflict probability with nodes Calculate lens-level controlled intensity parameters and accordingly assemble the corresponding control layers (layout, edge, pose, reference frame) and whiteboard vector operations; then, within the lens time window... The system synchronizes the whiteboard writing, narration, and subtitles, ensuring that key display, replacement, and emphasis actions are aligned with the timing of words. Finally, it outputs a shot-level frame sequence and records the keyframe index, providing direct positioning anchors for subsequent scoring and local regeneration.

[0116] Furthermore, to transform alignment quality and semantic risk into an adjustable knob during generation, a lens-level controlled intensity parameter is introduced. The intensity is determined by a weighted average of alignment weights and collision probabilities, and is used on the generation engine side to adjust the control layer weights and the acceptable range of pose offsets.

[0117]

[0118] Where: controlled strength parameter : A non-negative real number; the larger the value, the stronger the control. Baseline strength : A non-negative real number, representing the minimum controlled strength for all lenses; weight : Non-negative real numbers, respectively measuring the impact of inalignment and semantic risk; alignment weight : interval Alignment score from step three; node conflict probability : interval The result is from the coarse proofreading in step two.

[0119] Calculate the controlled intensity parameters for each shot. And based on this, select the control layer assembly strength; when the controlled strength parameter When the value is large, layout and edge control are given high weight, the pose reference frame prioritizes the unified reference frame, and the pause duration between emphasis and replacement in whiteboard operations is extended; when the controlled intensity parameter When the size is small, a small number of texture detail shots can be inserted to improve readability, but the order of whiteboard vector operations remains unchanged.

[0120] In use, this controlled intensity transforms image alignment and semantic risk into a tangible compositing knob, giving high-risk shots stronger constraints before they enter the scoring process and reducing the probability of subsequent regeneration; at the same time, it retains the limited freedom of low-risk shots to avoid stiff images.

[0121] To ensure that the display of words and formulas is consistent with the rhythm of the shots, the whiteboard vector operation sequence is used. The atomic actions serve as anchor frames, and the narration text collection... The sentence breaks serve as language anchors, and the subtitle timeline set... The timestamp serves as the display anchor point, and the three elements are within the camera's time window. The alignment process involves first determining key anchor points based on the priority of emphasis-pause-replacement, then snapping the remaining atomic actions and secondary clauses to both sides of the anchor points according to the principle of proximity; then allocating silence and pause budgets to ensure that at least one minimum gaze window is reserved before each replacement; finally, marking cross-shot continuity in the frame sequence to maintain the continuity of the whiteboard.

[0122] After aligning the three lines, generate the shot. The keyframe index list contains the start and end frame numbers corresponding to each atomic action, the text range corresponding to the narration sentence, the timestamp range corresponding to the subtitle, and the controlled intensity parameter. Write back to generate logs. When using it, the three-line anchor point alignment of action-language-display can replace the experience-based checkpoints, which can significantly reduce the misalignment of words when words are not seen or words have been replaced and words are lagging behind; the keyframe index list provides a locatable and replayable basis for subsequent scoring.

[0123] If factual consistency and security are postponed entirely, it will lead to error propagation, overall rework, and soaring costs. Therefore, a closed loop of scoring-triggering-regeneration should be established during the generation phase to limit errors or out-of-bounds issues to resolution at the scene level, and to leave traces in the generation log for playback and auditing.

[0124] First, let's look at the collection of evidence. The evidence summary corresponding to the semantics of the shot is extracted and compared with the narration text and whiteboard display of the shot to obtain a fact score; then the shot frame sequence and language load are tested for safety and age appropriateness to obtain a safety score; if any score is lower than the threshold, the controlled intensity is strengthened first, then the language load is adjusted, and if it still does not meet the standard, the order of the whiteboard main perspective is switched to perform local regeneration until it passes or reaches the regeneration limit.

[0125] Furthermore, the fact scoring employs a combined approach of evidence coverage, sign back substitution, and residual suppression, aggregating the results into a lens-level confidence score using a single differentiable function.

[0126]

[0127] In the formula: fact score : interval The larger the value, the better the consistency of the facts; Sigmoid function Defined as Mapping real numbers to Weight : Non-negative real numbers, reflecting the relative importance of the three factors; evidence coverage : interval , indicating the lens The narration text and blackboard writing are presented as evidence. The proportion of key elements covered;

[0128] Back-substitution consistency rate : interval This represents the proportion of the results that match the conclusion after substituting the formula and values ​​from the blackboard; normalized residual. : A non-negative real number representing the maximum normalized residual between the derivation and the conclusion.

[0129] Calculate evidence coverage Using a standardized glossary of evidence summaries as the thesaurus, the hit ratio between the video text and the blackboard writing was statistically analyzed; the back-substitution consistency rate was calculated. The key formulas appearing in the footage are then substituted and verified; the normalized residuals are calculated. Normalize the difference between the conclusion and the derivation; then aggregate the three factors according to their weights to form a fact score. If the fact score Less than the threshold If so, the recorded facts trigger and enter the local regeneration process.

[0130] When used, this aggregation function simultaneously rewards evidence coverage and consistency with backsubstitution, and penalizes derivation residuals. It can expose factual risks and trigger point-to-point corrections during the shot synthesis stage, preventing errors from spreading to subsequent shots.

[0131] Furthermore, the security score integrates the percentage of sensitive items, language load exceeding limits, and image alignment stability, using a monotonically increasing function to provide a shot-level security confidence level:

[0132]

[0133] Where: Safety score : interval A higher value indicates greater security; weight Non-negative real numbers reflect the importance of the three signals; the proportion of sensitive terms. : interval The normalized values ​​of the proportions of sensitive markers for video text, images, and audio; the intensity of language load exceeding limits. : interval The alignment weight is obtained from the normalized value of the language penalty in step three within the camera window. : interval Lens-level image alignment stability comes from step three.

[0134] Handling method: If the security score is... Local regeneration is performed in three stages: the first stage increases the controlled intensity parameters. And fix a unified reference frame; the second segment, without changing the semantics, is based on... Reduce the proportion of terminology and the number of information points per unit time, and redistribute subtitle pauses; in the third segment, switch the camera style to a whiteboard-centric perspective and replace the illustration with vector whiteboard text. Recalculate the factual score after each segment. With safety rating The process continues until both thresholds are met or the regeneration limit is reached. In use, a progressive strategy of controlled enhancement, language reduction, and style fallback ensures that security issues are addressed promptly at the scene level while maintaining factual accuracy. Regeneration logs can be used for post-event auditing and parameter playback.

[0135] Step 5: The initial media aggregation shot-level score forms a work-level a posteriori (APS) determination. If it fails, the trigger shot is used to revert to the compilation or controlled compositing stage. If it passes, a source credential fingerprint and generation chain summary are generated and written. It is then released and distributed according to the minor mode, and viewing and evaluation data are collected to generate learning feedback. The mastery vector and task worksheet are updated as input for the next task alignment and evidence readiness. Release records and version information are retained to support subsequent auditing and retrieval.

[0136] Shot-level scoring can expose local risks, but the release process requires clear conclusions at the work level. Simply averaging scores across different shots can easily dilute the results with long takes or high-risk shots, thus overlooking critical errors. Furthermore, if the impact of shot duration and semantic risk on weighting is not considered, the aggregated results lack interpretability. Therefore, a method is needed that can impose stronger penalties on low-scoring shots while reflecting the importance of duration and risk, and that uses clear gating variables to achieve a binary pass / fail determination.

[0137] Therefore, let's start with the length of the shot. Conflict probability with nodes Construct weights, then score the scene-level facts. With safety rating Power-mean aggregation is performed separately to obtain the work-level fact score and the work-level safety score; then, the relationship between the two is determined by gating variables; if it fails, the process is backed up to step three or four according to the shot list to avoid rework of the entire process.

[0138] To enhance sensitivity to low-resolution shots and retain the weighting effect of duration-risk, a power-law average was used for aggregation.

[0139]

[0140] In the formula: Work-level factual rating Work-level safety rating : interval A higher value indicates a better fit; lens-level factual rating Lens-level safety score : interval From step four; power exponent : Real number, usually taken Used to increase penalties for low scores; weighting : non-negative and Due to the length of the shot Conflict probability with nodes Joint decision (e.g., by) Normalization (This is a non-negative adjustment coefficient).

[0141] Calculate for each shot The formula of power-means for descendant input yields a work-level factual score. With work-level safety rating At the same time, the lens indexes and corresponding weights participating in the aggregation will be written back to the task worksheet for easy auditing and review.

[0142] When used, the power mean is... It is more sensitive to low-resolution shots, which can effectively prevent key shots from being masked by low-resolution shots; weighting By introducing two types of information—duration and risk—the aggregation results become interpretable and traceable.

[0143] Furthermore, in obtaining a work-level factual rating With work-level safety rating Then, define a binary gated variable:

[0144]

[0145] In the formula: published gating variable : Value Or 1, This indicates that publication is permitted; work-level fact threshold. Work-level security threshold : interval Preset by compliance policy; indicator function Return if condition is met Otherwise, return 0.

[0146] When publishing gated variables When the release process begins; when the gated variable is released... When triggering a shot, select the trigger shot from the shot-level scoring list and prioritize reverting to step four for shot-level regeneration. If the trigger shots are clustered in the same derivation segment, the system prompts to revert to step three to adjust the storyboard and whiteboard operation, and records the reason for reverting.

[0147] In use, binary gating ensures that both fact-checking and safety-checking gates are indispensable, and limits the rollback range to the triggering segment, reducing rework costs. Gating and rollback logic are logged in the work order for easy auditing later. Measurable effects and conditions: Under a fixed threshold, the gating pass rate-task volume curve and rollback path distribution are recorded, along with statistics at the shot level / compilation level, with software version and hardware environment kept consistent.

[0148] Furthermore, without source documentation, schools and platforms cannot verify the origin, basis, and generation chain of content. Without learning feedback, the system cannot absorb effective information density and rhythm feedback from viewing and assessment results, making subsequent generation difficult to adapt. Therefore, source documentation needs to be written into the media ontology, and viewing and assessment elements need to be incorporated into the mastery vector for updating. First, the evidence, generation chain, and scoring summary are organized and fingerprints are calculated to form source documentation; then, the media is distributed under the minor mode strategy; simultaneously, viewing behavior and in-class assessments are aggregated, the mastery vector is updated, and written back to the task worksheet for use in the next step.

[0149] Furthermore, to compress the evidence, generation chain, and score summary into a unique and verifiable identifier, a source credential fingerprint is constructed:

[0150]

[0151] In the formula: source credential fingerprint Fixed-length hash values ​​used for media-level verification; hash function Mathematical representation of collision-resistant hash functions; evidence set summary : On the set of evidence An ordered digest string; generate an instruction digest. : An ordered summary string of the generation control instruction set; a summary of the fact scoring sequence. Safety scoring sequence summary : A summary of ratings in shot order; timestamps : Monotonically increasing publication time marker; connector symbol : Indicates a concatenation operation.

[0152] The five elements are concatenated in a unified order with the coding procedure and then fed into a hash function to obtain the source credential fingerprint. Then, along with the model version, approved version, age-appropriate tag, and task work order number, write it into the metadata area of ​​the media file; after writing, display a copyable verification code and verification entry instructions on the page. When using it, the source credential fingerprint... By irreversibly binding the basis, link, rating, and time, media can be verified in distribution and implementation scenarios; metadata writing enables works to have their own credentials, reducing the cost of external evidence storage.

[0153] Furthermore, to robustly inject viewing and evaluation feedback into the mastery vector, a convergent update is adopted:

[0154]

[0155] In the formula: the updated master vector Before updating, understand vectors. : Defined in column vectors, Knowledge point dimension; learning rate : interval Control the magnitude of a single write-back; normalization factor : Positive real number, equal to the sum of weights Lens Feedback Vector : Defined in From the lens The viewing completion rate, dwell time, number of rewinds, and the correctness and error type of in-class assessments are mapped to the data.

[0156] in, ,in Project the above signals according to the dimensions of the knowledge points covered by the lens and linearly normalize them. (e.g., completion rate and correctness are positively correlated, while rewind and error type are negatively correlated). The projection matrix comes from the problem-solving trajectory diagram - knowledge point reference table.

[0157] Generate feedback vector for each shot (For example, mapping the correctness of in-class quizzes to the corresponding knowledge point dimensions, and mapping high-frequency rewinds to signals of unmastered knowledge points covered by that scene), weighted according to the same principles as the published aggregation. Calculate the weighted average, then use the learning rate. Before updating the known vector Perform a convergent update and then update the master vector. Write the task work order back along with the feedback elements, weights, and timestamps.

[0158] When in use, convergent updates can avoid the oscillations caused by one-time feedback, and weight sharing ensures that the attention given to the release decision and the impact of learning and writing are consistent, thus forming a parameter connection between decision-distribution-learning-re-decision.

[0159] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0160] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0162] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating a multimedia file based on a multi-modal large model, characterized in that: comprising, Receiving homework image or knowledge point theme and grade information, calculating course mapping score based on the homework image or knowledge point theme and grade information, determining target course item and evidence set according to preset threshold, setting age-appropriate target vector and upper limit of cognitive load, performing safety pre-audit, and outputting initial task plan containing learning situation summary and risk prompt; Structuring the homework image and generating candidate step chain, and arranging into problem solving trajectory graph, wherein the nodes of the problem solving trajectory graph carry board writing elements, oral broadcast anchor points, evidence abstracts and age-appropriate grades, the edges of the problem solving trajectory graph include precedence edges, dependency edges, replacement edges and emphasis edges, and time sequence numbers and nodes requiring review are given; According to the problem solving trajectory graph and the age-appropriate target vector and the upper limit of cognitive load, a compilation is obtained to obtain a split sequence, a whiteboard vector operation, a subtitle time axis and a control layer, and an alignment weight is calculated, and the split sequence, the whiteboard vector operation, the subtitle time axis, the control layer and the alignment weight are uniformly numbered and packaged into a generated control instruction set; specifically, the compilation to obtain the split sequence and the whiteboard vector operation includes assigning shot length and order according to time sequence number, evidence fitting degree and node conflict probability, and the whiteboard vector operation is atomized mapped from edge relationship to gradual display, erasing, highlighting, item-by-item replacement and graphical annotation, and is sliced and labeled with cross-shot continuous identification within the shot time window; The compilation to obtain the subtitle time axis and the control layer and the calculation of the alignment weight include generating the subtitle time axis and aligning it with the whiteboard atomic operation, constructing the control layer, calculating the intersection ratio of the main focus bounding box, the gradient direction histogram similarity and the unified reference frame indication, merging into the alignment weight and writing into the generated control instruction set for execution; Calling the generated control instruction set and the evidence set to perform controlled synthesis, calculating shot-level fact score and safety score for each shot number in the split sequence, wherein the shot-level pointer is calculated for each shot number; when any score is lower than the corresponding preset threshold, only the corresponding shot is locally regenerated, and controlled intensity enhancement, language reduction and whiteboard main perspective regeneration are performed in turn, and an initial version of the media is output; Performing work-level post-determination on the initial version of the media, performing point backtracking if it fails, writing source credential fingerprint if it passes, collecting watching and evaluation to generate learning backwriting, and updating learning situation for subsequent calling.

2. The multi-modal large model-based multimedia file generation method of claim 1, wherein: The calculation of the course mapping score includes fusion calculation of the course mapping score using three indicators of text embedding, element coverage and structure consistency; the determination of the target course item and the evidence set includes determination of the course item and the evidence set according to the threshold; and the safety score and the copyright score constitute a gate quantity, and an initial task plan is generated after passing, and the number and handover information are recorded.

3. The multi-modal large model-based multimedia file generation method of claim 2, wherein: The setting of the age-appropriate target vector and the upper limit of cognitive load includes setting the age-appropriate target vector as the four-dimensional threshold of the proportion of terms, sentence length, explanation level, and information points per unit time, and generating the upper limit of cognitive load and time length budget according to the four-dimensional threshold. The upper limit of cognitive load and time length budget are written into the task work sheet together with the learning situation mastering vector and the course mapping score as the unified input for the subsequent steps.

4. The multi-modal large model-based multimedia file generation method of claim 1, wherein: The structuring of the assignment image and the generation of the candidate step chain include performing layout positioning, optical character recognition, and formula structure tree analysis on the assignment image, calculating layout semantic combination scores and filtering effective regions, extracting stems, conditions, variables, and implicit constraints, forming a compilable candidate step chain input, and archiving the candidate step chain input to support subsequent generation.

5. The multi-modal large model-based multimedia file generation method of claim 4, wherein: The generation of the candidate step chain includes generating the candidate step chain by path solving based on concept matching, evidence matching, and sequence type constraints, generating oral anchor points and board writing element drafts for each step, giving time sequence numbers, and writing the time sequence numbers into the task work sheet for direct reference by the graph structure organization and subsequent compilation process.

6. The multi-modal large model-based multimedia file generation method of claim 5, wherein: The arrangement of the problem-solving trajectory graph includes arranging the candidate step chain into a problem-solving trajectory graph, establishing four types of edges, i.e., precedence, dependency, replacement, and emphasis, calculating node conflict probabilities based on evidence fitting degrees and symbolic back-substitution results, and marking nodes with conflict probabilities exceeding a threshold as requiring review and simultaneously rewriting them into the task work sheet.

7. The multi-modal large model-based multimedia file generation method of claim 1, wherein: The controlled synthesis further includes calculating a controlled intensity parameter based on alignment weights and node conflict probabilities, assembling layout, edge, and pose controls, implementing board writing, narration, and subtitle three-line synchronization according to key anchor points, generating shot frame sequences and key frame indexes, and establishing a generation log and a shot list.

8. The multi-modal large model-based multimedia file generation method of claim 7, wherein: The calculation of the shot-level fact score and the safety score includes that the fact score is composed of evidence coverage, consistent back-substitution, and residual error, and the safety score is composed of sensitive item proportion, language boundary crossing intensity, and alignment weight; when any score is lower than the corresponding preset threshold, controlled intensity enhancement, language reduction, and whiteboard main perspective regeneration are performed in sequence.

9. The multi-modal large model-based multimedia file generation method of claim 8, wherein: The work-level posteriori judgment includes power-averaging aggregation of the fact score and the safety score by weighting the shot length and the node conflict probability to obtain a work-level score, and determining whether to enter the release by a binary gating variable. If not, the process is backtracked to the processing of the focused area of the trigger shot.

10. The multi-modal large model-based multimedia file generation method of claim 9, wherein: The write source credential fingerprint comprises generating a source credential fingerprint, concatenating an evidence set digest, a generated control instruction set digest, a lens level score sequence and a timestamp in a fixed field order and anti-collision hashing, writing media file metadata, and page display verification code and verification entry.

11. The multi-modal large model-based multimedia file generation method of claim 10, characterized in that: The learning backwriting comprises constructing a lens feedback vector, including viewing completion, staying time, fast backward times, in-class test correctness and error types, weighted average according to lens weight and convergence update to mastery vector with learning rate, generating learning backwriting and storing in task worklist.

Citation Information

Patent Citations

  • Intelligent auxiliary teacher lesson preparation system and method based on large education model

    CN120407769A

  • Online resource adaptive recommendation method for multi-modal learning behavior analysis

    CN120561380A