A method, apparatus, and instruction editing system for dynamic video synthesis based on instruction sequences.

By using a dynamic video synthesis method based on instruction sequences, and leveraging an AI-assisted subsystem and an instruction editing system, video operation instruction sequences are generated and adjusted in real time. This solves the problems of insufficient flexibility and interactivity in traditional video synthesis systems, achieving efficient video generation and real-time interaction. It is suitable for AI-assisted teaching, meetings, and unattended online interactive classrooms.

CN120676223BActive Publication Date: 2026-01-30PEKING UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511191170.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-01-30
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing video synthesis technologies rely on fixed templates, which cannot dynamically respond to complex scenarios, resulting in low efficiency and failing to meet users' needs for customization and real-time interaction. In particular, they lack flexibility and interactivity in live streaming and online teaching.

Method used

By using a dynamic video synthesis method based on instruction sequences, and leveraging an AI-assisted subsystem and an instruction editing system, video operation instruction sequences are generated and adjusted in real time, including digital human broadcasting, material playback, scene switching, etc. Combined with voice intent recognition and dynamic instruction insertion, flexible customization and real-time interaction of video content can be achieved.

Benefits of technology

It enables flexible customization and real-time interaction in the video generation process, improves video generation efficiency, supports AI-assisted teaching, meetings, and unattended online interactive classrooms, and enhances user experience and system flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676223B_ABST
    Figure CN120676223B_ABST
Patent Text Reader

Abstract

This invention provides a dynamic video synthesis method, apparatus, and instruction editing system based on instruction sequences, applied in the field of real-time video generation technology. It is used for AI-assisted teaching, meetings, and unattended online interactive classrooms. Based on user-input prompts, an initial instruction sequence is generated through an AI-assisted subsystem and an instruction editor. An instruction parser then sequentially executes the instructions by calling execution plugins according to their type. During the execution of the initial instruction sequence, the AI-assisted subsystem performs real-time recognition of the intent of the input speech. Based on the recognition results, the input speech is converted into digital human-proclaimed instructions or new video operation instructions and inserted into the current execution position for execution. The video and audio are then rendered and synthesized based on the instruction execution results. This invention achieves dynamic video synthesis by issuing instruction sequences with AI assistance and enables real-time adjustment of video content through dynamic instruction insertion, meeting the needs of customized video synthesis and human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of real-time video generation technology, and in particular to a dynamic video synthesis method, apparatus, and instruction editing system based on instruction sequences. Background Technology

[0002] Video production is a complex process, especially when it involves live-action filming. Equipment, scenes, personnel, and post-production all require significant work. This has led to the development of simplified, automated, and efficient production techniques. These methods collect, sort, and match source material, then apply templates to automatically synthesize videos. However, the template-based video generation method lacks flexibility in material organization and arrangement. When the video has many customized details, requires frequent scene changes, or needs dynamic adjustments to the video content based on image signal changes—such as the dynamic changes required for interactive classroom videos—the template-based approach becomes less practical, significantly increasing the number of post-production steps.

[0003] In the fields of live streaming and online teaching, traditional live streaming relies on pre-prepared content and scripts, making it difficult to adjust content based on real-time audience feedback. Interaction between viewers and hosts is typically limited to text chat, lacking richer interactive formats. Traditional live streaming content is relatively monotonous, primarily relying on the host's explanations and demonstrations, lacking diverse visual and auditory experiences, and prone to stuttering and latency issues in high-concurrency scenarios, impacting the viewer experience. Traditional online teaching relies mainly on one-way lectures by teachers, with low student participation and a lack of real-time interaction. Once the course content is determined, it is difficult to adjust it according to students' actual needs, lacking flexibility. Traditional online teaching requires teachers to spend a significant amount of time preparing teaching content, including recording videos and creating PowerPoint presentations, increasing their workload. Furthermore, the relatively fixed teaching content makes it difficult to meet the personalized learning needs of different students. Summary of the Invention

[0004] This invention provides a dynamic video synthesis method, apparatus, and instruction editing system based on instruction sequences, which solves the shortcomings of existing technologies in video synthesis that rely on fixed templates, cannot dynamically respond to complex scenes, and are inefficient, thereby enabling flexible customization of the video generation process.

[0005] This invention provides a dynamic video synthesis method based on instruction sequences for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes:

[0006] The system obtains prompt words input by the user, and based on the prompt words, generates an initial instruction sequence for video synthesis through an AI-assisted subsystem and an instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plugin.

[0007] The initial instruction sequence is loaded by the instruction parser, and the corresponding execution plugin is called to execute them sequentially according to the instruction type of the video operation instructions in the initial instruction sequence.

[0008] During the execution of the initial instruction sequence, if user voice input is received, the AI-assisted subsystem performs real-time recognition of the intent of the voice input. If the intent of the voice input is recognized as a question, an answer text is generated and converted into a digital human to broadcast the instruction; if the intent of the voice input is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence continues to be executed.

[0009] Based on the execution results of the instructions, render and composite the visuals and audio.

[0010] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, the video operation instructions include at least digital human broadcast instructions, material playback instructions, scene switching instructions, PPT page turning instructions, wait-for-response instructions, and recording / live streaming instructions.

[0011] The digital human broadcasting instruction is configured with broadcast text and associated digital human parameters;

[0012] The media playback command is configured with the target media path and media configuration parameters;

[0013] The scene switching command is configured with scene layout parameters;

[0014] The PPT page-turning command is configured with page-turning control parameters;

[0015] The waiting response instruction is configured to pause at a specified knowledge point and listen for user questions.

[0016] The recording / live streaming command is configured to control the start and end of recording or to switch the push traffic.

[0017] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, the step of generating an initial instruction sequence for video synthesis based on the prompt words through an AI-assisted subsystem and an instruction editor specifically includes:

[0018] The AI-assisted subsystem calls the large AI model to parse the prompt words into multiple text paragraphs;

[0019] Each text paragraph is converted into a digital human broadcast instruction using an instruction editor; based on the converted digital human broadcast instructions and preset editing operation instructions, an initial instruction sequence for video synthesis is generated, wherein the editing operation instructions include: adding instructions, deleting instructions, modifying instructions, and inserting other video operation instructions between digital human broadcast instructions;

[0020] If other video operation commands inserted are media playback commands or scene switching commands, the command editor receives the user's configuration operations for the media and scenes.

[0021] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, the initial instruction sequence is loaded by an instruction parser, and the corresponding execution plugin is called sequentially according to the instruction type, specifically including:

[0022] When executing the digital human broadcast command, the digital human command plugin is invoked, the digital human module sends the broadcast text to the digital human platform, receives the audio and video stream data callback from the digital human platform and submits the audio and video stream data to the video rendering engine; the broadcast text string is distributed to the scene editor, and the scene editor draws the text on the transparent layer and overlays it onto the composite screen;

[0023] When the media playback command is executed, the media playback command plugin calls the media module to query the video file address, and calls the scene module to pass the file address to the video rendering engine, which then generates a video instance ID.

[0024] When executing a PPT page-turning command, the PPT page-turning command plugin sends page-turning control parameters to the PPT presentation window according to the window message mechanism.

[0025] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, the step of rendering and synthesizing images and audio based on the instruction execution result specifically includes:

[0026] Based on the command execution result and preset rendering parameters, extract the original video stream, audio stream, or image data, and perform leaf node rendering;

[0027] Multiple video streams or image data are superimposed to render the branch nodes of the image, and the audio stream is mixed.

[0028] Integrate scene layout parameters, subtitle layer, and transparent decoration layer for root node rendering;

[0029] Output the composite visuals and audio in Z-Order stacking order.

[0030] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, before inserting a new video operation instruction into the current execution position of the instruction sequence, the method further includes:

[0031] Determine the status of the currently executing plugin. If the currently executing plugin is in a blocked waiting state, immediately interrupt it and execute a new video operation command.

[0032] If the currently executing plugin is running, mark the interruption flag and insert the new video operation instruction when the current video operation instruction finishes execution.

[0033] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, during the execution of the initial instruction sequence, the method further includes:

[0034] The system monitors the execution status of video operation commands in real time. When the execution status is abnormal, the AI-assisted subsystem calls the AI ​​big model to analyze the abnormality type and context, generates an abnormality feedback command, and drives the digital human to broadcast it. The abnormal execution status includes at least one of command execution error, material loading failure, and rendering engine timeout.

[0035] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, when the method is used for AI-assisted teaching and meetings, it further includes:

[0036] Generate digital human broadcast instructions for each slide's narration;

[0037] Insert PPT page turning and scene switching commands between adjacent digital human broadcast commands;

[0038] Insert a wait-for-response instruction at the specified knowledge point to listen for student questions;

[0039] When the method is used in unattended online interactive classrooms, it further includes:

[0040] Pre-embed conditional jump instructions in the initial instruction sequence;

[0041] When the timeout period for waiting for a response command expires, the program will automatically redirect to the next teaching content.

[0042] When a student asks a question, a digital human is dynamically inserted to broadcast the answer.

[0043] This invention also provides a dynamic video synthesis device based on instruction sequences for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser. The device includes:

[0044] An initial instruction sequence generation module is used to obtain prompt words input by the user, and based on the prompt words, generate an initial instruction sequence for video synthesis through an AI-assisted subsystem and an instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plugin.

[0045] The initial instruction sequence execution module is used to load the initial instruction sequence through the instruction parser, and call the corresponding execution plugin to execute it sequentially according to the instruction type of the video operation instructions in the initial instruction sequence;

[0046] The dynamic instruction insertion module is used to, during the execution of the initial instruction sequence, if user voice input is received, to perform real-time recognition of the intent of the voice input through the AI-assisted subsystem. If the intent of the voice input is recognized as a question, an answer text is generated and converted into a digital human to broadcast the instruction; if the intent of the voice input is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence continues to be executed.

[0047] The compositing module is used to render and composite visuals and audio based on the results of command execution.

[0048] This invention also provides an instruction editing system, including an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, a video rendering engine, and device I / O. The instruction module includes an instruction editor and an instruction parser.

[0049] The instruction editor is used to add, delete, and modify video operation instructions according to user needs, so as to adjust the operation of the video operation instructions;

[0050] The instruction parser is used to determine the type of video operation instruction, query the instruction plugin list based on the type determination result, and call the execution plugin that matches the video operation instruction to perform the execution. The instruction plugin list is obtained by dynamically registering the execution plugin corresponding to each instruction type in the instruction parser in advance.

[0051] The AI-assisted subsystem is used to assist in generating a sequence of instructions based on prompts entered by the user when creating a new instruction list; during instruction execution, it recognizes the user's voice input, and if an instruction is recognized, it executes the instruction accordingly; if it is a question, it provides an answer; if it is a user answer, it provides feedback; when answering user questions, it combines an expert knowledge base and a large language model to generate an answer that conforms to the style of a specific role.

[0052] The digital human module is used to connect to a remote digital human platform, receive broadcast text, and obtain synchronous video and audio streams;

[0053] The material module is used to store various types of materials and provide a unified interface for calling;

[0054] The scene module includes a scene editor, which is used to store and manage scene layout parameters and material arrangement rules;

[0055] The video rendering engine is used to render and output the composite image and composite audio in a hierarchical tree structure based on preset video compositing rules, scene layout parameters and material arrangement rules.

[0056] The device I / O includes a device input module and a device output module. The device input module is used to capture user voice or user images, and the device output module is used to output presentation or data interaction.

[0057] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the dynamic video synthesis method based on instruction sequence as described above.

[0058] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the dynamic video synthesis method based on instruction sequences as described above.

[0059] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the dynamic video synthesis method based on instruction sequences as described above.

[0060] This invention provides a dynamic video synthesis method, apparatus, and instruction editing system based on instruction sequences, for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. By acquiring user-input prompts, and based on these prompts, the AI-assisted subsystem and the instruction editor generate an initial instruction sequence for video synthesis. This initial instruction sequence includes multiple types of video operation instructions, each corresponding to an independent execution plugin. The instruction parser then... The initial instruction sequence is loaded, and the corresponding execution plugins are called sequentially to execute the video operation instructions in the initial instruction sequence according to their instruction types. During the execution of the initial instruction sequence, if user voice input is received, the AI-assisted subsystem performs real-time recognition of the intent of the voice input. If the intent of the voice input is recognized as a question, an answer text is generated and converted into a digital human broadcast instruction; if the intent of the voice input is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence continues to be executed. Based on the instruction execution result, the image and audio are rendered and synthesized. This invention achieves dynamic video synthesis by issuing instruction sequences; through the dynamic insertion mechanism of instructions, the image content is adjusted in real time according to the state changes, and the flexibility of instructions meets the needs of customized video synthesis and human-computer interaction; with the assistance of AI, the generation, recognition, and insertion of instructions are realized, which greatly improves the efficiency of video generation, facilitates the interactive display of images and sounds, and is applicable to more application scenarios. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0062] Figure 1 This is a flowchart illustrating the dynamic video synthesis method based on instruction sequences provided by the present invention.

[0063] Figure 2 This is a schematic diagram of the structure of an instruction editing system provided by the present invention.

[0064] Figure 3 This is an interface demonstration of the dynamic video synthesis method based on instruction sequences provided by the present invention.

[0065] Figure 4 A schematic diagram of the structure of the dynamic video synthesis device based on instruction sequences provided by the present invention.

[0066] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0068] The present invention will now be described in detail with reference to the accompanying drawings. The specific operation methods in the method embodiments can also be applied to the device embodiments or system embodiments. In the description of the present invention, unless otherwise stated, "at least one" includes one or more. "Multiple" refers to two or more. For example, at least one of A, B, and C includes: A existing alone, B existing alone, A and B existing simultaneously, A and C existing simultaneously, B and C existing simultaneously, and A, B, and C existing simultaneously. In the present invention, " / " means "or". For example, A / B can mean A or B. "And / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0069] The present invention will now be described in detail with reference to specific embodiments.

[0070] In some specific embodiments of the present invention, such as Figure 1 As shown, this solution provides a dynamic video synthesis method based on instruction sequences for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes:

[0071] Step 100: Obtain the prompt words input by the user. Based on the prompt words, generate an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plugin.

[0072] Step 200: Load the initial instruction sequence through the instruction parser, and call the corresponding execution plugin to execute sequentially according to the instruction type of the video operation instructions in the initial instruction sequence;

[0073] Step 300: During the execution of the initial instruction sequence, if user voice input is received, the AI-assisted subsystem performs real-time recognition of the intent of the voice input. If the intent of the voice input is recognized as a question, an answer text is generated and converted into a digital human to broadcast the instruction; if the intent of the voice input is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence continues to be executed.

[0074] Step 400: Render and composite the visuals and audio based on the execution result of the instructions.

[0075] It should be noted that existing video generation solutions rely on fixed templates, cannot dynamically respond to complex scenarios, and can only adjust the arrangement of materials and timeline manually, which is cumbersome and inefficient. Furthermore, the content cannot be modified in real time according to the user's voice commands during the video generation process (such as interrupting the teaching process during Q&A), thus failing to meet the user's needs for customization and real-time performance.

[0076] Therefore, this invention provides a dynamic video synthesis method based on instruction sequences for AI-assisted teaching, meetings, and unattended online interactive classroom scenarios. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. By acquiring user prompts, an initial instruction sequence is generated based on the prompts, enabling rapid video script construction. A large AI model is used to transform user natural language needs into executable instructions, lowering the operational threshold. Instruction types are mapped to independent plugins, providing modular expansion capabilities. A plugin mechanism supports diverse instructions such as digital human broadcasting and PPT page turning, enhancing system flexibility. Real-time interactive control is achieved through voice intent recognition and dynamic insertion. AI recognizes questions or commands and dynamically inserts instructions, overcoming the limitation of traditional templates that cannot be dynamically adjusted. Layered synthesis of video and audio ensures efficient integration of multi-source data, and layered rendering processes the synthesis logic of materials, digital humans, and subtitles.

[0077] In some possible embodiments of the present invention, such as Figure 2As shown, an instruction editing system is provided, comprising an instruction module 1, an AI-assisted subsystem 2, a material module 3, a scene module or scene editor 4, a digital human module 5, and a video rendering engine 6. Instruction module 1 includes an instruction editor 11 and an instruction parser 12. The instruction editor 11 is used to add, delete, and modify video operation instructions according to user needs to adjust the operation of the video operation instructions. The instruction parser is used to determine the type of the video operation instruction, query the list of registered instruction plugins based on the type determination result, and call the instruction plugin matching the video operation instruction for execution. The AI-assisted subsystem is used to assist in generating a broadcast instruction sequence based on user-input prompts when creating a new instruction list. During instruction execution, it recognizes user voice input; if an instruction is recognized, it executes the corresponding instruction; if it is a question, it provides an answer; if it is a user answer, it provides feedback. When answering user questions, it combines an expert knowledge base and a language model to generate an answer that conforms to the style of a specific character. The material module 3 is used to store various types of materials and provide a unified interface for calling; the scene module 4 includes a scene editor, which is used to store and manage scene layout parameters and material arrangement rules; the digital human module 5 is used to connect to a remote digital human platform, receive broadcast text, and obtain synchronous video and audio streams; the video rendering engine 6 is used to render and output composite images and composite audio in a hierarchical tree structure based on preset video compositing rules, scene layout parameters, and material arrangement rules; the device I / O includes a device input module and a device output module, wherein the device input module is used to capture user voice or user images, and the device output module is used to output presentation or data interaction.

[0078] Specifically, the device input module of the device I / O is used for input control, receiving user voice commands such as capturing voice through a microphone to trigger real-time AI recognition and dynamic command insertion. The device input module can also be used to capture user images, such as user gestures. The device output module of the device I / O is used for output presentation and data interaction. Output presentation includes displaying composite images, playing composite audio, etc., with the display showing rendered images and the speaker playing digital human voice / mixed audio, etc. Data interaction supports live streaming and external signal access, specifically network streaming of camera signals, NDI / live signals, screen projection, and other external inputs, which are input from the device I / O to the material module. It can also be used for control interfaces, such as network streaming of remote control commands.

[0079] In some possible embodiments of the present invention, the dynamic video synthesis method based on instruction sequences provided by the present invention further includes:

[0080] The system monitors the execution status of video operation commands in real time. When the execution status is abnormal, the AI-assisted subsystem calls the AI ​​big model to analyze the abnormality type and context, generates an abnormality feedback command, and drives the digital human to broadcast it. The abnormal execution status includes at least one of command execution error, material loading failure, and rendering engine timeout.

[0081] Specifically, this embodiment provides an implementation method for handling execution status anomalies. By using an AI-assisted subsystem to assist in execution status monitoring and anomaly feedback, the robustness of the system is improved (extended functionality). By using AI to diagnose problems such as rendering timeouts and loading failures, the digital human is driven to broadcast anomaly prompts to avoid process interruption.

[0082] In some possible embodiments of the present invention, the video operation instructions include at least digital human broadcast instructions, material playback instructions, scene switching instructions, PPT page turning instructions, wait for response instructions, and recording / live streaming instructions.

[0083] The digital human broadcasting instruction is configured with broadcast text and associated digital human parameters;

[0084] The media playback command is configured with the target media path and media configuration parameters;

[0085] The scene switching command is configured with scene layout parameters;

[0086] The PPT page-turning command is configured with page-turning control parameters;

[0087] The waiting response instruction is configured to pause at a specified knowledge point and listen for user questions.

[0088] The recording / live streaming command is configured to control the start and end of recording or to switch the push traffic.

[0089] Specifically, this embodiment provides an implementation method for video operation instructions. The examples of video operation instructions described above only list some commonly used instruction types. In actual applications, other specific settings can be made according to different user needs, and it is not limited to the examples listed in this embodiment. It is worth noting that, in addition to obtaining user prompts and generating video operation instructions based on those prompts, video operation instructions, such as narration instructions, signal switching instructions, scene switching instructions, and voice question-and-answer instructions, can also be set before or during the video synthesis process.

[0090] In some possible embodiments of the present invention, the statement based on the prompt word,

[0091] Based on the prompts, an initial instruction sequence for video synthesis is generated through an AI-assisted subsystem and an instruction editor, specifically including:

[0092] The AI-assisted subsystem calls the large AI model to parse the prompt words into multiple text paragraphs;

[0093] Each text paragraph is converted into a digital human broadcast instruction using an instruction editor; based on the converted digital human broadcast instructions and preset editing operation instructions, an initial instruction sequence for video synthesis is generated, wherein the editing operation instructions include: adding instructions, deleting instructions, modifying instructions, and inserting other video operation instructions between digital human broadcast instructions;

[0094] If other video operation commands inserted are media playback commands or scene switching commands, the command editor receives the user's configuration operations for the media and scenes.

[0095] Specifically, this embodiment provides an implementation method for generating an initial instruction sequence. By calling an AI language model to parse the prompt words input by the user, the instruction sequence is optimized. For example, the prompt word "explain the four great inventions" is broken down into segmented broadcast instructions and bound to PPT page turning dependencies.

[0096] Specifically, first, the user inputs prompts, which are video generation requirements described in natural language. Examples include "explain the four great inventions of ancient China in three segments" or "use prompts to provide narration for each news video segment."

[0097] Furthermore, the prompt words are parsed into multiple text paragraphs using a large AI language model (such as a GPT-like model). For example, given the prompt words: "Explanation of the Four Great Inventions, 1 minute for each invention, followed by a summary," the parsed text paragraphs would look like this:

[0098] The invention of gunpowder;

[0099] The invention of papermaking;

[0100] The invention of printing;

[0101] The invention of the compass;

[0102] One minute per invention;

[0103] Summarize the significance of the Four Great Inventions.

[0104] The instruction editor converts each text paragraph into a digital human-generated instruction:

[0105] Broadcast instruction 1: Explanation of the invention of gunpowder (1 minute);

[0106] Broadcast instruction 2: Explanation of the invention of papermaking (1 minute);

[0107] Broadcast instruction 3: Explanation of the invention of printing (1 minute);

[0108] Broadcast instruction 4: Explanation of the invention of the compass (1 minute in length);

[0109] Broadcast instruction 5: Summarize the significance of the Four Great Inventions.

[0110] Each operation corresponds to a type of instruction (digital human broadcasting, PPT page turning, etc.).

[0111] Furthermore, the system dynamically binds scene parameters, automatically associating them with scene context parameters: the current scene layout (e.g., the position of the PPT and the digital human in a teacher's lecture scene) or material attributes (e.g., material position, PPT page number, video clip path). For example, upon receiving the prompt "Add narration to each PPT page," the system automatically binds the PPT page number to the broadcast command, ensuring that the narration is synchronized with the page.

[0112] Building upon this foundation, further supplementary non-broadcast commands such as scene switching commands, PPT page turning commands, and wait-for-response commands are added. Necessary dependent commands, such as delays between commands, are automatically added to ensure workflow continuity. For example, for a digital human broadcast command, the inserted command is a PPT page turning command (if associated with a PPT), ensuring synchronized page turning during broadcast. For a continuous video playback command, the inserted command is a scene switching command to change the layout of the materials. For a Q&A session, the inserted command is a wait-for-response command (timeout jump) to listen for user voice input. For instance, "Insert a wait-for-response command at the knowledge point to listen for student questions," automatically inserting the wait command after explaining key knowledge points.

[0113] Finally, an editable sequence of instructions is generated, and the output is a visual display of the generated instruction sequence in the instruction list editor. Examples of instruction sequences include: Digital Human's announcement instruction: Explanation of the invention of gunpowder (linked to PPT page 5); PPT page turning instruction: Jump to page 6; Digital Human's announcement instruction: Explanation of the invention of papermaking (linked to PPT page 6); Wait for response instruction (timeout 30 seconds).

[0114] Furthermore, users can manually add, delete, or modify commands such as adjusting the broadcast text and inserting special effects commands.

[0115] It's worth noting that traditional templates require pre-defining all materials, while the above-described settings in this invention allow for dynamic generation of instruction sequences via AI, adapting to any prompt requirements (such as from "Four Great Inventions" to "Industrial Revolution"), thus enabling flexible customization of dynamic videos. Simultaneously, instruction insertion (such as waiting for a response) directly supports real-time classroom interaction (such as the requirement for "dynamic adjustment of image signals"), achieving real-time dynamic adjustment of the video synthesis process. Furthermore, users do not need to understand the technical details of video synthesis; natural language descriptions are sufficient to generate professional instruction sequences, lowering the barrier to entry for dynamic video customization and synthesis.

[0116] For example, if the input prompt is: "Add narration to a 20-page PPT, using a teacher scenario on each page, and switching to a full-screen PPT scenario every 5 pages," the AI ​​will generate the following instruction sequence:

[0117] Loop 20 times: Digital human broadcast command (bound to the current PPT page) + PPT page turning command;

[0118] Insert every 5 times: Scene switching command (from teacher scene to full-screen PPT scene).

[0119] This process transforms user natural language needs into executable and editable instruction sequences through AI language model parsing, instruction editor conversion, scene parameter configuration, and instruction insertion. It fundamentally solves the problems of insufficient flexibility and lack of dynamic interaction in traditional video template systems, while significantly reducing the operational threshold.

[0120] In some possible embodiments of the present invention, the initial instruction sequence is loaded by an instruction parser, and the corresponding execution plugin is called sequentially according to the instruction type, specifically including:

[0121] When executing the digital human broadcast command, the digital human command plugin is invoked, the digital human module sends the broadcast text to the digital human platform, receives the audio and video stream data callback from the digital human platform and submits the audio and video stream data to the video rendering engine; the broadcast text string is distributed to the scene editor, and the scene editor draws the text on the transparent layer and overlays it onto the composite screen;

[0122] When the media playback command is executed, the media playback command plugin calls the media module to query the video file address, and calls the scene module to pass the file address to the video rendering engine, which then generates a video instance ID.

[0123] When executing a PPT page-turning command, the PPT page-turning command plugin sends page-turning control parameters to the PPT presentation window according to the window message mechanism.

[0124] Specifically, this embodiment provides an implementation method for calling an execution plugin to execute an initial instruction sequence. Through the execution flow of the digital human plugin, a digital human broadcasting pipeline is realized, connecting the entire chain from platform communication to audio and video stream reception and rendering compositing. Through the execution flow of the video playback plugin, dynamic on-screen display of materials is completed, and the scene module coordinates the rendering engine to manage the lifecycle of video instances.

[0125] In some possible embodiments of the present invention, the layered rendering and synthesis of visuals and audio are performed based on the result of instruction execution, specifically including:

[0126] Based on the command execution result and preset rendering parameters, extract the original video stream, audio stream, or image data, and perform leaf node rendering;

[0127] Multiple video streams or image data are superimposed to render the branch nodes of the image, and the audio stream is mixed.

[0128] Integrate scene layout parameters, subtitle layer, and transparent decoration layer for root node rendering;

[0129] Output the composite visuals and audio in Z-Order stacking order.

[0130] Specifically, this embodiment provides an implementation method for tree-structured layered rendering. By rendering in the order from leaves to branches to the root node, the workflow of the rendering engine is clarified, and the stacking priority is controlled by Z-Order (such as subtitles always being on the top layer). Layered processing ensures compositing efficiency.

[0131] In a possible embodiment, the CPU algorithm or GPU computing framework can be automatically selected to construct the rendering map based on the user's computer graphics card performance, and CPU / GPU acceleration can be switched according to the user's computer configuration to dynamically optimize resource usage; text can be drawn on the transparent layer through the scene module and superimposed onto the composite image.

[0132] In a possible implementation, layered rendering enables parallel processing, such as asynchronous execution of leaf node decoding and branch compositing. Adding new material types (such as projection signals) only requires extending the leaf node plugin; transparent layers / subtitles, as independent branch nodes, can be dynamically modified, allowing for flexible expansion.

[0133] In some possible embodiments of the present invention, before inserting a new video operation instruction into the current execution position of the instruction sequence, the method further includes:

[0134] Determine the status of the currently executing plugin. If the currently executing plugin is in a blocked waiting state, immediately interrupt it and execute a new video operation command.

[0135] If the currently executing plugin is running, mark the interruption flag and insert the new video operation instruction when the current video operation instruction finishes execution.

[0136] Specifically, this embodiment provides an implementation of an instruction insertion strategy. By blocking interrupts and running flags, the real-time nature of dynamic insertion is ensured. The plugin status is distinguished to determine whether to interrupt immediately or insert sequentially, thus avoiding instruction conflicts.

[0137] In some possible embodiments of the present invention, when the dynamic video synthesis method based on instruction sequences is used for AI-assisted teaching and meetings, it further includes:

[0138] Generate digital human broadcast instructions for each slide's narration;

[0139] Insert PPT page turning and scene switching commands between adjacent digital human broadcast commands;

[0140] Insert a wait-for-response instruction at the specified knowledge point to listen for student questions.

[0141] Specifically, this embodiment provides an implementation method for teaching scenarios, which uses AI to assist teaching, achieves automated course recording, and simulates the teacher's teaching process through command combinations (such as narration + PPT page turning + scene switching).

[0142] In some possible embodiments of the present invention, when the dynamic video synthesis method based on instruction sequences is used in unattended online interactive classrooms, it further includes:

[0143] Pre-embed conditional jump instructions in the initial instruction sequence;

[0144] When the timeout period for waiting for a response command expires, the program will automatically redirect to the next teaching content.

[0145] When a student asks a question, a digital human is dynamically inserted to broadcast the answer.

[0146] Specifically, this embodiment provides another implementation method for teaching scenarios. Through unattended interactive classrooms, it supports interactive classrooms, pre-embeds jump instruction timeout handling, dynamically inserts question and answer response instructions, and realizes fully automatic teaching.

[0147] In a possible embodiment, during the execution of the initial instruction sequence, user input is received in real time. The user input includes voice input and text input. If the user input is voice input, it is converted into text input and then transmitted to the AI ​​big model for intent recognition.

[0148] In possible implementations, user input includes not only voice and text input but also key input. Before the AI ​​model recognizes the user's intent, initial recognition of the user input allows the system to handle different forms of input more comprehensively. When the initial recognition result is a command, it is dynamically inserted into the current execution position and the registered plugin is invoked for execution. This improves the system's response speed and efficiency to explicit commands, providing a prerequisite for the AI ​​model to subsequently recognize the intent of non-command inputs and perfecting the overall logic of user input processing.

[0149] The dynamic video synthesis method based on instruction sequences provided in this invention abstracts the video generation process through instruction sequences and solves the flexibility problem of traditional template systems by using plug-in execution and AI dynamic interaction. It provides a practical technical solution, especially for the real-time interactive needs of educational scenarios (such as classroom Q&A).

[0150] In some possible embodiments of the present invention, see still Figure 2 The dynamic video generation system provided in this embodiment of the invention also includes a material module 3, a scene editor 4, a digital human module 5, a video rendering engine 6, and a device IO7.

[0151] In a possible embodiment, the instruction parser 11 supports plug-in extensions, with each instruction type corresponding to an independent plug-in, which is registered to the instruction parser 11 when the system starts.

[0152] Specifically, this embodiment provides one implementation of the instruction parser 11. The plug-in extension mechanism of the instruction parser 11 enhances the system's flexibility and scalability. By setting independent plug-ins for each instruction type and registering them with the instruction parser at system startup, this embodiment enables the system to easily add new instruction types without requiring large-scale modifications to the core system. This provides technical support for system function expansion and ensures that the system can adapt to future technological developments and changes in user needs.

[0153] In some possible embodiments of the present invention, the material module 3 supports plugin extensions, with each type of material corresponding to an independent plugin, which is registered to the material module 3 when the system starts.

[0154] Specifically, this embodiment provides one implementation method for the material module. The plug-in extension mechanism of the material module enhances the flexibility and scalability of the system. This embodiment sets up independent plug-ins for each type of material and registers them with the material module when the system starts. This allows the system to easily add new material types, enriching material resources and providing technical support for the system's material management. It ensures that the system can adapt to different types of material needs, improving the system's versatility and adaptability.

[0155] In some possible embodiments of the present invention, the material module 3 supports local and cloud storage, calls materials through a unified interface, and supports encryption and copyright protection of materials.

[0156] Specifically, this embodiment provides another implementation of the material module 3. Material module 3 supports not only local storage but also cloud storage, enabling the system to flexibly handle materials from different sources and formats, meeting user needs under varying network environments and resource conditions. The material module designs a unified plugin interface for each type, including functions such as opening, closing, and reading data. Calling materials through a unified interface ensures the system's compatibility and consistency when processing materials of different types and sources, reducing development and maintenance costs. The material module needs to handle various types of materials, which may involve copyright issues. Through encryption and copyright protection functions, the security and legality of the materials during storage and use are ensured.

[0157] In a specific embodiment, see still Figure 2 The video generation method provided by this invention uses... Figure 2 The architecture implementation includes a material module 3, a digital human module 5, scene modules such as a scene editor 4, an instruction module 1, a video rendering engine 6, an AI-assisted subsystem 2, and a device IO 7, wherein the device IO 7 includes a device input module 71 and a device output module 72.

[0158] The material module 3 stores various types of materials used for video generation, including videos, images, camera signals, live stream signals, NDI signals, Apple screen mirroring signals, PPTs, windows, and screens. To read these materials, the program has designed a unified plugin interface for each type. These interfaces include functions such as opening, closing, reading data, and reading parameter information. Basic material parameters include width and height, frame rate, pixel format, and playback status.

[0159] Digital Human Module 5 is used to connect to remote digital human platforms, such as Alibaba's Avatar digital human platform. By pushing broadcast commands to the digital human, it can be driven to speak and perform gestures. Digital humans can essentially be viewed as a type of content.

[0160] Specifically, the basic connection process for the digital human module is as follows:

[0161] Step 1: Access the user platform space through your digital human platform account to obtain a list of available digital human avatars;

[0162] Step two: The user selects a digital avatar;

[0163] Step 3: The program creates an instance on the digital human platform using the digital human ID and establishes a path for the transmission of image and sound signals.

[0164] Step four: During video generation, the program sends the broadcast text through the transmission channel; the digital human platform returns the calculated image and sound signals; the program receives these signals and sends them to the image engine for rendering and display.

[0165] Furthermore, the scene module or scene editor 4 is used to arrange materials, display subtitles, and decorate backgrounds. The scene module, used for managing and editing scenes, includes two sub-modules: a scene editor and a scene container. Through the scene editor, materials can be selected, and parameters such as their placement, size, cropping area, and transparency can be set. Scenes are more descriptive of specific situations than individual materials; for example, a teaching scene can be displayed simultaneously from the perspectives of a teacher lecturing and students listening.

[0166] Instruction Module 1 refers to a set of instructions related to video generation. Basic instruction types include broadcast instructions for digital human announcements, instructions for uploading materials, instructions for video playback, instructions for scene switching, instructions for turning PPT slides, operation instructions for live streaming and recording, and instructions for waiting. In addition to a list of display instructions, the system also includes an instruction editor and generator.

[0167] Specifically, specifying the generation and modification methods includes:

[0168] Step one: Based on the user's prompts, generate a set of text-based broadcast instructions, such as: dividing the story into several paragraphs to describe the four great inventions of ancient China. Then, call upon a large AI model to generate several text segments. Combining these texts, the program generates the broadcast instructions for the digital human.

[0169] Step two: The broadcast commands are displayed in the command list. The command list is also an editor, which allows users to add, delete, and modify commands to adjust their execution and meet more detailed user needs.

[0170] Step 3: When the instruction is being executed, the user can interrupt with language. The AI ​​will recognize the user's words. If it is recognized as a question, the AI ​​will generate an answer text using a large model and then send the broadcast instruction to the digital human to read it. If it is recognized as a command, it will be converted into a message and passed to the interface for response.

[0171] Correspondingly, the counterpart to instruction generation is the instruction parser, which refers to the mechanism that parses and executes each instruction in the instruction list. To improve the extensibility of instructions, each type of instruction is designed as a plugin, which is registered with the parser when the program starts. When the parser receives an instruction, it determines the type of the instruction and queries the list of registered instruction plugins to find the matching instruction to execute. For example, if the parser receives a digital human playback instruction, it will find the digital human plugin in the list of registered plugins and then drive the digital human plugin to complete the task.

[0172] The video rendering engine 6 refers to the underlying computing structure that synthesizes footage, digital humans, scenes, and sound together. The basic logic of the video rendering engine is as follows: Within a certain time sequence, it sequentially decodes one frame from each relevant piece of footage. Based on the placement, layering order, and effects parameters, it uses a CPU algorithm, or uploads data to a graphics card, applying a GPU computing framework to construct a graph. It renders the leaves first, then the branches and roots, ultimately forming a single image. Whether to use a CPU or GPU algorithm depends on the user's computer configuration; if the user has a powerful graphics card, the program defaults to using the GPU algorithm. For sound, it decodes and samples audio data, then uses a CPU synthesis algorithm to obtain synthesized sound.

[0173] AI-assisted subsystem 2 refers to using AI to assist users in creating command sequences. The specific functions of the AI-assisted subsystem include the following aspects:

[0174] Firstly, when creating a new instruction list, users can enter prompts, and AI will assist in generating a sequence of instructions to be broadcast.

[0175] Secondly, during the execution of instructions, AI recognizes the user's voice input. If an instruction is recognized, it is executed accordingly; if it is a question, an answer is given; if it is a user's response, feedback is provided.

[0176] Thirdly, when answering user questions, AI will combine expert knowledge bases and large language models to provide answers that match the style of a specific role.

[0177] Furthermore, the device I / O mentioned in this embodiment refers to a series of input / output devices, including device input module 71, such as microphones and cameras, and device output module 72, whose output objects include video files, live signals, screens, or speakers. For example, device I / O may include: a display for showing rendered images, speakers for supporting digital human broadcasts and video playback, a network stream for audio / video data transmission, and a sound input device, such as a microphone, for receiving user commands during instruction execution.

[0178] Based on the above embodiments, after video generation is initiated, the program will execute the following steps sequentially:

[0179] Step 201: The video rendering engine opens the video encoder and waits for each frame to be transmitted.

[0180] Step 202: The instruction parser obtains the instruction list and calls different instruction plugins to execute different instruction types.

[0181] In more detail, the basic execution of the digital human plugin and video playback plugin is as follows:

[0182] Step 202-1, Digital Human Instruction Plugin:

[0183] In this step, firstly, the plugin sends the broadcast text to the digital human platform (e.g., Alibaba Cloud avatar) via a network interface and then waits;

[0184] Based on this, the plugin receives a platform callback, which includes a video stream, an audio stream, and a text string to be played.

[0185] Furthermore, the plugin calls the rendering engine interface to complete (1) inserting video stream data, and the rendering engine combines the scene to synthesize a frame, which is then sent to the preview window, encoder and live interface respectively (Note: if open); (2) inserting audio stream, which is then synthesized and sent to the speaker, encoder and live interface respectively.

[0186] Furthermore, Figure 3 This is an interface demonstration of the dynamic video synthesis method based on instruction sequences provided by the present invention, such as... Figure 3 As shown, the text string obtained from the plugin callback is sent back to the interface through the instruction parser, and the interface distributes it to the instruction list as follows. Figure 3 The script list in the script is highlighted during broadcasting; at the same time, it is distributed to the scene editor, which draws text on a transparent layer and inputs it into the rendering engine. The rendering engine overlays the rendered image onto the composited image, which is also used for previewing, encoders and live streaming.

[0187] Finally, the plugin completes its task and sends a signal to the parser.

[0188] Step 202-2, Video Command Plugin:

[0189] In this step, the plugin first obtains the name of the input video file and calls the media module to query the file address of the video;

[0190] Next, the plugin sends the file address to the scene module, which then sends it to the rendering engine. The rendering engine returns the video instance ID, completing the video display.

[0191] Third, the plugin calls the rendering engine interface, passes in the video instance ID, starts playback, and waits;

[0192] Fourth, the rendering engine decodes the video, combines it with the scene, synthesizes the visual and audio data, and distributes it to the pre-monitoring, encoder, and live streaming interface;

[0193] Fifth, the plugin waits for the video playback to end signal returned by the rendering engine;

[0194] Sixth, the plugin calls the scene module, the scene module calls the rendering engine, closes the video instance ID, clears the video on-screen information, and completes the video closing;

[0195] Seventh, the plugin completes its task and sends a signal to the parser.

[0196] Step 203: The parser receives the plugin task completion signal and then executes the next instruction;

[0197] Step 204: If a user's verbal command is recognized from the audio input IO, the command is inserted into the current position of the command list for timely execution.

[0198] Step 205: The parser returns after executing all instructions;

[0199] Step 206: Close the file editor in the video rendering engine.

[0200] This embodiment categorizes and defines instructions, associating them with video generation. Supported by a material container, AI digital human, scene container, scene editor, instruction parser, and video rendering engine, it achieves instruction parsing and execution. With the support of the instruction editor and language model, it enables automatic generation and customized modification of instruction sequences. With the support of device I / O and the language model, it achieves dynamic recognition, insertion, parsing, and execution of instructions. Through the above settings of this embodiment, compared to existing technical solutions, using instruction sequences to describe the video generation process and leveraging AI digital human, language model, and proprietary rendering platform, it achieves efficient video generation. Describing video in a more straightforward way significantly reduces the difficulty of video editing and generation; it reduces video shooting steps, lowers video generation costs and complexity, and improves the flexibility and efficiency of video generation; by simply modifying instructions, it allows for easy intervention and dynamic adjustment of the video generation process; and by combining the instruction system with AI, it facilitates more interactive video generation.

[0201] In the specific embodiments of the present invention, see still Figure 2 The dynamic video synthesis method and dynamic video synthesis system based on instruction sequences provided by this invention are implemented through the following modules: material module 3, scene module or scene editor 4, digital human module 5, instruction module 1, video rendering engine 6, AI-assisted subsystem 2, and device IO 7.

[0202] In practical applications, this invention can be used to conveniently and efficiently test the following four types of application examples:

[0203] Application Example 1: News broadcast by a presenter;

[0204] Application Example 2: Course Recording;

[0205] Application Example 3: Large Screen Q&A;

[0206] Application Example 4: AI Classroom

[0207] For example application 1: Load all news clips into the material module, select a scene to display the digital human; in the instruction module, use prompts to add a narration to each news video clip, and the narration will be sent to the digital human as a broadcast instruction. Then, insert the instruction to play the corresponding news clip between each clip. In this way, a simple news broadcast instruction sequence is completed, and video generation can be started.

[0208] For application example two: Open and play a PowerPoint presentation. Design two scenarios. In the first scenario, place the PowerPoint presentation on the left and a digital human on the right, simulating a teacher lecturing at a blackboard. In the second scenario, display the entire PowerPoint presentation. With AI assistance, add narration to each slide as a broadcast command to the digital human. Set each slide to use either the first or second scenario, and insert "PowerPoint page turn" commands between the narrations. This creates a simple sequence of instructions for recording a lesson; simply start video generation.

[0209] For application example three: design a vertical screen scene, place the digital human in it, import the expert Q&A knowledge base in the instruction module, set a waiting Q&A instruction, and a simple Q&A instruction for the large screen is completed.

[0210] Application Example 4 is a combination of Application Examples 2 and 3. The specific process includes: first, opening a PowerPoint presentation to create a teacher-led teaching scenario, then creating a PowerPoint presentation scenario, and finally creating a scenario with only a digital human; importing an expert Q&A knowledge base into the instruction module, and designing a sequence of instructions as in Example 2 to enable the teaching, and setting the instruction module to accept questions and answers. The instructions are then started, and the digital human teaches normally without interruption. If someone's voice enters, the device I / O will recognize it. If a question is identified, the device first searches the expert Q&A knowledge base, then generates an answer text using a large model, which is read aloud by the digital human. If no question is asked and the timeout period expires, the digital human returns to normal teaching mode.

[0211] The following is the basic process of executing digital human broadcast commands, video playback commands, and PPT page turning commands:

[0212] A. The digital human broadcasts commands through the digital human command plugin, including the following steps:

[0213] Step 1: The plugin sends the broadcast text to the digital human platform (e.g., Alibaba Cloud avatar) via a network interface and waits;

[0214] Step 2: The plugin receives a platform callback, which includes a video stream, an audio stream, and a text string to be played.

[0215] Step 3: The plugin calls the rendering engine interface, inserts video stream data, and the rendering engine combines the scene to synthesize a frame, which is then sent to the preview window, encoder, and live streaming interface (Note: if enabled); the audio stream is inserted, synthesized, and then sent to the speakers, encoder, and live streaming interface.

[0216] Step 4: The text string obtained from the plugin callback is sent back to the interface through the parser. The interface then distributes it to the instruction list for highlighting during broadcast. At the same time, it is also distributed to the scene editor, which draws the text on the transparent layer and inputs it into the rendering engine. The rendering engine overlays it onto the composited image, which is also used for previewing, encoders, and live streaming.

[0217] B. Video playback commands are executed through the video playback command plugin, specifically including the following steps:

[0218] Step 1: The plugin obtains the name of the input video file and calls the media module to query the file address of the video.

[0219] Step 2: The plugin sends the file address to the scene module, which then sends it to the rendering engine. The rendering engine returns the video instance ID, completing the video display.

[0220] Step 3: The plugin calls the rendering engine interface, passes in the video instance ID, starts playback, and waits;

[0221] Step 4: The rendering engine decodes the video, combines it with the scene, synthesizes the visual and audio data, and distributes it to the pre-monitor (including speakers), encoder, and live streaming interface;

[0222] Step 5: The plugin waits for the video playback to end signal returned by the rendering engine;

[0223] Step 6: The plugin calls the scene module, which in turn calls the rendering engine to close the video instance ID, clear the video display information, and complete the video closing process.

[0224] C. PPT page-turning commands are executed through interface dialog box message responses, specifically including the following steps:

[0225] Step 1: Query the handle of the PowerPoint presentation window;

[0226] Step 2: Send the page-turning message parameters to the window.

[0227] The dynamic video synthesis method based on instruction sequences provided in this invention classifies and defines instructions, associating them with video generation; with the joint support of a material container, AI digital human, scene container, scene editor, instruction parser, and video rendering engine, it realizes the parsing and execution of instructions; with the support of an instruction editor and a large language model, it realizes the automatic generation and customized modification of instruction sequences; and with the support of device I / O and a large language model, it realizes the dynamic recognition, insertion, parsing, and execution of instructions.

[0228] The method provided in this embodiment of the invention is essentially a technical solution for defining video synthesis. Compared to existing technical solutions, this solution uses instruction sequences to describe the video generation process and leverages AI assistance, digital humans, large language models, and a proprietary rendering platform to achieve efficient video synthesis. The technical effects achievable by this solution can be summarized in at least the following aspects:

[0229] On the one hand, this invention describes video in a more straightforward way, significantly reducing the difficulty of video editing and generation;

[0230] On the other hand, the present invention can reduce the video shooting process, reduce the cost and complexity of video generation, and improve the flexibility and efficiency of video generation.

[0231] Thirdly, by simply modifying the instructions, the present invention can easily intervene in and dynamically adjust the video generation process;

[0232] Fourthly, by combining an instruction system with AI, this invention makes it easier to achieve interactive video synthesis.

[0233] In some specific embodiments of the present invention, such as Figure 4 As shown, this solution provides a dynamic video synthesis device based on instruction sequences for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser. The device includes:

[0234] The initial instruction sequence generation module 41 obtains the prompt words input by the user, and generates an initial instruction sequence for video synthesis based on the prompt words through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plugin.

[0235] The initial instruction sequence execution module 42 loads the initial instruction sequence through the instruction parser and calls the corresponding execution plugins to execute them sequentially according to the instruction type of the video operation instructions in the initial instruction sequence.

[0236] The dynamic instruction insertion module 43, during the execution of the initial instruction sequence, if it receives user input voice, it uses the AI-assisted subsystem to recognize the intent of the input voice in real time. If the intent of the input voice is recognized as a question, it generates an answer text and converts it into a digital human broadcast instruction; if the intent of the input voice is recognized as a control command, it converts it into a new video operation instruction and inserts the new video operation instruction into the current execution position of the instruction sequence; after executing the newly inserted video operation instruction through the instruction parser, it continues to execute the subsequent instruction sequence.

[0237] The compositing module 44 renders and composites the visuals and audio based on the results of the instruction execution.

[0238] The dynamic video synthesis apparatus based on instruction sequences provided in this embodiment of the invention has a similar implementation principle and beneficial effects to the dynamic video synthesis method based on instruction sequences shown in the above embodiments. Please refer to the implementation principle and beneficial effects of the dynamic video synthesis method based on instruction sequences shown in the above embodiments, which will not be repeated here.

[0239] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. The processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a dynamic video synthesis method based on instruction sequences. This method is used for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes: obtaining prompts input by the user; and based on the prompts, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution insert. The system loads the initial instruction sequence through an instruction parser and sequentially executes the corresponding execution plugins based on the instruction type of the video operation instructions in the initial instruction sequence. During the execution of the initial instruction sequence, if user voice input is received, the AI-assisted subsystem performs real-time recognition of the intent of the voice input. If the intent of the voice input is recognized as a question, an answer text is generated and converted into a digital human to broadcast the instruction. If the intent of the voice input is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After executing the newly inserted video operation instruction through the instruction parser, the subsequent instruction sequence continues to be executed. Based on the instruction execution result, the image and audio are rendered and synthesized.

[0240] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0241] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the dynamic video synthesis method based on instruction sequences provided by the above methods. This method is used for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes: obtaining prompt words input by the user; and based on the prompt words, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The instruction sequence includes multiple types of video operation instructions, each corresponding to an independent execution plugin. The initial instruction sequence is loaded via an instruction parser, and the corresponding execution plugins are called sequentially based on the instruction type of the video operation instructions within the initial instruction sequence. During the execution of the initial instruction sequence, if user voice input is received, the AI-assisted subsystem performs real-time recognition of the intent of the voice input. If the intent is recognized as a question, an answer text is generated and converted into a digital human to broadcast the instruction. If the intent is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After executing the newly inserted video operation instruction via the instruction parser, the subsequent instruction sequence continues to be executed. Based on the instruction execution results, the video and audio are rendered and synthesized.

[0242] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, this computer program implements the dynamic video synthesis method based on instruction sequences provided by the methods described above. This method is used for AI-assisted teaching, meetings, and unattended online interactive classrooms. It is implemented through an instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes: obtaining prompts input by a user; and based on the prompts, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operations. The system generates instructions, with each instruction type corresponding to an independent execution plugin. An initial instruction sequence is loaded via an instruction parser, and the corresponding execution plugin is called sequentially based on the instruction type of the video operation instructions in the initial instruction sequence. During the execution of the initial instruction sequence, if user voice input is received, the AI-assisted subsystem performs real-time recognition of the intent of the voice input. If the intent is recognized as a question, an answer text is generated and converted into a digital human to broadcast the instruction. If the intent is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After executing the newly inserted video operation instruction via the instruction parser, the subsequent instruction sequence continues to be executed. Based on the instruction execution results, the video and audio are rendered and synthesized.

[0243] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0244] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0245] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method of dynamic video synthesis based on instruction sequences, characterized in that, The application discloses an online interactive classroom for AI-assisted teaching, meeting and unattended teaching, which is realized through an instruction editing system, wherein the instruction editing system comprises an instruction module, an AI-assisted subsystem, a digital person module, a material module, a scene module and a video rendering engine; the instruction module comprises an instruction editor and an instruction parser; the scene module comprises a scene editor; the method comprises the following steps: acquiring a prompt word input by a user, generating an initial instruction sequence of video synthesis based on the prompt word through the AI-assisted subsystem and the instruction editor, wherein the initial instruction sequence comprises a plurality of types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; loading the initial instruction sequence through the instruction parser, and sequentially executing the corresponding execution plug-in according to the instruction type of the video operation instruction in the initial instruction sequence; during the execution of the initial instruction sequence, if an input voice of the user is received, the intention of the input voice is identified in real time through the AI-assisted subsystem, if the intention of the input voice is identified as a question, a reply text is generated and converted into a digital person broadcasting instruction, if the intention of the input voice is identified as a control command, a new video operation instruction is converted and inserted into the current execution position of the instruction sequence, and after the new inserted video operation instruction is executed through the instruction parser, the subsequent instruction sequence is continuously executed; rendering and synthesizing pictures and audios according to the instruction execution result; during the execution of the initial instruction sequence, the following steps are further included: monitoring the execution state of the video operation instruction in real time, when the execution state is abnormal, the AI big model is called through the AI-assisted subsystem to analyze the abnormal type and context, an abnormal feedback instruction is generated and a digital person is driven to broadcast, and the execution state abnormality includes at least one of instruction execution error, material loading failure and rendering engine timeout; when the method is used for AI-assisted teaching and meeting, the following steps are further included: generating a digital person broadcasting instruction for each page of PPT commentary; inserting a PPT page turning instruction and a scene switching instruction between adjacent digital person broadcasting instructions; inserting a waiting response instruction at a specified knowledge point to listen to student questions; when the method is used for an unattended online interactive classroom, the following steps are further included: pre-embedding a conditional jump instruction in the initial instruction sequence; when the waiting response instruction is timed out, automatically jumping to subsequent teaching content; when a student question is identified, a digital person broadcasting instruction is dynamically inserted to play an answer.

2. The method of claim 1, wherein, The video operation instruction at least comprises a digital person broadcasting instruction, a material playing instruction, a scene switching instruction, a PPT page turning instruction, a waiting response instruction and a recording / live broadcast instruction, wherein the digital person broadcasting instruction is configured with a broadcasting text and associated digital person parameters; the material playing instruction is configured with a target material path and material configuration parameters; the scene switching instruction is configured with scene layout parameters; the PPT page turning instruction is configured with page turning control parameters; the waiting response instruction is configured to pause at a specified knowledge point to listen to user questions; the recording / live broadcast instruction is configured to control the start, end or push flow switch of recording.

3. The method of claim 2, wherein, The initial instruction sequence of the video synthesis is generated based on the prompt word through an AI auxiliary subsystem and an instruction editor, and specifically includes: The AI auxiliary subsystem calls an AI large model to parse the prompt word into multiple text paragraphs; The instruction editor converts each text paragraph into a digital human broadcast instruction, and generates the initial instruction sequence of the video synthesis based on the converted digital human broadcast instruction and a preset editing operation instruction, wherein the editing operation instruction includes an adding instruction, a deleting instruction, a modifying instruction, and an inserting instruction of other video operation instructions between the digital human broadcast instructions; If the inserted other video operation instruction is a material playing instruction or a scene switching instruction, the instruction editor receives user configuration operations on the material and the scene.

4. The method of claim 3, wherein, The initial instruction sequence is loaded by the instruction parser, and corresponding execution plugins are called according to the instruction type for sequential execution, specifically including: When the digital human broadcast instruction is executed, the digital human instruction plugin is called, the digital human module sends the broadcast text to the digital human platform, receives the audio and video stream data returned by the callback of the digital human platform, and submits the audio and video stream data to the video rendering engine; the broadcast text string is distributed to the scene editor, and the text is drawn on the transparent layer and superimposed on the synthesis picture through the scene editor; When the material playing instruction is executed, the material playing instruction plugin is called, the material module queries the file address of the video, and the scene module delivers the file address to the video rendering engine, and the video instance ID is generated through the video rendering engine; When the PPT page turning instruction is executed, the PPT page turning instruction plugin sends the page turning control parameters to the PPT projection window according to the window message mechanism.

5. The method of claim 1, wherein, According to the instruction execution result, the picture and the audio are rendered and synthesized, specifically including: According to the instruction execution result and the preset rendering parameter, the original video stream, the audio stream or the image data are extracted for leaf node rendering; Multiple video streams or image data are superimposed for picture branch node rendering, and the audio stream is mixed; The scene layout parameter, the subtitle layer and the transparent decoration layer are integrated for root node rendering; The picture and the audio are synthesized according to the Z-Order layering order.

6. The method of claim 1, wherein, Before inserting the new video operation instruction into the current execution position of the instruction sequence, the method further includes: Judging the state of the current execution plugin, if the current execution plugin is in a blocking waiting state, the new video operation instruction is immediately interrupted and executed; If the current execution plugin is in a running state, a pause flag is marked and the new video operation instruction is inserted when the current video operation instruction is executed.

7. A dynamic video synthesis device based on instruction sequences, characterized in that, The device is applied to the dynamic video synthesis method based on the instruction sequence in any one of claims 1-6, and is used for AI-assisted teaching, conferences and unmanned online interactive classrooms, and is realized through an instruction editing system, the instruction editing system includes an instruction module, an AI auxiliary subsystem, a digital human module, a material module, a scene module and a video rendering engine, the instruction module includes an instruction editor and an instruction parser; the device includes: An initial instruction sequence generation module is configured to obtain a prompt word input by a user, and generate an initial instruction sequence for video synthesis based on the prompt word through an AI assistance subsystem and an instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in. An initial instruction sequence execution module is configured to load the initial instruction sequence through an instruction parser, and sequentially execute the corresponding execution plug-in according to the instruction type of the video operation instruction in the initial instruction sequence. A dynamic instruction insertion module is configured to, during execution of the initial instruction sequence, if an input voice of the user is received, identify the intention of the input voice in real time through the AI assistance subsystem, generate a reply text and convert it into a digital person broadcasting instruction if the intention of the input voice is identified as a question, and convert it into a new video operation instruction if the intention of the input voice is identified as a control command, and insert the new video operation instruction into the current execution position of the instruction sequence. After the new inserted video operation instruction is executed through the instruction parser, subsequent instruction sequences are continued to be executed. A synthesis module is configured to render and synthesize pictures and audio according to the instruction execution result.

8. An instruction editing system, characterized by, The instruction editing system applied to the dynamic video synthesis method based on the instruction sequence according to any one of claims 1-6 comprises an instruction module, an AI assistance subsystem, a digital person module, a material module, a scene module, a video rendering engine and a device IO. The instruction module comprises an instruction editor and an instruction parser. The instruction editor is configured to add, delete and modify the video operation instruction according to the user demand, so as to adjust the operation of the video operation instruction. The instruction parser is configured to judge the type of the video operation instruction, query the instruction plug-in list according to the type judgment result, and call the execution plug-in matched with the video operation instruction for execution. The instruction plug-in list is obtained by dynamically registering the execution plug-in corresponding to each instruction type in the instruction parser in advance. The AI assistance subsystem is configured to assist in generating a broadcasting instruction sequence according to the prompt word input by the user when a new instruction list is created, identify the voice input of the user during the execution of the instruction, execute the instruction if the instruction is identified, give an answer if it is a question, give a feedback if it is a user answer, and generate an answer conforming to the style of a specific role in combination with an expert knowledge base and a language large model when answering the question of the user. The digital person module is configured to connect a remote digital person platform, receive a broadcasting text and obtain a synchronous video stream and an audio stream. The material module is configured to store multiple types of materials and provide a unified interface for calling. The scene module comprises a scene editor, and is configured to store and manage scene layout parameters and material arrangement rules. The video rendering engine is configured to render and output a synthesized picture and a synthesized audio in a tree structure based on preset video synthesis rules, scene layout parameters and material arrangement rules. The device IO includes a device input module and a device output module, the device input module is used for capturing user voice or user image, and the device output module is used for outputting presentation or data interaction.

Citation Information

Patent Citations

  • Digital human live broadcast interaction method and system based on large model

    CN119967197A

  • MOOC video method and device based on AIGC digital human

    CN120378707A