Dynamic video synthesis method and device based on instruction sequence and instruction editing system

Through a dynamic video synthesis method based on instruction sequences, AI is used to generate the initial instruction sequence and recognize user input in real time, which solves the problem of insufficient flexibility in traditional video synthesis systems and realizes flexible customization and real-time interaction of video generation. It is suitable for AI-assisted teaching, conferences and online interactive classrooms.

CN120676223AActive Publication Date: 2025-09-19PEKING UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511191170.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-19
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

In existing technologies, video synthesis relies on fixed templates and cannot dynamically respond to complex scenes, resulting in low video generation efficiency, lack of flexibility and real-time interactive capabilities, especially in live broadcasts and online teaching, which makes it difficult to meet personalized needs.

Method used

A dynamic video synthesis method based on instruction sequences is adopted, with AI-assisted generation of the initial instruction sequence, real-time recognition of user input and dynamic insertion of video operation instructions. Combined with the digital human module, material module and scene module, flexible customization of video content and real-time interaction are achieved.

Benefits of technology

It realizes flexible customization and real-time interaction of the video generation process, improves the efficiency of video generation, is suitable for AI-assisted teaching, conferences and unattended online interactive classrooms, and enhances the interactive display capabilities of images and sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676223A_ABST
    Figure CN120676223A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic video synthesis method and device based on an instruction sequence and an instruction editing system, which are applied to the technical field of video real-time generation, are used for AI auxiliary teaching, conferences and unattended online interactive classes, and are used for generating an initial instruction sequence through an AI auxiliary subsystem and an instruction editor based on prompt words input by a user; calling execution plug-ins for sequential execution through an instruction parser according to the instruction types; and in the execution process of the initial instruction sequence, the AI auxiliary subsystem identifies the intention of the input voice in real time, converts the input voice into a digital human broadcast instruction or a new video operation instruction according to an identification result, inserts the instruction into a current execution bit for execution, and renders and synthesizes a picture and an audio according to an instruction execution result. According to the method, dynamic video synthesis is realized by issuing the instruction sequence under the assistance of the AI, and real-time adjustment of video picture content is realized by dynamically inserting the instruction, so that the requirements of video customization synthesis and human-computer interaction are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of real-time video generation, and in particular to a dynamic video synthesis method, device and instruction editing system based on instruction sequences. Background Art

[0002] Video production is a complex process, especially when it involves live action. Equipment, scenes, personnel, and post-processing all require significant effort. This has led to the development of a streamlined, automated, and efficient video production method. This method collects footage, sorts and matches it, and then applies a template to automatically synthesize the video. However, this template-based video generation method lacks flexibility in the organization and layout of footage. When video customization requires extensive detail, frequent scene switching, or dynamic adjustments to video content as the image signal changes, such as in classroom interactions, the practicality of the template-based approach decreases, and the number of post-production steps required increases significantly.

[0003] For example, in the areas of live streaming and online teaching, traditional live streaming relies on pre-prepared content and scripts, making it difficult to adjust content based on real-time viewer feedback. Interaction between viewers and hosts is typically limited to text chat, lacking more diverse interactive forms. Traditional live streaming content is relatively simple, relying primarily on the host's explanations and demonstrations, lacking a diverse visual and auditory experience. In high-concurrency scenarios, it is prone to issues such as lag and delay, which impact the viewer experience. Traditional online teaching relies primarily on one-way teacher explanations, resulting in low student engagement and a lack of real-time interaction. Once the course content is determined, it is difficult to adjust it to meet students' actual needs, resulting in a lack of flexibility. Traditional online teaching requires teachers to spend a considerable amount of time preparing teaching content, including recording videos and creating PowerPoint presentations, which increases the teacher's workload. Furthermore, the relatively fixed content makes it difficult to meet the personalized learning needs of different students. Summary of the Invention

[0004] The present invention provides a dynamic video synthesis method, device and instruction editing system based on instruction sequences, which are used to solve the defects of the existing technology that video synthesis relies on fixed templates, cannot dynamically respond to complex scenes and is inefficient, and realizes flexible customization of the video generation process.

[0005] The present invention provides a dynamic video synthesis method based on an instruction sequence, which is used for AI-assisted teaching, conferences, and unattended online interactive classrooms. The method is implemented through an instruction editing system. The instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes: Obtaining a prompt word input by the user, and based on the prompt word, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor, wherein the initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; The initial instruction sequence is loaded through an instruction parser, and corresponding execution plug-ins are called according to instruction types of video operation instructions in the initial instruction sequence to perform execution in sequence; During the execution of the initial instruction sequence, if a user's input voice is received, the AI ​​auxiliary subsystem will perform real-time recognition of the intention of the input voice. If the intention of the input voice is recognized as a question, an answer text will be generated and converted into a digital human broadcast instruction; if the intention of the input voice is recognized as a control command, it will be converted into a new video operation instruction and inserted into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence will continue to be executed; Render and synthesize images and audio based on the command execution results.

[0006] According to the dynamic video synthesis method based on instruction sequence provided by the present invention, the video operation instructions include at least digital human broadcast instructions, material playback instructions, scene switching instructions, PPT page turning instructions, waiting response instructions and recording / live broadcast instructions. The digital human broadcast instruction is configured with a broadcast text and associated digital human parameters; The material playback instruction is configured with a target material path and material configuration parameters; The scene switching instruction is configured with scene layout parameters; The PPT page turning instruction is configured with page turning control parameters; The wait-for-response instruction is configured to pause at a specified knowledge point and listen for user questions; The recording / live broadcast instruction is configured to control the start and end of recording or push traffic switch.

[0007] According to the method for dynamic video synthesis based on instruction sequences provided by the present invention, the method generates an initial instruction sequence for video synthesis based on the prompt word through the AI ​​auxiliary subsystem and the instruction editor, specifically including: The AI ​​auxiliary subsystem calls the AI ​​large model to parse the prompt word into multiple text paragraphs; Each text paragraph is converted into a digital human broadcast instruction through a command editor; based on the converted digital human broadcast instruction and preset editing operation instructions, an initial instruction sequence for video synthesis is generated, wherein the editing operation instructions include: adding instructions, deleting instructions, modifying instructions, and inserting other video operation instructions between the digital human broadcast instructions; If the other inserted video operation instructions are material playback instructions or scene switching instructions, the user's configuration operations on the materials and scenes are received through the instruction editor.

[0008] According to the dynamic video synthesis method based on instruction sequence provided by the present invention, the initial instruction sequence is loaded by an instruction parser, and the corresponding execution plug-in is called according to the instruction type for sequential execution, specifically including: When the Digihuman broadcast command is executed, the Digihuman command plug-in is called, the Digihuman module sends the broadcast text to the Digihuman platform, receives the audio and video stream data called back by the Digihuman platform and submits the audio and video stream data to the video rendering engine; the broadcast text string is distributed to the scene editor, and the scene editor draws the text on the transparent layer and overlays it on the composite picture; When the material playback instruction is executed, the material playback instruction plug-in is used to call the material module to query the file address of the video, and the scene module is called to pass the file address to the video rendering engine, and the video instance ID is generated by the video rendering engine; When the PPT page turning command is executed, the page turning control parameters are sent to the PPT projection window according to the window message mechanism through the PPT page turning command plug-in.

[0009] According to the dynamic video synthesis method based on the instruction sequence provided by the present invention, the rendering and synthesis of the picture and audio according to the instruction execution result specifically includes: Extract the original video stream, audio stream or image data and render the leaf nodes according to the command execution results and preset rendering parameters; Overlay multiple video streams or image data to render branches and nodes of the screen, and mix the audio streams; Integrate scene layout parameters, subtitle layer and transparent decoration layer for root node rendering; Output the composite image and composite audio in Z-Order stacking order.

[0010] According to the method for dynamic video synthesis based on an instruction sequence provided by the present invention, before inserting the new video operation instruction into the current execution position of the instruction sequence, the method further includes: Determine the status of the currently executing plug-in. If the currently executing plug-in is in a blocked waiting state, immediately interrupt and execute a new video operation instruction. If the currently executed plug-in is in a running state, an interrupt flag is marked and the new video operation instruction is inserted when the execution of the current video operation instruction ends.

[0011] According to the method for dynamic video synthesis based on an instruction sequence provided by the present invention, during the execution of the initial instruction sequence, the method further includes: Monitor the execution status of video operation instructions in real time. When the execution status is abnormal, the AI ​​auxiliary subsystem calls the AI ​​large model to analyze the abnormal type and context, generate abnormal feedback instructions and drive the digital human to broadcast. The abnormal execution status includes at least one of instruction execution error, material loading failure, and rendering engine timeout.

[0012] According to the dynamic video synthesis method based on instruction sequences provided by the present invention, when used for AI-assisted teaching and meetings, the method further includes: Generate digital human broadcast instructions for each page of PPT commentary; Insert PPT page turning instructions and scene switching instructions between adjacent digital human broadcast instructions; Insert wait-for-response instructions at designated knowledge points to listen for students' questions; When the method is used in an unattended online interactive classroom, it also includes: Pre-embedding a conditional jump instruction in the initial instruction sequence; When the waiting time for the response instruction times out, it will automatically jump to the subsequent teaching content; When a student asks a question, the digital human dynamically inserts instructions to play the answer.

[0013] The present invention also provides a dynamic video synthesis device based on instruction sequences, which is used for AI-assisted teaching, conferences, and unattended online interactive classrooms. The device is implemented through an instruction editing system. The instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser. The device includes: An initial instruction sequence generation module is used to obtain a prompt word input by the user and, based on the prompt word, generate an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; An initial instruction sequence execution module, configured to load the initial instruction sequence through an instruction parser, and call corresponding execution plug-ins according to instruction types of video operation instructions in the initial instruction sequence for sequential execution; A dynamic instruction insertion module is configured to, during the execution of the initial instruction sequence, if a user's voice input is received, identify the intent of the input voice in real time through the AI ​​auxiliary subsystem; if the input voice is identified as a question, generate an answer text and convert it into a digital human broadcast instruction; if the input voice is identified as a control command, convert it into a new video operation instruction and insert the new video operation instruction into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, continue to execute the subsequent instruction sequence; The synthesis module is used to render and synthesize images and audio according to the instruction execution results.

[0014] The present invention also provides an instruction editing system, including an instruction module, an AI auxiliary subsystem, a digital human module, a material module, a scene module, a video rendering engine and a device IO. The instruction module includes an instruction editor and an instruction parser. The command editor is used to add, delete and modify video operation commands according to user needs to adjust the operation of the video operation commands; The instruction parser is used to determine the type of the video operation instruction, query the instruction plug-in list according to the type determination result, and call the execution plug-in matching the video operation instruction for execution, wherein the instruction plug-in list is obtained by dynamically registering the execution plug-in corresponding to each instruction type in the instruction parser in advance; The AI-assisted subsystem is used to assist in generating a broadcast command sequence based on the prompt words entered by the user when creating a new command list; during the command execution process, it recognizes the user's voice input and executes the command if it is recognized; if it is a question, it gives a corresponding answer; if the user answers, it gives feedback; when answering user questions, it combines the expert knowledge base and the language model to generate answers that conform to the style of the specific character; The digital human module is used to connect to the remote digital human platform, receive the broadcast text and obtain the synchronized video and audio streams; The material module is used to store various types of materials and provide a unified interface call; The scene module includes a scene editor, and the scene module is used to store and manage scene layout parameters and material arrangement rules; The video rendering engine is used to render and output the synthesized images and synthesized audio in a hierarchical tree structure based on preset video synthesis rules, scene layout parameters and material arrangement rules; The device 10 includes a device input module and a device output module. The device input module is used to capture user voice or user image, and the device output module is used for output presentation or data interaction.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for dynamic video synthesis based on an instruction sequence as described above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements any of the above-mentioned dynamic video synthesis methods based on instruction sequences.

[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned dynamic video synthesis methods based on instruction sequences.

[0018] The present invention provides a method, device and instruction editing system for dynamic video synthesis based on instruction sequence, which are used for AI-assisted teaching, conferences and unattended online interactive classrooms. The method is implemented through the instruction editing system, which includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. By obtaining the prompt word input by the user, based on the prompt word, the initial instruction sequence for video synthesis is generated by the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; through the instruction parser The initial instruction sequence is loaded, and the corresponding execution plug-in is called according to the instruction type of the video operation instruction in the initial instruction sequence for sequential execution; during the execution of the initial instruction sequence, if the user's input voice is received, the intention of the input voice is recognized in real time by the AI ​​auxiliary subsystem. If the intention of the input voice is recognized as a question, an answer text is generated and converted into a digital human broadcast instruction; if the intention of the input voice is recognized as a control command, it is converted into a new video operation instruction, and the new video operation instruction is inserted into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence is continued to be executed; according to the instruction execution result, the picture and audio are rendered and synthesized. The present invention realizes dynamic video synthesis by issuing an instruction sequence; through the dynamic instruction insertion mechanism, the picture content is adjusted in real time with the state change, and the video customization synthesis and human-computer interaction requirements are met through the flexibility of the instruction; and with the assistance of AI, the generation, recognition and insertion of instructions are realized, which greatly improves the efficiency of video generation, is conducive to the interactive display of images and sounds, and is suitable for more application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 It is a flow chart of the dynamic video synthesis method based on instruction sequence provided by the present invention.

[0021] Figure 2It is a structural diagram of an instruction editing system provided by the present invention.

[0022] Figure 3 This is an interface display of the dynamic video synthesis method based on instruction sequence provided by the present invention.

[0023] Figure 4 A schematic structural diagram of a dynamic video synthesis device based on instruction sequences provided by the present invention.

[0024] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0026] The present invention is described in detail below with reference to the accompanying drawings. The specific operating methods in the method embodiments can also be applied to device embodiments or system embodiments. In the description of the present invention, unless otherwise specified, "at least one" includes one or more. "Multiple" refers to two or more. For example, at least one of A, B, and C includes: A exists alone, B exists alone, A and B exist at the same time, A and C exist at the same time, B and C exist at the same time, and A, B, and C exist at the same time. In the present invention, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0027] The present invention will be described in detail below with reference to specific embodiments.

[0028] In some specific embodiments of the present invention, Figure 1 As shown, this solution provides a dynamic video synthesis method based on instruction sequences, which is used for AI-assisted teaching, conferences, and unattended online interactive classrooms. It is implemented through an instruction editing system. The instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes: Step 100: Obtain a prompt word input by the user, and based on the prompt word, generate an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; Step 200: Load the initial instruction sequence through an instruction parser, and call corresponding execution plug-ins according to the instruction type of the video operation instruction in the initial instruction sequence to execute them in sequence; Step 300: During the execution of the initial instruction sequence, if a user's voice input is received, the AI ​​auxiliary subsystem performs real-time recognition of the intent of the input voice. If the input voice is recognized as a question, an answer text is generated and converted into a digital human broadcast instruction. If the input voice is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence continues to be executed. Step 400: Render and synthesize the image and audio according to the instruction execution result.

[0029] It should be noted that existing video generation solutions rely on fixed templates and cannot dynamically respond to complex scenarios. They can only adjust the material layout and timeline manually, which is cumbersome and inefficient to operate. In addition, during the video generation process, the content cannot be modified instantly according to user voice commands (such as interrupting the teaching process for questions and answers, etc.), so it cannot meet user customization and real-time requirements.

[0030] Therefore, the present invention provides a dynamic video synthesis method based on an instruction sequence, which is used for AI-assisted teaching, conferences and unattended online interactive classroom scenarios, and is implemented through an instruction editing system. The instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. By obtaining the user's prompt words, an initial instruction sequence is generated based on the prompt words to achieve rapid video script construction, and the user's natural language requirements are converted into executable instructions using an AI large model to lower the operating threshold. The instruction type is matched with an independent plug-in to provide modular expansion capabilities. A plug-in mechanism is used to support various instructions such as digital human broadcasting, PPT page turning, etc., thereby enhancing system flexibility. Real-time interactive control is achieved through voice intent recognition and dynamic insertion. Questions or commands are identified by AI, and instructions are dynamically inserted to solve the defect that traditional templates cannot be dynamically adjusted. By layering the synthesis of images and audio, efficient integration of multi-source data is guaranteed, and the synthesis logic of materials, digital humans, and subtitles is processed through cascade rendering.

[0031] In some possible embodiments of the present invention, Figure 2 As shown, a command editing system is provided, which includes a command module 1, an AI-assisted subsystem 2, a material module 3, a scene module or scene editor 4, a digital human module 5 and a video rendering engine 6. The command module 1 includes a command editor 11 and a command parser 12. The command editor 11 is used to add, delete and modify video operation commands according to user needs to adjust the operation of the video operation commands; the command parser is used to judge the type of the video operation command, query the registered command plug-in list according to the type judgment result, and call the command plug-in matching the video operation command for execution; the AI-assisted subsystem is used to assist in generating a broadcast command sequence according to the prompt word entered by the user when creating a new command list; during the command execution process, the user voice input is recognized, and if the command is recognized, the command is executed accordingly; if it is a question, an answer is given accordingly; if the user answers, feedback is given; when answering user questions, the expert knowledge base and the language model are combined to generate answers that conform to the style of a specific role. The material module 3 is used to store various types of materials and provide a unified interface call; the scene module 4 includes a scene editor for storing and managing scene layout parameters and material arrangement rules; the digital human module 5 is used to connect to the remote digital human platform, receive broadcast text and obtain synchronized video streams and audio streams; the video rendering engine 6 is used to render and output composite images and composite audio in a hierarchical tree structure based on preset video synthesis rules, scene layout parameters and material arrangement rules; the device IO includes a device input module and a device output module, the device input module is used to capture user voice or user images, and the device output module is used for output presentation or data interaction.

[0032] Specifically, the device input module of the device IO is used for input control, receiving user voice commands such as capturing voice through a microphone, triggering AI real-time recognition and dynamic command insertion. The device input module can also be used to capture user images, such as user gestures, etc. The device output module of the device IO is used for output presentation and data interaction. Output presentation includes displaying a synthesized image, playing synthesized audio, etc. The display outputs a rendered image, and the speaker plays digital human voice / mixed audio, etc.; data interaction includes supporting live streaming and external signal access, specifically network streaming to transmit camera signals, NDI / live signals, screen projection and other external inputs, which are input to the material module by the device IO. It can also be used for control interfaces, such as network streaming to transmit remote control commands.

[0033] In some possible implementations of the present invention, the method for dynamic video synthesis based on an instruction sequence provided by the present invention further includes: Monitor the execution status of video operation instructions in real time. When the execution status is abnormal, the AI ​​auxiliary subsystem calls the AI ​​large model to analyze the abnormal type and context, generate abnormal feedback instructions and drive the digital human to broadcast. The abnormal execution status includes at least one of instruction execution error, material loading failure, and rendering engine timeout.

[0034] Specifically, this embodiment provides an implementation method for handling execution status exceptions. It uses an AI-assisted subsystem to assist in execution status monitoring and exception feedback, improves system robustness (extended functions), and uses AI to diagnose rendering timeouts, loading failures and other problems, driving the digital human to broadcast exception prompts to avoid process interruptions.

[0035] In some possible implementations of the present invention, the video operation instructions include at least digital human broadcast instructions, material playback instructions, scene switching instructions, PPT page turning instructions, waiting for response instructions and recording / live broadcast instructions. The digital human broadcast instruction is configured with a broadcast text and associated digital human parameters; The material playback instruction is configured with a target material path and material configuration parameters; The scene switching instruction is configured with scene layout parameters; The PPT page turning instruction is configured with page turning control parameters; The wait-for-response instruction is configured to pause at a specified knowledge point and listen for user questions; The recording / live broadcast instruction is configured to control the start and end of recording or push traffic switch.

[0036] Specifically, this embodiment provides an implementation method for video operation instructions. The above examples for video operation instructions only list some commonly used instruction types. In actual applications, other specific settings can be made according to different user needs and are not limited to the above examples listed in this embodiment. It is worth noting that in addition to obtaining user prompts and generating video operation instructions based on the prompts, video operation instructions such as commentary instructions, signal switching instructions, scene switching instructions, and voice question and answer instructions can also be set before the video synthesis process is started or during operation.

[0037] In some possible implementations of the present invention, based on the prompt word, Based on the prompt words, an initial instruction sequence for video synthesis is generated by the AI-assisted subsystem and the instruction editor, specifically including: The AI ​​auxiliary subsystem calls the AI ​​large model to parse the prompt word into multiple text paragraphs; Each text paragraph is converted into a digital human broadcast instruction through a command editor; based on the converted digital human broadcast instruction and preset editing operation instructions, an initial instruction sequence for video synthesis is generated, wherein the editing operation instructions include: adding instructions, deleting instructions, modifying instructions, and inserting other video operation instructions between the digital human broadcast instructions; If the other inserted video operation instructions are material playback instructions or scene switching instructions, the user's configuration operations on the materials and scenes are received through the instruction editor.

[0038] Specifically, this embodiment provides an implementation method for generating an initial instruction sequence by calling the AI ​​language model to parse the prompt words input by the user and optimize the instruction sequence. For example, the prompt word "Explain the Four Great Inventions" is broken down into segmented broadcast instructions and bound to the PPT page turning dependency.

[0039] Specifically, first, the user inputs a prompt word, and the input form is a natural language description of the video generation requirements, for example, such as "divide into three paragraphs to explain the four great inventions of ancient China", "use prompt words to provide a commentary for each news video", etc.

[0040] Furthermore, the prompt word is parsed into multiple text paragraphs through an AI language model (such as a GPT model). For example, the prompt word "Explain the Four Great Inventions, each invention for 1 minute, and summarize at the end" is parsed into text paragraphs such as: the invention of gunpowder; The invention of papermaking; the invention of printing; The compass was invented; 1 minute for each invention; Summarize the significance of the Four Great Inventions.

[0041] The command editor converts each text paragraph into a digital human-reading command: Broadcast instruction 1: Explanation of the invention of gunpowder (1 minute); Broadcast instruction 2: Explanation of the invention of papermaking (1 minute); Broadcast instruction 3: Explanation of the invention of printing (1 minute); Broadcast instruction 4: Explanation of the invention of the compass (1 minute); Broadcast instruction 5: Summarize the significance of the four great inventions.

[0042] Each operation corresponds to a type of instruction (digital human broadcast, PPT page turning, etc.).

[0043] Furthermore, the system dynamically binds scene parameters, automatically associating them with scene context parameters: the current scene layout (e.g., the position of the PowerPoint presentation or the position of the digital human in a lecture scene) or material attributes (e.g., material location, PowerPoint page number, video clip path). For example, upon receiving the prompt "Add commentary to each PowerPoint page," the system automatically binds the PowerPoint page number to the presentation instruction, ensuring that the commentary is synchronized with the page.

[0044] On this basis, non-broadcasting instructions such as scene switching instructions, PPT page turning instructions, and wait response instructions are further supplemented, and necessary dependent instructions such as delay waiting between instructions are automatically supplemented to ensure process continuity. For example, if the digital human is giving a broadcast instruction, the inserted instruction is a PPT page turning instruction (if associated with PPT) to ensure that the PPT page turns synchronously during the broadcast. If it is a continuous video playback instruction, the inserted instruction is a scene switching instruction to switch the layout of the material. If it is a question and answer session, the inserted instruction is a wait response instruction (timeout jump) to monitor user voice input. For example, "Insert a wait response instruction at the knowledge point to monitor students' questions", and automatically insert a wait instruction after explaining the key knowledge points.

[0045] Finally, an editable command sequence is generated, output as a command list editor that visualizes the generated command sequence. Examples of command sequences include: the digital human broadcast command: "Explanation of the invention of gunpowder" (bound to PPT page 5); the PPT page turn command: "Jump to page 6"; the digital human broadcast command: "Explanation of the invention of papermaking" (bound to PPT page 6); and the wait for response command (timeout 30 seconds).

[0046] In addition, users can manually add, delete or modify instructions such as adjusting the broadcast text and inserting special effect instructions.

[0047] It's worth noting that traditional templates require pre-defining all the elements. However, with the aforementioned configuration in this embodiment of the present invention, command sequences are dynamically generated via AI and can adapt to any prompt word requirement (e.g., "Four Great Inventions" to "Industrial Revolution"). This allows for flexible customization of dynamic videos. Furthermore, command insertion (such as the wait-for-response command) directly supports real-time classroom interaction (e.g., the requirement for "dynamic adjustment of image signals"), enabling real-time dynamic adjustment of the video synthesis process. Furthermore, users do not need to understand the technical details of video synthesis; natural language descriptions can generate professional command sequences, lowering the barrier to dynamic video customization and synthesis.

[0048] For example, if the prompt is "Create a 20-page PPT with a narration, use a teacher scene for each page, and switch to a full-screen PPT scene every 5 pages," the AI ​​generates the following command sequence: Loop 20 times: Digital human broadcast command (bound to the current PPT page) + PPT page turning command; Insert every 5 times: scene switching instruction (teacher scene to full-screen PPT scene).

[0049] This process converts user natural language requirements into executable and editable instruction sequences through AI language large model parsing, instruction editor conversion, scene parameter configuration, and instruction insertion. It fundamentally solves the problems of insufficient flexibility and lack of dynamic interaction in traditional video template systems, while significantly lowering the operational threshold.

[0050] In some possible implementations of the present invention, the initial instruction sequence is loaded by an instruction parser, and corresponding execution plug-ins are called according to the instruction type for sequential execution, specifically including: When the Digihuman broadcast command is executed, the Digihuman command plug-in is called, the Digihuman module sends the broadcast text to the Digihuman platform, receives the audio and video stream data called back by the Digihuman platform and submits the audio and video stream data to the video rendering engine; the broadcast text string is distributed to the scene editor, and the scene editor draws the text on the transparent layer and overlays it on the composite picture; When the material playback instruction is executed, the material playback instruction plug-in is used to call the material module to query the file address of the video, and the scene module is called to pass the file address to the video rendering engine, and the video instance ID is generated by the video rendering engine; When the PPT page turning command is executed, the page turning control parameters are sent to the PPT projection window according to the window message mechanism through the PPT page turning command plug-in.

[0051] Specifically, this embodiment provides a method for invoking an execution plug-in to execute an initial instruction sequence. Through the digital human plug-in execution process, a digital human broadcast pipeline is implemented, bridging the entire chain from platform communication to audio and video stream reception and rendering synthesis. The video playback plug-in execution process enables dynamic on-screen playback of material, and the scene module coordinates the rendering engine to manage the video instance lifecycle.

[0052] In some possible implementations of the present invention, layered rendering and synthesizing images and audio according to the instruction execution result specifically includes: Extract the original video stream, audio stream or image data and render the leaf nodes according to the command execution results and preset rendering parameters; Overlay multiple video streams or image data to render branches and nodes of the screen, and mix the audio streams; Integrate scene layout parameters, subtitle layer and transparent decoration layer for root node rendering; Output the composite image and composite audio in Z-Order stacking order.

[0053] Specifically, this embodiment provides an implementation method for hierarchical rendering of a tree structure, which clarifies the workflow of the rendering engine through the rendering order from leaves to branches to root nodes, and controls the stacking priority through Z-Order (such as subtitles are always on the top layer), ensuring synthesis efficiency through hierarchical processing.

[0054] In a possible embodiment, the CPU algorithm or GPU computing framework can be automatically selected to build a rendering map based on the performance of the user's computer graphics card, and CPU / GPU acceleration can be switched according to the user's computer configuration to dynamically optimize resource usage; text can be drawn on a transparent layer through a scene module, and the drawn text can be superimposed on the composite picture.

[0055] In a possible implementation, layered rendering can enable parallel processing, such as asynchronous execution of leaf node decoding and branch synthesis. Adding new material types (such as projection signals) simply requires expanding the leaf node plugin. Transparency layers / subtitles are treated as independent branch nodes and can be dynamically modified, enabling flexible expansion.

[0056] In some possible implementations of the present invention, before inserting the new video operation instruction into the current execution position of the instruction sequence, the method further includes: Determine the status of the currently executing plug-in. If the currently executing plug-in is in a blocked waiting state, immediately interrupt and execute a new video operation instruction. If the currently executed plug-in is in a running state, an interrupt flag is marked and the new video operation instruction is inserted when the execution of the current video operation instruction ends.

[0057] Specifically, this embodiment provides an implementation method for an instruction insertion strategy, which ensures the real-time performance of dynamic insertion by blocking interrupts and running flags, distinguishes plug-in statuses to determine immediate interruption or sequential insertion, and avoids instruction conflicts.

[0058] In some possible implementations of the present invention, when the instruction sequence-based dynamic video synthesis method is used for AI-assisted teaching and conferencing, it also includes: Generate digital human broadcast instructions for each page of PPT commentary; Insert PPT page turning instructions and scene switching instructions between adjacent digital human broadcast instructions; Insert wait-for-response instructions at designated knowledge points to listen for students' questions.

[0059] Specifically, this embodiment provides an implementation method for teaching scenario applications, which uses AI-assisted teaching to achieve automated course recording and simulate the teacher's teaching process through a combination of instructions (such as commentary + PPT page turning + scene switching).

[0060] In some possible implementations of the present invention, when the instruction sequence-based dynamic video synthesis method is used in an unattended online interactive classroom, the method further includes: Pre-embedding a conditional jump instruction in the initial instruction sequence; When the waiting time for the response instruction times out, it will automatically jump to the subsequent teaching content; When a student asks a question, the digital human dynamically inserts instructions to play the answer.

[0061] Specifically, this embodiment provides another implementation method for teaching scenario applications, which supports interactive classrooms through unattended interactive classrooms, pre-embeds jump instructions to handle timeouts, and dynamically inserts question and answer response instructions to achieve fully automatic teaching.

[0062] In a possible embodiment, during the execution of the initial instruction sequence, user input is received in real time, and the user input includes voice input and text input. If the user input is voice input, the voice input is converted into text input and then passed to the AI ​​large model for intent recognition.

[0063] In a possible embodiment, user input includes not only voice and text input but also keystrokes. Before the AI ​​model recognizes user intent, initial recognition of user input is performed, enabling the system to more comprehensively process diverse forms of user input. When the initial recognition result is a command, it is dynamically inserted into the current execution position and executed by a registered plug-in. This improves the system's response speed and efficiency to explicit commands, paving the way for the AI ​​model to subsequently recognize intent for non-command input, and improves the overall logic of user input processing.

[0064] The dynamic video synthesis method based on instruction sequences provided by the embodiments of the present invention abstracts the video generation process through instruction sequences, utilizes plug-in execution and AI dynamic interaction to solve the flexibility problem of traditional template systems, and especially provides a feasible technical solution for the real-time interaction needs of educational scenarios (such as classroom question and answer).

[0065] In some possible embodiments of the present invention, see Figure 2 The dynamic video generation system provided by the embodiment of the present invention also includes a material module 3, a scene editor 4, a digital human module 5, a video rendering engine 6 and a device IO7.

[0066] In a possible embodiment, the instruction parser 11 supports plug-in extensions, and each instruction type corresponds to an independent plug-in, which is registered to the instruction parser 11 when the system starts.

[0067] Specifically, this embodiment provides an implementation of the instruction parser 11. The plug-in extension mechanism of the instruction parser 11 enhances the flexibility and scalability of the system. By providing a separate plug-in for each instruction type and registering it with the instruction parser at system startup, this embodiment enables the system to easily add new instruction types without requiring large-scale modifications to the core system. This provides technical support for system functionality expansion and ensures that the system can adapt to future technological developments and changes in user needs.

[0068] In some possible implementations of the present invention, the material module 3 supports plug-in extensions, and each type of material corresponds to an independent plug-in, which is registered to the material module 3 when the system is started.

[0069] Specifically, this embodiment provides an implementation of the material module and its plug-in extension mechanism, enhancing the system's flexibility and scalability. By providing independent plug-ins for each type of material and registering them with the material module upon system startup, this embodiment facilitates the addition of new material types, enriching material resources and providing technical support for the system's material management. This ensures the system can adapt to different material requirements and improves its versatility and adaptability.

[0070] In some possible implementations of the present invention, the material module 3 supports local and cloud storage, calls materials through a unified interface, and supports encryption and copyright protection of materials.

[0071] Specifically, this embodiment provides another implementation method of the material module 3. The material module 3 not only supports local storage, but also supports cloud storage, so that the system can flexibly process materials from different sources and forms to meet the needs of users under different network environments and resource conditions. The material module has designed a unified plug-in interface for each type, including functions such as opening, closing, and reading data. Calling materials through a unified interface ensures the compatibility and consistency of the system when processing materials of different types and sources, reducing development and maintenance costs. The material module needs to process various types of materials, which may involve copyright issues. Through the encryption and copyright protection functions of the materials, the security and legality of the materials during storage and use are ensured.

[0072] In a specific embodiment, see Figure 2 The video generation method provided by the present invention is Figure 2 The architecture is implemented, which includes a material module 3, a digital human module 5, a scene module such as a scene editor 4, an instruction module 1, a video rendering engine 6, an AI auxiliary subsystem 2, and a device IO 7, wherein the device IO 7 includes a device input module 71 and a device output module 72.

[0073] Material Module 3 is used to store the video, images, camera signals, live broadcast signals, NDI signals, Apple screencast signals, PPTs, windows, screens, and other types of materials used in video generation. To access these materials, the program has designed a unified plug-in interface for each type. These interfaces include opening, closing, reading data, and reading parameter information. Basic material parameters include width, height, frame rate, pixel format, and playback status.

[0074] Digital Human Module 5 is used to connect to remote digital human platforms, such as Alibaba's avatar platform. By pushing broadcast commands to the digital human, it can be driven to speak and perform gestures. Digital humans can essentially be considered a type of material.

[0075] Specifically, the basic connection process of the Digihuman module is as follows: Step 1: Open the user platform space through the Digital Human platform account and obtain the list of available Digital Human images; Step 2: The user selects a digital human image; Step 3: The program creates an instance on the Digital Human Platform using the Digital Human ID and establishes a path for transmitting image and sound signals; Step 4: During the video generation process, the program sends the broadcast text through the transmission channel; the digital human platform returns the calculated image and sound signals; after receiving them, the program sends them to the image engine for rendering and display.

[0076] Furthermore, the scene module or scene editor 4 is used to arrange materials, display subtitles, and decorate the background. The scene module, used to manage and edit scenes, includes two submodules: the scene editor and the scene container. Through the scene editor, you can select materials and set parameters such as their placement, size, cropping area, and transparency. A scene can better describe a specific situation than a single material, for example, showing a teaching scene from both the teacher's lecture and the students' listening angles.

[0077] Command module 1 refers to a set of commands related to video generation. Basic command types include broadcast commands for digital humans, commands for displaying materials on the screen, commands for video playback, commands for scene switching, commands for turning pages in PowerPoint presentations, commands for live streaming and recording, and commands for waiting. In addition to a list of display commands, the system also includes a command editor and generator.

[0078] Specifically, the generation and modification methods include: Step 1: Based on the user's prompt, a set of text broadcast instructions is generated, for example: "Divide the four great inventions of ancient China into several paragraphs." At this point, the AI ​​model is called to generate several paragraphs of text. Combined with these texts, the program generates the broadcast instructions for the digital human.

[0079] Step 2: The broadcast command is displayed in the command list. The command list is also an editor that can add, delete and modify commands to adjust the operation of the command to meet the user's more detailed needs.

[0080] Step three: When the instruction is being executed, the user can interrupt with words. The AI ​​will recognize what the user says. If it is recognized as a question, the AI ​​model will be used to generate the answer text, and then the broadcast instruction will be sent to the digital human. If it is recognized as a command, it will be converted into a message and passed to the interface for response.

[0081] Correspondingly, the instruction parser corresponds to instruction generation, which refers to the mechanism that parses and executes each instruction in the instruction list. To improve the scalability of instructions, each type of instruction is designed as a plug-in. When the program starts, it will be registered with the parser. When the parser receives an instruction, it will determine the instruction type and query the list of registered instruction plug-ins to find a matching instruction to execute. For example, if the parser receives a digital human play instruction, it will find the digital human plug-in from the registered plug-in list and then drive the digital human plug-in to complete the task.

[0082] Video Rendering Engine 6 refers to the underlying computing structure that synthesizes footage, digital humans, scenes, and sound. The basic logic behind the Video Rendering Engine is to sequentially decode a frame from each relevant source material within a specific time sequence. Based on placement, layering order, and special effects parameters, the engine then constructs a graph using CPU algorithms or uploads the data to a graphics card, applying a GPU computing framework. This graph is then rendered first, followed by branches and roots, ultimately forming a single image. Whether CPU or GPU algorithms are used depends on the user's computer configuration; if a high-performance graphics card is used, the program defaults to GPU algorithms. Sound is also synthesized by decoding sampled sound data and then using CPU synthesis algorithms.

[0083] AI-assisted subsystem 2 refers to the use of AI to assist users in creating instruction sequences. The specific functions of the AI-assisted subsystem include the following aspects: First, when creating a new command list, users can enter prompt words and AI will assist in generating a broadcast command sequence; Secondly, during the execution of commands, the AI ​​recognizes user voice input. If it recognizes a command, it executes it accordingly; if it is a question, it gives an answer; if the user answers, it provides feedback.

[0084] Third, when answering user questions, AI will combine the expert knowledge base and language model to give answers that are consistent with the style of a specific role.

[0085] Furthermore, the device IO mentioned in this embodiment refers to a series of input and output devices, including a device input module 71, such as a microphone and camera, and a device output module 72, whose output objects include video files, live broadcast signals, screens, or speakers. For example, the device IO may include: a display for displaying rendered images, speakers for supporting digital human broadcasts and video material playback, a network stream for transmitting audio and video data, and a voice input device, such as a microphone, for receiving user commands during command execution.

[0086] Based on the above embodiment, when video generation is started, the program will execute the following steps in sequence: Step 201: The video rendering engine opens the video encoder and waits for each frame to be input.

[0087] Step 202: The instruction parser obtains an instruction list and calls different instruction plug-ins to execute different instruction types.

[0088] In more detail, the basic implementation of the Digital Human plug-in and the video playback plug-in is as follows: Step 202-1, Digital Human Command Plug-in: In this step, first, the plug-in sends the broadcast text to the digital human platform (for example, Alibaba Cloud Avatar) through the network interface and waits; Based on this, the plug-in receives a platform callback, which includes a video stream, an audio stream, and a text string to be broadcast.

[0089] Furthermore, the plug-in calls the rendering engine interface to complete (1) inserting video stream data, the rendering engine combines the scene, synthesizes a frame, and sends it to the preview window, encoder, and live broadcast interface (Note: if it is turned on); (2) inserting audio stream, and after synthesis, it is sent to the speaker, encoder, and live broadcast interface.

[0090] Furthermore, Figure 3 This is the interface display of the dynamic video synthesis method based on instruction sequence provided by the present invention, such as Figure 3 As shown, the text string obtained by the plugin callback is returned to the interface through the instruction parser, and the interface distributes it to the instruction list as shown in Figure 3 The script list in the file is highlighted during broadcasting; at the same time, it is distributed to the scene editor, which draws the text on a transparent layer and feeds it into the rendering engine, which then overlays it into the synthesized picture. It is also used for preview, encoder, and live broadcast.

[0091] Finally, the plugin completes the task and sends a signal to the parser. Step 202-2, video instruction plug-in: In this step, first, the plug-in obtains the name of the incoming video file and calls the material module to query the file address of the video; Next, the plug-in passes the file address to the scene module, which then passes it to the rendering engine. The rendering engine returns the instance ID of the video, completing the video display. Third, the plug-in calls the rendering engine interface, passes in the video instance ID, starts playback, and waits; Fourth, the rendering engine decodes the video, combines the scene, synthesizes the picture and audio data, and distributes it to the preview, encoder, and live broadcast interface; Fifth, the plug-in waits until the rendering engine returns a signal that the video playback has ended; Sixth, the plug-in calls the scene module, the scene module calls the rendering engine, closes the video instance ID, clears the video on-screen information, and completes the video closing; Seventh, the plugin completes the task and sends a signal to the parser.

[0092] Step 203: The parser receives the plug-in task completion signal and then executes the next instruction; Step 204: If a command sent by the user through speech is recognized from the audio input IO, the command is inserted into the current position of the command list for timely execution; Step 205: The parser completes all instruction executions and returns; Step 206: The video rendering engine closes the file editor.

[0093] This embodiment classifies and defines instructions, and associates instructions with video generation; with the joint support of material container, AI digital human, scene container, scene editor, instruction parser, and video rendering engine, the parsing and execution of instructions are realized; with the support of instruction editor and language large model, the automatic generation and customized modification of instruction sequence are realized; with the support of device IO and language large model, the dynamic recognition, insertion and parsing execution of instructions are realized. Through the above-mentioned settings of this embodiment, compared with the existing technical solutions, the video generation process is described using instruction sequence, and the efficient generation of video is achieved by utilizing AI digital human, language large model and proprietary rendering platform. By describing the video in a more straightforward way, the difficulty of video editing and generation is significantly reduced; the video shooting link is reduced, the cost and complexity of video generation are reduced, and the flexibility and efficiency of video generation are improved; by simply modifying the instructions, the video generation process can be easily intervened and dynamically adjusted; by combining the instruction system and AI, interactive video generation is more easily achieved.

[0094] In the specific embodiment of the present invention, see Figure 2The dynamic video synthesis method and dynamic video synthesis system based on instruction sequence provided by the present invention are specifically implemented through the following modules, namely material module 3, scene module or scene editor 4, digital human module 5, instruction module 1, video rendering engine 6, AI auxiliary subsystem 2, and device IO7.

[0095] In practical applications, the present invention can be used to conveniently and efficiently test the following four types of application examples: Application example 1: host news broadcast; Application Example 2: Course Recording; Application Example 3: Large Screen Q&A; Application example 4: AI classroom.

[0096] For application example one: load all news clips into the material module and select a scene to display the digital human; in the instruction module, use prompt words to provide a commentary for each news video. The commentary will be transmitted to the digital human as a broadcast instruction, and then insert the instruction to play the corresponding news clip between each segment. In this way, a simple news broadcast instruction sequence is ready, and video generation can be started.

[0097] For application example 2: Open a PowerPoint presentation and play it. Two scenes were designed. The first scene featured the PowerPoint presentation on the left and the digital human on the right, simulating a teacher lecturing at a blackboard. The second scene displayed the entire PowerPoint presentation. With AI assistance, each page of the presentation was accompanied by a commentary, which was then transmitted to the digital human as a broadcast command. Each page was set to use either the first or second scene, and a "turn the PowerPoint page" command was inserted between the commentary. This created a simple course recording command sequence, allowing video generation to begin.

[0098] For application example three: design a vertical screen scene and place the digital human in it. In the instruction module, import the expert Q&A knowledge base and set a waiting Q&A instruction. In this way, a simple large-screen Q&A instruction is ready.

[0099] Application Example 4 is a combination of Application Example 2 and Application Example 3. The specific process includes: first, opening a PowerPoint presentation, establishing a teacher teaching scene, creating a PowerPoint presentation scene, and then establishing a scene with only the digital human. The instruction module imports the expert Q&A knowledge base. To facilitate teaching, a sequence of instructions is designed according to the method of the second example, and the instruction module is configured to accept Q&A. At this point, the instructions are started, and the digital human teaches normally without interruption. If a voice is heard, the device IO will recognize it. If a question is identified, it will first search the expert Q&A knowledge base, then generate a text answer using the large model, and the digital human will read it out. If no questions are asked and the timeout period is reached, the digital human will return to normal teaching.

[0100] The following is the basic process of executing the digital human broadcast command, video playback command, and PPT page turning command: A. The Digihuman broadcast command is executed through the Digihuman command plug-in, including the following steps: Step 1: The plug-in sends the broadcast text to the digital human platform (for example, Alibaba Cloud Avatar) through the network interface and waits; Step 2: The plug-in receives a platform callback, which includes a video stream, an audio stream, and a text string to be broadcast; Step 3: The plug-in calls the rendering engine interface, inserts the video stream data, and the rendering engine combines the scene to synthesize a frame, which is sent to the preview window, encoder, and live broadcast interface (note: if enabled). The plug-in also inserts the audio stream, and after synthesis, it is sent to the speaker, encoder, and live broadcast interface. Step 4: The text string obtained by the plugin callback is returned to the interface through the parser, and the interface distributes it to the instruction list for highlighting during broadcast; at the same time, it is distributed to the scene editor, and the scene editor draws the text on the transparent layer and enters it into the rendering engine. The rendering engine overlays it into the synthesized picture, which is also used for preview, encoder and live broadcast.

[0101] B. The video playback command is executed through the video playback command plug-in, which specifically includes the following steps: Step 1: The plug-in obtains the name of the incoming video file and calls the material module to query the file address of the video; Step 2: The plug-in passes the file address to the scene module, which then passes it to the rendering engine. The rendering engine returns the instance ID of the video, and the video is displayed on the screen. Step 3: The plug-in calls the rendering engine interface, passes in the video instance ID, starts playback, and waits; Step 4: The rendering engine decodes the video, combines the scene, synthesizes the picture and audio data, and distributes it to the preview (including speakers), encoder, and live broadcast interface; Step 5: The plug-in waits until the rendering engine returns a signal indicating the video has finished playing. Step 6. The plug-in calls the scene module, the scene module calls the rendering engine, closes the video instance ID, clears the video on-screen information, and completes the video closing.

[0102] C. The PPT page turning command is executed through the interface dialog box message response, which specifically includes the following steps: Step 1: Query the PPT presentation window handle; Step 2: Send page turning message parameters to the window.

[0103] The dynamic video synthesis method based on instruction sequences provided by the embodiment of the present invention associates instructions with video generation by classifying and defining instructions; with the joint support of material containers, AI digital humans, scene containers, scene editors, instruction parsers, and video rendering engines, it realizes the parsing and execution of instructions; with the support of instruction editors and language big models, it realizes the automatic generation and customized modification of instruction sequences; with the support of device IO and language big models, it realizes the dynamic recognition, insertion, parsing and execution of instructions.

[0104] The above-mentioned method provided by the embodiments of the present invention is essentially a technical solution for defining video synthesis. Compared with existing technical solutions, this solution uses an instruction sequence to describe the video generation process, utilizing AI assistance, digital humans, large language models, and a proprietary rendering platform to achieve efficient video synthesis. The technical effects achieved by this technical solution can be summarized as follows: On the one hand, the present invention significantly reduces the difficulty of video editing and generation by describing the video in a more straightforward way; On the other hand, the present invention can reduce the video shooting process, reduce the cost and complexity of video generation, and improve the flexibility and efficiency of video generation; Thirdly, the present invention can conveniently intervene in and dynamically adjust the video generation process by simply modifying the instructions; Fourthly, the present invention makes it easier to realize interactive video synthesis by combining the instruction system and AI.

[0105] In some specific embodiments of the present invention, Figure 4 As shown, this solution provides a dynamic video synthesis device based on instruction sequences, which is used for AI-assisted teaching, conferences, and unattended online interactive classrooms. It is implemented through an instruction editing system. The instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The instruction module includes an instruction editor and an instruction parser. The device includes: The initial instruction sequence generation module 41 obtains a prompt word input by the user and, based on the prompt word, generates an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; The initial instruction sequence execution module 42 loads the initial instruction sequence through the instruction parser and calls the corresponding execution plug-in according to the instruction type of the video operation instruction in the initial instruction sequence to execute it in sequence; The dynamic instruction insertion module 43, during the execution of the initial instruction sequence, if a user's input voice is received, the AI ​​auxiliary subsystem performs real-time recognition of the intent of the input voice. If the input voice is recognized as a question, an answer text is generated and converted into a digital human broadcast instruction. If the input voice is recognized as a control command, it is converted into a new video operation instruction and inserted into the current execution position of the instruction sequence. After the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence continues to be executed. The synthesis module 44 renders and synthesizes the image and audio according to the instruction execution result.

[0106] The implementation principle and beneficial effects of the dynamic video synthesis device based on instruction sequence provided in the embodiment of the present invention are similar to the implementation principle and beneficial effects of the dynamic video synthesis method based on instruction sequence shown in the above embodiment. Please refer to the implementation principle and beneficial effects of the dynamic video synthesis method based on instruction sequence shown in the above embodiment, and no further details will be given here.

[0107] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communications interface 520 and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute a dynamic video synthesis method based on an instruction sequence. The method is used for AI-assisted teaching, conferences and unattended online interactive classrooms, and is implemented through an instruction editing system. The instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor. The method includes: obtaining a prompt word input by the user, and based on the prompt word, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in. The invention relates to a method for realizing a video manipulation system, comprising: loading the initial instruction sequence through an instruction parser, and calling the corresponding execution plug-in to execute the instruction sequence in sequence according to the instruction type of the video operation instruction in the initial instruction sequence; during the execution of the initial instruction sequence, if the user's input voice is received, the AI ​​auxiliary subsystem recognizes the intention of the input voice in real time, and if the intention of the input voice is recognized as a question, generates an answer text and converts it into a digital human broadcast instruction; if the intention of the input voice is recognized as a control command, converts it into a new video operation instruction, and inserts the new video operation instruction into the current execution position of the instruction sequence; after executing the newly inserted video operation instruction through the instruction parser, the subsequent instruction sequence is continued to be executed; and according to the instruction execution result, the picture and audio are rendered and synthesized.

[0108] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the dynamic video synthesis method based on the instruction sequence provided by the above methods. The method is used for AI-assisted teaching, meetings and unattended online interactive classrooms, and is implemented through an instruction editing system. The instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module and a video rendering engine. The instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor; the method includes: obtaining a prompt word input by the user, and based on the prompt word, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor, the initial The instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; the initial instruction sequence is loaded through the instruction parser, and the corresponding execution plug-in is called in sequence according to the instruction type of the video operation instruction in the initial instruction sequence; during the execution of the initial instruction sequence, if the user's input voice is received, the AI ​​auxiliary subsystem recognizes the intention of the input voice in real time. If the intention of the input voice is recognized as a question, an answer text is generated and converted into a digital human broadcast instruction; if the intention of the input voice is recognized as a control command, it is converted into a new video operation instruction, and the new video operation instruction is inserted into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence continues to be executed; according to the instruction execution result, the picture and audio are rendered and synthesized.

[0110] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the dynamic video synthesis method based on an instruction sequence provided by the above-mentioned methods, which is used for AI-assisted teaching, meetings and unattended online interactive classrooms, and is implemented through an instruction editing system, wherein the instruction editing system includes an instruction module, an AI-assisted subsystem, a digital human module, a material module, a scene module and a video rendering engine, wherein the instruction module includes an instruction editor and an instruction parser, and the scene module includes a scene editor; the method includes: obtaining a prompt word input by the user, and based on the prompt word, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor, wherein the initial instruction sequence includes multiple types of video operations instructions, and each instruction type corresponds to an independent execution plug-in; the initial instruction sequence is loaded through the instruction parser, and the corresponding execution plug-in is called in sequence according to the instruction type of the video operation instruction in the initial instruction sequence; during the execution of the initial instruction sequence, if the user's input voice is received, the intention of the input voice is recognized in real time through the AI ​​auxiliary subsystem, and if the intention of the input voice is recognized as a question, an answer text is generated and converted into a digital human broadcast instruction; if the intention of the input voice is recognized as a control command, it is converted into a new video operation instruction, and the new video operation instruction is inserted into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence is continued to be executed; according to the instruction execution result, the picture and audio are rendered and synthesized.

[0111] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0112] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A dynamic video synthesis method based on instruction sequence, characterized in that: An online interactive classroom for AI-assisted teaching, conferencing, and unattended learning is implemented through a command editing system, which includes a command module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The command module includes a command editor and a command parser, and the scene module includes a scene editor. The method includes: Obtaining a prompt word input by the user, and based on the prompt word, generating an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor, wherein the initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; The initial instruction sequence is loaded through an instruction parser, and corresponding execution plug-ins are called according to instruction types of video operation instructions in the initial instruction sequence to perform execution in sequence; During the execution of the initial instruction sequence, if a user's input voice is received, the AI ​​auxiliary subsystem will perform real-time recognition of the intention of the input voice. If the intention of the input voice is recognized as a question, an answer text will be generated and converted into a digital human broadcast instruction; if the intention of the input voice is recognized as a control command, it will be converted into a new video operation instruction and inserted into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, the subsequent instruction sequence will continue to be executed; Render and synthesize images and audio based on the command execution results.

2. The method for dynamic video synthesis based on instruction sequence according to claim 1, characterized in that: The video operation instructions include at least digital human broadcast instructions, material playback instructions, scene switching instructions, PPT page turning instructions, waiting for response instructions and recording / live broadcast instructions. The digital human broadcast instruction is configured with a broadcast text and associated digital human parameters; The material playback instruction is configured with a target material path and material configuration parameters; The scene switching instruction is configured with scene layout parameters; The PPT page turning instruction is configured with page turning control parameters; The wait-for-response instruction is configured to pause at a specified knowledge point and listen for user questions; The recording / live broadcast instruction is configured to control the start and end of recording or push traffic switch.

3. The method for dynamic video synthesis based on instruction sequence according to claim 2, characterized in that: The initial instruction sequence for video synthesis is generated based on the prompt word by the AI ​​auxiliary subsystem and the instruction editor, specifically including: The AI ​​auxiliary subsystem calls the AI ​​large model to parse the prompt word into multiple text paragraphs; Each text paragraph is converted into a digital human broadcast instruction through a command editor; based on the converted digital human broadcast instruction and preset editing operation instructions, an initial instruction sequence for video synthesis is generated, wherein the editing operation instructions include: adding instructions, deleting instructions, modifying instructions, and inserting other video operation instructions between the digital human broadcast instructions; If the other inserted video operation instructions are material playback instructions or scene switching instructions, the user's configuration operations on the materials and scenes are received through the instruction editor.

4. The method for dynamic video synthesis based on instruction sequence according to claim 3, characterized in that: The initial instruction sequence is loaded through the instruction parser, and the corresponding execution plug-in is called according to the instruction type for sequential execution, specifically including: When the Digihuman broadcast command is executed, the Digihuman command plug-in is called, the Digihuman module sends the broadcast text to the Digihuman platform, receives the audio and video stream data called back by the Digihuman platform and submits the audio and video stream data to the video rendering engine; the broadcast text string is distributed to the scene editor, and the scene editor draws the text on the transparent layer and overlays it on the composite picture; When the material playback instruction is executed, the material playback instruction plug-in is used to call the material module to query the file address of the video, and the scene module is called to pass the file address to the video rendering engine, and the video instance ID is generated by the video rendering engine; When the PPT page turning instruction is executed, the page turning control parameters are sent to the PPT projection window according to the window message mechanism through the PPT page turning instruction plug-in.

5. The method for dynamic video synthesis based on instruction sequence according to claim 1, characterized in that: Rendering and synthesizing the image and audio according to the instruction execution result specifically includes: Extract the original video stream, audio stream or image data and render the leaf nodes according to the command execution results and preset rendering parameters; Overlay multiple video streams or image data to render branches and nodes of the screen, and mix the audio streams; Integrate scene layout parameters, subtitle layer and transparent decoration layer for root node rendering; Composite images and audio in Z-Order stacking order.

6. The method for dynamic video synthesis based on instruction sequence according to claim 1, characterized in that: Before inserting the new video operation instruction into the current execution position of the instruction sequence, the method further includes: Determine the status of the currently executing plug-in. If the currently executing plug-in is in a blocked waiting state, immediately interrupt and execute a new video operation instruction. If the currently executed plug-in is in a running state, an interrupt flag is marked and the new video operation instruction is inserted when the execution of the current video operation instruction ends.

7. The method for dynamic video synthesis based on instruction sequence according to any one of claims 1 to 6, characterized in that: During execution of the initial instruction sequence, the method further includes: Monitor the execution status of video operation instructions in real time. When the execution status is abnormal, the AI ​​auxiliary subsystem calls the AI ​​large model to analyze the abnormal type and context, generate abnormal feedback instructions and drive the digital human to broadcast. The abnormal execution status includes at least one of instruction execution error, material loading failure, and rendering engine timeout.

8. The method for dynamic video synthesis based on instruction sequence according to claim 7, characterized in that: When the method is used for AI-assisted teaching and meetings, it also includes: Generate digital human broadcast instructions for each page of PPT commentary; Insert PPT page turning instructions and scene switching instructions between adjacent digital human broadcast instructions; Insert wait-for-response instructions at designated knowledge points to listen for students' questions; When the method is used in an unattended online interactive classroom, it also includes: Pre-embedding a conditional jump instruction in the initial instruction sequence; When the waiting time for the response instruction times out, it will automatically jump to the subsequent teaching content; When a student asks a question, the digital human dynamically inserts instructions to play the answer.

9. A dynamic video synthesis device based on an instruction sequence, characterized in that: The system is used for AI-assisted teaching, conferencing, and unattended online interactive classrooms. The system is implemented through a command editing system, which includes a command module, an AI-assisted subsystem, a digital human module, a material module, a scene module, and a video rendering engine. The command module includes a command editor and a command parser. The device includes: An initial instruction sequence generation module is used to obtain a prompt word input by the user and, based on the prompt word, generate an initial instruction sequence for video synthesis through the AI-assisted subsystem and the instruction editor. The initial instruction sequence includes multiple types of video operation instructions, and each instruction type corresponds to an independent execution plug-in; An initial instruction sequence execution module, configured to load the initial instruction sequence through an instruction parser, and call corresponding execution plug-ins according to instruction types of video operation instructions in the initial instruction sequence for sequential execution; A dynamic instruction insertion module is configured to, during the execution of the initial instruction sequence, if a user's voice input is received, identify the intent of the input voice in real time through the AI ​​auxiliary subsystem; if the input voice is identified as a question, generate an answer text and convert it into a digital human broadcast instruction; if the input voice is identified as a control command, convert it into a new video operation instruction and insert the new video operation instruction into the current execution position of the instruction sequence; after the newly inserted video operation instruction is executed by the instruction parser, continue to execute the subsequent instruction sequence; The synthesis module is used to render and synthesize images and audio according to the instruction execution results.

10. An instruction editing system, characterized in that: It includes an instruction module, an AI auxiliary subsystem, a digital human module, a material module, a scene module, a video rendering engine, and device IO. The instruction module includes an instruction editor and an instruction parser. The command editor is used to add, delete and modify video operation commands according to user needs to adjust the operation of the video operation commands; The instruction parser is used to determine the type of the video operation instruction, query the instruction plug-in list according to the type determination result, and call the execution plug-in matching the video operation instruction for execution, wherein the instruction plug-in list is obtained by dynamically registering the execution plug-in corresponding to each instruction type in the instruction parser in advance; The AI-assisted subsystem is used to assist in generating a broadcast command sequence based on the prompt words entered by the user when creating a new command list; during the command execution process, it recognizes the user's voice input and executes the command if it is recognized; if it is a question, it gives a corresponding answer; if the user answers, it gives feedback; when answering user questions, it combines the expert knowledge base and the language model to generate answers that conform to the style of the specific character; The digital human module is used to connect to the remote digital human platform, receive the broadcast text and obtain the synchronized video and audio streams; The material module is used to store various types of materials and provide a unified interface call; The scene module includes a scene editor, and the scene module is used to store and manage scene layout parameters and material arrangement rules; The video rendering engine is used to render and output the synthesized images and synthesized audio in a hierarchical tree structure based on preset video synthesis rules, scene layout parameters and material arrangement rules; The device 10 includes a device input module and a device output module. The device input module is used to capture user voice or user image, and the device output module is used for output presentation or data interaction.

Citation Information

Patent Citations

  • Digital human live broadcast interaction method and system based on large model

    CN119967197A

  • MOOC video method and device based on AIGC digital human

    CN120378707A

  • System

    JP2025050761A

  • Automatic generation of videos for digital products

    US20210319781A1