Design process accelerated video generation method based on semantic perception and multi-modal operation

CN122824952APending Publication Date: 2026-09-25BEIJING CHUANGZUOMEIHAO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610807746.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

传统方法通常需要设计师手动录制整个设计过程,不仅耗时耗力,还可能影响设计师的创作思路

Benefits of technology

[0043]本发明上述实施例提供的方案,实时采集用户在画布页上的全量设计交互事件,得到结构化的原始操作日志流,其中,所述原始操作日志流包含操作ID、时间戳、操作语义类型、目标图层或组件的标识符以及操作前后状态的增量差异数据;将所述原始操作日志流输入至视频关键帧识别模型,输出带有置信度评分的关键操作序列并标注视觉叙事中的信息熵权重,其中,所述视频关键帧识别模型包括时序上下文模块、操作语义类别模块、状态变更幅度模块、自适应时序压缩模块以及视觉连贯性增强模块;所述自适应时序压缩模块根据所述关键操作序列及其信息熵权重动态构建非均匀时间映射函数,对低信息密度区间进行帧聚合压缩以及对高价值操作区间保留原始粒度或进行微插值增强以生成初步优化帧调度表;所述视觉连贯性增强模块根据所述优化帧调度表及相邻帧间的增量差异数据进行局部运动矢量估计,合成具有中间插值帧的高保真帧序列;通过多模态同步渲染器对所述高保真帧序列中包含时序媒体元素的时段进行播放速率对齐与循环边界优化,得到加速创作过程视频。本发明通过视频关键帧识别模型生成高保真帧序列,在创作过程中自动记录和生成有吸引力的加速视频。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824952A_ABST
    Figure CN122824952A_ABST
Patent Text Reader

Abstract

The application discloses a design process accelerated video generation method and device based on semantic perception and multi-modal operation, wherein the method comprises the following steps: collecting all the design interaction events of a user on a canvas page in real time to obtain an original operation log stream; inputting the original operation log stream into a video key frame identification model to output a key operation sequence and label the information entropy weight in visual narration, wherein the video key frame identification model comprises a time sequence context, an operation semantic category, a state change amplitude, an adaptive time sequence compression module and a visual coherence enhancement module; the adaptive time sequence compression module generates a preliminary optimization frame scheduling table; the visual coherence enhancement module synthesizes a high-fidelity frame sequence with intermediate interpolation frames; and the high-fidelity frame sequence is subjected to playing rate alignment and cycle boundary optimization to obtain an accelerated creation process video. The application generates a high-fidelity frame sequence through a video key frame identification model, and automatically records and generates an attractive accelerated video in the creation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software design technology, and specifically to a method and apparatus for accelerating video generation in a design process based on semantic awareness and multimodal operation. Background Technology

[0002] In the design industry, showcasing the design process is invaluable for teaching, presentation, and marketing. Traditional methods typically require designers to manually record the entire design process, which is not only time-consuming and labor-intensive but can also hinder their creative flow. While some design software offers simple recording functions, they often lack intelligent processing, failing to effectively capture key design steps or automatically generate engaging accelerated videos. Therefore, developing a system capable of automatically recording the design process and intelligently generating accelerated videos is crucial for improving design efficiency and enhancing value presentation.

[0003] To address the aforementioned issues, this invention develops a method and apparatus for accelerating video generation during the design process based on semantic awareness and multimodal operation. It generates high-fidelity frame sequences through a video keyframe recognition model, enabling the automatic recording and generation of attractive accelerated videos during the creation process. Summary of the Invention

[0004] In view of the above problems, a method and apparatus for accelerating video generation based on semantic awareness and multimodal operation design process is proposed to facilitate the automatic recording and generation of attractive accelerated videos during the creation process.

[0005] According to one aspect of the present invention, a method for accelerating video generation based on semantic awareness and multimodal operation design process is provided, comprising:

[0006] Real-time collection of all design interaction events of users on the canvas page to obtain a structured raw operation log stream, wherein the raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of target layer or component, and incremental difference data of state before and after operation.

[0007] The original operation log stream is input into a video keyframe recognition model, which outputs a key operation sequence with confidence scores and labels the information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change magnitude module, an adaptive temporal compression module, and a visual coherence enhancement module. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequence and its information entropy weights, performs frame aggregation compression on low-information-density intervals, and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule table. The visual coherence enhancement module performs local motion vector estimation based on the optimized frame schedule table and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. The information entropy weights are dynamically labeled based on visual saliency, structural change complexity, and color perception differences. The confidence score is obtained by weighting the operation semantic type and the state change magnitude.

[0008] By using a multimodal synchronous renderer to align the playback rate and optimize the loop boundaries of the time segments containing temporal media elements in the high-fidelity frame sequence, a video with an accelerated creation process can be obtained.

[0009] In one alternative approach, real-time collection of all user design interaction events on the canvas page yields a structured raw operation log stream that further includes:

[0010] Inject an event delegate listener layer into the canvas page to capture the raw event stream, including all explicit user interactions and implicit triggering operations, by intercepting the canvas component's native DOM events and state change hooks.

[0011] The original event stream is passed to the operation semantic parser, which outputs a standardized event object with an operation semantic type. The operation semantic parser maps the low-order input signals in the original event stream to high-level design semantic actions based on the layer tree structure model and the component instantiation context.

[0012] Before and after each effective operation, the incremental difference data of the current canvas page is automatically recorded and bound to the corresponding event; the operation ID, timestamp, unique path identifier of the target layer or component, standardized event object and incremental difference data are encapsulated to form a structured log record; the structured log record is sorted according to the operation time sequence and aggregated into the original operation log stream.

[0013] In one alternative approach, the adaptive temporal compression module generating the preliminary optimized frame schedule table further includes:

[0014] Based on the timestamps of the key operation sequences and their corresponding information entropy weights, the entire design process is divided into multiple consecutive time intervals with adjacent key operations as boundaries, and the average information entropy weight of each consecutive time interval is calculated.

[0015] According to the preset weighted compression ratio mapping curve, a target time compression coefficient is dynamically assigned to each continuous time interval. Among them, the interval with the target time compression coefficient greater than the preset threshold is judged as a low information density interval and frame aggregation compression is performed to obtain an aggregated key frame sequence. The interval with the target time compression coefficient equal to or close to 1 is judged as a high-value operation interval and the original operation granularity is retained. Micro-interpolation enhancement is performed on the high semantic value operation according to the semantic model of UI components to obtain micro-interpolation enhanced frames.

[0016] The keyframe sequence is integrated and aggregated in the order of the target video timeline, along with the original frames and micro-interpolation enhancement frames. Each frame is associated with the original design timestamp, the target presentation timestamp, the operation ID list, and the frame source type identifier, and a preliminary optimized frame scheduling table is output. The preliminary optimized frame scheduling table is used to accelerate the mapping relationship between each frame in the video and the original design behavior and canvas state.

[0017] In an alternative approach, the execution of frame aggregation compression to obtain the aggregated keyframe sequence further includes:

[0018] Based on the incremental difference data of the state before and after the operation in the original operation log stream corresponding to the low information density interval, all intermediate state snapshots are extracted as candidate frames.

[0019] By using a motion energy minimization model, a subset of frames that maximize visual coherence or minimize visual abruptness is selected while satisfying the target number of frames, and an aggregated keyframe sequence for this low information density range is generated.

[0020] In an alternative approach, obtaining the micro-interpolated enhanced frame from the high semantic value operation based on the UI component semantic model further includes:

[0021] Based on the attribute differences in the corresponding incremental difference data, insert 1 to 3 intermediate state frames on the target time axis after the preset weight compression ratio mapping;

[0022] Enhanced frames are synthesized using easing functions and micro-interpolation to improve the visual smoothness and expressiveness of key design actions;

[0023] The high semantic value operations include layout adjustment, vector path editing, and primary color replacement; the attribute differences include coordinate offset, control point changes, and color channel gradients.

[0024] In one alternative approach, the visual coherence enhancement module synthesizing a high-fidelity frame sequence with intermediate interpolated frames further includes:

[0025] Extract the sequence of frame entries sorted by the target video timeline from the preliminary optimized frame scheduling table;

[0026] Based on the operation ID list associated with any two adjacent frame entries in the frame entry sequence, backtrack to the original operation log stream to obtain the corresponding start state snapshot and end state snapshot;

[0027] Analyze the incremental difference data between the initial state snapshot and the final state snapshot. The incremental difference data is the attribute change path and numerical difference of each node in the layer tree, including position, size, rotation, fill color, stroke, transparency and component instance parameters.

[0028] The incremental difference data is semantically grouped according to the layer semantic topology model, wherein the same visual motivation operation is divided into the same motion semantic unit, and the same visual motivation operation includes the same group of alignment adjustments and the same color theme switch.

[0029] Based on the time ratio of the frames to be inserted on the target video timeline, attribute-level interpolation is performed in each motion semantic unit to generate a complete layer tree description of the intermediate state and call the off-screen rendering engine to draw it as a bitmap frame in real time to obtain the intermediate interpolated frame.

[0030] All original frames and intermediate interpolated frames are sorted by the target rendering timestamp and merged into a high-fidelity frame sequence with natural motion transitions.

[0031] In an alternative approach, the method further includes:

[0032] The local motion vectors of the same motion semantic unit are calculated independently. The position change is simulated by the Bezier easing trajectory curve, the color gradient is displayed by the shortest perceptual path interpolation in the CIELAB color space, and the vector path editing is interpolated by the quadratic / cubic splines of the control points.

[0033] In one alternative approach, the formula for calculating the average information entropy weight for each consecutive time interval is:

[0034]

[0035] in,[ [] represents the start and end times of the k-th time interval; ; The visual information entropy corresponding to the operation at time t; For the temporal attention decay function, , To control the contribution intensity coefficient of recent operations to the interval entropy; ; This represents the number of nodes whose layer tree structure has changed. A normalized difference vector for the position, size, and rotation properties of all affected layers; The perceived color difference between fill and outline colors in the CIELAB color space; , These are the three-dimensional vector representations of the color attributes of the layer or component before and after the same operation in the CIELAB color space. .

[0036] In one alternative approach, the expression for selecting the subset of frames that maximizes visual coherence or minimizes visual abruptness is:

[0037]

[0038] in, This is a set of snapshots of all candidate intermediate states within a low information density range. For the target number of frames, ; This is the compressed interval duration. For the first A time interval; To output video frame rate; Let be the kinetic energy function. ; The set of layers that change between two frames; , Layers In frame and The state vector in; This is a non-linear perception mapping function used to map UI properties to a human eye sensitivity weighted space; The Mahalanobis distance; For layers Importance weights.

[0039] According to another aspect of this application, a video generation apparatus based on semantic awareness and multimodal operation for accelerating the design process is provided, comprising:

[0040] The full-volume interaction event collection module is used to collect all design interaction events of users on the canvas page in real time and obtain a structured raw operation log stream. The raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of target layer or component, and incremental difference data of state before and after operation.

[0041] A keyframe recognition and compression module is used to input the original operation log stream into a video keyframe recognition model, output key operation sequences with confidence scores, and label information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change magnitude module, an adaptive temporal compression module, and a visual coherence enhancement module. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequences and their information entropy weights, performs frame aggregation compression on low information density intervals, and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule table. The visual coherence enhancement module performs local motion vector estimation based on the optimized frame schedule table and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. The information entropy weights are dynamically labeled based on visual saliency, structural change complexity, and color perception differences. The confidence score is obtained by weighted calculation based on operation semantic type and state change magnitude.

[0042] The synchronous rendering module is used to perform playback rate alignment and loop boundary optimization on the time periods containing temporal media elements in the high-fidelity frame sequence through a multimodal synchronous renderer, so as to obtain a video that accelerates the creation process.

[0043] The solution provided in the above embodiments of the present invention collects all design interaction events of the user on the canvas page in real time to obtain a structured raw operation log stream. The raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of the target layer or component, and incremental difference data of the state before and after the operation. The raw operation log stream is input into a video keyframe recognition model, which outputs a key operation sequence with confidence scores and labels the information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change amplitude module, an adaptive temporal compression module, and a visual coherence module. The invention comprises a video keyframe recognition model to generate high-fidelity frame sequences and automatically records and generates attractive accelerated videos during the creation process. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequence and its information entropy weights. It performs frame aggregation compression on low-information-density intervals and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule. The visual coherence enhancement module estimates local motion vectors based on the optimized frame schedule and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. A multimodal synchronous renderer aligns the playback rate and optimizes the loop boundaries of the time periods containing temporal media elements in the high-fidelity frame sequence, resulting in an accelerated creation process video.

[0044] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above description and other objects, features and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0045] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0046] Figure 1 A flowchart illustrating the design process of the video generation method based on semantic awareness and multimodal operation according to an embodiment of the present invention is shown.

[0047] Figures 2a to 2f An illustration of accelerated video according to an embodiment of the present invention is shown. Figure 1 ;

[0048] Figures 3a to 3i A second schematic diagram of accelerated video according to an embodiment of the present invention is shown;

[0049] Figure 4 A schematic diagram illustrating the temporal compression and visual enhancement process according to an embodiment of the present invention is shown;

[0050] Figure 5 The diagram shows a functional structure of a device for accelerating video generation based on semantic awareness and multimodal operation design process according to an embodiment of the present invention. Detailed Implementation

[0051] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0052] Before implementing the embodiments of the present invention, the technical terms or custom terms used below are uniformly explained as follows:

[0053] Components: They allow the UI to be divided into independent, reusable parts. In practice, components are often organized into a nested tree structure.

[0054] Instance: An instance refers to a specific reference to a component. Any component that a user drags out of the component list or copies and pastes from a component is an "instance", not the component itself.

[0055] The following specific embodiments will provide a detailed description of the video generation method and apparatus for accelerating the design process based on semantic awareness and multimodal operation proposed in this invention.

[0056] Example 1:

[0057] Figure 1 This diagram illustrates the functional structure of a video generation method based on semantic awareness and multimodal operation, according to an embodiment of the present invention. Specifically, as shown... Figure 1 As shown, it includes the following steps:

[0058] Step S101: Collect all design interaction events of the user on the canvas page in real time to obtain a structured raw operation log stream, wherein the raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of target layer or component, and incremental difference data of state before and after operation.

[0059] In this embodiment, user design interaction behaviors are captured in real time, whether explicit (such as mouse clicks and keyboard input) or implicit (such as component auto-alignment), and all are recorded in detail. The captured raw event stream is converted into a standardized format containing operation ID, timestamp, operation semantic type, target layer or component identifier, and incremental difference data, which facilitates subsequent data analysis and processing, eliminating the need for designers to manually record or organize workflows.

[0060] In one alternative approach, real-time collection of all user design interaction events on the canvas page yields a structured raw operation log stream that further includes:

[0061] Inject an event delegate listener layer into the canvas page to capture the raw event stream, including all explicit user interactions and implicit triggering operations, by intercepting the canvas component's native DOM events and state change hooks.

[0062] The original event stream is passed to the operation semantic parser, which outputs a standardized event object with an operation semantic type. The operation semantic parser maps the low-order input signals in the original event stream to high-level design semantic actions based on the layer tree structure model and the component instantiation context.

[0063] Before and after each effective operation, the incremental difference data of the current canvas page is automatically recorded and bound to the corresponding event; the operation ID, timestamp, unique path identifier of the target layer or component, standardized event object and incremental difference data are encapsulated to form a structured log record; the structured log record is sorted according to the operation time sequence and aggregated into the original operation log stream.

[0064] In this embodiment, by hijacking native DOM events and state change hooks through an event proxy listening layer, it not only captures explicit user-initiated interactions such as clicks and drags, but also detects implicit operations automatically triggered by the system (such as component alignment and snapping, automatic layout adjustment, and style inheritance updates), achieving true full recording and avoiding the omission of key design intentions. Utilizing a layer tree structure model and component context information, the original low-level signals such as coordinate offsets and attribute changes are abstracted into semantic actions with design significance (such as aligning to the top, changing the main color, and ungrouping containers), significantly improving log readability and the accuracy of subsequent analysis. Incremental differences in the canvas state are automatically recorded before and after each effective operation and strongly bound to the event object, enabling accurate reconstruction of the design state at any given time, providing reliable data support for subsequent keyframe recognition, interpolation synthesis, and playback rendering.

[0065] Step S102: The original operation log stream is input into the video keyframe recognition model, which outputs a key operation sequence with confidence scores and labels the information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change magnitude module, an adaptive temporal compression module, and a visual coherence enhancement module. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequence and its information entropy weights, performs frame aggregation compression on low information density intervals, and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule table. The visual coherence enhancement module performs local motion vector estimation based on the optimized frame schedule table and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. The information entropy weights are dynamically labeled based on visual saliency, structural change complexity, and color perception differences. The confidence score is obtained by weighting the operation semantic type and the state change magnitude.

[0066] In this embodiment, by labeling each key operation with an information entropy weight in the visual narrative, its information value in the viewer's perception can be quantified. Based on this, a non-uniform temporal mapping function is constructed to achieve intelligent rhythm control that is "fast when it should be fast, slow when it should be slow," efficiently compressing redundant operations while fully showcasing high-value operations, significantly improving the video's information density and watchability. The adaptive temporal compression module not only determines which frames to retain but also how to enhance key frames, performing micro-interpolation (such as smooth movement and gradual transitions) on high semantic value intervals. Even under a significantly accelerated overall rhythm, key design actions remain smooth, natural, and expressive. The visual coherence enhancement module estimates local motion vectors based on incremental difference data and interpolates independently by semantic unit, eliminating visual imperfections such as flickering and jumps caused by frame skipping, generating a high-fidelity frame sequence approaching professional animation standards.

[0067] For example, when designing the login page, the designer repeatedly adjusted the font size of the "Buy" button (12px → 13px → 12.5px → 12px) in Phase 1, taking 30 seconds; in Phase 2, the entire form area was changed from centered to left-aligned and the button color was simultaneously changed from blue to the brand's green. Figure 4 As shown, the keyframe recognition model determines that the operation semantic value of stage one is low and the state change amplitude is small, so it is assigned a low information entropy weight; while stage two involves layout reconstruction and theme color change, which are marked as high-confidence key operations and assigned a high weight. The adaptive temporal compression module compresses the 30 seconds of stage one into 1 second, retaining only the initial and final state frames (selected by minimizing motion energy); stage two retains all operation details and inserts 2 micro-interpolation frames during the movement and color change processes, so that the movement trajectory is a gentle curve and the color gradient transitions along a uniform path perceived by the human eye. The visual coherence enhancement module recognizes that the form container movement and button color change are two independent motion semantic units, and applies position Bezier interpolation and CIELAB color difference interpolation respectively to ensure that the two animations do not interfere with each other and are smooth. In the output high-fidelity frame sequence, redundant fine-tuning is extremely compressed, and key designs are clearly presented with cinematic motion effects. Viewers can intuitively understand the entire creative intent within 10 seconds, far exceeding the experience of traditional uniform acceleration or simple fast forward.

[0068] In one alternative approach, the adaptive temporal compression module generating the preliminary optimized frame schedule table further includes:

[0069] Based on the timestamps of the key operation sequences and their corresponding information entropy weights, the entire design process is divided into multiple consecutive time intervals with adjacent key operations as boundaries, and the average information entropy weight of each consecutive time interval is calculated.

[0070] According to the preset weighted compression ratio mapping curve, a target time compression coefficient is dynamically assigned to each continuous time interval. Among them, the interval with the target time compression coefficient greater than the preset threshold is judged as a low information density interval and frame aggregation compression is performed to obtain an aggregated key frame sequence. The interval with the target time compression coefficient equal to or close to 1 is judged as a high-value operation interval and the original operation granularity is retained. Micro-interpolation enhancement is performed on the high semantic value operation according to the semantic model of UI components to obtain micro-interpolation enhanced frames.

[0071] The keyframe sequence is integrated and aggregated in the order of the target video timeline, along with the original frames and micro-interpolation enhancement frames. Each frame is associated with the original design timestamp, the target presentation timestamp, the operation ID list, and the frame source type identifier, and a preliminary optimized frame scheduling table is output. The preliminary optimized frame scheduling table is used to accelerate the mapping relationship between each frame in the video and the original design behavior and canvas state.

[0072] In this embodiment, by dividing the time intervals of key operation sequences and their information entropy weights, it is possible to automatically identify which time periods in the design process contain important creative behaviors (such as layout adjustments, color replacements, etc.) and which time periods are only for fine-tuning or waiting, thereby avoiding redundant recording of all operations and significantly improving the information transmission efficiency of the video. Unlike fixed speed acceleration or simple frame extraction, differentiated compression coefficients are assigned to different intervals according to a preset weighted compression ratio mapping curve to achieve non-uniform time compression. Among them, low-value periods are compressed efficiently, while high-value periods retain or even enhance details, shortening the total duration without losing key visual semantics. In the high-value operation intervals, not only is the original granularity preserved, but also 1 to 3 intermediate frames are selectively inserted using the UI component semantic model and synthesized using easing functions, making key actions smoother and more expressive, improving the video's viewing experience. At the same time, frame aggregation compression reduces the number of frames in low-information intervals, reducing rendering and storage overhead.

[0073] In one alternative approach, the formula for calculating the average information entropy weight for each consecutive time interval is:

[0074]

[0075] in,[ [] represents the start and end times of the k-th time interval; ; The visual information entropy corresponding to the operation at time t; Let be the temporal attention decay function. , To control the contribution intensity coefficient of recent operations to the interval entropy; ; This represents the number of nodes whose layer tree structure has changed. A normalized difference vector for the position, size, and rotation properties of all affected layers; The perceived color difference between fill and outline colors in the CIELAB color space; , These are the three-dimensional vector representations of the color attributes of the layer or component before and after the same operation in the CIELAB color space. .

[0076] In this embodiment, by measuring changes in the number of layer tree structure changes, position, size, rotation, and color differences, the impact of operations at each moment in the design process on the final visual result can be comprehensively and accurately measured. This allows for the identification of time periods containing important and significant design changes. The temporal attention decay function assigns higher weight to more recently occurring operations when calculating the information entropy of the interval, thereby more accurately simulating the distribution of user memory and focus on the design process. It considers not only the geometric transformations of UI components, such as position, size, and rotation, but also differences in color perception. In particular, it uses the CIELAB color space to calculate the perceived color difference before and after adjustments, ensuring that even subtle color changes are appropriately reflected. For example, a designer creates a webpage prototype, and the entire design process lasts 60 minutes. Between the 10th and 20th minute, a large-scale adjustment to the page layout is made, involving the movement and resizing of multiple elements, but the colors remain unchanged. At this point, and The value is relatively high, while The result was zero. In the following 20 to 30 minutes, detailed optimizations were performed, particularly fine-tuning the text color inside the text boxes. Although there were no major layout changes during this stage, color adjustments were frequent. Using the above formula, the different characteristics of the two time periods can be distinguished and appropriate information entropy weights can be assigned. In the first stage, due to numerous geometric transformations, a high [information entropy] was achieved. In addition, it is close to the current time. The average information entropy weight is also relatively high during this period. In the second stage, although there are fewer geometric transformations, the frequent and significant color adjustments also result in a good weight score. The adaptive temporal compression module then determines which strategy to adopt in each stage (e.g., the former is more suitable for retaining more original frames to showcase the complex layout adjustment process, while the latter uses appropriate micro-interpolation to smoothly transition color changes).

[0077] In an alternative approach, the execution of frame aggregation compression to obtain the aggregated keyframe sequence further includes:

[0078] Based on the incremental difference data of the state before and after the operation in the original operation log stream corresponding to the low information density interval, all intermediate state snapshots are extracted as candidate frames.

[0079] By using a motion energy minimization model, a subset of frames that maximize visual coherence or minimize visual abruptness is selected while satisfying the target number of frames, and an aggregated keyframe sequence for this low information density range is generated.

[0080] In this embodiment, in low information density intervals (such as repeated fine-tuning, drag-and-drop trial and error, and empty operations without substantial changes), the original operation log may contain a large number of intermediate state snapshots. Directly discarding some frames can easily lead to abrupt distortion, while by optimizing and selecting the fewest but smoothest subset of frames, a natural visual transition can be maintained while significantly compressing the duration. Motion energy is not simply calculated by pixel differences or attribute differences, but rather by mapping UI attribute changes to a human eye sensitivity weighted space through nonlinear perceptual mapping functions, making the differences between frames more consistent with the laws of human visual perception. For example, a slight shift in the position of a small icon is more easily noticed than a slight change in the background color of a large area, and the weights can be dynamically adjusted accordingly. The compressed interval needs to be adapted to the total duration and frame rate of the target video. The optimal subset is solved under a fixed frame limit to ensure that the output video does not stutter or lag due to excessive or insufficient local compression.

[0081] In one alternative approach, the expression for selecting the subset of frames that maximizes visual coherence or minimizes visual abruptness is:

[0082]

[0083] in, This is a set of snapshots of all candidate intermediate states within a low information density range. For the target number of frames, ; This is the compressed interval duration. For the first A time interval; To output video frame rate; Let be the kinetic energy function. ; The set of layers that change between two frames; , Layers In frame and The state vector in; This is a non-linear perception mapping function used to map UI properties to a human eye sensitivity weighted space; The Mahalanobis distance; For layers Importance weights.

[0084] In this embodiment, for example, when a designer creates a dashboard interface, the starting angle of a circular progress bar is repeatedly adjusted between 10:05 and 10:10 (5 seconds), with a total of 8 fine-tuning adjustments (each +2° to 5°), while other elements remain unchanged. The candidate frame set contains 9 state snapshots (initial + 8 operations). This interval has low information entropy and is compressed to 0.5 seconds. = ⌊0.5 × 20⌋ = 10, but in practice, due to the simplicity of the content, a maximum of 4 frames are retained to avoid redundancy. If snapshots {1, 3, 6, 9} are selected, the angle changes evenly (0°→10°→25°→40°), with each step approximately 10°~15°. The sum is low. If {1, 2, 8, 9} is selected, the first two frames show almost no change (0°→2°), while the last two frames show a sudden change (35°→40°). The sudden jump after a long period of stillness in between leads to… The higher the sensitivity (especially to the last jump, as the human eye is more sensitive to sudden changes after a period of stillness), the better. (Differences amplified after mapping). The optimal solution {1, 4, 7, 9} was selected, making the angle change approximately linear, visually presenting a uniform rotation effect. The original 5-second operation was compressed to 0.2 seconds (4 frames @ 20fps), while viewers can still clearly understand the design intent of the gradual adjustment of the progress bar's starting angle, without any sense of jumps or pauses. While ensuring extreme compression, the best balance between information retention and viewing experience was achieved through perceptual optimization and global planning.

[0085] In an alternative approach, obtaining the micro-interpolated enhanced frame from the high semantic value operation based on the UI component semantic model further includes:

[0086] Based on the attribute differences in the corresponding incremental difference data, insert 1 to 3 intermediate state frames on the target time axis after the preset weight compression ratio mapping;

[0087] Enhanced frames are synthesized using easing functions and micro-interpolation to improve the visual smoothness and expressiveness of key design actions;

[0088] The high semantic value operations include layout adjustment, vector path editing, and primary color replacement; the attribute differences include coordinate offset, control point changes, and color channel gradients.

[0089] In this embodiment, only operations with genuine design intent value (such as layout adjustment, vector path editing, and primary color replacement) are enhanced, avoiding wasting computational resources on irrelevant operations. This allows the accelerated video to significantly shorten its duration while clearly presenting key creative logic. Unlike video interpolation algorithms such as optical flow, the interpolation logic is customized based on the semantic type of UI components and the physical meaning of attribute changes (such as coordinate offset, control point movement, and color gradient). For example, perceptibly uniform easing transitions are used for colors, and spline smoothing is used for paths, making the animation more in line with designer expectations and user cognition. A small number of intermediate frames (1-3 frames) are inserted on the target timeline, which avoids the original operations appearing abrupt due to compression and prevents excessive interpolation from causing the video to become lengthy, maximizing visual smoothness within a limited frame budget. Non-linear easing functions (such as Bézier curves) are used to synthesize intermediate states, simulating the acceleration changes of real UI interaction effects, making processes such as element movement and color switching more rhythmic, significantly better than the mechanical feel of linear interpolation. Among them, micro-interpolation is only applied to the high-value operation range (compression coefficient ≈ 1), complementing the frame aggregation in the low-information range (the former enhances fidelity, the latter compresses efficiently) to jointly construct an optimized frame scheduling table that takes into account information density, visual coherence and viewing rhythm.

[0090] In one alternative approach, the visual coherence enhancement module synthesizing a high-fidelity frame sequence with intermediate interpolated frames further includes:

[0091] Extract the sequence of frame entries sorted by the target video timeline from the preliminary optimized frame scheduling table;

[0092] Based on the operation ID list associated with any two adjacent frame entries in the frame entry sequence, backtrack to the original operation log stream to obtain the corresponding start state snapshot and end state snapshot;

[0093] Analyze the incremental difference data between the initial state snapshot and the final state snapshot. The incremental difference data is the attribute change path and numerical difference of each node in the layer tree, including position, size, rotation, fill color, stroke, transparency and component instance parameters.

[0094] The incremental difference data is semantically grouped according to the layer semantic topology model, wherein the same visual motivation operation is divided into the same motion semantic unit, and the same visual motivation operation includes the same group of alignment adjustments and the same color theme switch.

[0095] Based on the time ratio of the frames to be inserted on the target video timeline, attribute-level interpolation is performed in each motion semantic unit to generate a complete layer tree description of the intermediate state and call the off-screen rendering engine to draw it as a bitmap frame in real time to obtain the intermediate interpolated frame.

[0096] All original frames and intermediate interpolated frames are sorted by the target rendering timestamp and merged into a high-fidelity frame sequence with natural motion transitions.

[0097] In this embodiment, instead of relying on pixel-level optical flow or black-box AI models, attribute-level reconstruction and interpolation are performed based on structured incremental difference data in the original operation log. This ensures that the intermediate frames are logically completely faithful to the user's true design intent, avoiding layer misalignment, color distortion, or inconsistent component states caused by blind frame interpolation. Multiple layer changes are merged into unified motion semantic units through a layer semantic topology model (e.g., an alignment operation involves the synchronous movement of 5 icons), ensuring that these elements maintain consistent movement and rhythmic synchronization during interpolation, avoiding element drift and asynchrony issues caused by independent interpolation. High-fidelity transitions of complex UI attributes are supported, handling full-dimensional attribute changes including position, size, rotation, color, stroke, transparency, and even component instance parameters. Combined with interpolation methods such as Bezier trajectory, CIELAB color space gradient, and spline path, smoother, accelerated videos with transitions more consistent with human visual perception are generated. The interpolation operation is integrated with the initial compression, directly affecting adjacent keyframes in the preliminary optimized frame schedule table, further improving visual coherence on the already compressed timeline.

[0098] In an alternative approach, the method further includes:

[0099] The local motion vectors of the same motion semantic unit are calculated independently. The position change is simulated by the Bezier easing trajectory curve, the color gradient is displayed by the shortest perceptual path interpolation in the CIELAB color space, and the vector path editing is interpolated by the quadratic / cubic splines of the control points.

[0100] In this embodiment, position changes use Bezier easing trajectories to simulate acceleration changes in real physical motion, avoiding the mechanical feel of linear movement; color gradients are interpolated along the shortest perceptual path in the CIELAB color space to ensure natural color transitions in human vision, avoiding the common graying or abrupt changes in intermediate colors; vector paths are controlled by quadratic / cubic spline interpolation points to maintain curve smoothness and topological structure, preventing path jitter or self-intersection. All layers belonging to the same visual motive (such as a single alignment or a single theme switch) are grouped into a single motion unit and share the same interpolation logic and time function, ensuring rhythmic uniformity and visual coordination during multi-element collaborative movement. Figures 2a to 2f , Figures 3a to 3i As shown, accelerated videos not only demonstrate changes in results, but also convey how and why changes occur through smooth transition animations, enabling viewers to clearly understand the designer's operational logic and aesthetics, greatly enhancing the teaching, demonstration, or review value of the videos.

[0101] Step S103: The playback rate of the time segments containing temporal media elements in the high-fidelity frame sequence is aligned and the loop boundary is optimized by a multimodal synchronous renderer to obtain a video that accelerates the creation process.

[0102] In this embodiment, time-series media elements such as videos, GIFs, or automatic carousel components are often embedded during the design process, each with its own independent timeline. Playing these embedded media directly at the accelerated main video's rhythm would cause fast-forward distortion or loop breaks. If the complete retention time of the embedded time-series media causes the total duration of the accelerated video to exceed a preset target, weighted compression or truncation is performed based on the interaction priority of the media elements to ensure the final video duration meets the acceleration expectations. Playback rate alignment ensures that the embedded media maintains its original rhythm in the accelerated video (e.g., a 2-second animation still plays completely once after compression), avoiding information loss or visual confusion. For looping media (such as loading animations or background particle effects), truncating at any frame would result in noticeable flickering back to the starting point in the video. Loop boundary optimization identifies the loop cycle and aligns the video clip point to the loop's start / end frame, achieving seamless transitions and improving visual smoothness. Therefore, it ensures that even if the overall design process is significantly accelerated, these dynamic elements can still be presented completely and naturally. For example, when a designer is creating a prototype of an app login page, a 4-second Lottie login success animation is inserted within a 10-second timeframe and set to play once and then stop. A 1-second looping background particle effect is then added. If the entire design process is compressed into a 2-minute video, this 10-second interval might be compressed to only 1 second. If the video is sped up proportionally, the Lottie animation will flash at 4x speed, making it impossible for viewers to see the animation details; the particle effect might only play 0.7 loops within 1 second, suddenly jumping back to the starting point at the end, causing flickering. In this embodiment, for the Lottie effect, a full 4 seconds of playback is retained in the sped-up video (without local compression), and the main design operation is paused during this period; for the particle effect, loop boundary optimization is used. Instead of playing 0.8 loops within the allocated 0.8 seconds, it is adjusted to play a full loop (1 second), and the time is compensated by fine-tuning the duration of the static frames before and after, ensuring that the last frame is the loop start point; in the final video, the Lottie animation is clear and complete, the particle background loops seamlessly, and the overall rhythm remains tight. Viewers can quickly browse the entire design process without missing key animation details, meeting the needs of high-requirement scenarios such as design reviews and portfolio presentations.

[0103] The solution provided in the above embodiments of the present invention collects all design interaction events of the user on the canvas page in real time to obtain a structured raw operation log stream. The raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of the target layer or component, and incremental difference data of the state before and after the operation. The raw operation log stream is input into a video keyframe recognition model, which outputs a key operation sequence with confidence scores and labels the information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change amplitude module, an adaptive temporal compression module, and a visual coherence module. The invention comprises a video keyframe recognition model to generate high-fidelity frame sequences and automatically records and generates attractive accelerated videos during the creation process. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequence and its information entropy weights. It performs frame aggregation compression on low-information-density intervals and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule. The visual coherence enhancement module estimates local motion vectors based on the optimized frame schedule and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. A multimodal synchronous renderer aligns the playback rate and optimizes the loop boundaries of the time periods containing temporal media elements in the high-fidelity frame sequence, resulting in an accelerated creation process video.

[0104] Example 2:

[0105] Figure 5 This diagram illustrates the functional structure of a device for accelerating video generation based on semantic awareness and multimodal operation design processes, according to an embodiment of the present invention. Figure 5 As shown, the device includes:

[0106] The full-volume interactive event acquisition module 501 is used to collect all design interaction events of the user on the canvas page in real time and obtain a structured raw operation log stream. The raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of target layer or component, and incremental difference data of state before and after operation.

[0107] The keyframe recognition and compression module 502 is used to input the original operation log stream into the video keyframe recognition model, output key operation sequences with confidence scores, and label the information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change magnitude module, an adaptive temporal compression module, and a visual coherence enhancement module. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequences and their information entropy weights, performs frame aggregation compression on low information density intervals, and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule table. The visual coherence enhancement module performs local motion vector estimation based on the optimized frame schedule table and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. The information entropy weights are dynamically labeled based on visual saliency, structural change complexity, and color perception differences. The confidence score is obtained by weighted calculation based on the operation semantic type and state change magnitude.

[0108] The synchronous rendering module 503 is used to perform playback rate alignment and loop boundary optimization on the time periods containing temporal media elements in the high-fidelity frame sequence through a multimodal synchronous renderer, so as to obtain a video that accelerates the creation process.

[0109] The solution provided in the above embodiments of the present invention collects all design interaction events of the user on the canvas page in real time to obtain a structured raw operation log stream. The raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of the target layer or component, and incremental difference data of the state before and after the operation. The raw operation log stream is input into a video keyframe recognition model, which outputs a key operation sequence with confidence scores and labels the information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change amplitude module, an adaptive temporal compression module, and a visual coherence module. The invention comprises a video keyframe recognition model to generate high-fidelity frame sequences and automatically records and generates attractive accelerated videos during the creation process. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequence and its information entropy weights. It performs frame aggregation compression on low-information-density intervals and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule. The visual coherence enhancement module estimates local motion vectors based on the optimized frame schedule and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. A multimodal synchronous renderer aligns the playback rate and optimizes the loop boundaries of the time periods containing temporal media elements in the high-fidelity frame sequence, resulting in an accelerated creation process video.

[0110] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of the present invention are not directed to any particular programming language. It should be understood that the content of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the invention.

[0111] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0112] Similarly, it should be understood that, in order to simplify the invention and aid in understanding one or more of the various inventive aspects, features of the embodiments of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the above description of exemplary embodiments of the invention. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.

[0113] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0114] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0115] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0116] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A method for accelerating video generation based on semantic awareness and multimodal operation design process, characterized in that, include: Real-time collection of all design interaction events of users on the canvas page to obtain a structured raw operation log stream, wherein the raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of target layer or component, and incremental difference data of state before and after operation. The original operation log stream is input into a video keyframe recognition model, which outputs a key operation sequence with confidence scores and labels the information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change magnitude module, an adaptive temporal compression module, and a visual coherence enhancement module. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequence and its information entropy weights, performs frame aggregation compression on low-information-density intervals, and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule table. The visual coherence enhancement module performs local motion vector estimation based on the optimized frame schedule table and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. The information entropy weights are dynamically labeled based on visual saliency, structural change complexity, and color perception differences. The confidence score is obtained by weighting the operation semantic type and the state change magnitude. By using a multimodal synchronous renderer to align the playback rate and optimize the loop boundaries of the time segments containing temporal media elements in the high-fidelity frame sequence, a video with an accelerated creation process can be obtained.

2. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 1, characterized in that, Real-time collection of all user design interaction events on the canvas page yields a structured raw operation log stream, which further includes: Inject an event delegate listener layer into the canvas page to capture the raw event stream, including all explicit user interactions and implicit triggering operations, by intercepting the canvas component's native DOM events and state change hooks. The original event stream is passed to the operation semantic parser, which outputs a standardized event object with an operation semantic type. The operation semantic parser maps the low-order input signals in the original event stream to high-level design semantic actions based on the layer tree structure model and the component instantiation context. Before and after each effective operation, the incremental difference data of the current canvas page is automatically recorded and bound to the corresponding event; the operation ID, timestamp, unique path identifier of the target layer or component, standardized event object and incremental difference data are encapsulated to form a structured log record; the structured log record is sorted according to the operation time sequence and aggregated into the original operation log stream.

3. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 1 or 2, characterized in that, The adaptive temporal compression module generates a preliminary optimized frame scheduling table, which further includes: Based on the timestamps of the key operation sequences and their corresponding information entropy weights, the entire design process is divided into multiple consecutive time intervals with adjacent key operations as boundaries, and the average information entropy weight of each consecutive time interval is calculated. According to the preset weighted compression ratio mapping curve, a target time compression coefficient is dynamically assigned to each continuous time interval. Among them, the interval with the target time compression coefficient greater than the preset threshold is judged as a low information density interval and frame aggregation compression is performed to obtain an aggregated key frame sequence. The interval with the target time compression coefficient equal to or close to 1 is judged as a high-value operation interval and the original operation granularity is retained. Micro-interpolation enhancement is performed on the high semantic value operation according to the semantic model of UI components to obtain micro-interpolation enhanced frames. The keyframe sequence is integrated and aggregated in the order of the target video timeline, along with the original frames and micro-interpolation enhancement frames. Each frame is associated with the original design timestamp, the target presentation timestamp, the operation ID list, and the frame source type identifier, and a preliminary optimized frame scheduling table is output. The preliminary optimized frame scheduling table is used to accelerate the mapping relationship between each frame in the video and the original design behavior and canvas state.

4. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 3, characterized in that, The aggregated keyframe sequence obtained by the execution frame aggregation and compression further includes: Based on the incremental difference data of the state before and after the operation in the original operation log stream corresponding to the low information density interval, all intermediate state snapshots are extracted as candidate frames. By using a motion energy minimization model, a subset of frames that maximize visual coherence or minimize visual abruptness is selected while satisfying the target number of frames, and an aggregated keyframe sequence for this low information density range is generated.

5. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 3, characterized in that, The step of obtaining micro-interpolation enhanced frames by performing high semantic value operations based on the semantic model of UI components further includes: Based on the attribute differences in the corresponding incremental difference data, insert 1 to 3 intermediate state frames on the target time axis after the preset weight compression ratio mapping; Enhanced frames are synthesized using easing functions and micro-interpolation to improve the visual smoothness and expressiveness of key design actions; The high semantic value operations include layout adjustment, vector path editing, and primary color replacement; the attribute differences include coordinate offset, control point changes, and color channel gradients.

6. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 1 or 2, characterized in that, The visual coherence enhancement module synthesizes a high-fidelity frame sequence with intermediate interpolated frames, which further includes: Extract the sequence of frame entries sorted by the target video timeline from the preliminary optimized frame scheduling table; Based on the operation ID list associated with any two adjacent frame entries in the frame entry sequence, backtrack to the original operation log stream to obtain the corresponding start state snapshot and end state snapshot; Analyze the incremental difference data between the initial state snapshot and the final state snapshot. The incremental difference data is the attribute change path and numerical difference of each node in the layer tree, including position, size, rotation, fill color, stroke, transparency and component instance parameters. The incremental difference data is semantically grouped according to the layer semantic topology model, wherein the same visual motivation operation is divided into the same motion semantic unit, and the same visual motivation operation includes the same group of alignment adjustments and the same color theme switch. Based on the time ratio of the frames to be inserted on the target video timeline, attribute-level interpolation is performed in each motion semantic unit to generate a complete layer tree description of the intermediate state and call the off-screen rendering engine to draw it as a bitmap frame in real time to obtain the intermediate interpolated frame. All original frames and intermediate interpolated frames are sorted by the target rendering timestamp and merged into a high-fidelity frame sequence with natural motion transitions.

7. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 6, characterized in that, The method further includes: The local motion vectors of the same motion semantic unit are calculated independently. The position change is simulated by the Bezier easing trajectory curve, the color gradient is displayed by the shortest perceptual path interpolation in the CIELAB color space, and the vector path editing is interpolated by the quadratic / cubic splines of the control points.

8. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 3, characterized in that, The formula for calculating the average information entropy weight for each continuous time interval is: ,in,[ [] represents the start and end times of the k-th time interval; ; The visual information entropy corresponding to the operation at time t; Let be the temporal attention decay function. , To control the contribution intensity coefficient of recent operations to the interval entropy; ; This represents the number of nodes whose layer tree structure has changed. A normalized difference vector for the position, size, and rotation properties of all affected layers; The perceived color difference between fill and outline colors in the CIELAB color space; , These are the three-dimensional vector representations of the color attributes of the layer or component before and after the same operation in the CIELAB color space. .

9. The method for accelerating video generation based on semantic awareness and multimodal operation design process according to claim 4, characterized in that, The expression for finding the subset of frames that maximizes visual coherence or minimizes visual abruptness is as follows: ,in, This is a set of snapshots of all candidate intermediate states within a low information density range. For the target number of frames, ; This is the compressed interval duration. For the first A time interval; To output video frame rate; Let be the kinetic energy function. ; The set of layers that change between two frames; , Layers In frame and The state vector in; This is a non-linear perception mapping function used to map UI properties to a human eye sensitivity weighted space; The Mahalanobis distance; For layers Importance weights.

10. A video generation device with accelerated design process based on semantic awareness and multimodal operation, characterized in that, include: The full-volume interaction event collection module is used to collect all design interaction events of users on the canvas page in real time and obtain a structured raw operation log stream. The raw operation log stream includes operation ID, timestamp, operation semantic type, identifier of target layer or component, and incremental difference data of state before and after operation. A keyframe recognition and compression module is used to input the original operation log stream into a video keyframe recognition model, output key operation sequences with confidence scores, and label information entropy weights in the visual narrative. The video keyframe recognition model includes a temporal context module, an operation semantic category module, a state change magnitude module, an adaptive temporal compression module, and a visual coherence enhancement module. The adaptive temporal compression module dynamically constructs a non-uniform temporal mapping function based on the key operation sequences and their information entropy weights, performs frame aggregation compression on low information density intervals, and retains the original granularity or performs micro-interpolation enhancement on high-value operation intervals to generate a preliminary optimized frame schedule table. The visual coherence enhancement module performs local motion vector estimation based on the optimized frame schedule table and incremental difference data between adjacent frames, synthesizing a high-fidelity frame sequence with intermediate interpolated frames. The information entropy weights are dynamically labeled based on visual saliency, structural change complexity, and color perception differences. The confidence score is obtained by weighted calculation based on operation semantic type and state change magnitude. The synchronous rendering module is used to perform playback rate alignment and loop boundary optimization on the time periods containing temporal media elements in the high-fidelity frame sequence through a multimodal synchronous renderer, so as to obtain a video that accelerates the creation process.