Method and device for generating first screen video background based on page elements

By acquiring the multimodal interface features and user natural language commands of the target website's homepage, a structured visual control description is generated using a multimodal alignment and fusion model and a cross-modal attention model. This description is decoupled into a control subspace of content, style, and motion, solving the problem of video background conflict with page elements in existing technologies and realizing the generation of dynamic video backgrounds that are highly coordinated with page elements.

CN121725084APending Publication Date: 2026-03-24BEIJING CHUANGZUOMEIHAO TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing video background generation methods rely on fixed video files pre-made by designers, which are inflexible, cannot be adapted to page content, and the generated video backgrounds conflict with page elements, making it impossible to generate customized dynamic visual content.

Method used

By acquiring the multimodal interface features of the target website's homepage and the user's natural language commands, a structured visual control description is generated using a multimodal alignment and fusion model and a cross-modal attention model. This description is decoupled into control subspaces of content, style, and motion, and then used to synthesize a dynamic video background.

Benefits of technology

The generated video background is highly coordinated with the page elements, ensuring that the composition, color tone, and theme are consistent with the original content of the page, thus achieving a customized dynamic visual effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725084A_ABST
    Figure CN121725084A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for generating a first-screen video background based on page elements, and the method comprises the steps: obtaining a multi-mode interface feature of a first screen of a target website, and receiving a natural language instruction inputted by a user through an interaction dialog box; analyzing the visual intention, style preference and dynamic demand in the natural language instruction to generate a structured visual control description; extracting layout, color and semantic entities from the multi-modal interface features through a cross-modal attention model to form and generate condition vectors; and inputting the generated condition vector into a condition decoupling video diffusion model, decoupling the generated condition vector into a control subspace of content, style and motion, and synthesizing a video background sequence according to the control subspace and the structured visual control description so as to synthesize a dynamic background in the first screen of the target website. According to the method, the dynamic video background is generated under the condition of the webpage interface and the guidance of the user language, so that the defect of rough control of a general text video model is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software design technology, and specifically to a method and apparatus for generating a first-screen video background based on page elements. Background Technology

[0002] In the internet experience, the website's homepage serves as the first window users see, and its visual appeal plays a crucial role in user retention, brand awareness, and interaction guidance. While static text and image layouts can convey information, they are ill-suited for scenarios that prioritize experience and high interactivity (such as product launch websites, brand image sites, game promotional pages, and art and creative platforms). Dynamic video backgrounds, with their rich expressiveness, strong emotional communication capabilities, and visual impact, have become an important means of enhancing the visual hierarchy of the homepage and improving the user experience.

[0003] Traditional video background applications primarily rely on pre-made and uploaded fixed video files by designers, resulting in high production costs and requiring professional visual design and video production skills. They also lack flexibility, as once the video content is determined, it's difficult to adapt it to the specific content of the page (such as color, layout, and text) or real-time user preferences. Furthermore, they cannot generate customized dynamic visual content for different users or different access scenarios. With the rapid development of AIGC technology, especially Text-to-Image and Text-to-Video models, while they can generate dynamic backgrounds or videos, they lack analysis of the existing visual structure, color system, and component functions of the page. This leads to the generated background images or videos clashing with foreground elements in terms of composition, tone, and visual weight, resulting in a very uncoordinated appearance. In addition, most can only generate static images; the quality and consistency of the generated dynamic images are poor, and they cannot be finely controlled according to user instructions. Adjusting some page elements and regenerating often leads to other unpredictable changes.

[0004] To address the aforementioned issues, this invention proposes a method and apparatus for generating a first-screen video background based on page elements. This method generates a dynamic video background using the webpage interface itself as a condition and user language as a guide, ensuring that the generated video background is highly coordinated with the original page content in terms of composition, color scheme, and theme, thus avoiding the shortcomings of existing general-purpose text-based video models in terms of their rough control. Summary of the Invention

[0005] In view of the above problems, a method and apparatus for generating the first screen video background based on page elements are proposed to facilitate this process.

[0006] According to one aspect of the present invention, a method for generating a first-screen video background based on page elements is provided, comprising:

[0007] The system acquires the multimodal interface features of the target website's home screen and receives natural language commands input by the user through an interactive dialog box; wherein, the multimodal interface features include layout structure, visual colors, text semantics, and interactive components;

[0008] The natural language instructions and multimodal interface features are input into a multimodal alignment and fusion model to analyze the visual intent, style preferences and dynamic requirements in the natural language instructions to generate a structured visual control description; layout, color and semantic entities are extracted from the multimodal interface features through a cross-modal attention model to form a generation condition vector;

[0009] The generated conditional vector is input into the conditional decoupled video diffusion model, and the generated conditional vector is decoupled into a control subspace of content, style and motion. A video background sequence is synthesized based on the control subspace and the structured visual control description.

[0010] A dynamic background is synthesized based on the video background sequence and then composited onto the homepage of the target website.

[0011] In one alternative approach, the multimodal alignment and fusion model includes an instruction parsing and knowledge retrieval module, a heterogeneous feature encoding module, a shared semantic space optimal transmission projection module, a multi-stage alignment and adaptive fusion module, and a conditional vector refinement output module.

[0012] The instruction parsing and knowledge retrieval module is used to receive the natural language instruction and decompose and reason it, and simultaneously perform multi-hop retrieval from the Web visual design pattern library and motion graphics parameter library, and output an executable visual description script including hierarchical structure, abstract intent, specific attributes and constraints.

[0013] The heterogeneous feature encoding module, including a layout encoder, a color encoder, and a unified encoder, runs in parallel with the instruction parsing and knowledge retrieval module and is used to receive the multimodal interface features. The layout encoder constructs a directed heterogeneous graph of interface elements and their spatial relationships, encoding global typographic semantics through message passing and graph attention mechanisms. The color encoder calculates color harmony and prominence indices in the CIELAB color space and generates sentiment semantic embeddings. The unified encoder jointly encodes textual semantics and interactive component images to extract their joint functional and thematic representations.

[0014] The shared semantic space optimal transport projection module receives the executable visual description script and the modal features encoded by the heterogeneous feature encoding module, respectively; it treats the description script from the language domain and the page features from the visual domain as two distributions, and finds the minimum cost projection matrix that maps them to the shared semantic space by calculating the regularized optimal transport plan.

[0015] The multi-stage alignment and adaptive fusion module includes a multi-granularity semantic routing network, used to receive the minimum cost projection matrix and filter relevant page feature modal subsets according to the abstract intent in the description script;

[0016] The condition vector refinement output module is used to output the generated condition vector based on the subset of page feature modes.

[0017] In one alternative approach, the cross-modal attention model includes a multimodal tensor quantized feature encoder, a hierarchical semantic routing and attention network, a feature unwrapping and fusion module based on causal intervention, and an adaptive generative conditional projector.

[0018] The multimodal tensor quantization feature encoder includes a layout neural coding submodule, a perception and emotion quantization coding submodule, and a multi-granularity visual alignment coding submodule. The layout neural coding submodule constructs a dynamic hypergraph from the geometric attributes, spatial relationships, and visual hierarchy of interface elements. The hyperedges of the dynamic hypergraph dynamically represent visual balance groups, reading flow sequences, or interactive function clusters, and are encoded and output as a layout tensor with explicit structural semantics through high-order message passing. The perception and emotion quantization coding submodule simulates human eye perception in the ICtCp color space based on a color adaptation transformation model, calculates the visual saliency, emotional arousal, and value of colors, and decomposes the overall color distribution into three factor matrices—base, emotion, and emphasis—through non-negative tensor decomposition to output a structured color emotion tensor. The multi-granularity visual alignment coding submodule establishes the association between text descriptions, component visual screenshots, and general vision, and outputs a semantic consistency tensor.

[0019] The hierarchical semantic routing and attention network includes a semantic-aware routing layer and an intra-channel modulation attention layer; wherein, the semantic-aware routing layer analyzes the meta-semantics of each input tensor and queries a learnable cross-modal interaction prior codebook to dynamically generate a set of routing weight vectors; the intra-channel modulation attention layer performs conditional multi-head cross-attention in each channel;

[0020] The feature untangling fusion module based on causal intervention is connected to the hierarchical semantic routing and attention network. It performs causal graph discovery on the fused feature tensor and infers the potential causal relationships between different semantic dimensions. Through do-calculus intervention, it forcibly cuts off non-causal statistical correlation paths and strengthens the transmission of causal paths at the feature level, and outputs a set of untangled causal semantic features.

[0021] The adaptive generating conditional projector receives the unwrapped causal semantic feature set and includes a multi-head conditional linear projection layer.

[0022] In one alternative approach, the conditional decoupled video diffusion model includes a conditional decoupled encoder, a temporal fusion module, a spatiotemporal decoupled injection module, a staged scheduling controller, and a video latent space synthesis network connected in sequence.

[0023] The conditional decoupling encoder receives the generated conditional vector, which includes a content path, a style path, and a motion path set in parallel. The content path consists of a residual convolutional network and a self-attention layer. It maximizes the mutual information between the content and the entities and composition in the generated conditional vector to impose orthogonality constraints on the style and motion information, and outputs a content latent code representing the main subject and static layout of the scene. The style path uses an adaptive multilayer perceptron to extract information strongly related to appearance attributes and optimizes it by contrast decoupling loss with the output of the content path, outputting a time-independent style latent code. The motion path consists of a one-dimensional temporal convolutional network and a gated recurrent unit. It captures the dynamic patterns and evolutionary trends contained in the generated conditional vector and outputs a motion latent code representing the temporal change pattern.

[0024] The timing fusion module is connected to the conditional decoupling encoder, receives the motion latent code and dynamic parameters in the structured visual control description, and outputs the enhanced timing condition signal.

[0025] The spatiotemporal decoupling injection module adopts a U-ViT structure, starting with the latent representation of random Gaussian noise video, and receives feature maps from the previous layer in each level of the U-ViT structure.

[0026] The staged scheduling controller, as a global control logic unit, receives the time step signal of the denoising process and dynamically adjusts the relative weights of the feature maps at different stages of denoising.

[0027] The video latent space synthesis network outputs a video background latent sequence, which is then mapped by a decoder to obtain a video background sequence in pixel space.

[0028] In an alternative approach, synthesizing the video background sequence based on the control subspace and the structured visual control description further includes:

[0029] Using the content latent code and style latent code as conditions, an initial video noise latent sequence is generated by combining randomly sampled Gaussian noise. The scene subject and static layout information carried by the content latent code are injected into the U-ViT network with high weights through the spatiotemporal decoupling injection module, driving the network to predict the key frame static composition of the video sequence, so as to ensure that the generated background matches the page layout intention and the abstract intention specified by the user, thereby outputting a basic video latent sequence draft.

[0030] The basic video latent sequence draft is used for dynamic effect synthesis. In this process, the phased scheduling controller transfers the control of the generation conditions to the enhanced temporal condition signal. The spatiotemporal decoupling injection module synthesizes dynamic elements that conform to the description on the basic video latent sequence draft based on the enhanced temporal condition signal. At the same time, the dynamic parameters in the description are transformed into temporal target constraints through the explicit motion trajectory constraint module, and the motion path in the synthesis is corrected in real time through gradient feedback to generate an intermediate video latent sequence with dynamics.

[0031] Style latent codes are injected into the intermediate video latent sequence and visual consistency optimization is performed. A multi-attribute joint optimization loop is started to analyze the differences between adjacent latent frames and apply smoothing constraints to reduce inter-frame flicker and jitter. The optimized video latent sequence is then output.

[0032] The optimized video latent sequence is input into the high-definition video latent space decoder, which converts it to pixel space to obtain the initial video background sequence. The sequence is then matched with the overall color tone of the target website's homepage and the video edges are feathered to ensure smooth integration with the foreground page elements, resulting in the final video background sequence.

[0033] In an alternative approach, decoupling the generation of conditional vectors into control subspaces of content, style, and motion further includes:

[0034] The generated conditional vector is input into a shared basic feature analysis network to perform global analysis on the mixed semantic information in the generated conditional vector, and to initially identify and separate the feature responses related to scene entities and spatial relationships, visual appearance and texture attributes, and temporal change patterns.

[0035] The initial latent code is obtained by deep extraction of the feature response. Features strongly correlated with object category, layout position, and geometric structure are extracted through the content pathway. Simultaneously, by using online comparison decoupling loss with style and motion pathways, the representation of appearance texture and dynamic information in the features of this pathway is actively suppressed, and a content latent code representing the static scene composition is output. A style latent code representing the visual appearance is output through the style pathway. A pure temporal dynamic pattern is distilled from the mixed conditions through the motion pathway, and a motion latent code representing the motion law is output.

[0036] For any initial latent code, the properties of other initial latent codes are predicted by a multi-head discriminator network, wherein the properties of other initial latent codes cannot be inferred from any given initial latent code, thus forcing the three initial latent codes to be mutually orthogonal in the feature space.

[0037] The orthogonalized initial latent codes are spliced ​​together and fine-tuned and scaled through a conditional output projection layer to output decoupled and task-adapted content, style, and motion control subspace signals.

[0038] In one alternative approach, incorporating the dynamic background into the homepage of the target website further includes:

[0039] Create a full-screen WebGL canvas or Canvas element as a rendering container, and render the dynamic background into the rendering container;

[0040] Identify and track the visual salience of key interactive elements on the page, analyze the local motion intensity, brightness change frequency and texture complexity of the dynamic background in the key interactive element area, and process the background area of ​​the dynamic background through an adaptive visual noise reduction filter.

[0041] In an alternative approach, computing the regularized optimal transport plan to find the minimum cost projection matrix that maps it to the shared semantic space further includes:

[0042] A cost matrix is ​​formed based on the feature matrix of the executable visual description script in the language domain and the multimodal feature matrix of the page in the visual domain. The cost of the cost matrix is:

[0043]

[0044] Where, x i ,y j Let X and Y be the row vectors of the feature matrix X and the page multimodal feature matrix Y, respectively, for the executable visual description script; α is the balancing weight; ∑ is the covariance matrix estimated from all features; X∈R m×d , Y∈R n×d m and n represent the number of language and visual features, respectively, and d represents the feature dimension.

[0045] Find a transmission plan matrix that minimizes the total transmission cost and is constrained by entropy regularization and mode structure preservation; the objective function of this transmission plan matrix is:

[0046]

[0047] in,<P,C> F For the total transmission cost, The proportion by which the quality of language feature i is allocated to visual feature j; Let a be the feasible region of the transmission plan; a and b be the normalized quality distributions of language and visual features, respectively; H(P) = -∑ i,j P ij (logP ij -1), which is the entropy regularization term of the transmission plan; L X L Y These are the graph Laplacian matrices constructed based on the k-nearest neighbor graphs of features X and Y, respectively. To maintain constraints on the modal structure; P is the transport plan matrix; C is the cost matrix; <.,> F This is the Frobenius inner product;

[0048] The minimum cost projection matrix is ​​obtained by solving the optimization problem of the objective function using an accelerated approximation point algorithm based on Bregman iteration.

[0049] In one alternative approach, a multi-attribute joint optimization loop is initiated to analyze the differences between adjacent latent frames and apply smoothing constraints to reduce inter-frame flicker and jitter. The output optimized video latent sequence further includes:

[0050] Based on the intermediate video latent sequence and the low-resolution pixel frames decoded from the sequence, calculate the forward and backward optical flows between adjacent frames and their respective optical flow confidence masks.

[0051] Based on the forward and backward optical flows and their respective optical flow confidence masks, as well as the consistency loss function of the multi-attribute joint optimization cycle, inter-frame flicker and jitter are reduced and an optimized video latent sequence is output; the consistency loss function includes latent spatial cycle consistency loss, pixel-level optical flow consistency loss, and motion trajectory smoothness loss.

[0052] According to another aspect of this application, an apparatus for generating a first-screen video background based on page elements is provided, comprising:

[0053] The feature and instruction acquisition module is used to acquire the multimodal interface features of the target website's home screen and receive natural language instructions input by the user through an interactive dialog box; wherein, the multimodal interface features include layout structure, visual color, text semantics, and interactive components;

[0054] The feature understanding and condition generation module is used to input the natural language instructions and multimodal interface features into the multimodal alignment and fusion model, analyze the visual intent, style preference and dynamic requirements in the natural language instructions to generate a structured visual control description; and extract layout, color and semantic entities from the multimodal interface features through a cross-modal attention model to form a generated condition vector.

[0055] The video background sequence generation module is used to input the generation condition vector into the decoupled video diffusion model, decouple the generation condition vector into a control subspace of content, style and motion, and synthesize a video background sequence based on the control subspace and the structured visual control description.

[0056] The background compositing module is used to compose a dynamic background based on the video background sequence and to composite the dynamic background onto the homepage of the target website.

[0057] The solution provided by the above embodiments of the present invention acquires the multimodal interface features of the target website's homepage and receives natural language commands input by the user through an interactive dialog box. The multimodal interface features include layout structure, visual color, text semantics, and interactive components. The natural language commands and multimodal interface features are input into a multimodal alignment and fusion model to parse the visual intent, style preferences, and dynamic requirements in the natural language commands to generate a structured visual control description. A cross-modal attention model extracts layout, color, and semantic entities from the multimodal interface features to form a generation condition vector. The generation condition vector is input into a conditional decoupling video diffusion model, decoupling the generation condition vector into a control subspace of content, style, and motion. A video background sequence is synthesized based on the control subspace and the structured visual control description. A dynamic background is synthesized based on the video background sequence and then composited into the target website's homepage. This invention generates dynamic video backgrounds based on the webpage interface itself and guided by user language, avoiding the coarse control defects of general text-based video models. Specifically, through multimodal alignment and fusion models and cross-modal attention models, the system not only understands the abstract intent of users' natural language commands but also deeply integrates the layout structure, visual colors, textual semantics, and interactive component features of the target website's homepage. This ensures that the generated video background is highly coordinated with the original page content in terms of composition, tone, and theme. A conditional decoupling video diffusion model is employed to decouple the generation conditions into three independent control subspaces: content, style, and motion. This enables the system to meet complex dynamic requirements in user commands while maintaining visual quality. By using optimal transport projection in a shared semantic space and a feature untangling fusion method based on causal intervention, the system achieves efficient conversion from heterogeneous page features to standardized generation conditions, significantly lowering the technical threshold for creating high-quality dynamic backgrounds.

[0058] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above description and other objects, features and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0059] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0060] Figure 1 A flowchart illustrating a method for generating a first-screen video background based on page elements according to an embodiment of the present invention is shown.

[0061] Figure 2A schematic diagram of the multimodal alignment and fusion process according to an embodiment of the present invention is shown;

[0062] Figure 3 A flowchart illustrating the cross-modal attention and causal untangling module according to an embodiment of the present invention is shown;

[0063] Figure 4 This diagram illustrates the process of generating conditionally decoupled video diffusion according to an embodiment of the present invention.

[0064] Figures 5a to 5b This illustration shows an embodiment of the present invention for generating a video background based on page elements. Figure 1 ;

[0065] Figures 6a to 6c This illustration shows an embodiment of the present invention for generating a video background based on page elements. Figure 2 ;

[0066] Figure 7 The diagram illustrates the functional structure of a method device for generating a first-screen video background based on page elements according to an embodiment of the present invention. Detailed Implementation

[0067] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0068] The following detailed embodiments illustrate the method and apparatus for generating a first-screen video background based on page elements proposed in this invention.

[0069] Example 1:

[0070] Figure 1 This diagram illustrates the functional structure of a method for generating a first-screen video background based on page elements, according to an embodiment of the present invention. Specifically, as shown... Figure 1 As shown, it includes the following steps:

[0071] Step S101: Obtain the multimodal interface features of the target website's home screen and receive natural language commands input by the user through an interactive dialog box; wherein, the multimodal interface features include layout structure, visual colors, text semantics, and interactive components.

[0072] In this embodiment, by acquiring the layout structure, visual colors, text semantics, and interactive components, a structured understanding of the target website's homepage can be achieved. The generated video background automatically adapts to the overall page design, avoiding information gaps caused by relying solely on screenshots or HTML structures. Furthermore, natural language commands ensure that the generated results conform to both objective page constraints and the user's subjective needs.

[0073] Specifically, the page's DOM tree is parsed, calculating the absolute / relative position, size (width and height), hierarchy, and arrangement (Flexbox / Grid layout) of each visible element (such as div, section, button) to construct the page's spatial topology. Pixel-level analysis is performed on the first screen screenshot to extract the distribution of primary, secondary, and contrasting colors, the main color palette, and color hues (such as warm / cool tones). All visible text content (titles, paragraphs, link text) within the first screen is extracted, and keyword extraction, theme summarization, and sentiment analysis are performed to understand the page's core theme and atmosphere. UI elements such as buttons, forms, navigation bars, and carousels are identified. Natural language commands input by the user (typed or voice input) are received through floating dialog boxes or sidebar input boxes.

[0074] Step S102: Input the natural language instructions and multimodal interface features into the multimodal alignment and fusion model to analyze the visual intent, style preferences and dynamic requirements in the natural language instructions to generate a structured visual control description; extract layout, color and semantic entities from the multimodal interface features through a cross-modal attention model to form a generation condition vector.

[0075] In this embodiment, the user's vague and abstract natural language descriptions (such as "full of technological feel" and "dynamic and natural") are parsed and decomposed into structured, executable visual control descriptions, transforming subjective user preferences into an objective set of instructions containing specific attributes, constraints, and hierarchical structures. The cross-modal attention model does not simply stack all page features (layout, color, semantics, etc.) together, but rather extracts the layout, color, and semantic entities most relevant to the generation task from a massive amount of features. For example, when generating a background, it focuses more on the page's main color tone and overall composition, ensuring that the generated conditional vector filters out irrelevant details, significantly improving the fit between the generated content and the page's theme. The generated conditional vector contains the video's intent, style, and dynamics, and provides specific color values, spatial coordinates, and semantic labels, ensuring that the user's natural language instructions and the page's presentation (multimodal features) are associated in the same semantic space.

[0076] In one alternative approach, the multimodal alignment and fusion model includes an instruction parsing and knowledge retrieval module, a heterogeneous feature encoding module, a shared semantic space optimal transmission projection module, a multi-stage alignment and adaptive fusion module, and a conditional vector refinement output module.

[0077] The instruction parsing and knowledge retrieval module is used to receive the natural language instruction and decompose and reason it, and simultaneously perform multi-hop retrieval from the Web visual design pattern library and motion graphics parameter library, and output an executable visual description script including hierarchical structure, abstract intent, specific attributes and constraints.

[0078] The heterogeneous feature encoding module, including a layout encoder, a color encoder, and a unified encoder, runs in parallel with the instruction parsing and knowledge retrieval module and is used to receive the multimodal interface features. The layout encoder constructs a directed heterogeneous graph of interface elements and their spatial relationships, encoding global typographic semantics through message passing and graph attention mechanisms. The color encoder calculates color harmony and prominence indices in the CIELAB color space and generates sentiment semantic embeddings. The unified encoder jointly encodes textual semantics and interactive component images to extract their joint functional and thematic representations.

[0079] The shared semantic space optimal transport projection module receives the executable visual description script and the modal features encoded by the heterogeneous feature encoding module, respectively; it treats the description script from the language domain and the page features from the visual domain as two distributions, and finds the minimum cost projection matrix that maps them to the shared semantic space by calculating the regularized optimal transport plan.

[0080] The multi-stage alignment and adaptive fusion module includes a multi-granularity semantic routing network, used to receive the minimum cost projection matrix and filter relevant page feature modal subsets according to the abstract intent in the description script;

[0081] The condition vector refinement output module is used to output the generated condition vector based on the subset of page feature modes.

[0082] In this embodiment, as Figure 2As shown, by using optimal transport projection in a shared semantic space, user language commands and page visual features are treated as distributions of two different languages. Optimal transport seeks the minimum-cost solution to map both to the same semantic space, fundamentally avoiding misunderstandings caused by modal differences. The command parsing and knowledge retrieval module performs multi-hop searches from external libraries (Web design pattern library, motion graphics parameter library), significantly enriching the generated details and ensuring a more professional effect. The local encoder understands the spatial relationships and visual hierarchy between elements, acquiring typographic semantics. The color encoder (based on CIELAB) operates in a color space more consistent with human visual perception, quantifying harmony and sentiment rather than just RGB values. The unified encoder (multimodal joint encoding) links text and component images. The multi-stage alignment and adaptive fusion module dynamically selects the most relevant subset of page features (such as the brand's primary color) based on the abstract intent of the current command, avoiding information overload and noise interference, making the generated conditional vector highly relevant to the task.

[0083] Specifically, user commands and page features are fed in parallel into two processing branches: one branch (command parsing) to understand what is wanted, and the other branch (heterogeneous encoding) to analyze what is currently available. A parser based on a large language model decomposes commands into structured elements and simultaneously queries a knowledge base. For example, when parsing for a "tech-savvy" feel, it retrieves halo parameters to supplement the description script. Layout information is transformed into a graph data structure, and deep relationships between elements are learned through a graph neural network. Color information is analyzed in the CIELAB space to determine hue harmony, which color is most prominent, and mapped to tags such as "professional." Text and component images are jointly encoded using a visual-language model. Cross-modal optimal transmission alignment uses the processed language description feature matrix and visual feature matrix to solve an optimal transmission problem with regularization terms, calculating how to project them into the same shared semantic space with minimal semantic distortion to generate an alignment mapping relationship. The routing network decides to primarily adopt color and texture features based on the command intent, weakening layout features, and finally outputs a generation condition vector that integrates user intent and page characteristics.

[0084] In one alternative approach, the cross-modal attention model includes a multimodal tensor quantized feature encoder, a hierarchical semantic routing and attention network, a feature unwrapping and fusion module based on causal intervention, and an adaptive generative conditional projector.

[0085] The multimodal tensor quantization feature encoder includes a layout neural coding submodule, a perception and emotion quantization coding submodule, and a multi-granularity visual alignment coding submodule. The layout neural coding submodule constructs a dynamic hypergraph from the geometric attributes, spatial relationships, and visual hierarchy of interface elements. The hyperedges of the dynamic hypergraph dynamically represent visual balance groups, reading flow sequences, or interactive function clusters, and are encoded and output as a layout tensor with explicit structural semantics through high-order message passing. The perception and emotion quantization coding submodule simulates human eye perception in the ICtCp color space based on a color adaptation transformation model, calculates the visual saliency, emotional arousal, and value of colors, and decomposes the overall color distribution into three factor matrices—base, emotion, and emphasis—through non-negative tensor decomposition to output a structured color emotion tensor. The multi-granularity visual alignment coding submodule establishes the association between text descriptions, component visual screenshots, and general vision, and outputs a semantic consistency tensor.

[0086] The hierarchical semantic routing and attention network includes a semantic-aware routing layer and an intra-channel modulation attention layer; wherein, the semantic-aware routing layer analyzes the meta-semantics of each input tensor and queries a learnable cross-modal interaction prior codebook to dynamically generate a set of routing weight vectors; the intra-channel modulation attention layer performs conditional multi-head cross-attention in each channel;

[0087] The feature untangling fusion module based on causal intervention is connected to the hierarchical semantic routing and attention network. It performs causal graph discovery on the fused feature tensor and infers the potential causal relationships between different semantic dimensions. Through do-calculus intervention, it forcibly cuts off non-causal statistical correlation paths and strengthens the transmission of causal paths at the feature level, and outputs a set of untangled causal semantic features.

[0088] The adaptive generating conditional projector receives the unwrapped causal semantic feature set and includes a multi-head conditional linear projection layer.

[0089] In this embodiment, as Figure 3As shown, through the layout neural coding submodule, interface elements are transformed into dynamic hypergraphs, capable of capturing complex geometric attributes, spatial relationships, and visual hierarchies. Higher-order message passing helps the model understand visual balance groups, reading flow sequences, or clusters of interactive functions, thereby generating layout tensors that better align with human cognitive habits. The perception and emotion quantization coding submodule operates in the ICtCp color space and utilizes a color adaptation transformation model to simulate the human eye's perception process, improving the accuracy of color analysis and calculating the emotional value of colors, thus contributing to the generation of video backgrounds that resonate with users. The multi-granularity visual alignment coding submodule ensures strong correlation between text descriptions, component visual screenshots, and general vision, making the generated video background consistent and coherent with the actual page content. Hierarchical semantic routing and semantic perception routing in the attention network dynamically adjust the importance of different modal features, and the intra-channel modulation attention layer further refines the feature interaction methods within each channel, enabling adaptive performance optimization in different application scenarios. The feature untangling fusion module based on causal intervention discovers causal graphs on the fused features, identifies and strengthens truly meaningful causal paths, reduces the influence of irrelevant variables, and thus improves the quality of the generated results. The adaptive generating conditional projector ultimately integrates all the above information into a generating conditional vector that is highly adapted to the specific task requirements, providing precise control signals for video generation.

[0090] Specifically, layout neural encoding abstracts page elements (such as div, img, and button) into graph nodes, and spatial relationships between elements (such as top / bottom, containment, and alignment) constitute edges. It identifies elements that visually form a whole (such as a card and its internal elements), forming hyperedges. Through message passing in the graph neural network, it outputs a layout tensor representing how the page is organized and how the eye moves. Perceptual sentiment quantization encoding analyzes page screenshots in the ICtCp color space, evaluating the saliency of color regions, the emotional arousal of the overall tone, and the value of colors. Tensor decomposition separates this information into a primary tone matrix, a sentiment matrix, and an emphasis matrix, which together constitute the color sentiment tensor. Multi-granularity visual alignment encoding inputs text (such as purchase information) and component screenshots from the page into a pre-trained visual-language model (such as CLIP), outputting a semantic consistency tensor that represents the joint embedding of text descriptions and their corresponding visual forms in a general semantic space. After inputting three tensors (layout, color, and semantics), the semantic-aware routing network analyzes their respective characteristics (e.g., the current color tensor exhibits high "emotional arousal") and queries the codebook learned from a large amount of design data to return multiple sets of weights, indicating the emotions and layouts that should be focused on in the current task. The intra-channel modulated attention layer performs cross-calculation at the feature channel level based on the multiple sets of weights, allowing the weighted features to interact and fuse into a comprehensive feature tensor. Causal discovery is performed on the initially fused comprehensive feature tensor to establish a causal graph between feature dimensions. Subsequently, do-calculus intervention is performed, whereby a certain feature is fixed or intervened within the model (e.g., setting "theme" to "natural"), and other features (e.g., "hue") are observed to change, thereby severing statistically related but non-causal associations. After intervention, several unentangled causal semantic feature groups are output, such as: [theme concept group], [spatial structure group], [emotional hue group], and [texture material group]. The unwrapped feature set is mapped to the conditional vector space desired by the video generation model through a learnable linear projection layer, and scale and dimension are calibrated to form the final generation conditional vector.

[0091] In an alternative approach, computing the regularized optimal transport plan to find the minimum cost projection matrix that maps it to the shared semantic space further includes:

[0092] A cost matrix is ​​formed based on the feature matrix of the executable visual description script in the language domain and the multimodal feature matrix of the page in the visual domain. The cost of the cost matrix is:

[0093]

[0094] Where, x i ,y jLet X and Y be the row vectors of the feature matrix X and the page multimodal feature matrix Y, respectively, for the executable visual description script; α is the balancing weight; ∑ is the covariance matrix estimated from all features; X∈R m×d , Y∈R n×d m and n represent the number of language and visual features, respectively, and d represents the feature dimension.

[0095] Find a transmission plan matrix that minimizes the total transmission cost and is constrained by entropy regularization and mode structure preservation; the objective function of this transmission plan matrix is:

[0096]

[0097] in,<P,C> F For the total transmission cost, The proportion by which the quality of language feature i is allocated to visual feature j; Let a be the feasible region of the transmission plan; a and b be the normalized quality distributions of language and visual features, respectively; H(P) = -∑ i,j P i,j (logP ij -1), which is the entropy regularization term of the transmission plan; L X L Y These are the graph Laplacian matrices constructed based on the k-nearest neighbor graphs of features X and Y, respectively. To maintain constraints on the modal structure; P is the transport plan matrix; C is the cost matrix; <.,> F This is the Frobenius inner product;

[0098] The minimum cost projection matrix is ​​obtained by solving the optimization problem of the objective function using an accelerated approximation point algorithm based on Bregman iteration.

[0099] In this embodiment, instead of performing black-box feature fusion, the feature alignment problem between language and visual modalities is modeled as an optimal transmission problem with a well-defined mathematical framework. The goal is to find the optimal solution for transferring the language feature distribution to the visual feature distribution with minimal semantic cost, making the alignment process measurable and interpretable. Minimizing the total transmission cost aims for globally optimal alignment, while entropy regularization prevents the transmission plan from becoming too concentrated or deterministic (i.e., avoiding a language feature matching only one visual feature). Modal structure preservation constraints require that the aligned feature relationships should, as far as possible, maintain the structure within their respective original modalities. That is, if "luxury" and "refined" are synonyms in the language description (closely related in the graph Laplacian matrix), then the matched visual features should also be closely related in the visual space. This ensures that cross-modal alignment does not disrupt the inherent logic of each modality, achieving structure-preserving semantic mapping. The accelerated approximation point algorithm based on Bregman iteration can effectively solve large-scale feature matrices, ensuring real-time or near-real-time interactive scenarios.

[0100] Step S103: Input the generated conditional vector into the conditional decoupled video diffusion model, decouple the generated conditional vector into a control subspace of content, style and motion, and synthesize a video background sequence based on the control subspace and the structured visual control description.

[0101] In this embodiment, the decoupling of content, style, and motion ensures that the background of the generated video is consistent with the layout, colors, and user intent of the target website's homepage. Figure 1 The content subspace ensures spatial and semantic alignment between the generated scene subject and page elements; the style subspace ensures that visual attributes such as color and texture are coordinated with the overall visual style of the page; the motion subspace ensures that dynamic effects meet the user's described needs, thereby improving the overall integration of the video background with the page environment. The motion subspace can more accurately capture and generate temporal dynamic patterns (such as motion trajectories and speed changes) and perform real-time corrections in conjunction with the motion trajectory constraint module, effectively reducing common problems in video generation such as inter-frame flickering and jitter. Multi-attribute joint optimization can further improve the temporal smoothness and visual coherence of the video. The decoupled control subspace can be combined with a staged scheduling controller to dynamically adjust the injection weights of each subspace at different generation stages (such as keyframe static composition, dynamic effect compositing, and visual consistency optimization), thereby generating higher-quality video sequences that better meet expectations. Figures 5a to 5b As shown, dynamic backgrounds (videos) are generated based on web page elements and natural language instructions.

[0102] In one alternative approach, the conditional decoupled video diffusion model includes a conditional decoupled encoder, a temporal fusion module, a spatiotemporal decoupled injection module, a staged scheduling controller, and a video latent space synthesis network connected in sequence.

[0103] The conditional decoupling encoder receives the generated conditional vector, which includes a content path, a style path, and a motion path set in parallel. The content path consists of a residual convolutional network and a self-attention layer. It maximizes the mutual information between the content and the entities and composition in the generated conditional vector to impose orthogonality constraints on the style and motion information, and outputs a content latent code representing the main subject and static layout of the scene. The style path uses an adaptive multilayer perceptron to extract information strongly related to appearance attributes and optimizes it by contrast decoupling loss with the output of the content path, outputting a time-independent style latent code. The motion path consists of a one-dimensional temporal convolutional network and a gated recurrent unit. It captures the dynamic patterns and evolutionary trends contained in the generated conditional vector and outputs a motion latent code representing the temporal change pattern.

[0104] The timing fusion module is connected to the conditional decoupling encoder, receives the motion latent code and dynamic parameters in the structured visual control description, and outputs the enhanced timing condition signal.

[0105] The spatiotemporal decoupling injection module adopts a U-ViT structure, starting with the latent representation of random Gaussian noise video, and receives feature maps from the previous layer in each level of the U-ViT structure.

[0106] The staged scheduling controller, as a global control logic unit, receives the time step signal of the denoising process and dynamically adjusts the relative weights of the feature maps at different stages of denoising.

[0107] The video latent space synthesis network outputs a video background latent sequence, which is then mapped by a decoder to obtain a video background sequence in pixel space.

[0108] In this embodiment, as Figure 4 As shown, by setting three dedicated paths—content, style, and motion—in parallel within the conditional decoupling encoder and applying orthogonality constraints (such as maximizing mutual information and contrastive decoupling loss), the hybrid generative conditional vector can be decoupled into semantically clear content latent codes, style latent codes, and motion latent codes, as shown. Figures 6a to 6c As shown, this allows users to control the static composition, visual appearance, and dynamic mode of the video separately, improving the flexibility and controllability of the generation process. Combining the spatiotemporal decoupling injection module of the U-ViT structure with a staged scheduling controller, the injection weights of content, style, motion, and other conditions can be dynamically adjusted at different time steps in the denoising (generation) process, making the generation process more stable. This results in a smooth dynamic overlay after a static composition is formed, reducing global / local flicker and motion distortion issues in the video.

[0109] In this embodiment, synthesizing a video background sequence based on the control subspace and the structured visual control description further includes:

[0110] Using the content latent code and style latent code as conditions, an initial video noise latent sequence is generated by combining randomly sampled Gaussian noise. The scene subject and static layout information carried by the content latent code are injected into the U-ViT network with high weights through the spatiotemporal decoupling injection module, driving the network to predict the key frame static composition of the video sequence, so as to ensure that the generated background matches the page layout intention and the abstract intention specified by the user, thereby outputting a basic video latent sequence draft.

[0111] The basic video latent sequence draft is used for dynamic effect synthesis. In this process, the phased scheduling controller transfers the control of the generation conditions to the enhanced temporal condition signal. The spatiotemporal decoupling injection module synthesizes dynamic elements that conform to the description on the basic video latent sequence draft based on the enhanced temporal condition signal. At the same time, the dynamic parameters in the description are transformed into temporal target constraints through the explicit motion trajectory constraint module, and the motion path in the synthesis is corrected in real time through gradient feedback to generate an intermediate video latent sequence with dynamics.

[0112] Style latent codes are injected into the intermediate video latent sequence and visual consistency optimization is performed. A multi-attribute joint optimization loop is started to analyze the differences between adjacent latent frames and apply smoothing constraints to reduce inter-frame flicker and jitter. The optimized video latent sequence is then output.

[0113] The optimized video latent sequence is input into the high-definition video latent space decoder, which converts it to pixel space to obtain the initial video background sequence. The sequence is then matched with the overall color tone of the target website's homepage and the video edges are feathered to ensure smooth integration with the foreground page elements, resulting in the final video background sequence.

[0114] In one alternative approach, a multi-attribute joint optimization loop is initiated to analyze the differences between adjacent latent frames and apply smoothing constraints to reduce inter-frame flicker and jitter. The output optimized video latent sequence further includes:

[0115] Based on the intermediate video latent sequence and the low-resolution pixel frames decoded from the sequence, calculate the forward and backward optical flows between adjacent frames and their respective optical flow confidence masks.

[0116] Based on the forward and backward optical flows and their respective optical flow confidence masks, as well as the consistency loss function of the multi-attribute joint optimization cycle, inter-frame flicker and jitter are reduced and an optimized video latent sequence is output; the consistency loss function includes latent spatial cycle consistency loss, pixel-level optical flow consistency loss, and motion trajectory smoothness loss.

[0117] In this embodiment, the latent spatial cycle consistency loss ensures the temporal smoothness of deep learning features, reducing semantic incoherence at its source. The pixel-level optical flow consistency loss eliminates visible flickering and jumps at the visual level, while the motion trajectory smoothness loss avoids unnatural abrupt changes. Forward and backward optical flow mutually verify each other, improving the accuracy of motion estimation. Confidence masks distinguish reliable and unreliable optical flow regions, avoiding erroneous optimization in occluded or blurred areas.

[0118] In an alternative approach, decoupling the generation of conditional vectors into control subspaces of content, style, and motion further includes:

[0119] The generated conditional vector is input into a shared basic feature analysis network to perform global analysis on the mixed semantic information in the generated conditional vector, and to initially identify and separate the feature responses related to scene entities and spatial relationships, visual appearance and texture attributes, and temporal change patterns.

[0120] The initial latent code is obtained by deep extraction of the feature response. Features strongly correlated with object category, layout position, and geometric structure are extracted through the content pathway. Simultaneously, by using online comparison decoupling loss with style and motion pathways, the representation of appearance texture and dynamic information in the features of this pathway is actively suppressed, and a content latent code representing the static scene composition is output. A style latent code representing the visual appearance is output through the style pathway. A pure temporal dynamic pattern is distilled from the mixed conditions through the motion pathway, and a motion latent code representing the motion law is output.

[0121] For any initial latent code, the properties of other initial latent codes are predicted by a multi-head discriminator network, wherein the properties of other initial latent codes cannot be inferred from any given initial latent code, thus forcing the three initial latent codes to be mutually orthogonal in the feature space.

[0122] The orthogonalized initial latent codes are spliced ​​together and fine-tuned and scaled through a conditional output projection layer to output decoupled and task-adapted content, style, and motion control subspace signals.

[0123] In this embodiment, the semantic information required for video background generation is explicitly divided into three subspaces: content (static layout), style (visual appearance), and motion (dynamic mode). Each subspace can be adjusted independently without interfering with other dimensions. For example, only the dynamic rhythm can be modified while maintaining the overall composition and color. By comparing the decoupling loss with a multi-head discriminator online, it is ensured that the content latent code does not carry style or motion features, thereby improving the semantic consistency of the generated results. The decoupled subspace signals are more structured and task-adaptable, better adapting to natural language instructions and page features, avoiding generation ambiguity or distortion caused by modal mixing.

[0124] Step S104: Synthesize a dynamic background based on the video background sequence, and composite the dynamic background onto the homepage of the target website.

[0125] In one alternative approach, incorporating the dynamic background into the homepage of the target website further includes:

[0126] Create a full-screen WebGL canvas or Canvas element as a rendering container, and render the dynamic background into the rendering container;

[0127] Identify and track the visual salience of key interactive elements on the page, analyze the local motion intensity, brightness change frequency and texture complexity of the dynamic background in the key interactive element area, and process the background area of ​​the dynamic background through an adaptive visual noise reduction filter.

[0128] In this embodiment, by identifying key interactive elements (such as buttons, input boxes, navigation bars, etc.) on the page and analyzing their visual salience, visual obscuring or distraction caused by overly active backgrounds is avoided. For key interactive areas, the motion intensity, brightness flicker frequency, and texture complexity of the background are dynamically analyzed, and an adaptive visual noise reduction filter is applied to reduce visual noise and improve interface clarity. By differentiating the background areas (staticating interactive areas and retaining dynamics in non-interactive areas), a natural transition between dynamic backgrounds and static foreground elements is achieved, enhancing overall visual unity.

[0129] The solution provided by the above embodiments of the present invention acquires the multimodal interface features of the target website's homepage and receives natural language commands input by the user through an interactive dialog box. The multimodal interface features include layout structure, visual color, text semantics, and interactive components. The natural language commands and multimodal interface features are input into a multimodal alignment and fusion model to parse the visual intent, style preferences, and dynamic requirements in the natural language commands to generate a structured visual control description. A cross-modal attention model extracts layout, color, and semantic entities from the multimodal interface features to form a generation condition vector. The generation condition vector is input into a conditional decoupling video diffusion model, decoupling the generation condition vector into a control subspace of content, style, and motion. A video background sequence is synthesized based on the control subspace and the structured visual control description. A dynamic background is synthesized based on the video background sequence and then composited into the target website's homepage. This invention generates dynamic video backgrounds based on the webpage interface itself and guided by user language, avoiding the coarse control defects of general text-based video models. Specifically, through multimodal alignment and fusion models and cross-modal attention models, the system not only understands the abstract intent of users' natural language commands but also deeply integrates the layout structure, visual colors, textual semantics, and interactive component features of the target website's homepage. This ensures that the generated video background is highly coordinated with the original page content in terms of composition, tone, and theme. A conditional decoupling video diffusion model is employed to decouple the generation conditions into three independent control subspaces: content, style, and motion. This enables the system to meet complex dynamic requirements in user commands while maintaining visual quality. By using optimal transport projection in a shared semantic space and a feature untangling fusion method based on causal intervention, the system achieves efficient conversion from heterogeneous page features to standardized generation conditions, significantly lowering the technical threshold for creating high-quality dynamic backgrounds.

[0130] Example 2:

[0131] Figure 7 This diagram illustrates the functional structure of a method apparatus for generating a first-screen video background based on page elements, according to an embodiment of the present invention. Figure 7 As shown, the device includes:

[0132] The feature and instruction acquisition module 701 is used to acquire the multimodal interface features of the target website's home screen and receive natural language instructions input by the user through an interactive dialog box; wherein, the multimodal interface features include layout structure, visual color, text semantics, and interactive components;

[0133] The feature understanding and condition generation module 702 is used to input the natural language instructions and multimodal interface features into the multimodal alignment and fusion model, analyze the visual intent, style preference and dynamic requirements in the natural language instructions to generate a structured visual control description; and extract layout, color and semantic entities from the multimodal interface features through a cross-modal attention model to form a generated condition vector.

[0134] The video background sequence generation module 703 is used to input the generation condition vector into the conditional decoupling video diffusion model, decouple the generation condition vector into a control subspace of content, style and motion, and synthesize a video background sequence based on the control subspace and the structured visual control description.

[0135] Background compositing module 704 is used to compose a dynamic background based on the video background sequence and to composite the dynamic background in the homepage of the target website.

[0136] The solution provided by the above embodiments of the present invention acquires the multimodal interface features of the target website's homepage and receives natural language commands input by the user through an interactive dialog box. The multimodal interface features include layout structure, visual color, text semantics, and interactive components. The natural language commands and multimodal interface features are input into a multimodal alignment and fusion model to parse the visual intent, style preferences, and dynamic requirements in the natural language commands to generate a structured visual control description. A cross-modal attention model extracts layout, color, and semantic entities from the multimodal interface features to form a generation condition vector. The generation condition vector is input into a conditional decoupling video diffusion model, decoupling the generation condition vector into a control subspace of content, style, and motion. A video background sequence is synthesized based on the control subspace and the structured visual control description. A dynamic background is synthesized based on the video background sequence and then composited into the target website's homepage. This invention generates dynamic video backgrounds based on the webpage interface itself and guided by user language, avoiding the coarse control defects of general text-based video models. Specifically, through multimodal alignment and fusion models and cross-modal attention models, the system not only understands the abstract intent of users' natural language commands but also deeply integrates the layout structure, visual colors, textual semantics, and interactive component features of the target website's homepage. This ensures that the generated video background is highly coordinated with the original page content in terms of composition, tone, and theme. A conditional decoupling video diffusion model is employed to decouple the generation conditions into three independent control subspaces: content, style, and motion. This enables the system to meet complex dynamic requirements in user commands while maintaining visual quality. By using optimal transport projection in a shared semantic space and a feature untangling fusion method based on causal intervention, the system achieves efficient conversion from heterogeneous page features to standardized generation conditions, significantly lowering the technical threshold for creating high-quality dynamic backgrounds.

[0137] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of the present invention are not directed to any particular programming language. It should be understood that the content of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the invention.

[0138] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0139] Similarly, it should be understood that, in order to simplify the invention and aid in understanding one or more of the various inventive aspects, features of the embodiments of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the above description of exemplary embodiments of the invention. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.

[0140] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0141] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0142] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0143] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A method for generating a first-screen video background based on page elements, characterized in that, include: The system acquires the multimodal interface features of the target website's home screen and receives natural language commands input by the user through an interactive dialog box; wherein, the multimodal interface features include layout structure, visual colors, text semantics, and interactive components; The natural language instructions and multimodal interface features are input into a multimodal alignment and fusion model to analyze the visual intent, style preferences and dynamic requirements in the natural language instructions to generate a structured visual control description; layout, color and semantic entities are extracted from the multimodal interface features through a cross-modal attention model to form a generation condition vector; The generated conditional vector is input into the conditional decoupled video diffusion model, and the generated conditional vector is decoupled into a control subspace of content, style and motion. A video background sequence is synthesized based on the control subspace and the structured visual control description. A dynamic background is synthesized based on the video background sequence and then composited onto the homepage of the target website.

2. The method for generating a first-screen video background based on page elements according to claim 1, characterized in that, The multimodal alignment and fusion model includes an instruction parsing and knowledge retrieval module, a heterogeneous feature encoding module, a shared semantic space optimal transmission projection module, a multi-stage alignment and adaptive fusion module, and a conditional vector refinement output module. The instruction parsing and knowledge retrieval module is used to receive the natural language instruction and decompose and reason it, and simultaneously perform multi-hop retrieval from the Web visual design pattern library and motion graphics parameter library, and output an executable visual description script including hierarchical structure, abstract intent, specific attributes and constraints. The heterogeneous feature encoding module, including a layout encoder, a color encoder, and a unified encoder, runs in parallel with the instruction parsing and knowledge retrieval module and is used to receive the multimodal interface features. The layout encoder constructs a directed heterogeneous graph of interface elements and their spatial relationships, encoding global typographic semantics through message passing and graph attention mechanisms. The color encoder calculates color harmony and prominence indices in the CIELAB color space and generates sentiment semantic embeddings. The unified encoder jointly encodes textual semantics and interactive component images to extract their joint functional and thematic representations. The shared semantic space optimal transport projection module receives the executable visual description script and the modal features encoded by the heterogeneous feature encoding module, respectively; it treats the description script from the language domain and the page features from the visual domain as two distributions, and finds the minimum cost projection matrix that maps them to the shared semantic space by calculating the regularized optimal transport plan. The multi-stage alignment and adaptive fusion module includes a multi-granularity semantic routing network, used to receive the minimum cost projection matrix and filter relevant page feature modal subsets according to the abstract intent in the description script; The condition vector refinement output module is used to output the generated condition vector based on the subset of page feature modes.

3. The method for generating a first-screen video background based on page elements according to claim 1, characterized in that, The cross-modal attention model includes a multimodal tensorized feature encoder, a hierarchical semantic routing and attention network, a feature unwrapping and fusion module based on causal intervention, and an adaptive generative conditional projector. The multimodal tensor quantization feature encoder includes a layout neural coding submodule, a perception and emotion quantization coding submodule, and a multi-granularity visual alignment coding submodule. The layout neural coding submodule constructs a dynamic hypergraph from the geometric attributes, spatial relationships, and visual hierarchy of interface elements. The hyperedges of the dynamic hypergraph dynamically represent visual balance groups, reading flow sequences, or interactive function clusters, and are encoded and output as a layout tensor with explicit structural semantics through high-order message passing. The perception and emotion quantization coding submodule simulates human eye perception in the ICtCp color space based on a color adaptation transformation model, calculates the visual saliency, emotional arousal, and value of colors, and decomposes the overall color distribution into three factor matrices—base, emotion, and emphasis—through non-negative tensor decomposition to output a structured color emotion tensor. The multi-granularity visual alignment coding submodule establishes the association between text descriptions, component visual screenshots, and general vision, and outputs a semantic consistency tensor. The hierarchical semantic routing and attention network includes a semantic-aware routing layer and an intra-channel modulation attention layer; wherein, the semantic-aware routing layer analyzes the meta-semantics of each input tensor and queries a learnable cross-modal interaction prior codebook to dynamically generate a set of routing weight vectors; the intra-channel modulation attention layer performs conditional multi-head cross-attention in each channel; The feature untangling fusion module based on causal intervention is connected to the hierarchical semantic routing and attention network. It performs causal graph discovery on the fused feature tensor and infers the potential causal relationships between different semantic dimensions. Through do-calculus intervention, it forcibly cuts off non-causal statistical correlation paths and strengthens the transmission of causal paths at the feature level, and outputs a set of untangled causal semantic features. The adaptive generating conditional projector receives the unwrapped causal semantic feature set and includes a multi-head conditional linear projection layer.

4. The method for generating a first-screen video background based on page elements according to claim 1, characterized in that, The conditional decoupled video diffusion model includes a conditional decoupled encoder, a temporal fusion module, a spatiotemporal decoupled injection module, a staged scheduling controller, and a video latent space synthesis network connected in sequence. The conditional decoupling encoder receives the generated conditional vector, which includes a content path, a style path, and a motion path set in parallel. The content path consists of a residual convolutional network and a self-attention layer. It maximizes the mutual information between the content and the entities and composition in the generated conditional vector to impose orthogonality constraints on the style and motion information, and outputs a content latent code representing the main subject and static layout of the scene. The style path uses an adaptive multilayer perceptron to extract information strongly related to appearance attributes and optimizes it by contrast decoupling loss with the output of the content path, outputting a time-independent style latent code. The motion path consists of a one-dimensional temporal convolutional network and a gated recurrent unit. It captures the dynamic patterns and evolutionary trends contained in the generated conditional vector and outputs a motion latent code representing the temporal change pattern. The timing fusion module is connected to the conditional decoupling encoder, receives the motion latent code and dynamic parameters in the structured visual control description, and outputs the enhanced timing condition signal. The spatiotemporal decoupling injection module adopts a U-ViT structure, starting with the latent representation of random Gaussian noise video, and receives feature maps from the previous layer in each level of the U-ViT structure. The staged scheduling controller, as a global control logic unit, receives the time step signal of the denoising process and dynamically adjusts the relative weights of the feature maps at different stages of denoising. The video latent space synthesis network outputs a video background latent sequence, which is then mapped by a decoder to obtain a video background sequence in pixel space.

5. The method for generating a first-screen video background based on page elements according to claim 4, characterized in that, The synthesized video background sequence based on the control subspace and the structured visual control description further includes: Using the content latent code and style latent code as conditions, an initial video noise latent sequence is generated by combining randomly sampled Gaussian noise. The scene subject and static layout information carried by the content latent code are injected into the U-ViT network with high weights through the spatiotemporal decoupling injection module, driving the network to predict the key frame static composition of the video sequence, so as to ensure that the generated background matches the page layout intention and the abstract intention specified by the user, thereby outputting a basic video latent sequence draft. The basic video latent sequence draft is used for dynamic effect synthesis. In this process, the phased scheduling controller transfers the control of the generation conditions to the enhanced temporal condition signal. The spatiotemporal decoupling injection module synthesizes dynamic elements that conform to the description on the basic video latent sequence draft based on the enhanced temporal condition signal. At the same time, the dynamic parameters in the description are transformed into temporal target constraints through the explicit motion trajectory constraint module, and the motion path in the synthesis is corrected in real time through gradient feedback to generate an intermediate video latent sequence with dynamics. Style latent codes are injected into the intermediate video latent sequence and visual consistency optimization is performed. A multi-attribute joint optimization loop is started to analyze the differences between adjacent latent frames and apply smoothing constraints to reduce inter-frame flicker and jitter. The optimized video latent sequence is then output. The optimized video latent sequence is input into the high-definition video latent space decoder, which converts it to pixel space to obtain the initial video background sequence. The sequence is then matched with the overall color tone of the target website's homepage and the video edges are feathered to ensure smooth integration with the foreground page elements, resulting in the final video background sequence.

6. The method for generating a first-screen video background based on page elements according to claim 4, characterized in that, The decoupling of the generated conditional vectors into control subspaces for content, style, and motion further includes: The generated conditional vector is input into a shared basic feature analysis network to perform global analysis on the mixed semantic information in the generated conditional vector, and to initially identify and separate the feature responses related to scene entities and spatial relationships, visual appearance and texture attributes, and temporal change patterns. The initial latent code is obtained by deep extraction of the feature response. Features strongly correlated with object category, layout position, and geometric structure are extracted through the content pathway. Simultaneously, by using online comparison decoupling loss with style and motion pathways, the representation of appearance texture and dynamic information in the features of this pathway is actively suppressed, and a content latent code representing the static scene composition is output. A style latent code representing the visual appearance is output through the style pathway. A pure temporal dynamic pattern is distilled from the mixed conditions through the motion pathway, and a motion latent code representing the motion law is output. For any initial latent code, the properties of other initial latent codes are predicted by a multi-head discriminator network, wherein the properties of other initial latent codes cannot be inferred from any given initial latent code, thus forcing the three initial latent codes to be mutually orthogonal in the feature space. The orthogonalized initial latent codes are spliced ​​together and fine-tuned and scaled through a conditional output projection layer to output decoupled and task-adapted content, style, and motion control subspace signals.

7. The method for generating a first-screen video background based on page elements according to claim 5, characterized in that, Compositing the dynamic background onto the homepage of the target website further includes: Create a full-screen WebGL canvas or Canvas element as a rendering container, and render the dynamic background into the rendering container; Identify and track the visual salience of key interactive elements on the page, analyze the local motion intensity, brightness change frequency and texture complexity of the dynamic background in the key interactive element area, and process the background area of ​​the dynamic background through an adaptive visual noise reduction filter.

8. The method for generating a first-screen video background based on page elements according to claim 2, characterized in that, Computing the regularized optimal transport plan to find the minimum cost projection matrix that maps it to the shared semantic space further includes: A cost matrix is ​​formed based on the feature matrix of the executable visual description script in the language domain and the multimodal feature matrix of the page in the visual domain. The cost of the cost matrix is: Where, x i ,y j Let X and Y be the row vectors of the feature matrix X and the page multimodal feature matrix Y, respectively, for the executable visual description script; α is the balancing weight; ∑ is the covariance matrix estimated from all features; X∈R m×d , Y∈R n×d m and n represent the number of language and visual features, respectively, and d represents the feature dimension. Find a transmission plan matrix that minimizes the total transmission cost and is constrained by entropy regularization and mode structure preservation; the objective function of this transmission plan matrix is: in,<P,C> F For the total transmission cost, P ij The proportion by which the quality of language feature i is allocated to visual feature j; Let a be the feasible region of the transmission plan; a and b be the normalized quality distributions of language and visual features, respectively; H(P) = -∑ i,j P ij (logP ij -1), which is the entropy regularization term of the transmission plan; L X L Y These are the graph Laplacian matrices constructed based on the k-nearest neighbor graphs of features X and Y, respectively. To maintain constraints on the modal structure; P is the transport plan matrix; C is the cost matrix; <.,> F This is the Frobenius inner product; The minimum cost projection matrix is ​​obtained by solving the optimization problem of the objective function using an accelerated approximation point algorithm based on Bregman iteration.

9. The method for generating a first-screen video background based on page elements according to claim 5, characterized in that, A multi-attribute joint optimization loop is initiated to analyze the differences between adjacent latent frames and apply smoothing constraints to reduce inter-frame flicker and jitter. The optimized video latent sequence output further includes: Based on the intermediate video latent sequence and the low-resolution pixel frames decoded from the sequence, calculate the forward and backward optical flows between adjacent frames and their respective optical flow confidence masks. Based on the forward and backward optical flows and their respective optical flow confidence masks, as well as the consistency loss function of the multi-attribute joint optimization cycle, inter-frame flicker and jitter are reduced and an optimized video latent sequence is output; the consistency loss function includes latent spatial cycle consistency loss, pixel-level optical flow consistency loss, and motion trajectory smoothness loss.

10. A device for generating a first-screen video background based on page elements, characterized in that, include: The feature and instruction acquisition module is used to acquire the multimodal interface features of the target website's home screen and receive natural language instructions input by the user through an interactive dialog box; wherein, the multimodal interface features include layout structure, visual color, text semantics, and interactive components; The feature understanding and condition generation module is used to input the natural language instructions and multimodal interface features into the multimodal alignment and fusion model, analyze the visual intent, style preference and dynamic requirements in the natural language instructions to generate a structured visual control description; and extract layout, color and semantic entities from the multimodal interface features through a cross-modal attention model to form a generated condition vector. The video background sequence generation module is used to input the generation condition vector into the decoupled video diffusion model, decouple the generation condition vector into a control subspace of content, style and motion, and synthesize a video background sequence based on the control subspace and the structured visual control description. The background compositing module is used to compose a dynamic background based on the video background sequence and to composite the dynamic background onto the homepage of the target website.

Citation Information

Cited By

  • Adaptive semantic recombination and noise reduction method oriented to large language model retrieval enhancement

    CN122065849A