A three-branch architecture and method for text-to-image personalized generation

CN122820892APending Publication Date: 2026-09-25HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610433075.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]针对现有技术中文本到图像个性化生成时,主体结构与风格表征耦合、文本提示贴合度不足、复杂场景鲁棒性差等问题,本发明提出一种用于文本到图像个性化生成的三分支架构及方法,通过四大核心模块的协同工作,实现主体结构与风格特征的解耦与精准融合,生成既保留参考主体固有特征,又精准匹配文本风格与场景要求的个性化图像

Benefits of technology

[0063]本发明提出的用于文本到图像个性化生成的三分支架构及方法,在融合深度学习扩散模型技术的基础上,创新设计了三分支特征交互架构及同步匹配引导优化机制,实现了参考主体结构保真与文本风格精准适配的双重目标。该方法通过分工明确的重构分支、融合分支与生成分支,有效解耦结构与风格特征,结合循环一致性校验与指数加权引导策略,显著提升了生成图像的主体完整性、风格贴合度及复杂场景适应性,成功解决了传统方法中结构与风格耦合、文本响应不精准、鲁棒性不足等问题。其可广泛应用于图像编辑、数字内容创作、虚拟形象生成等多个领域,为用户提供高效、高质量的个性化图像生成服务,助力数字创作行业的智能化升级,同时降低个性化图像生成的技术门槛与时间成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820892A_ABST
    Figure CN122820892A_ABST
Patent Text Reader

Abstract

The application discloses a diffusion model-based three-branch architecture and method for personalized generation of text-to-image, and relates to the field of generative artificial intelligence and image editing. The architecture includes image inversion, three-branch, semantic matching consistency modeling and synchronous matching guide modules. First, a DDIM inverse transform is used to map the reference image to standard Gaussian noise. Second, reconstruction, fusion and generation branches are constructed to realize subject structure extraction, text style integration and structure style alignment fusion, respectively. Third, a matching cost function is calculated based on semantic information to maximize the optical flow field and realize subject structure conversion. Finally, the prediction noise is dynamically corrected by an exponential weighted MSE score function to strengthen the style guide effect. The application decouples the structure style representation, improves the text fit, and generates images that not only retain the inherent appearance structure of the reference subject, but also accurately match the target style, making it suitable for creative design, digital content creation, virtual image generation and other scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of generative artificial intelligence, diffusion models, and image editing technology, specifically to a three-branch architecture and method for personalized text-to-image (T2I) generation, applicable to scenarios requiring customized style image generation such as image editing, digital content creation, and virtual character generation. Background Technology

[0002] Personalized text-to-image generation is a core application area in generative AI. Its core objective is to generate customized images that retain the inherent appearance and structural features of the subject while accurately matching the style and scene of the text description, based on user-provided reference images and stylized text prompts. With the explosive growth in demand for digital content creation, the industry has placed higher demands on the accuracy, efficiency, and robustness of this technology, but existing technical solutions still have significant shortcomings.

[0003] Current mainstream technical approaches can be broadly categorized into two types: optimization-based methods and plug-in methods. Optimization-based methods (such as Textual Inversion, DreamBooth, and Custom Diffusion) optimize text pseudowords or local model parameters to enable pre-trained diffusion models to learn the visual concepts of specific reference subjects. However, these methods compress complex visual information into a low-dimensional semantic space, making it difficult to capture the fine-grained spatial structure and texture details of the subject. The generated images are prone to color shifts and structural distortions, and they have poor adaptability to complex stylized text prompts. The optimization process is time-consuming and labor-intensive, and plug-and-play functionality cannot be achieved.

[0004] Plug-in methods (such as MasaCtrl and DreamMatcher) do not require additional training and directly inject the spatial structure features of the reference image into the diffusion model denoising process. However, most methods do not have a specific design for the transfer of stylistic information. Directly injecting the V features of the reconstruction branch into the generation branch, while in the self-attention mechanism, V features dominate the appearance attributes, this operation easily covers the style information encoded in the text, resulting in the generated image only retaining the main structure and failing to meet the needs of changing the main style (such as "Van Gogh style") or the main color attributes.

[0005] Furthermore, stylized generation faces the core challenge of "structure-style decoupling": First, it is difficult to balance the weights of subject fidelity and style effectiveness, easily leading to extreme cases of "style covering the subject" or "subject rejecting style"; second, style information representation is insufficient, and existing text encoders struggle to transform abstract style features into visual features that the model can recognize; third, stylized information is unstable during diffusion, easily covered by noise or subject features, resulting in inconsistent style presentation. Simultaneously, in complex scenarios such as large displacements, occlusions, and new perspectives, the semantic matching accuracy of existing methods drops significantly, further exacerbating the instability of the generated results.

[0006] Traditional methods can no longer meet the triple requirements of "high subject fidelity, high text fit, and adaptability to complex scenarios" in practical application scenarios. A brand-new technical architecture is needed to achieve precise decoupling and collaborative generation of subject structure, inherent appearance and text style. Summary of the Invention

[0007] To address the problems of coupling between subject structure and style representation, insufficient text prompt fit, and poor robustness in complex scenes in existing text-to-image personalized generation technologies, this invention proposes a three-branch architecture and method for text-to-image personalized generation. Through the collaborative work of four core modules, it achieves decoupling and precise fusion of subject structure and style features, generating personalized images that retain the inherent features of the reference subject while accurately matching the text style and scene requirements. The architecture includes an image inverse transformation module, a three-branch feature interaction module, a semantic matching consistency modeling module, and a synchronous matching guidance module. The specific implementation steps are as follows:

[0008] Step 1, the inverse image transformation process converts the reference subject image into standard Gaussian noise that meets the input requirements of the diffusion model, providing a unified input carrier for subsequent three-branch feature interaction. The specific steps are as follows:

[0009] 1.1 The reference image is preprocessed and encoded into the latent space. The specific steps are as follows:

[0010] 1.1.1 Input reference subject image I x The pixel values ​​are normalized.

[0011] 1.1.2 Based on the preset input resolution of the Latent Diffusion Model (LDM) encoder, a bilinear interpolation algorithm is used to process the image to ensure that the image size matches the encoder input dimension;

[0012] 1.1.3 Detect the number of channels in the input image to ensure it matches the three-channel input requirement of the LDM encoder;

[0013] 1.1.4 Load the pre-trained LDM diffusion model;

[0014] 1.1.5 The preprocessed image I x The input encoder progressively compresses the spatial dimension of the image and improves the level of feature abstraction to obtain the latent feature tensor z0, which retains the core structure, texture and appearance features of the reference subject.

[0015] 1.2 Perform inverse DDIM transform on the preprocessed tensor z0 to generate standard Gaussian noise. The specific steps are as follows:

[0016] 1.2.1 Set the total number of steps T in the diffusion model and initialize the noise scheduling parameter α.t Sequence, calculate cumulative noise reduction coefficient Generate α t and The lookup table;

[0017] 1.2.2 Initialize the current latent noise z using the latent feature tensor z0 as the initial input at time step t=0. t =z0, set the iteration counter step = 0;

[0018] 1.2.3 Perform the inverse transformation operation sequentially from t=0 to t=T-1: First, convert the current potential noise z t With time step encoding EMB t Input the LDM pre-trained U-Net to predict the noise tensor ∈ at the current time step. θ (z t ,t), representing the noise component in the current potential features; secondly, according to the formula Calculate the potential noise z at the next time step t+1 ,in, To preserve the original characteristic components, This formula is used to gradually enhance the noise component in each iteration, adding new noise components. Finally, after T iterations, the potential noise at time step t = T is obtained. At this point, the noise follows a standard Gaussian distribution.

[0019] Step 2, three-branch feature interaction construction: through the collaborative work of reconstructing branches, merging branches, and generating branches, the extraction of reference subject structural features, the integration of text style features, and the precise alignment of structural and style features are completed. The specific steps are as follows:

[0020] 2.1 Create a refactoring branch, the specific steps are as follows:

[0021] 2.1.1 Standard Gaussian noise output from the image inverse transform module As the initial input for the refactoring branch;

[0022] 2.1.2 The noise scheduling parameter α from the inverse image transform stage is used. t With cumulative noise reduction coefficient Ensure the consistency of parameters between the denoising process and the inverse transform process to avoid feature reconstruction distortion caused by parameter differences;

[0023] 2.1.3 Set the diffusion denoising step number t;

[0024] 2.1.4 A sinusoidal positional encoding method is used to encode the time steps, transforming the time step information into features that the model can recognize and embedding them into the EMB. t ;

[0025] 2.1.5 Initialize the current potential noise Iteration counter iter = 0;

[0026] 2.1.6 Perform forward denoising operation step by step, the specific steps are as follows:

[0027] 2.1.6.1 The current potential noise z t Input LDM encoder;

[0028] 2.1.6.2 Basic features are extracted through residual blocks, and the resulting feature map contains high-order semantic information of the reference subject;

[0029] 2.1.6.3 Input the feature map into the cross-attention layer to complete cross-attention aggregation;

[0030] 2.1.6.3 Input the output results into the self-attention layer to complete self-attention aggregation;

[0031] 2.1.6.4 Extract self-attention features corresponding to each time step

[0032] 2.1.6.5 Synchronize the structural features to the self-attention layer of the fusion branch according to time steps. At the corresponding time step t of the fusion branch, directly replace its original features. and Ensure that the fusion branch always takes the spatial structure of the reference subject as a constraint when integrating text style features, so as to avoid problems such as subject deformation and component misalignment during the stylization process;

[0033] 2.1.6.6 The deep features enhanced by self-attention are input into the U-Net decoder, and the output is the prediction noise at the current time step. θ (z t ,t);

[0034] 2.1.6.7 Update latent features according to the DDIM denoising formula. Updated Order t = t-1, iter = iter+1;

[0035] 2.1.7 Repeat step 2.1.6 until t=0 to obtain the latent features of the reconstructed reference image.

[0036] 2.2 Construct the fusion branch, the specific steps are as follows:

[0037] 2.2.1 Sampling independent of the reconstruction branch using standard Gaussian noise To avoid interference with the original style of the reference image through noise transmission;

[0038] 2.2.2 Load the pre-trained CLIP text encoder, input the text prompt into the encoder, and generate the semantic embedding p;

[0039] 2.2.3 Convert the initial semantic embedding p into text style features p emb ;

[0040] 2.2.4 Independent noise Input LDM encoder;

[0041] 2.2.5 The text style feature p emnb As cross attention and Using image features as Calculate cross attention;

[0042] 2.2.6 The self-attention layer and Directly replace with the branch passed by the refactoring branch and

[0043] 2.2.7 Inject the output of the cross-attention layer with the reconstructed branch. Input from the attention layer, and complete the attention aggregation according to the formula: During the polymerization process, and Provide spatial structural constraints, It carries stylistic semantics, achieving an initial fusion of structure and style;

[0044] 2.3 Create and generate branches, the specific steps are as follows:

[0045] 2.3.1 The same initial standard Gaussian noise as the fusion branch is used. Ensure that the initial feature distribution of the generated branch is consistent with that of the fused branch; avoid alignment failure caused by differences in feature distribution.

[0046] 2.3.2 Following step 2.2.3, embed the target text p emb ;

[0047] 2.3.3 Integrating Features Injecting branch generation from the attention layer, and generating branches... According to the formula After completing self-attention aggregation, the output result is input into the LDM decoder for subsequent denoising.

[0048] Step 3: Construct the semantic matching consistency modeling module to estimate the optical flow field of the semantic structure of images through semantic matching between them.

[0049] 3.1 Initialize confidence thresholds γ and λc ;

[0050] 3.2 At each time step t, extract intermediate features ψ from the fusion branch respectively. Y Extracting intermediate features ψ from the generated branches Y ′;

[0051] 3.3 Constructing the matching loss function By maximizing C t+1 (i, j) are obtained

[0052] 3.4 On the displacement field Perform the inverse mapping to obtain the inverse displacement field.

[0053] 3.5 Generate confidence mask U t ,

[0054] 3.6 Extract the attention weight matrix from the cross-attention layer of the generated branch, label pixels with weight values ​​above the threshold as foreground, pixels below the threshold as background, and the intermediate region as transition region, to obtain the foreground mask M. t ;

[0055] 3.7 Applying the Hadamard product to the foreground mask M t With confidence mask U t Fusion, formula M′ t =M t ⊙U t ;

[0056] 3.8 Based on displacement field Stylized features of the fusion branch Perform the Warp operation to achieve spatial alignment between stylized features and the generated branch target structure;

[0057] 3.9 Extracting Value Features from the Self-Attention Layer for Generating Branches According to the formula The fusion is completed, and the result is input into the self-attention layer that generates the branch.

[0058] Step 4: Correct the prediction results using exponentially weighted MSE, and force the generation of stylized features that fit the fusion branch throughout the diffusion process, while preserving the integrity of the main structure. The specific steps are as follows:

[0059] 4.1 Constructing the bootstrap optimization function

[0060] 4.2 Constructing the Guiding Optimization Function

[0061] 4.3 Setting the weight parameter w of the guided optimization function t = [1 + (1 - t / T)] 2 ;

[0062] Beneficial effects

[0063] This invention proposes a three-branch architecture and method for personalized text-to-image generation. Based on the integration of deep learning diffusion model technology, it innovatively designs a three-branch feature interaction architecture and a synchronous matching guided optimization mechanism, achieving the dual goals of faithfully preserving the reference subject structure and accurately adapting the text style. This method effectively decouples structural and style features through clearly defined reconstruction, fusion, and generation branches. Combined with cyclic consistency verification and exponential weighting guidance strategies, it significantly improves the subject integrity, style fit, and adaptability to complex scenes in the generated images, successfully solving problems such as structure-style coupling, inaccurate text response, and insufficient robustness in traditional methods. It can be widely applied in multiple fields such as image editing, digital content creation, and virtual avatar generation, providing users with efficient and high-quality personalized image generation services, assisting in the intelligent upgrade of the digital creation industry, and reducing the technical threshold and time cost of personalized image generation. Attached Figure Description

[0064] Figure 1 Three-branch model architecture diagram

[0065] Figure 2 Figure showing the qualitative comparison results with the benchmark model

[0066] Figure 3 The results are shown in the figure for qualitative comparison with related models.

[0067] Figure 4 Figure showing the results of quantitative comparison with the baseline model.

[0068] Figure 5 The results are shown in the figure for quantitative comparison with the model that does not require training.

[0069] Figure 6 The graph shows the quantitative comparison results with the untrained model on a challenging dataset.

[0070] Figure 7 A graph showing the quantitative comparison results with optimization-based models on challenging datasets. Detailed Implementation

[0072] Taking the personalized text-to-image generation of "A Van Gogh style sks cat wearing a Santa hat" as an example, the technical solution of the present invention will be further described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0073] Step 1, Data Preprocessing:

[0074] 1.1 Select a clear frontal image with a resolution of 1024×1024 as the input reference image:

[0075] 1.2 The pixel values ​​of the reference image are normalized to the [0, 1] interval, and then scaled to 512×512 resolution using bilinear interpolation to verify that it is in RGB three-channel format;

[0076] 1.3 Convert the image from (512, 512, 3) to the standard input of the PyTorch framework (3, 512, 512); add a batch dimension to finally generate an input tensor with shape (1, 3, 512, 512) to adapt to the batch processing logic of the encoder.

[0077] 1.4 Load the pre-trained LDM diffusion model;

[0078] 1.5 Self-attention layer and cross-attention layer feature dimension d = 512, multi-head attention head number head = 8, head dimension d head =64, dropout=0.1;

[0079] 1.6 Set the total number of diffusion steps T = 50, and the time step t ∈ [0, 50];

[0080] 1.7 A linearly increasing strategy is used to generate the weights β at each time step. t ;

[0081] 1.8 Calculate the noise reduction coefficient α for each time step t ;

[0082] 1.9 Calculate α t The cumulative coefficient characterizes the cumulative noise reduction effect from t=0 to time t; it is pre-calculated for all time steps. And store it as an array to avoid repeated calculations during iteration and improve efficiency;

[0083] 1.10 For each time step t, a 512-dimensional time step embedding (emb) is generated using sinusoidal position encoding. t ;

[0084] 1.11 Set text hints ("") for the inverse image transformation and reconstruction branches to indicate empty text hints;

[0085] 1.12 Set the merge branch and generate branch text hints ("A Van Gogh style sks cat wearing a santa hat") to indicate the target text hints;

[0086] 1.13 Input the two text prompts into the CLIP-B / 32 encoder to generate an initial 768-dimensional semantic embedding, which is then transformed into a 512-dimensional text style feature p through a projection layer. emb and p′ emb ;

[0087] Step 2, Image Reverse:

[0088] 2.1 Input the preprocessed input tensor into the LDM encoder and output the latent feature tensor z0;

[0089] 2.2 Initialize the current noisy feature z using the latent feature tensor z0 as the initial input at time t=0. t =z0.

[0090] 2.3 Execute the noise-adding process in a loop: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] t With emb t Input LDM model ε θ Output the prediction noise at the current time step ∈ θ (z t emb t p emb ), calculate the noisy feature z at time t+1. t+1 Finally, standard Gaussian noise is obtained.

[0091] Step 3, construct the three-branch architecture, the specific steps are as follows:

[0092] 3.1 Standard Gaussian noise generated by inverse image transformation As initial input;

[0093] 3.2 Set the current potential noise The time step decreases from t=50 to t=0;

[0094] 3.3 Time-step denoising: z t Time embedding in EMB t and text prompts embedded in p emb Input the LDM noise prediction network and extract the intermediate data from the attention layer. Output the current time step prediction noise ∈ θ (z t emb t ), calculate the denoised features Iterative noise reduction.

[0095] 3.4 A random number generator is used to sample standard Gaussian noise. The tensor has dimensions (1, 4, 64, 64), and is related to the reconstruction branch. Irrelevant, to avoid interference with the original style of the reference image;

[0096] 3.5 Set the current potential noise The time step decreases from t=50 to t=0;

[0097] 3.6 Time-step denoising: z t Time embedding in EMB t and text prompt embed p′ emb Input the LDM noise prediction network and add the self-attention layer Replace with the one extracted from the refactoring branch. and Complete multi-head attention aggregation and extract intermediate data from the attention layer. The second and third layers of the decoder extract features ψ. Y Input is fed into the semantic matching and consistency modeling module;

[0098] 3.7 Set the current potential noise The time step decreases from t=50 to t=0;

[0099] 3.8 Time-step denoising: z t Time embedding in EMB t and text prompt embed p′ emb Input the LDM noise prediction network and add the self-attention layer The result of replacing it with the semantic matching consistency modeling module Perform multi-head attention aggregation and extract features ψ from layers 2 and 3 of the decoder. Y ′, input to the semantic matching consistency modeling module;

[0100] Step 4: Construct the semantic matching consistency modeling module. The specific steps are as follows:

[0101] 4.1 Initialize confidence thresholds γ = 0.8, λ c =0.1;

[0102] 4.2 Constructing the matching loss function C t+1 (i, j), by maximizing C t+1 (i, j) are obtained

[0103] 4.3 Based on Constructing the inverse displacement field

[0104] 4.4 Calculate the consistency mask Ut (x);

[0105] 4.5 Extracting the attention weight matrix Attn from the generative branch cross attention layer weight Set the foreground threshold (thres) fg =0.7, background threshold (thres) bg =0.3; Label weight ≥thres fg Foreground, ≤thres bg Using the background as the background, a coarse foreground mask M is obtained. t ;

[0106] 4.6 Calculate the final semantically consistent mask M′ t ;

[0107] 4.7 pairs Executing the warp operation yields

[0108] 4.8 Calculation using a mask The input is fed into the self-attention layer that generates the branch;

[0109] Step 5: Construct the synchronization matching bootstrap module. The specific steps are as follows:

[0110] 5.1 Constructing the Guided Optimization Function L MSE ;

[0111] 5.2 Setting the weight parameter w of the guided optimization function t ;

[0112] 5.3 Correct the prediction result ε at each time step t guided ;

[0113] Step 6: Iterate for T time steps and output the final result image.

[0114] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principle of the present invention. For example, in the improved method, the selection of the denoising network is not limited to the network mentioned, and other networks suitable for the image data to be processed can also be selected according to the method proposed in the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A three-branch architecture and method for personalized text-to-image generation, characterized in that, The system comprises four modules: an image inversion module, a three-branch architecture, a semantic matching consistency modeling module, and a synchronous matching guidance module. These four modules work together to achieve personalized text-to-image generation. The specific steps are as follows: 1.1 The reference image is mapped to standard Gaussian noise using the inverse DDIM transform; 1.2 Spatial structural features of the reference image are extracted through the reconstruction branch, text style features are integrated through the fusion branch, and alignment and fusion of structural and style features are completed through the generation branch; 1.3 Construct a semantically consistent mask based on cycle consistency to ensure the accuracy of feature fusion; 1.4 An exponentially weighted MSE score function is used to dynamically guide and correct the predicted noise of the diffusion process, and output a personalized image that combines subject fidelity and style adaptation.

2. The method according to claim 1, characterized in that, The steps for implementing the inverse image transformation include: 2.1 The reference image is preprocessed by standardization, and the pixel values ​​are normalized to the [0, 1] interval. The resolution is adjusted and the channel adaptation is completed according to the requirements of the Latent Diffusion Model (LDM) encoder. 2.2 Input the preprocessed image into the LDM pre-trained encoder and map it to the low-dimensional latent space to obtain the latent feature tensor. The dimension of this tensor is lower than that of the original pixel image and retains the core features of the reference subject. 2.3 Initialize the diffusion time step t = T (T is the total number of diffusion steps) and the noise scheduling parameter α t According to the formula The DDIM diffusion process is executed in reverse iteration to generate standard Gaussian noise with a mean of 0 and a variance of 1.

3. The method according to claim 1, characterized in that, In the three-branch feature interaction module, the steps for reconstructing branches include: 3.1 Standard Gaussian noise As the initial input, the diffusion denoising step number t∈[0,T] is set, and the weights of the LDM diffusion model U-Net are loaded; 3.2 Starting from t=T, denoising is performed step by step, and the basic features are projected into query features in the self-attention layer. Key features Value characteristics According to the formula Complete self-attention aggregation; 3.3 After completing T-step denoising, the latent features of the reconstructed reference image are obtained. 3.4 Extracting from the attention layer by time step and Synchronize to the self-attention layer at the corresponding time step of the fusion branch.

4. The method according to claim 1, characterized in that, In the three-branch feature interaction module, the steps for implementing the fusion branch include: 4.1 Sampling independent of the reconstruction branch: random standard Gaussian noise As initial input, avoid interference from the original style of the reference image; 4.2 Input the stylized text cue into the CLIP text encoder to generate semantic embeddings, and obtain the text style features p emb ; 4.3 At each time step t, receive the refactoring branch passed to you. Replace the native self-attention layer of the fusion branch 4.4 Projecting the basic features of the fusion branch into value features Through cross-attention layer and text style feature p emb Interactive, integrating the style and semantics specified in the text; 4.5 According to the formula Complete self-attention aggregation to generate a fusion feature that combines reference structure and text style.

5. The method according to claim 1, characterized in that, In the three-branch feature interaction module, the steps for generating branches include: 5.1 Use the same initial noise as the fusion branch. As input, the target text prompt is accessed and the target text embedding is generated. emb ; 5.2 Load the semantic matching consistency modeling module and initialize the confidence thresholds γ and λ. c ; 5.3 Extracting features from the intermediate layer between the fusion branch and the generation branch and 5.4 Constructing the matching loss function By maximizing C t+1 (i, j) are obtained 5.5 Based on Constructing the inverse displacement field 5.6 Calculate the consistency mask 5.7 Extracting the attention weight matrix Attn from the generative branch cross attention layer weight Set the foreground threshold (thres) fg Background threshold (thres) bg ; Label weight ≥ thres fg Foreground, ≤thres bg Using the background as the background, a coarse foreground mask M is obtained. t ; 5.8 Calculate the final semantically consistent mask M′ t =M t ⊙U t , where ⊙ is the Hadamard accumulation; 5.9 Based on displacement field pairs Execute the Warp operation according to the formula. Achieve spatial alignment between stylized features and the generated branch target structure; 5.10 According to the formula Complete the fusion of foreground and background features; 5.11 will integrate features Injecting a self-attention layer that generates branches, according to the formula Complete attention aggregation.

6. The method according to claim 1, characterized in that, The implementation steps of the synchronization matching guidance module include: 6.1 Constructing the bootstrap optimization function 6.2 Setting the weight parameter w of the guided optimization function t = [1 + (1 - t / T)] 2 ; 6.3 Correct the prediction results for each time step t.