Image local redrawing optimization method and system based on branch network
Through the branch network-based image local redrawing optimization method, the problems of insufficient learning of mask area features and poor adaptability in complex backgrounds in the prior art are solved, and efficient and controllable image local redrawing is achieved. The generation results are highly consistent with the original image and are suitable for image editing with multi-resolution input.
Patent Information
- Application Number
- CN202510526417.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-12
AI Technical Summary
The existing image local redrawing technology has problems such as insufficient learning of mask area features, poor adaptability of complex backgrounds and limited controllability, resulting in poor consistency between the generated content and the original image, stiff edge transitions, distorted details, and it is difficult to accurately control the generation of results through multimodal input.
The local image redrawing optimization method based on branch network is adopted, and the multi-modal mask input and hierarchical feature fusion mechanism is combined with multi-task loss design to build a lightweight branch network, focusing on learning mask regional features, and supporting multi-resolution input through dynamic downsampling and multi-scale feature extraction, improving the visual coherence and controllability of the generated results.
The generated content is highly consistent with the original image in the airspace and frequency domain, which significantly improves the generation quality in complex scenarios, reduces training costs, and supports seamless adaptation from low resolution to 4K level inputs. Users can accurately control the generation results through multimodal inputs.
Smart Images

Figure CN120472025A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image optimization, and in particular to a branch network-based image local redrawing optimization method and system. Background Art
[0002] Existing image local redrawing technologies mainly rely on traditional manual editing tools (such as the Clone Healing Brush) or deep learning-based generative models (such as generative adversarial networks and diffusion models). Although diffusion models can achieve local content generation by combining text descriptions with mask areas, they still have the following technical drawbacks:
[0003] Insufficient learning of mask region features: Existing methods do not specifically model the contextual information of the mask region, resulting in poor consistency between the texture, lighting, and geometric structure of the generated content and the original image. This manifests as abrupt edge transitions, distorted details (such as blurred textures and mismatched shadows), and noticeable artifacts.
[0004] Poor adaptability to complex backgrounds: For highly complex backgrounds (such as dense textures and dynamic lighting scenes), existing models have difficulty effectively capturing the semantic association between local and global contexts, and visual dissonance is prone to occur in the generated content.
[0005] Limited controllability: The lack of a refined control mechanism makes it difficult for users to guide the generation process through multimodal input (such as sketches and color markings), resulting in large deviations from expected redrawing results. Summary of the Invention
[0006] The purpose of the present invention is to solve the shortcomings mentioned in the above background technology and to propose an image local redrawing optimization method and system based on a branch network.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The branch network-based image local redrawing optimization method includes the following steps:
[0009] S1. Input preprocessing: Generate multimodal mask input that fits the model’s latent space;
[0010] S2. Branch network architecture: Build a lightweight branch network to focus on learning mask region features;
[0011] S3, hierarchical feature fusion: injecting local features learned by the branch network into the original model;
[0012] S4, training strategy: Efficiently fine-tune the branch network to balance generation quality and training cost;
[0013] S5, Resolution Adaptation: Supports generation consistency from low resolution to 4K input;
[0014] S6. Reasoning optimization: Improve the visual coherence and controllability of generated results.
[0015] Preferably, in said S1, the specific steps are as follows:
[0016] S101. Generate RGB mask: retain the original pixel values of the target area of the original image, blacken the non-target area, generate an RGB three-channel mask image, and compress the RGB mask to the latent space dimension through the VAE encoder of the original model to ensure alignment with the noise input dimension;
[0017] S102, generating a binary mask: setting the target area to black and the background to white to generate a single-channel binary mask image, and using the nearest neighbor interpolation method to downsample the binary mask to the potential spatial resolution to avoid edge blurring caused by bilinear interpolation;
[0018] S103, noise input processing: retain the noise input of the original model as the basic condition of the generation process.
[0019] Preferably, in said S2, the specific steps are as follows:
[0020] S201, parameter reuse initialization: copy the network structure and weight parameters from the cross-attention layer of the original generative model as the initialization parameters of the branch network to ensure that the branch network is compatible with the original model architecture and reduce training instability;
[0021] S202, multimodal input concatenation: concatenate the noise, RGB mask features, and binary mask along the channel dimension to form an input tensor of dimension 64×64×9;
[0022] S203. Remove text input interference: Delete the input channels related to text embedding in the branch network, retaining only the mask and noise conditions, forcing the network to autonomously learn the texture and structure of the local area.
[0023] Preferably, in S3, the specific steps are as follows:
[0024] S301, feature weighted summation: adding the feature map output by the branch network to the feature map of the corresponding layer of the original model element by element;
[0025] S302, Channel Attention Mechanism: Before fusion, the SENet module dynamically adjusts the weights of each channel of branch features to enhance the contribution of important features;
[0026] S303, cross-layer consistency constraint: impose consistency loss on the multi-layer fusion results to ensure that the feature distributions between different layers are aligned.
[0027] Preferably, in S301, the weighted summation formula is: output feature = original feature + α × branch feature, where α is a learnable weight coefficient and α is initialized to 0.1.
[0028] Preferably, in said S4, the specific steps are as follows:
[0029] S401, parameter freezing strategy: fix all parameters of the original generative model and only allow the cross-attention layer parameters of the branch network to be updated;
[0030] S402, Multi-task Loss Design: Combines pixel-level L1 loss, VGG-16 perceptual loss, and frequency domain consistency loss to improve detail fidelity and visual consistency of the generated area through multi-scale constraints from local to global, spatial to frequency domain;
[0031] S403, optimizer configuration: Use the AdamW optimizer, set the initial learning rate to 1e-4, and dynamically adjust the learning rate with the cosine annealing scheduler.
[0032] Preferably, the step S402 includes the following steps:
[0033] S4021, pixel-level L1 loss: constrains the absolute pixel error between the generated area and the real image;
[0034] S4022, Perceptual Loss: Extract multi-scale features through the pre-trained VGG network and calculate the feature similarity between the generated region and the real image;
[0035] S4023, frequency domain consistency loss: Perform fast Fourier transform (FFT) on the generated area and the original image to constrain the matching of low-frequency and high-frequency components.
[0036] Preferably, in S5, the specific steps are as follows:
[0037] S501, dynamic downsampling strategy: for the high-resolution input binary mask, use the nearest neighbor interpolation method to downsample to the latent space dimension, preserving the hard edge characteristics;
[0038] S502, multi-scale feature extraction: introduce hollow volumes in the branch network to expand the receptive field to capture cross-resolution context information;
[0039] S503, adaptive normalization: perform instance normalization on input noise and mask features to eliminate distribution offset caused by resolution change.
[0040] Preferably, in said S6, the specific steps are as follows:
[0041] S601. Multimodal Conditional Guidance
[0042] The input text hint and mask features jointly guide the generation process, aligning semantic and spatial information through the cross-attention mechanism;
[0043] S602, edge post-processing: performing adaptive Gaussian blur on the boundary between the generated area and the original image to achieve a natural transition;
[0044] S603, super-resolution reconstruction: For high-resolution input, enhance local details through a pre-trained super-resolution model after generation.
[0045] The present invention further provides a branch network-based image local redrawing optimization system, which is applied to the above-mentioned branch network-based image local redrawing optimization method, and includes:
[0046] The input preprocessing module generates a multimodal mask input that fits the model's latent space, highlighting the features of the target area and accurately locating its boundaries, while preserving the randomness and creativity of the noise input;
[0047] The branch network architecture module builds a lightweight branch network, reuses parameter initialization, splices multimodal inputs, removes text interference, and focuses on learning mask area features;
[0048] The hierarchical feature fusion module injects the local features learned by the branch network into the original model to achieve hierarchical feature fusion through feature weighted summation, channel attention mechanism and cross-layer consistency constraints;
[0049] The training strategy module freezes the original model parameters, updates the branch network parameters, combines multi-task losses, configures the optimizer, and efficiently fine-tunes the branch network to balance generation quality and training cost;
[0050] The resolution adaptation module uses strategies such as dynamic downsampling, multi-scale feature extraction, and adaptive normalization to ensure that the system supports consistent generation from low-resolution to 4K inputs.
[0051] The inference optimization module guides the generation process with multimodal conditions, performs edge post-processing, and performs super-resolution reconstruction when necessary to improve the visual coherence and controllability of the generated results.
[0052] Compared with the prior art, the present invention provides a method and system for optimizing local image redrawing based on a branch network, which has the following beneficial effects:
[0053] (1) Through multimodal mask input (RGB mask retains the color of the target area, binary mask accurately locates the boundary) and hierarchical feature fusion mechanism (weighted sum + channel attention), the model can accurately capture the texture, lighting and geometric structure characteristics of the local area. Combined with the multi-task loss design (L1 constrained pixels, VGG-16 perceptual loss alignment semantics, frequency domain loss matching lighting and edges), the generated content is highly consistent with the original image in both the spatial and frequency domains. Experiments show that in complex scenes such as animation and portraits, the LPIPS (learning-perceptual image patch similarity) index of the generated area is improved compared with the baseline model, and is significantly better than the traditional diffusion model;
[0054] (2) By freezing the original model parameters and fine-tuning only the branch network (approximately 5% of the parameters), the training cost is significantly reduced while maintaining the generation quality. The parameter reuse mechanism (cross-attention layer initialization) shortens the model convergence time and reduces the number of training iterations by 40%. In addition, the dynamic downsampling and adaptive normalization strategy enables the model to support multi-resolution input, eliminating the need for repeated training for different resolutions, further saving resources.
[0055] (3) Based on dynamic downsampling (nearest neighbor interpolation preserves hard edges), multi-scale feature extraction (atrous convolution expands the receptive field), and instance normalization (eliminating resolution offset), the system can seamlessly adapt to input resolutions from 512×512 to 4K. In high-resolution test sets (such as 1600×1264), there are no artifacts or blurring. This feature gives it a unique advantage in fields such as film and television post-production and advertising design that require processing ultra-high-definition images, avoiding the problem of detail loss caused by resolution scaling in traditional methods.
[0056] (4) After removing the text input interference in the branch network, the model relies entirely on mask and noise conditions to generate content. Combined with the multimodal guidance (text prompt + mask positioning) in the reasoning stage, users can accurately control the generation results by adjusting the mask shape and text semantics.
[0057] (5) This method is compatible with mainstream generative architectures such as UNet and Dit, and its modular design facilitates rapid adaptation to emerging models. Its effectiveness has been verified in applications such as animation generation, e-commerce portrait retouching, and historical photo restoration. Its lightweight and high-precision features provide a technical foundation for the large-scale application of AI-driven image editing tools. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flow chart of the branch network-based image local redrawing optimization method proposed by the present invention;
[0059] Figure 2 This is a step diagram of the branch network-based image local redrawing optimization method proposed by the present invention. DETAILED DESCRIPTION
[0060] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0061] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.
[0062] Reference Figure 1-2 , an image local redrawing optimization method based on a branch network, comprising the following steps:
[0063] S1. Input preprocessing: Generate multimodal mask input that fits the model’s latent space;
[0064] S2. Branch network architecture: Build a lightweight branch network to focus on learning mask region features;
[0065] S3, hierarchical feature fusion: injecting local features learned by the branch network into the original model;
[0066] S4, training strategy: Efficiently fine-tune the branch network to balance generation quality and training cost;
[0067] S5, Resolution Adaptation: Supports generation consistency from low resolution to 4K input;
[0068] S6. Reasoning optimization: Improve the visual coherence and controllability of generated results.
[0069] In this embodiment, in S1, the specific steps are as follows:
[0070] S101. Generate RGB mask: retain the original pixel values of the target area of the original image, blacken the non-target area, generate an RGB three-channel mask image, and compress the RGB mask to the latent space dimension through the VAE encoder of the original model to ensure alignment with the noise input dimension;
[0071] S102, generating a binary mask: setting the target area to black and the background to white to generate a single-channel binary mask image, and using the nearest neighbor interpolation method to downsample the binary mask to the potential spatial resolution to avoid edge blurring caused by bilinear interpolation;
[0072] S103, noise input processing: retain the noise input of the original model as the basic condition of the generation process.
[0073] In this embodiment, in S2, the specific steps are as follows:
[0074] S201, parameter reuse initialization: copy the network structure and weight parameters from the cross-attention layer of the original generative model as the initialization parameters of the branch network to ensure that the branch network is compatible with the original model architecture and reduce training instability;
[0075] S202, multimodal input concatenation: concatenate the noise, RGB mask features, and binary mask along the channel dimension to form an input tensor of dimension 64×64×9;
[0076] S203. Remove text input interference: Delete the input channels related to text embedding in the branch network, retaining only the mask and noise conditions, forcing the network to autonomously learn the texture and structure of the local area.
[0077] In this embodiment, in S3, the specific steps are as follows:
[0078] S301, feature weighted summation: adding the feature map output by the branch network to the feature map of the corresponding layer of the original model element by element;
[0079] S302, Channel Attention Mechanism: Before fusion, the SENet module dynamically adjusts the weights of each channel of branch features to enhance the contribution of important features;
[0080] S303, cross-layer consistency constraint: impose consistency loss on the multi-layer fusion results to ensure that the feature distributions between different layers are aligned.
[0081] In this embodiment, in S301 , the weighted summation formula is: output feature = original feature + α×branch feature, where α is a learnable weight coefficient and is initialized to 0.1.
[0082] In this embodiment, in S4, the specific steps are as follows:
[0083] S401, parameter freezing strategy: fix all parameters of the original generative model and only allow the cross-attention layer parameters of the branch network to be updated;
[0084] S402, Multi-task Loss Design: Combines pixel-level L1 loss, VGG-16 perceptual loss, and frequency domain consistency loss to improve detail fidelity and visual consistency of the generated area through multi-scale constraints from local to global, spatial to frequency domain;
[0085] S403, optimizer configuration: Use the AdamW optimizer, set the initial learning rate to 1e-4, and dynamically adjust the learning rate with the cosine annealing scheduler.
[0086] In this embodiment, S402 includes the following steps:
[0087] S4021, pixel-level L1 loss: constrains the absolute pixel error between the generated area and the real image;
[0088] S4022, Perceptual Loss: Extract multi-scale features through the pre-trained VGG network and calculate the feature similarity between the generated region and the real image;
[0089] S4023, frequency domain consistency loss: Perform fast Fourier transform (FFT) on the generated area and the original image to constrain the matching of low-frequency and high-frequency components.
[0090] In this embodiment, in S5, the specific steps are as follows:
[0091] S501, dynamic downsampling strategy: for the high-resolution input binary mask, use the nearest neighbor interpolation method to downsample to the latent space dimension, preserving the hard edge characteristics;
[0092] S502, multi-scale feature extraction: introduce hollow volumes in the branch network to expand the receptive field to capture cross-resolution context information;
[0093] S503, adaptive normalization: perform instance normalization on input noise and mask features to eliminate distribution offset caused by resolution change.
[0094] In this embodiment, in S6, the specific steps are as follows:
[0095] S601. Multimodal Conditional Guidance
[0096] The input text hint and mask features jointly guide the generation process, aligning semantic and spatial information through the cross-attention mechanism;
[0097] S602, edge post-processing: performing adaptive Gaussian blur on the boundary between the generated area and the original image to achieve a natural transition;
[0098] S603, super-resolution reconstruction: For high-resolution input, enhance local details through a pre-trained super-resolution model after generation.
[0099] The present invention further provides a branch network-based image local redrawing optimization system, which is applied to the above-mentioned branch network-based image local redrawing optimization method, and includes:
[0100] The input preprocessing module generates a multimodal mask input that fits the model's latent space, highlighting the features of the target area and accurately locating its boundaries, while preserving the randomness and creativity of the noise input;
[0101] The branch network architecture module builds a lightweight branch network, reuses parameter initialization, splices multimodal inputs, removes text interference, and focuses on learning mask area features;
[0102] The hierarchical feature fusion module injects the local features learned by the branch network into the original model to achieve hierarchical feature fusion through feature weighted summation, channel attention mechanism and cross-layer consistency constraints;
[0103] The training strategy module freezes the original model parameters, updates the branch network parameters, combines multi-task losses, configures the optimizer, and efficiently fine-tunes the branch network to balance generation quality and training cost;
[0104] The resolution adaptation module uses strategies such as dynamic downsampling, multi-scale feature extraction, and adaptive normalization to ensure that the system supports consistent generation from low-resolution to 4K inputs.
[0105] The inference optimization module guides the generation process with multimodal conditions, performs edge post-processing, and performs super-resolution reconstruction when necessary to improve the visual coherence and controllability of the generated results.
[0106] In this embodiment, starting from input preprocessing, an RGB mask map is generated for the target area of the original image and a blackening operation is performed on the non-target area. The RGB mask is compressed to the latent space dimension using a model encoder. At the same time, a single-channel binary mask map is generated and downsampled to the latent space resolution through the nearest neighbor interpolation method to retain edge sharpness. The noise input of the original model is retained as the basic generation condition to maintain generation diversity. Then, the branch network construction phase is entered. The network structure and parameters are copied from the cross-attention layer of the original generation model to initialize the branch network and maintain architectural consistency. The text input channel in the branch network is removed to avoid interference and make it focus on learning the local features of the mask area. The noise input, RGB mask features and binary mask are spliced into a multimodal input tensor along the channel dimension. Then, the output features of the branch network and the cross-attention layer features of the original model are weighted and summed layer by layer through a hierarchical feature fusion mechanism. The channel attention mechanism is introduced to dynamically adjust the fusion weights to enhance the contribution of key features, and the branch network is constructed. Add cross-layer consistency constraints to ensure feature distribution alignment; freeze the original model parameters during the training phase and only fine-tune the branch network, jointly optimize the pixel-level L1 loss, VGG-16 perceptual loss, and frequency domain consistency loss to improve the detail fidelity and visual consistency of the generated area, and configure the AdamW optimizer and cosine annealing scheduler to accelerate convergence; adopt a dynamic downsampling strategy and binary mask hard edge retention for multi-resolution input, expand the receptive field through void convolution to capture cross-resolution contextual information, and combine instance normalization to eliminate resolution offset; finally, in the inference phase, guide the generation process through multimodal conditions, input text prompts and mask features jointly guide semantic and spatial alignment, perform adaptive Gaussian blur or Poisson fusion on the boundaries of the generated area to achieve natural transition, and enhance local details through super-resolution reconstruction when necessary to ensure that the generated results are highly consistent with the original image in texture, lighting, and edge transition. It is suitable for complex scenes with inputs from low resolution to 4K level, significantly improving generation efficiency and industrial landing value.
[0107] The standard parts used in the present invention can all be purchased from the market, and special-shaped parts can be customized according to the description in the specification and the drawings. The specific connection methods of each part adopt conventional means such as mature bolts, rivets, welding, etc. in the existing technology. The machinery, parts and equipment all adopt conventional models in the existing technology, and the circuit connection adopts the conventional connection method in the existing technology, which will not be described in detail here.
Claims
1. A branch network-based image local redrawing optimization method, characterized in that: The following steps are involved: S1. Input preprocessing: Generate multimodal mask input that fits the model’s latent space; S2. Branch network architecture: Build a lightweight branch network to focus on learning mask region features; S3, hierarchical feature fusion: injecting local features learned by the branch network into the original model; S4, training strategy: Efficiently fine-tune the branch network to balance generation quality and training cost; S5, Resolution Adaptation: Supports generation consistency from low resolution to 4K input; S6. Reasoning optimization: Improve the visual coherence and controllability of generated results.
2. The image local redrawing optimization method based on branch network according to claim 1, characterized in that: In S1, the specific steps are as follows: S101. Generate RGB mask: retain the original pixel values of the target area of the original image, blacken the non-target area, generate an RGB three-channel mask image, and compress the RGB mask to the latent space dimension through the VAE encoder of the original model to ensure alignment with the noise input dimension; S102, generating a binary mask: setting the target area to black and the background to white to generate a single-channel binary mask image, and using the nearest neighbor interpolation method to downsample the binary mask to the potential spatial resolution to avoid edge blurring caused by bilinear interpolation; S103, noise input processing: retain the noise input of the original model as the basic condition of the generation process.
3. The image local redrawing optimization method based on branch network according to claim 1, characterized in that: In S2, the specific steps are as follows: S201, parameter reuse initialization: copy the network structure and weight parameters from the cross-attention layer of the original generative model as the initialization parameters of the branch network to ensure that the branch network is compatible with the original model architecture and reduce training instability; S202, multimodal input concatenation: concatenate the noise, RGB mask features, and binary mask along the channel dimension to form an input tensor of dimension 64×64×9; S203. Remove text input interference: Delete the input channels related to text embedding in the branch network, retaining only the mask and noise conditions, forcing the network to autonomously learn the texture and structure of the local area.
4. The image local redrawing optimization method based on branch network according to claim 1, characterized in that: In S3, the specific steps are as follows: S301, feature weighted summation: adding the feature map output by the branch network to the feature map of the corresponding layer of the original model element by element; S302, Channel Attention Mechanism: Before fusion, the SENet module dynamically adjusts the weights of each channel of branch features to enhance the contribution of important features; S303, cross-layer consistency constraint: impose consistency loss on the multi-layer fusion results to ensure that the feature distributions between different layers are aligned.
5. The image local redrawing optimization method based on branch network according to claim 1, characterized in that: In S301 , the weighted summation formula is: output feature = original feature + α×branch feature, where α is a learnable weight coefficient and is initialized to 0.
1.
6. The image local redrawing optimization method based on branch network according to claim 1, characterized in that: In said S4, the specific steps are as follows: S401, parameter freezing strategy: fix all parameters of the original generative model and only allow the cross-attention layer parameters of the branch network to be updated; S402, Multi-task Loss Design: Combines pixel-level L1 loss, VGG-16 perceptual loss, and frequency domain consistency loss to improve detail fidelity and visual consistency of the generated area through multi-scale constraints from local to global, spatial to frequency domain; S403, optimizer configuration: Use the AdamW optimizer, set the initial learning rate to 1e-4, and dynamically adjust the learning rate with the cosine annealing scheduler.
7. The image local redrawing optimization method based on branch network according to claim 1, characterized in that: The S402 includes the following steps: S4021, pixel-level L1 loss: constrains the absolute pixel error between the generated area and the real image; S4022, Perceptual Loss: Extract multi-scale features through the pre-trained VGG network and calculate the feature similarity between the generated region and the real image; S4023, frequency domain consistency loss: Perform fast Fourier transform (FFT) on the generated area and the original image to constrain the matching of low-frequency and high-frequency components.
8. The method for optimizing local image redrawing based on a branch network according to claim 1, characterized in that: In said S5, the specific steps are as follows: S501, dynamic downsampling strategy: for the high-resolution input binary mask, use the nearest neighbor interpolation method to downsample to the latent space dimension, preserving the hard edge characteristics; S502, multi-scale feature extraction: introduce hollow volumes in the branch network to expand the receptive field to capture cross-resolution context information; S503, adaptive normalization: perform instance normalization on input noise and mask features to eliminate distribution offset caused by resolution change.
9. The method for optimizing local image redrawing based on a branch network according to claim 1, characterized in that: In said S6, the specific steps are as follows: S601. Multimodal Conditional Guidance The input text hint and mask features jointly guide the generation process, aligning semantic and spatial information through the cross-attention mechanism; S602, edge post-processing: performing adaptive Gaussian blur on the boundary between the generated area and the original image to achieve a natural transition; S603, super-resolution reconstruction: For high-resolution input, enhance local details through a pre-trained super-resolution model after generation.
10. A branch network-based image local redrawing optimization system, applied to the branch network-based image local redrawing optimization method according to claim 9, characterized in that: include: The input preprocessing module generates a multimodal mask input that fits the model's latent space, highlighting the features of the target area and accurately locating its boundaries, while preserving the randomness and creativity of the noise input; The branch network architecture module builds a lightweight branch network, reuses parameter initialization, splices multimodal inputs, removes text interference, and focuses on learning mask area features; The hierarchical feature fusion module injects the local features learned by the branch network into the original model to achieve hierarchical feature fusion through feature weighted summation, channel attention mechanism and cross-layer consistency constraints; The training strategy module freezes the original model parameters, updates the branch network parameters, combines multi-task losses, configures the optimizer, and efficiently fine-tunes the branch network to balance generation quality and training cost; The resolution adaptation module uses strategies such as dynamic downsampling, multi-scale feature extraction, and adaptive normalization to ensure that the system supports consistent generation from low-resolution to 4K inputs. The inference optimization module guides the generation process with multimodal conditions, performs edge post-processing, and performs super-resolution reconstruction when necessary to improve the visual coherence and controllability of the generated results.
Citation Information
Cited By
Priori feature assisted directional coordinate attention remote sensing road extraction method
CN120976762A
Non-paired low-light real image enhancement method based on multi-mode guidance
CN121860873A