Video restoration data set construction and restoration method combined with generation of large model
Through the hierarchical cross-attention mechanism and gradient locking strategy, combined with the generation of a large model for video restoration, the problem of poor local restoration effect in existing technologies is solved, and the video restoration effect of global naturalness and local refinement is achieved, reducing computational costs.
Patent Information
- Application Number
- CN202510601443.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video restoration techniques fail to effectively utilize the potential of generating large models, resulting in poor restoration effects in local areas, especially in terms of object scratches and color distortion.
A hierarchical cross-attention mechanism and gradient locking strategy are adopted to perform video restoration by generating a large model, injecting global and local features in stages, combining text descriptions and object masks, optimizing the collaboration between local and global features, and only training the cross-attention layer parameters to reduce computational costs.
It achieves global naturalness and local refinement of video restoration, significantly improves the restoration effect, reduces computing resource requirements, and ensures consistency and smoothness of the restored content with the original video.
Smart Images

Figure CN120634908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video restoration technology, and in particular to a video restoration data set construction and restoration method combined with generating a large model. Background Art
[0002] Video restoration technology aims to enhance low-quality videos through algorithms, including tasks such as denoising, deblurring, super-resolution, and frame rate improvement. Existing technologies often use small-scale neural networks (such as ESRGAN) to process videos frame by frame and then merge the results. However, due to the limited number of model parameters, it is difficult to capture the temporal consistency of videos and the details of complex scenes.
[0003] In recent years, large generative models (such as Sora and Wan) have demonstrated powerful temporal modeling and detail generation capabilities in video generation. However, their enormous number of parameters (in the billions to hundreds of billions) has prevented them from being applied to video restoration tasks. Existing video restoration methods fail to leverage the generative potential of large generative models, resulting in limited restoration results, particularly in the detailed restoration of localized areas (such as scratches and color distortion).
[0004] In the above technical solutions; therefore, we propose a video restoration dataset construction and restoration method that combines the generation of a large model to solve this problem. Summary of the Invention
[0005] The purpose of the present invention is to provide a video restoration dataset construction and restoration method combined with the generation of a large model to solve the problems raised in the above background technology.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A video restoration dataset construction and restoration method combined with generation of a large model includes the following steps:
[0008] S1, slicing: Select high-quality original video and perform scene-consistent slicing using the PySceneDetect algorithm;
[0009] S2, screening: aesthetic scoring and motion detection are performed on the slices to select and retain high-dynamic and high-quality clips;
[0010] S3. Core object segmentation: Use a multimodal model to accurately segment the core objects in each frame and generate object masks;
[0011] S4. Extract text features: Extract the text features of the video clip description and the segmented objects through the text generation model to form a single description containing the original video slice, object mask, object text description and the overall video description;
[0012] S5, model input: input random noise, object mask and mask video;
[0013] S6, Feature Fusion: Mask video is compressed by VAE and concatenated with the downsampled object mask. The input is used to generate a large model. The video description is injected into the first 80% layers of the model through the cross-attention layer, and the object description is injected into the last 20% layers to achieve collaborative optimization of local and global features.
[0014] S7, model training: only train the cross-attention layer parameters and lock the gradients of other layers to reduce computational cost;
[0015] S8, loss function calculation: The model output is decoded by VAE and compared with the complete video slice to enhance the local restoration capability;
[0016] S9, video input: input the video to be repaired and generate its object mask and text description;
[0017] S10, repaired video output: input the mask video, object mask and text description into the improved generative model, and output the repaired video.
[0018] Preferably, in S2, the aesthetic score adopts a pre-trained visual quality assessment model, and the screening threshold is set to aesthetic score>0.8 and motion amplitude>10 pixels / frame.
[0019] Preferably, in S3, the core object segmentation uses a SAM2 model to generate a frame-by-frame binary mask, and the mask area accounts for 5%-50% of each frame.
[0020] Preferably, in S4, the object text description includes the object material, motion state and spatial relationship.
[0021] Preferably, in S5, in the Mask video, Gaussian blur transition is used for the edge of the blackened area, and the blur radius is ≤5 pixels.
[0022] Preferably, in S7, the gradient locking adopts the LayerFreeze algorithm, and only updates the parameters of the Query-Key matrix of the cross-attention layer.
[0023] Preferably, the specific steps of S6 are:
[0024] S601, VAE compression processing of mask video: The mask video to be repaired is input into the pre-trained variational autoencoder frame by frame. The resolution of each frame of video is H times W RGB format. After compression by the encoder, the output feature dimension of each frame is reduced to one-eighth of the original resolution. The number of channels is fixed to 4, and finally a low-dimensional potential feature sequence containing the time dimension T is obtained to preserve the spatiotemporal information of the video.
[0025] S602, downsampling alignment of object masks: adjust the resolution of the binary object mask of each frame to match the feature size after VAE compression, and use bilinear interpolation to downsample the mask to ensure a smooth transition of the binary information of each pixel. The size of the downsampled mask is consistent with the VAE output feature, which is convenient for subsequent stitching operations;
[0026] S603, Feature Concatenation and Model Input: Concatenate the VAE-compressed Mask video features and the downsampled object Mask along the channel dimension to form a fused feature tensor. At the same time, generate the original input of the large model—the random noise matrix—keeping the original dimension and value range unchanged. The concatenated fused features and random noise are input together into the backbone network of the large model, such as the Unet or Dit structure, as the core input of the model.
[0027] S604, hierarchical cross-attention mechanism injection of text features: In the first 80% of the network layers of the large model, the overall description text of the video is converted into a high-dimensional feature vector through the pre-trained text encoder. This feature vector interacts with the spatial features of the current layer through the cross-attention mechanism, focusing on the generation of global scene information, such as background lighting and overall composition. In the last 20% of the network layers of the model, the specific description text of the segmented objects is injected in the same way, but the attention weight is significantly increased, forcing the model to focus on the restoration of local details, such as object edges and textures.
[0028] S605. Dynamic collaborative optimization of local and global features: During the forward propagation of the model, the first 80% layers use high-weight video description features to guide the generation of globally consistent content. For example, the overall color of the repaired area matches the surrounding environment. The last 20% layers enhance local details through object description features. For example, the motion trajectory of the repaired object is consistent with the physical laws of the original video. At the same time, the downsampled object mask serves as a spatial constraint to ensure that the generated content is strictly limited to the target area, avoiding excessive modification of the background. Ultimately, the model outputs repair features that have both global naturalness and local refinement.
[0029] Preferably, in S10, when the repaired video is output, a timing consistency check is performed on the repaired area, and abnormal frames with adjacent frame PSNR < 30dB are removed.
[0030] Preferably, in said S7, the specific steps are as follows:
[0031] S701. Definition of model parameter gradient locking range: In generating a large model, except for the cross-attention layer, the parameters of all other network layers are marked as non-trainable;
[0032] S702. Gradient locking implementation method: Utilize the automatic gradient calculation function of the deep learning framework to traverse all model parameters. For parameters of non-cross attention layers, set their requires_grad attribute to False to prevent gradient calculation and update.
[0033] S703, optimizer configuration and parameter binding: select an adaptive optimizer (such as AdamW) and pass only the trainable cross-attention layer parameters into the optimizer;
[0034] S704. Dynamic identification of cross-attention layers: Different strategies are used to locate cross-attention layers based on the structural type of the generated large model. The Unet structure identifies the cross-attention module in the skip connection, and the Dit structure locates the cross-attention layer in the Transformer block that interacts with the text condition.
[0035] S705. Monitoring and parameter adjustment during training: Verify that the gradients of the non-cross-attention layers are all zero at the beginning of training to ensure successful locking. If the training loss does not decrease within 5 epochs, gradually increase the learning rate of the cross-attention layer by 10%, with an upper limit of 1e-4. When the PSNR indicator of the validation set fluctuates less than 0.1dB for 3 consecutive epochs, terminate the training early.
[0036] The beneficial effects of the present invention are:
[0037] 1. In the present invention, a video restoration dataset construction and restoration method combined with the generation of a large model is described. Through the layered cross-attention mechanism, the model realizes refined restoration in the coordinated optimization of global and local features. The first 80% of the network layers inject the overall description features of the video to ensure that the background lighting and color distribution of the repaired area are physically consistent with the original scene; the last 20% of the layers introduce object-level text descriptions to enhance the restoration of details such as material texture and motion trajectory. For example, the object material description guides the model to generate semantically consistent reflection or transparency features through cross-attention, while the motion state description constrains the dynamic rationality of the repaired object to avoid conflict with the physical laws of the original video. At the same time, the object mask acts as a spatial constraint to accurately limit the generation range and prevent the background area from being modified by mistake. This staged feature injection strategy takes into account both global naturalness and local realism, significantly improving the visual consistency between the repaired content and the original video.
[0038] 2. In the present invention, a video restoration dataset construction and restoration method combined with the generation of a large model adopts a gradient locking strategy to freeze the non-critical layer parameters in the generated large model, and only trains the cross-attention layer parameters, which greatly reduces the video memory usage and computing cost. The LayerFreeze algorithm identifies and locks the gradient updates of non-essential layers, such as the basic convolutional layer or the Transformer backbone layer, and only allows the Query-Key matrix parameters of the cross-attention layer to participate in the optimization. This design reduces the number of trainable parameters by about 70% while ensuring model performance, allowing the training process to be completed on consumer-grade graphics cards. In addition, the dynamic learning rate adjustment mechanism automatically adjusts the learning rate of the cross-attention layer according to the change in loss in the early stage of training, avoiding invalid iterations and further shortening the training cycle. The combination of gradient locking and optimizer parameter binding technology achieves a balance between efficient resource utilization and model performance.
[0039] 3. In the present invention, the video restoration dataset construction and restoration method combined with the generation of a large model, based on the frame-by-frame object mask generated by the multimodal segmentation model, combines semantic information such as material and spatial relationship in the text description to accurately guide the generation of local details, ensuring that the core object is reasonably proportional in the picture and avoiding distortion caused by being too small or too large. In addition, the mask video edge uses Gaussian blur transition to eliminate the abrupt boundary between the restoration area and the background. This "semantic + spatial" dual constraint mechanism effectively solves the problem of blurred details or logical contradictions in traditional methods;
[0040] 4. In the present invention, the construction and restoration method of a video restoration dataset combined with a large generation model ensures the dynamic coherence of the restored video in the temporal dimension through motion amplitude screening and abnormal frame verification mechanisms. In the training phase, high-dynamic segments are screened to force the model to learn complex temporal features such as motion blur and deformation. In the inference phase, adjacent frame PSNR detection is performed on the restoration results, mutation frames are removed and interpolation is performed to complete them. At the same time, the low-dimensional latent features after VAE compression retain the spatiotemporal information of the video, allowing the generation model to implicitly learn the relationship between frames. For example, the motion trajectory of an object is modeled through the temporal continuity of the latent features to avoid jumping or jittering of the restored object. In addition, the injection of spatial relationships in text descriptions further constrains the temporal rationality of multi-object interactions and enhances the overall smoothness of the video.
[0041] 5. In the present invention, a video restoration dataset construction and restoration method combined with the generation of a large model, dynamic learning rate adjustment and early stopping mechanism intelligently balance the model convergence speed and overfitting risk, verify the non-cross-attention layer gradient lock state at the beginning of training to ensure that the parameters are frozen correctly; adaptively increase the cross-attention layer learning rate according to the gradient amplitude to accelerate model convergence, and automatically terminate training when the verification set PSNR index fluctuation is less than the threshold to avoid invalid iterations. At the same time, the layered text feature injection strategy clarifies the optimization goals of each stage and reduces training oscillations caused by feature conflicts. This "monitoring-feedback-adjustment" closed-loop mechanism enables the model to achieve optimal performance in a shorter time, which is more efficient than full-parameter training. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flowchart of a video restoration dataset construction and restoration method proposed by the present invention in combination with generating a large model. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0044] Reference Figure 1 ,A video restoration dataset construction and restoration method combined with generating a large model, comprising the following steps:
[0045] S1, slicing: Select high-quality original video and perform scene-consistent slicing using the PySceneDetect algorithm;
[0046] S2, screening: aesthetic scoring and motion detection are performed on the slices to select and retain high-dynamic and high-quality clips;
[0047] S3. Core object segmentation: Use a multimodal model to accurately segment the core objects in each frame and generate object masks;
[0048] S4. Extract text features: Extract the text features of the video clip description and the segmented objects through the text generation model to form a single description containing the original video slice, object mask, object text description and the overall video description;
[0049] S5, model input: input random noise, object mask and mask video;
[0050] S6, Feature Fusion: Mask video is compressed by VAE and concatenated with the downsampled object mask. The input is used to generate a large model. The video description is injected into the first 80% layers of the model through the cross-attention layer, and the object description is injected into the last 20% layers to achieve collaborative optimization of local and global features.
[0051] S7, model training: only train the cross-attention layer parameters and lock the gradients of other layers to reduce computational cost;
[0052] S8, loss function calculation: The model output is decoded by VAE and compared with the complete video slice to enhance the local restoration capability;
[0053] S9, video input: input the video to be repaired and generate its object mask and text description;
[0054] S10, repaired video output: input the mask video, object mask and text description into the improved generative model, and output the repaired video.
[0055] In this embodiment, in S2, the aesthetic score adopts a pre-trained visual quality assessment model, and the screening threshold is set to aesthetic score>0.8 and motion amplitude>10 pixels / frame.
[0056] In this embodiment, in S3, the core object segmentation uses the SAM2 model to generate a frame-by-frame binary mask, and the mask area accounts for 5%-50% of each frame.
[0057] In this embodiment, in S4, the object text description includes the object material, motion state and spatial relationship.
[0058] In this embodiment, in S5, in the Mask video, the edge of the blackened area adopts Gaussian blur transition, and the blur radius is ≤5 pixels.
[0059] In this embodiment, in S7, the gradient locking adopts the LayerFreeze algorithm, and only the Query-Key matrix of the cross attention layer is updated.
[0060] In this embodiment, the specific steps of S6 are:
[0061] S601, VAE compression processing of mask video: The mask video to be repaired is input into the pre-trained variational autoencoder frame by frame. The resolution of each frame of video is H times W RGB format. After compression by the encoder, the output feature dimension of each frame is reduced to one-eighth of the original resolution. The number of channels is fixed to 4, and finally a low-dimensional potential feature sequence containing the time dimension T is obtained to preserve the spatiotemporal information of the video.
[0062] S602, downsampling alignment of object masks: adjust the resolution of the binary object mask of each frame to match the feature size after VAE compression, and use bilinear interpolation to downsample the mask to ensure a smooth transition of the binary information of each pixel. The size of the downsampled mask is consistent with the VAE output feature, which is convenient for subsequent stitching operations;
[0063] S603, Feature Concatenation and Model Input: Concatenate the VAE-compressed Mask video features and the downsampled object Mask along the channel dimension to form a fused feature tensor. At the same time, generate the original input of the large model—the random noise matrix—keeping the original dimension and value range unchanged. The concatenated fused features and random noise are input together into the backbone network of the large model, such as the Unet or Dit structure, as the core input of the model.
[0064] S604, hierarchical cross-attention mechanism injection of text features: In the first 80% of the network layers of the large model, the overall description text of the video is converted into a high-dimensional feature vector through the pre-trained text encoder. This feature vector interacts with the spatial features of the current layer through the cross-attention mechanism, focusing on the generation of global scene information, such as background lighting and overall composition. In the last 20% of the network layers of the model, the specific description text of the segmented objects is injected in the same way, but the attention weight is significantly increased, forcing the model to focus on the restoration of local details, such as object edges and textures.
[0065] S605. Dynamic collaborative optimization of local and global features: During the forward propagation of the model, the first 80% layers use high-weight video description features to guide the generation of globally consistent content. For example, the overall color of the repaired area matches the surrounding environment. The last 20% layers enhance local details through object description features. For example, the motion trajectory of the repaired object is consistent with the physical laws of the original video. At the same time, the downsampled object mask serves as a spatial constraint to ensure that the generated content is strictly limited to the target area, avoiding excessive modification of the background. Ultimately, the model outputs repair features that have both global naturalness and local refinement.
[0066] In this embodiment, in S10, when the repaired video is output, a timing consistency check is performed on the repaired area, and abnormal frames with adjacent frame PSNR less than 30dB are eliminated.
[0067] In this embodiment, in S7, the specific steps are as follows:
[0068] S701. Definition of model parameter gradient locking range: In generating a large model, except for the cross-attention layer, the parameters of all other network layers are marked as non-trainable;
[0069] S702. Gradient locking implementation method: Utilize the automatic gradient calculation function of the deep learning framework to traverse all model parameters. For parameters of non-cross attention layers, set their requires_grad attribute to False to prevent gradient calculation and update.
[0070] S703, optimizer configuration and parameter binding: select an adaptive optimizer (such as AdamW) and pass only the trainable cross-attention layer parameters into the optimizer;
[0071] S704. Dynamic identification of cross-attention layers: Different strategies are used to locate cross-attention layers based on the structural type of the generated large model. The Unet structure identifies the cross-attention module in the skip connection, and the Dit structure locates the cross-attention layer in the Transformer block that interacts with the text condition.
[0072] S705. Monitoring and parameter adjustment during training: Verify that the gradients of the non-cross-attention layers are all zero at the beginning of training to ensure successful locking. If the training loss does not decrease within 5 epochs, gradually increase the learning rate of the cross-attention layer by 10%, with an upper limit of 1e-4. When the PSNR indicator of the validation set fluctuates less than 0.1dB for 3 consecutive epochs, terminate the training early.
[0073] In this embodiment, a scene detection algorithm is first used to automatically slice high-quality original videos, extracting video segments with consistent scenes. High-dynamic and high-quality segments are then screened and retained using a pre-trained aesthetic evaluation model and motion detection algorithm. A multimodal segmentation model is then used to segment the core objects in each frame at the pixel level and generate frame-by-frame binary masks. A text generation model is then used to extract the overall video description and object-level attribute features. When constructing a dataset, the mask video is compressed using a variational autoencoder and then spliced along the channel dimension with the downsampled aligned object mask. Text features are then injected into different network layers of the generative model using a hierarchical cross-attention mechanism. The global description of the video guides the front network layer to generate background lighting and composition, while the object material and motion description drives the back network layer to optimize local details. During the model training phase, the parameters of the non-cross-attention layers are frozen and only the attention matrix is fine-tuned. Dynamic learning rate adjustment and early stopping are used to accelerate convergence. During actual restoration, the video to be restored undergoes the same preprocessing and is input into the improved generative model. The output is subjected to a temporal consistency check to remove abnormal frames, resulting in a restored video that combines global naturalness with local refinement.
[0074] The above is a detailed introduction to the video restoration dataset construction and restoration method provided by the present invention in combination with the generation of a large model. This article uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A video restoration dataset construction and restoration method combined with the generation of a large model, characterized in that: The following steps are involved: S1, slicing: Select high-quality original video and perform scene-consistent slicing using the PySceneDetect algorithm; S2, screening: aesthetic scoring and motion detection are performed on the slices to select and retain high-dynamic and high-quality clips; S3. Core object segmentation: Use a multimodal model to accurately segment the core objects in each frame and generate object masks; S4. Extract text features: Extract the text features of the video clip description and the segmented objects through the text generation model to form a single description containing the original video slice, object mask, object text description and the overall video description; S5, model input: input random noise, object mask and mask video; S6, Feature Fusion: Mask video is compressed by VAE and concatenated with the downsampled object mask. The input is used to generate a large model. The video description is injected into the first 80% layers of the model through the cross-attention layer, and the object description is injected into the last 20% layers to achieve collaborative optimization of local and global features. S7, model training: only train the cross-attention layer parameters and lock the gradients of other layers to reduce computational cost; S8, loss function calculation: The model output is decoded by VAE and compared with the complete video slice to enhance the local restoration capability; S9, video input: input the video to be repaired and generate its object mask and text description; S10, repaired video output: input the mask video, object mask and text description into the improved generative model, and output the repaired video.
2. The video restoration dataset construction and restoration method combined with the generation of a large model according to claim 1 is characterized by: In S2, the aesthetic score adopts a pre-trained visual quality assessment model, and the screening threshold is set as aesthetic score>0.8 and motion amplitude>10 pixels / frame.
3. The video restoration dataset construction and restoration method combined with the generation of a large model according to claim 1 is characterized by: In S3, the core object segmentation uses the SAM2 model to generate a frame-by-frame binary mask, and the mask area accounts for 5%-50% of each frame.
4. The video restoration dataset construction and restoration method combined with the generation of a large model according to claim 1 is characterized by: In S4, the object text description includes the object material, motion state and spatial relationship.
5. The video restoration dataset construction and restoration method combined with the generation of a large model according to claim 1 is characterized by: In S5, in the Mask video, the edge of the blackened area adopts Gaussian blur transition, and the blur radius is ≤5 pixels.
6. The video restoration dataset construction and restoration method combined with the generation of a large model according to claim 1 is characterized by: In S7, the gradient locking adopts the LayerFreeze algorithm, and only updates the parameters of the Query-Key matrix of the cross-attention layer.
7. The method for constructing and restoring a video restoration dataset in combination with generating a large model according to claim 1, characterized in that: The specific steps of S6 are: S601, VAE compression processing of mask video: The mask video to be repaired is input into the pre-trained variational autoencoder frame by frame. The resolution of each frame of video is H times W RGB format. After compression by the encoder, the output feature dimension of each frame is reduced to one-eighth of the original resolution. The number of channels is fixed to 4, and finally a low-dimensional potential feature sequence containing the time dimension T is obtained to preserve the spatiotemporal information of the video. S602, downsampling alignment of object masks: adjust the resolution of the binary object mask of each frame to match the feature size after VAE compression, and use bilinear interpolation to downsample the mask to ensure a smooth transition of the binary information of each pixel. The size of the downsampled mask is consistent with the VAE output feature, which is convenient for subsequent stitching operations; S603, Feature Concatenation and Model Input: Concatenate the VAE-compressed Mask video features and the downsampled object Mask along the channel dimension to form a fused feature tensor. At the same time, generate the original input of the large model—the random noise matrix—keeping the original dimension and value range unchanged. The concatenated fused features and random noise are input together into the backbone network of the large model, such as the Unet or Dit structure, as the core input of the model. S604, hierarchical cross-attention mechanism injection of text features: In the first 80% of the network layers of the large model, the overall description text of the video is converted into a high-dimensional feature vector through the pre-trained text encoder. This feature vector interacts with the spatial features of the current layer through the cross-attention mechanism, focusing on the generation of global scene information, such as background lighting and overall composition. In the last 20% of the network layers of the model, the specific description text of the segmented objects is injected in the same way, but the attention weight is significantly increased, forcing the model to focus on the restoration of local details, such as object edges and textures. S605. Dynamic collaborative optimization of local and global features: During the forward propagation of the model, the first 80% layers use high-weight video description features to guide the generation of globally consistent content. For example, the overall color of the repaired area matches the surrounding environment. The last 20% layers enhance local details through object description features. For example, the motion trajectory of the repaired object is consistent with the physical laws of the original video. At the same time, the downsampled object mask serves as a spatial constraint to ensure that the generated content is strictly limited to the target area, avoiding excessive modification of the background. Ultimately, the model outputs repair features that have both global naturalness and local refinement.
8. The video restoration dataset construction and restoration method combined with the generation of a large model according to claim 1 is characterized by: In S10, when the repaired video is output, a timing consistency check is performed on the repaired area, and abnormal frames with adjacent frame PSNR < 30dB are removed.
9. The method for constructing and restoring a video restoration dataset in combination with generating a large model according to claim 1, characterized in that: In said S7, the specific steps are as follows: S701. Definition of model parameter gradient locking range: In generating a large model, except for the cross-attention layer, the parameters of all other network layers are marked as non-trainable; S702. Gradient locking implementation method: Utilize the automatic gradient calculation function of the deep learning framework to traverse all model parameters. For parameters of non-cross attention layers, set their requires_grad attribute to False to prevent gradient calculation and update. S703, optimizer configuration and parameter binding: select an adaptive optimizer (such as AdamW) and pass only the trainable cross-attention layer parameters into the optimizer; S704, Dynamic Identification of Cross-Attention Layers: Different strategies are used to locate cross-attention layers based on the structural type of the generated large model. The Unet structure identifies the cross-attention modules in the skip connection, and the Dit structure locates the cross-attention layers in the Transformer block that interact with the text condition. S705. Monitoring and parameter adjustment during training: Verify that the gradients of the non-cross-attention layers are all zero at the beginning of training to ensure successful locking. If the training loss does not decrease within 5 epochs, gradually increase the learning rate of the cross-attention layer by 10%, with an upper limit of 1e-4. When the PSNR indicator of the validation set fluctuates less than 0.1dB for 3 consecutive epochs, terminate the training early.
Citation Information
Cited By
Energy-saving and consumption-reducing method for air blower of sewage treatment plant based on machine learning
CN121541475A
Virtual reality data acquisition and restoration method
CN121707869A