Deep learning enhancement model updating method and system based on image quality evaluation

By introducing the ControlNet architecture and image quality assessment model and dynamically adjusting the weight of the image enhancement model, the problem of insufficient model generalization in power equipment inspection is solved, the image enhancement effect and fault detection accuracy are improved, and it can adapt to complex environments.

CN120672765AActive Publication Date: 2025-09-19STATE GRID HUBEI ELECTRIC POWER CO LTD +3
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511189646.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-19
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing deep learning models have insufficient generalization performance in power equipment inspections due to inconsistencies between training data and actual scenarios. In addition, existing technologies make it difficult to achieve real-time updates and obtain high-quality reference images, affecting the accuracy and automation level of fault detection.

Method used

The ControlNet architecture and image quality assessment model are introduced, and the image enhancement model is improved through auxiliary modules and conditional perception cross entropy loss function. The weights are dynamically adjusted, and the model parameters are optimized in real time in combination with the multimodal loss function to ensure the adaptability and stability of the model in complex environments.

Benefits of technology

It improves the generalization ability and effect of the image enhancement model in actual scenarios, realizes the automation and stability of images, reduces the dependence on manually labeled data, and improves the accuracy of fault detection and the ability to adapt to complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672765A_ABST
    Figure CN120672765A_ABST
Patent Text Reader

Abstract

The invention provides a deep learning enhancement model updating method and system based on image quality assessment, and the method comprises the steps: integrating a ControlNet architecture to an image enhancement model, and improving the image enhancement model through adding an auxiliary module and a conditional perception cross entropy loss function. The auxiliary module is used for introducing an additional condition to generate a control signal so as to dynamically adjust a generation strategy, and the improved image enhancement model generates an enhanced image according to an input additional condition and an input image; inputting the enhanced image into an image quality evaluation model, and outputting a quality score and text description; and if the quality score is lower than a preset threshold value, designing a multi-modal loss function according to the quality score output by the image quality evaluation model and the description text, and adjusting a weight parameter of the improved image enhancement model in real time through back propagation. The weight information of the image enhancement model is adjusted in real time according to the evaluation result, so that the generalization ability and the image enhancement effect of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image quality assessment, and in particular to a method and system for updating a deep learning enhancement model based on image quality assessment. Background Art

[0002] The power system is a vital component of national infrastructure, and its safe and stable operation relies on efficient equipment inspection technology. In recent years, with the widespread use of intelligent inspection equipment such as drones and robots, power inspection has gradually shifted from manual to automated and intelligent modes. However, inspection images of power equipment often exhibit problems such as blurring, low contrast, or noise interference due to complex environments, which directly affects the accuracy of subsequent fault detection. At the same time, with the widespread application of deep learning technology in the field of image enhancement, existing deep learning models often exhibit significant discrepancies between training and actual use scenarios. This is because the dataset used during training is inconsistent with the image distribution in actual application scenarios, resulting in insufficient generalization performance of the model.

[0003] To address this issue, research over the years has primarily focused on static image enhancement algorithms and deep learning models with fixed parameters. For example, enhancement models based on convolutional neural networks (CNNs) improve image quality through end-to-end training, but their weights are fixed after deployment and cannot be dynamically adjusted according to actual scenarios. While adaptive methods that introduce reinforcement learning can partially address the generalization problem, they rely on manually designed reward functions and large trial-and-error costs, resulting in low efficiency. Furthermore, existing technologies often rely on manually labeled "true values" as optimization targets, but this approach is not only time-consuming and labor-intensive, but also difficult to achieve real-time updates and the acquisition of high-quality reference images, resulting in a lack of effective basis for model updates. These shortcomings make it difficult for existing technologies to balance real-time, robustness, and automation requirements in complex power scenarios, limiting the further application of intelligent inspection technologies. Summary of the Invention

[0004] The purpose of the present invention is to address the above-mentioned existing problems and propose a real-time updating method and system for a deep learning enhancement model based on image quality assessment. For images of power equipment in different environments and fault types, the method aims to automatically evaluate the quality of the enhanced image by introducing an image quality assessment model, and adjust the weight information of the image enhancement model in real time according to the evaluation results, thereby improving the generalization ability of the model and the image enhancement effect.

[0005] A deep learning enhancement model updating method based on image quality assessment includes the following steps:

[0006] Integrating a ControlNet architecture into an image enhancement model. The ControlNet architecture improves the image enhancement model by adding an auxiliary module and a condition-aware cross-entropy loss function. The auxiliary module is used to introduce additional conditions to generate a control signal to dynamically adjust the generation strategy. The improved image enhancement model generates an enhanced image based on the additional conditions and the input image.

[0007] The enhanced image is input into the image quality assessment model, which outputs a quality score and text description;

[0008] If the quality score is lower than the preset threshold, a multimodal loss function is designed based on the quality score and descriptive text output by the image quality assessment model, and the weight parameters of the improved image enhancement model are adjusted in real time through back propagation.

[0009] Furthermore, the auxiliary module includes a locked copy and a trainable copy. The locked copy is used to copy the structure and weights of the image enhancement model. The parameters remain fixed during the training process, retaining the generation capability of the image enhancement model. The trainable copy has the same structure as the image enhancement model, and the parameters can be fine-tuned. The control signal is generated through additional conditions to dynamically adjust the generation strategy.

[0010] Furthermore, the trainable replica generates a control signal through additional conditions, specifically including:

[0011]

[0012] in, is the generated control signal tensor, are the height, width and number of channels of the feature map, respectively. To fine-tune the model, is the input image;

[0013] control signal Intermediate feature map with the main model Interacting through adaptive feature fusion enables the model to dynamically adjust the generation process based on additional conditions when generating images:

[0014]

[0015] in, is a learnable scaling factor that dynamically adjusts the strength of the control signal.

[0016] Furthermore, the loss function of the image enhancement model is For pixel-level loss or perceptual loss, we introduce the conditional perceptual cross entropy loss function Forming a mixed loss function:

[0017]

[0018] in, It is a dynamically adjusted weight coefficient, which is larger in the initial stage and gradually decreases in the later stage;

[0019] The design of the conditional perception cross entropy loss function is as follows: Assume that the control signal output by ControlNet is It is necessary to guide the model to generate images that meet the quality conditions and define the target distribution and the model predicts distribution The cross entropy loss between:

[0020]

[0021] in, The probability map output by ControlNet is obtained by Softmax normalization , dynamically generated according to the quality assessment results;

[0022] During the pre-training main model phase, only Train the original image enhancement model to ensure basic generation capabilities; secondly, freeze the main model and fix the main model weights when training the ControlNet model. Optimize the trainable copy of ControlNet and learn to generate control signals that adapt to additional conditions; finally, in the joint fine-tuning stage, unfreeze the main model and use a hybrid loss Jointly optimize the main model and ControlNet parameters to improve the overall performance of the image enhancement model.

[0023] Furthermore, the step of inputting the enhanced image into an image quality assessment model and outputting a quality score and a text description includes:

[0024] The introduction of a text-guided mechanism enables the image quality assessment model to simultaneously output an image quality score and a detailed text description, helping to understand the problems in the image.

[0025] The texture perception module based on the Transformer structure performs weight analysis on each token in the image, determines the contribution weight of each token in the image quality assessment, and evaluates the local features of the image;

[0026] The global quality assessment module fuses all tokens to generate global weights to evaluate the local features of the image;

[0027] Combining local and global weights to generate image quality scores and text descriptions.

[0028] Furthermore, the additional conditions are multi-dimensional external information used to dynamically control the image generation process, and are used to guide the model to generate enhanced images that are more in line with the target according to specific scene requirements.

[0029] Furthermore, the multimodal loss function is designed based on the quality score and description text output by the image quality assessment model, specifically including:

[0030] Combining the quality score loss and text generation loss, a multimodal loss function is designed to guide model updates:

[0031]

[0032] in, For quality evaluation loss, for text generation loss, and is a dynamic weight coefficient used to balance the optimization direction of quality score and text description;

[0033]

[0034] in, is the model prediction score, For true quality labels, and are their means respectively;

[0035] The text description loss converts the text description Encoded as a feature vector , and with image features Alignment:

[0036]

[0037] Update the weights of the improved image enhancement model according to the multimodal loss function, and calculate the parameters of the improved image enhancement model by the chain rule The gradient of , uses the optimizer to update the model weights:

[0038] .

[0039] A real-time updating system for a deep learning enhancement model based on image quality assessment, comprising:

[0040] An image enhancement module, configured to integrate a ControlNet architecture into the original image enhancement model. The ControlNet architecture improves the original image enhancement model by adding an auxiliary module and a cross-entropy-based loss function. The auxiliary module is configured to introduce additional conditions to generate a control signal to dynamically adjust the generation strategy. The improved image enhancement model generates an enhanced image based on the additional conditions and the input image.

[0041] The quality assessment module is used to input the enhanced image into the image quality assessment model and output the quality score and text description;

[0042] The dynamic update module is used to design a multimodal loss function based on the quality score and description text output by the image quality assessment model if the quality score is lower than the preset threshold, and to adjust the weight parameters of the improved image enhancement model in real time through back propagation.

[0043] Furthermore, the auxiliary module includes a locked copy and a trainable copy. The locked copy is used to replicate the structure and weights of the image enhancement model. The parameters remain unchanged during the training process, retaining the generation capability of the image enhancement model. The trainable copy has the same structure as the image enhancement model and can be fine-tuned. The control signal is generated through additional conditions to dynamically adjust the generation strategy.

[0044] The trainable replica generates a control signal through additional conditions, specifically including:

[0045]

[0046] in, is the generated control signal tensor, are the height, width and number of channels of the feature map, respectively. To fine-tune the model, is the input image;

[0047] control signal Intermediate feature map with the main model Interacting through adaptive feature fusion enables the model to dynamically adjust the generation process based on additional conditions when generating images:

[0048]

[0049] in, is a learnable scaling factor that dynamically adjusts the strength of the control signal.

[0050] Furthermore, the loss function of the image enhancement model is For pixel-level loss or perceptual loss, we introduce the conditional perceptual cross entropy loss function Forming a mixed loss function:

[0051]

[0052] in, It is a dynamically adjusted weight coefficient, which is larger in the initial stage and gradually decreases in the later stage;

[0053] The design of the conditional perception cross entropy loss function is as follows: Assume that the control signal output by ControlNet is It is necessary to guide the model to generate images that meet the quality conditions and define the target distribution and the model predicts distribution The cross entropy loss between:

[0054]

[0055] in, The probability map output by ControlNet is obtained by Softmax normalization , dynamically generated according to the quality assessment results;

[0056] During the pre-training main model phase, only Train the original image enhancement model to ensure basic generation capabilities; secondly, freeze the main model and fix the main model weights when training the ControlNet model. Optimize the trainable copy of ControlNet and learn to generate control signals that adapt to additional conditions; finally, in the joint fine-tuning stage, unfreeze the main model and use a hybrid loss Jointly optimize the main model and ControlNet parameters to improve the overall performance of the image enhancement model.

[0057] Compared with the existing technology, the present invention improves the effect and generalization ability of image enhancement by optimizing the weights of the enhancement model in real time in actual usage scenarios. The beneficial effects include:

[0058] 1. Introducing the ControlNet architecture into an existing image enhancement model and locking the weights of the replica from updates ensures the core capabilities of the original model remain intact, improving model stability. Furthermore, an improved hybrid loss function further enhances model performance. This solution dynamically adjusts model weights in real-world applications, improving image enhancement while reducing reliance on manually labeled data, significantly increasing the model's automation and adaptability.

[0059] 2. The image quality assessment model evaluates enhanced images in real time, outputs a quality score and text description, and dynamically adjusts model weights based on the evaluation results. This enables continuous optimization of the model for practical use scenarios, improving the model's generalization capabilities and image enhancement effects.

[0060] 3. The image enhancement model can be continuously optimized in real-world scenarios, gradually improving its adaptability to complex environments (such as low light and high noise), and more meticulously assessing the local and global quality of images, ensuring the model's adaptability in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1Schematic diagram of the structure of an image quality assessment model in an embodiment of the present invention;

[0062] Figure 2 Schematic diagram of the structure of the global quality assessment module in an embodiment of the present invention;

[0063] Figure 3 This is a flowchart of a method for updating a deep learning enhancement model based on image quality assessment according to an embodiment of the present invention. DETAILED DESCRIPTION

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0065] See also Figure 1-3 , an embodiment of the present invention provides a method for updating a deep learning enhancement model based on image quality assessment, comprising the following steps:

[0066] Step 1: Integrate the ControlNet architecture into the image enhancement model. The ControlNet architecture improves the image enhancement model by adding auxiliary modules and a condition-aware cross-entropy loss function. The auxiliary modules are used to introduce additional conditions to generate control signals to dynamically adjust the generation strategy. The improved image enhancement model generates an enhanced image based on the input additional conditions and the input image. The specific implementation process is as follows:

[0067] ControlNet is an auxiliary neural network structure that controls the image generation process by adding auxiliary modules to the model and introducing additional conditions. It mainly consists of two parts: the first part is a locked copy that replicates the structure and weights of the original image enhancement model, but the parameters remain unchanged during the training process to retain the generation ability of the original model; the second part is a trainable copy, which has the same structure as the original model but with fine-tunable parameters, and generates control signals through additional conditions (such as quality scores or text descriptions).

[0068] After inputting the image and additional conditions, the output control signal is generated after being processed by the trainable copy:

[0069]

[0070] in, is the generated control signal tensor, are the height, width and number of channels of the feature map, respectively. To fine-tune the model, is the input image;

[0071] control signal Intermediate feature map with the main model Interacting through adaptive feature fusion allows the model to dynamically adjust the generation process based on additional conditions (such as image quality feedback) when generating images:

[0072]

[0073] in, is a learnable scaling factor used to dynamically adjust the strength of the control signal. The primary model refers to the original, unmodified basic image enhancement model using the ControlNet architecture; the improved image enhancement model is a combination of the primary model and the ControlNet architecture.

[0074] In order to further improve the performance of the model, a loss function based on cross entropy is used for improvement. The loss function of the original model is Typically pixel-level loss (e.g. Loss) or perceptual loss, the present invention introduces conditional perceptual cross entropy loss , forming a mixed loss function:

[0075]

[0076] in, It is a dynamically adjusted weight coefficient, which is larger in the initial stage and gradually decreases in the later stage.

[0077] The design of conditional-aware cross entropy loss is as follows: Assume that the control signal output by ControlNet is It is necessary to guide the model to generate images that meet the quality conditions and define the target distribution (such as the distribution of high-quality images) and the model's predicted distribution The cross entropy loss between:

[0078]

[0079] in, The probability map output by ControlNet is obtained by Softmax normalization , which is dynamically generated based on the quality assessment results (e.g. the probability of a high-scoring area is 1, and the probability of a low-scoring area is 0).

[0080] During the pre-training main model phase, only Train the original image enhancement model to ensure basic generation capabilities; secondly, when freezing the main model and training the ControlNet model, fix the main model weights, through Optimize the trainable copy of ControlNet and learn to generate control signals that adapt to additional conditions; finally, in the joint fine-tuning stage, unfreeze the main model and use a hybrid loss Jointly optimize the main model and ControlNet parameters to improve the overall performance of the image enhancement model.

[0081] Step 2: Input the enhanced image into the image quality assessment model, which outputs a quality score and text description to provide more comprehensive quality feedback. The specific implementation process of Step 2 is as follows:

[0082] Step 2.1: Introduce a text-guided mechanism so that the image quality assessment model can not only output a quality score, but also generate natural language text describing the image problem, achieving dual output of image quality score and text description:

[0083] The text generation branch is introduced into the image quality assessment model. The main body of the model consists of two branches: the scoring branch is based on the CLS Token in the ViT model and outputs the quality score through the fully connected layer regression. S ∈ [ 0 , 1 ] The text generation branch is based on the Transformer decoder, which takes the ViT token sequence as input to generate natural language descriptions. .

[0084] Use the pre-trained vision-language model CLIP to align image features with text features.

[0085] Define the contrast loss function:

[0086]

[0087] in, is the image feature, is the text feature, is the temperature coefficient.

[0088] Use the autoregressive decoder to generate text descriptions and define the text-guided loss function, where the loss function is cross entropy:

[0089]

[0090] in, For the generated words, It is the Token sequence feature output by ViT.

[0091] Step 2.2: Based on the texture perception module of the Transformer structure, perform weight analysis on each token in the image to determine the local contribution weight of each token in the image quality assessment. :

[0092] Enter the ViT token sequence , calculate the weight of each Token through the self-attention mechanism :

[0093]

[0094] in, and For the The query vector and key vector of a Token, is the feature dimension;

[0095] The multi-scale features extracted by the pre-training model are compressed through the convolution layer to obtain local texture feature information. , and weighted fusion with ViT Token:

[0096]

[0097] Among them, MLP is used to project local features to the dimension of ViT.

[0098] Step 2.3: Introduce a global quality assessment module to integrate all token information (such as Figure 2 As shown), generate a global weight Assess the overall quality of the image;

[0099] Perform average pooling on the Token sequence of ViT to obtain the global feature vector :

[0100]

[0101] Generate global weights through fully connected layers and Sigmoid activation function :

[0102]

[0103] in, 、 is a learnable parameter, is the Sigmoid function.

[0104] Step 2.4: Combine the local weights and global weights for analysis and output the final image quality score and text description:

[0105] Use global weights Adjust the scoring branch and output the final quality score:

[0106]

[0107] in, is the historical average score, used for stable training;

[0108] The update formula is:

[0109]

[0110] in, For all Number.

[0111] The fused Token sequence As input, generate the problem description text through the Transformer decoder:

[0112]

[0113] Step 3: Dynamically adjust the weight parameters of the image enhancement model based on the quality score and text description output by the image quality assessment model.

[0114] Set quality score thresholds based on application scenarios (such as power inspection and safety supervision) S th ∈ [ 0 , 1 ] , to determine whether the enhanced image quality reaches the preset threshold. If the final quality score , then keep the current enhanced model weight; if , a multimodal loss function is designed based on the score and text description output by the image quality assessment model, and the weights of the image enhancement model are updated through back propagation.

[0115] Combining the quality score loss and text generation loss, a multimodal loss function is designed to guide model updates:

[0116]

[0117] in, For quality evaluation loss, for text generation loss, , is a dynamic weight coefficient that balances the optimization direction of quality score and text description. It is higher to strengthen the text guidance and gradually decreases in the later stage.

[0118] For example, when a keyword triggers a defect label, the corresponding sub-loss weight will increase, causing the enhanced model to be more severely penalized in this regard, thereby improving these deficiencies in a targeted manner.

[0119] The quality score loss PLCC (Pearson linear correlation coefficient) loss measures the consistency between the predicted score and the manual annotation and is defined as:

[0120]

[0121] in, is the model prediction score, For true quality labels, and are their mean values ​​respectively.

[0122] The text description loss converts the text description Encoded as a feature vector , and with image features Alignment:

[0123]

[0124] Update the image enhancement model weights according to the loss function and calculate the image enhancement model parameters using the chain rule The gradient of , uses the optimizer to update the model weights:

[0125]

[0126] The weights of ControlNet's "locked copy" are fixed, and only the backbone parameters of the "trainable copy" and image enhancement model are updated to ensure that the basic generation capabilities are not damaged.

[0127] An embodiment of the present invention further provides a real-time update system for a deep learning enhancement model based on image quality assessment, comprising:

[0128] An image enhancement module, configured to integrate a ControlNet architecture into the original image enhancement model. The ControlNet architecture improves the original image enhancement model by adding an auxiliary module and a cross-entropy-based loss function. The auxiliary module is configured to introduce additional conditions to generate a control signal to dynamically adjust the generation strategy. The improved image enhancement model generates an enhanced image based on the additional conditions and the input image.

[0129] The quality assessment module is used to input the enhanced image into the image quality assessment model and output the quality score and text description;

[0130] The dynamic update module is used to design a multimodal loss function based on the quality score and description text output by the image quality assessment model if the quality score is lower than the preset threshold, and to adjust the weight parameters of the improved image enhancement model in real time through back propagation.

[0131] The present invention has the following characteristics and effects:

[0132] Dynamic adjustment technology based on the ControlNet architecture: This technology introduces auxiliary modules consisting of locked and trainable replicas. The locked replica fixes the original model weights to preserve basic generation capabilities, while the trainable replica generates dynamic control signals based on additional conditions (such as quality scores) and adjusts the generation strategy through adaptive feature fusion. This technology enables real-time optimization of model weights based on actual scenarios after deployment, significantly improving the adaptability and stability of image enhancement while reducing reliance on manually annotated data.

[0133] Hybrid loss function optimization technology: This technology combines the conditional-aware cross-entropy loss function with traditional pixel-level / perceptual loss, dynamically adjusting weight coefficients (initially focusing on conditional guidance and later enhancing generative capabilities). The model is optimized in stages (pre-training the main model, training the ControlNet, and then joint fine-tuning). This technology enhances the model's ability to fit high-quality image distributions, addresses the generalization issues inherent in traditional methods due to fixed loss functions, and significantly improves enhancement in complex scenes (such as low light and high noise).

[0134] Quality Assessment and Text-Guided Technology: The image quality assessment model uses a Transformer architecture to simultaneously output a quality score and text description. It combines local texture weight analysis with global feature fusion, and designs a multimodal loss function (quality score loss + text generation loss) to guide model updates. This technology uses text feedback to accurately locate image defects (such as blur and low contrast) and optimize model parameters in a targeted manner, achieving more fine-grained quality control and improving the accuracy and automation of fault detection in scenarios such as power inspections.

[0135] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A deep learning enhancement model updating method based on image quality assessment, characterized in that: The following steps are involved: Integrating a ControlNet architecture into an image enhancement model. The ControlNet architecture improves the image enhancement model by adding an auxiliary module and a condition-aware cross-entropy loss function. The auxiliary module is used to introduce additional conditions to generate a control signal to dynamically adjust the generation strategy. The improved image enhancement model generates an enhanced image based on the additional conditions and the input image. The enhanced image is input into the image quality assessment model, which outputs a quality score and text description; If the quality score is lower than the preset threshold, a multimodal loss function is designed based on the quality score and descriptive text output by the image quality assessment model, and the weight parameters of the improved image enhancement model are adjusted in real time through back propagation.

2. The method according to claim 1, characterized in that The auxiliary module includes a locked copy and a trainable copy. The locked copy is used to replicate the structure and weights of the image enhancement model. The parameters remain unchanged during the training process, retaining the generation capability of the image enhancement model. The trainable copy has the same structure as the image enhancement model, and the parameters can be fine-tuned. Control signals are generated through additional conditions to dynamically adjust the generation strategy.

3. The method according to claim 2, characterized in that The trainable replica generates a control signal through additional conditions, specifically including: ; in, is the generated control signal tensor, are the height, width and number of channels of the feature map, respectively. To fine-tune the model, is the input image; control signal Intermediate feature map with the main model Interacting through adaptive feature fusion enables the model to dynamically adjust the generation process based on additional conditions when generating images: ; in, is a learnable scaling factor that dynamically adjusts the strength of the control signal.

4. The method according to claim 1, wherein Loss function for image enhancement model For pixel-level loss or perceptual loss, we introduce the conditional perceptual cross entropy loss function Forming a mixed loss function: ; in, It is a dynamically adjusted weight coefficient, which is larger in the initial stage and gradually decreases in the later stage; The design of the conditional perception cross entropy loss function is as follows: Assume that the control signal output by ControlNet is It is necessary to guide the model to generate images that meet the quality conditions and define the target distribution and the model predicts distribution The cross entropy loss between: ; in, The probability map output by ControlNet is obtained by Softmax normalization , dynamically generated according to the quality assessment results; During the pre-training main model phase, only Train the original image enhancement model to ensure basic generation capabilities; secondly, freeze the main model and fix the main model weights when training the ControlNet model. Optimize the trainable copy of ControlNet and learn to generate control signals that adapt to additional conditions; finally, in the joint fine-tuning stage, unfreeze the main model and use a hybrid loss Jointly optimize the main model and ControlNet parameters to improve the overall performance of the image enhancement model.

5. The method according to claim 1, wherein The enhanced image is input into the image quality assessment model, and the quality score and text description are output, including: The introduction of a text-guided mechanism enables the image quality assessment model to simultaneously output an image quality score and a detailed text description, helping to understand the problems in the image. The texture perception module based on the Transformer structure performs weight analysis on each token in the image, determines the contribution weight of each token in the image quality assessment, and evaluates the local features of the image; The global quality assessment module fuses all tokens to generate global weights to evaluate the local features of the image; Combining local and global weights to generate image quality scores and text descriptions.

6. The method according to claim 1, characterized in that The additional conditions are multi-dimensional external information used to dynamically control the image generation process, and are used to guide the model to generate enhanced images that are more in line with the target according to specific scene requirements.

7. The method according to claim 1, characterized in that The multimodal loss function is designed based on the quality score and description text output by the image quality assessment model, specifically including: Combining the quality score loss and text generation loss, a multimodal loss function is designed to guide model updates: ; in, For quality evaluation loss, for text generation loss, and is a dynamic weight coefficient used to balance the optimization direction of quality score and text description; ; in, is the model prediction score, For true quality labels, and are their means respectively; The text description loss converts the text description Encoded as a feature vector , and with image features Alignment: ; Update the weights of the improved image enhancement model according to the multimodal loss function, and calculate the parameters of the improved image enhancement model by the chain rule The gradient of , uses the optimizer to update the model weights: 。 8. A real-time update system for a deep learning enhancement model based on image quality assessment, characterized in that: include: An image enhancement module, configured to integrate a ControlNet architecture into the original image enhancement model. The ControlNet architecture improves the original image enhancement model by adding an auxiliary module and a cross-entropy-based loss function. The auxiliary module is configured to introduce additional conditions to generate a control signal to dynamically adjust the generation strategy. The improved image enhancement model generates an enhanced image based on the additional conditions and the input image. The quality assessment module is used to input the enhanced image into the image quality assessment model and output the quality score and text description; The dynamic update module is used to design a multimodal loss function based on the quality score and description text output by the image quality assessment model if the quality score is lower than the preset threshold, and to adjust the weight parameters of the improved image enhancement model in real time through back propagation.

9. The system according to claim 8, characterized in that The auxiliary module includes a locked copy and a trainable copy. The locked copy is used to replicate the structure and weights of the image enhancement model. The parameters remain unchanged during the training process, preserving the generation capability of the image enhancement model. The trainable copy has the same structure as the image enhancement model and can be fine-tuned. The control signal is generated through additional conditions to dynamically adjust the generation strategy. The trainable replica generates a control signal through additional conditions, specifically including: ; in, is the generated control signal tensor, are the height, width and number of channels of the feature map, respectively. To fine-tune the model, is the input image; control signal Intermediate feature map with the main model Interacting through adaptive feature fusion enables the model to dynamically adjust the generation process based on additional conditions when generating images: ; in, is a learnable scaling factor that dynamically adjusts the strength of the control signal.

10. The system according to claim 8, wherein: Loss function for image enhancement model For pixel-level loss or perceptual loss, we introduce the conditional perceptual cross entropy loss function Forming a mixed loss function: ; in, It is a dynamically adjusted weight coefficient, which is larger in the initial stage and gradually decreases in the later stage; The design of the conditional perception cross entropy loss function is as follows: Assume that the control signal output by ControlNet is It is necessary to guide the model to generate images that meet the quality conditions and define the target distribution and the model predicts distribution The cross entropy loss between: ; in, The probability map output by ControlNet is obtained by Softmax normalization , dynamically generated according to the quality assessment results; During the pre-training main model phase, only Train the original image enhancement model to ensure basic generation capabilities; secondly, freeze the main model and fix the main model weights when training the ControlNet model. Optimize the trainable copy of ControlNet and learn to generate control signals that adapt to additional conditions; finally, in the joint fine-tuning stage, unfreeze the main model and use a hybrid loss Jointly optimize the main model and ControlNet parameters to improve the overall performance of the image enhancement model.

Citation Information

Patent Citations

  • Image processing method and device and electronic equipment

    CN112801536A

  • Underwater image enhancement quality evaluation method based on quality perception domain adaptation

    CN119313661A

  • Pathological section virtual staining image generation method based on diffusion model

    CN119810254A

  • Human body image restoration method and product based on diffusion model

    CN119919315A

  • Training and deployment of image generation models

    US11995803B1