Image enhancement method and device based on on-board acquisition, electronic equipment and storage medium
Patent Information
- Application Number
- CN202510747102.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-06-05
AI Technical Summary
[0003]然而,相关技术中,基于机载采集的图像增强方法存在明显的性能与资源消耗矛盾
[0050]The image enhancement method, apparatus, electronic device, and storage medium proposed in this application based on airborne acquisition acquire multiple frames of target images to be enhanced; perform image enhancement on the multiple frames of target images based on a first enhancement image model to obtain multiple enhanced target images; wherein, the training process of the first enhancement image model includes: acquiring a first pair of multi-frame sample images, the first pair of multi-frame sample images including two sets of multi-frame sample images with different image qualities; calling the first enhancement image model to perform image enhancement on the first pair of multi-frame sample images to obtain a first enhanced image, and determining an enhancement loss based on the first enhanced image and the first pair of multi-frame sample images; calling a pre-trained second enhancement image model to perform image enhancement on the first pair of multi-frame sample images to obtain a second enhanced image, and determining a distillation loss based on the first and second enhanced images; performing enhancement image source detection based on the first enhanced image, the second enhanced image, and the multi-frame sample images to obtain enhancement image source detection results, and updating the parameters of the first enhancement image model according to the enhancement image source detection results, enhancement loss, and distillation loss.
Smart Images

Figure CN120782648B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image enhancement method and apparatus, electronic device and storage medium based on airborne acquisition. Background Technology
[0002] Image enhancement based on airborne acquisition refers to the process of improving the quality of images acquired by airborne devices in embedded systems through technical means. For example, noise removal, detail enhancement, and environmental restoration can be performed on images acquired by vehicle-mounted or airborne image acquisition systems to improve image quality. In embedded systems, the application of image enhancement methods can improve the system's environmental awareness capabilities, thereby enhancing the practicality of the embedded system.
[0003] However, in related technologies, image enhancement methods based on airborne acquisition present a significant performance-resource trade-off. High-performance image enhancement methods have high computational complexity and consume large amounts of resources, making them difficult to apply in embedded systems with limited computing resources. Conversely, image enhancement methods with lower computational complexity have limited performance and cannot meet the image enhancement requirements of embedded systems. Summary of the Invention
[0004] The main objective of this application is to propose an image enhancement method, apparatus, electronic device, and storage medium based on airborne acquisition, which can improve image enhancement performance and reduce computational complexity.
[0005] To achieve the above objectives, a first aspect of this application proposes an image enhancement method based on airborne acquisition, the method comprising:
[0006] Acquire multiple frames of the target image to be enhanced;
[0007] Image enhancement is performed on the multi-frame target images based on the first enhanced image model to obtain multi-frame target enhanced images;
[0008] The training process of the first enhanced image model includes the following steps:
[0009] Obtain a first multi-frame sample image pair, which includes two sets of multi-frame sample images with different image qualities;
[0010] The first image enhancement model is invoked to enhance the image based on the first multi-frame sample image pair to obtain the first enhanced image, and the enhancement loss is determined based on the first enhanced image and the first multi-frame sample image pair.
[0011] The pre-trained second image enhancement model is invoked to perform image enhancement based on the first multi-frame sample image pair to obtain the second enhanced image, and the distillation loss is determined based on the first enhanced image and the second enhanced image;
[0012] Enhanced image source detection is performed based on the first enhanced image, the second enhanced image, and the multi-frame sample images to obtain enhanced image source detection results. The parameters of the first enhanced image model are then updated based on the enhanced image source detection results, the enhancement loss, and the distillation loss.
[0013] In some embodiments, the step of performing enhanced image source detection based on the first enhanced image, the second enhanced image, and the multi-frame sample images to obtain enhanced image source detection results includes:
[0014] Based on the first enhanced image and the second enhanced image, perform enhanced image source detection to obtain the first sub-source detection result;
[0015] Based on the first enhanced image and the multi-frame sample images, the enhanced image source detection is performed to obtain the second sub-source detection result;
[0016] The enhanced image source detection result is determined based on the first sub-source detection result and the second sub-source detection result.
[0017] In some embodiments, the training process of the second enhanced image model includes the following steps:
[0018] Obtain a second multi-frame sample image pair, which includes two sets of multi-frame sample images with different image qualities;
[0019] The second enhanced image model is invoked to perform image enhancement based on the second multi-frame sample image pair to obtain a third enhanced image, and the diffusion loss is determined based on the third enhanced image and the second multi-frame sample image pair;
[0020] The parameters of the second enhanced image model are updated based on the diffusion loss.
[0021] In some embodiments, determining the diffusion loss based on the third enhanced image and the second multi-frame sample image pair includes:
[0022] Based on the third enhanced image and the second multi-frame sample image pair, the feature consistency loss is calculated to obtain the feature consistency loss.
[0023] Based on the third enhanced image and the second multi-frame sample image pair, the image spatial consistency loss is calculated to obtain the image spatial consistency loss.
[0024] Based on the third enhanced image and the second multi-frame sample image, the foreground object consistency loss is calculated to obtain the foreground object consistency loss.
[0025] The diffusion loss is determined based on the feature consistency loss, the image spatial consistency loss, and the foreground object consistency loss.
[0026] In some embodiments, the second enhanced image model includes an encoding layer, a diffusion layer, a decoding layer, and a gradient embedding layer. The step of calling the second enhanced image model to perform image enhancement based on the second multi-frame sample image pair to obtain a third enhanced image includes:
[0027] The coding layer is invoked to perform feature encoding on the second multi-frame sample image pair to obtain multi-frame encoded features;
[0028] The gradient embedding layer is invoked to determine the gradient embedding features, and the diffusion layer is invoked to perform image enhancement on the gradient embedding features and the multi-frame coding features to obtain multi-frame enhanced coding features;
[0029] The decoding layer is invoked to perform feature decoding on the multi-frame enhanced coding features to obtain the third enhanced image.
[0030] In some embodiments, the step of performing image enhancement on the multi-frame target images based on the first enhanced image model to obtain multi-frame enhanced target images includes:
[0031] The multi-frame target images are subjected to quality detection to obtain a first quality score;
[0032] When the first quality score is less than a preset threshold, the multi-frame target images are enhanced based on the first enhanced image model to obtain multi-frame preliminary enhanced images.
[0033] The quality of the multiple pre-enhanced images is tested to obtain a second quality score;
[0034] When the second quality score is greater than or equal to the preset threshold, the multi-frame target enhancement image is determined based on the multi-frame preliminary enhancement image.
[0035] In some embodiments, quality detection is performed on the multiple frames of target images to obtain the first quality score, including:
[0036] The pre-trained discriminator model is invoked to perform quality detection on the multi-frame target images to obtain the first quality score, wherein the discriminator model includes a sampling layer and a prediction layer;
[0037] The process of calling a pre-trained discriminator model to perform quality detection on the multi-frame target images to obtain the first quality score includes:
[0038] The sampling layer is invoked to perform feature sampling on the multi-frame target images to obtain multi-frame target features;
[0039] The prediction layer is invoked to perform quality detection on the target features of the multi-frame dataset to obtain the first quality score.
[0040] To achieve the above objectives, a second aspect of this application provides an image enhancement device based on airborne acquisition, the device comprising:
[0041] The image acquisition module is used to acquire multiple frames of target images to be enhanced;
[0042] The image enhancement module is used to enhance the multi-frame target images based on the first enhanced image model to obtain multi-frame enhanced target images;
[0043] The training process of the first enhanced image model includes the following steps:
[0044] Obtain a first multi-frame sample image pair, which includes two sets of multi-frame sample images with different image qualities;
[0045] The first image enhancement model is invoked to enhance the image based on the first multi-frame sample image pair to obtain the first enhanced image, and the enhancement loss is determined based on the first enhanced image and the first multi-frame sample image pair.
[0046] The pre-trained second image enhancement model is invoked to perform image enhancement based on the first multi-frame sample image pair to obtain the second enhanced image, and the distillation loss is determined based on the first enhanced image and the second enhanced image;
[0047] Enhanced image source detection is performed based on the first enhanced image, the second enhanced image, and the multi-frame sample images to obtain enhanced image source detection results. The parameters of the first enhanced image model are then updated based on the enhanced image source detection results, the enhancement loss, and the distillation loss.
[0048] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0049] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0050] The image enhancement method, apparatus, electronic device, and storage medium proposed in this application based on airborne acquisition acquire multiple frames of target images to be enhanced; perform image enhancement on the multiple frames of target images based on a first enhancement image model to obtain multiple enhanced target images; wherein, the training process of the first enhancement image model includes: acquiring a first pair of multi-frame sample images, the first pair of multi-frame sample images including two sets of multi-frame sample images with different image qualities; calling the first enhancement image model to perform image enhancement on the first pair of multi-frame sample images to obtain a first enhanced image, and determining an enhancement loss based on the first enhanced image and the first pair of multi-frame sample images; calling a pre-trained second enhancement image model to perform image enhancement on the first pair of multi-frame sample images to obtain a second enhanced image, and determining a distillation loss based on the first and second enhanced images; performing enhancement image source detection based on the first enhanced image, the second enhanced image, and the multi-frame sample images to obtain enhancement image source detection results, and updating the parameters of the first enhancement image model according to the enhancement image source detection results, enhancement loss, and distillation loss.
[0051] This application trains a first image enhancement model using a first set of multi-frame sample images. During training, enhancement loss and distillation loss are used, along with the source detection results of the enhanced images, to update the model parameters of the first image enhancement model. Finally, the updated first image enhancement model is used to enhance multiple target images, resulting in multi-frame enhanced target images. This allows updating the model parameters of the first image enhancement model using a pre-trained second image enhancement model, avoiding the computational resources required to train the first image enhancement model from scratch. Therefore, this application improves image enhancement performance and reduces computational complexity. Attached Figure Description
[0052] Figure 1 This is a flowchart of the image enhancement method based on airborne acquisition provided in the embodiments of this application;
[0053] Figure 2 This is a flowchart of the first enhanced image model training process provided in the embodiments of this application;
[0054] Figure 3 yes Figure 2 The flowchart of step S204 in the process;
[0055] Figure 4 This is a flowchart of the training process of the second enhanced image model provided in the embodiments of this application;
[0056] Figure 5 yes Figure 4 The flowchart of step S402 in the document;
[0057] Figure 6 yes Figure 4 Another flowchart for step S402 in the process;
[0058] Figure 7 yes Figure 1 The flowchart of step S102 in the document;
[0059] Figure 8 yes Figure 7 The flowchart of step S701 in the process;
[0060] Figure 9 This is a schematic diagram of the structure of the image enhancement device based on airborne acquisition provided in the embodiments of this application;
[0061] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0063] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0065] First, let's analyze some of the terms used in this application:
[0066] Environmental perception capability refers to the ability of an embedded system to monitor, identify, and analyze its surrounding environment through sensors, image processing, and other technologies. Environmental perception capability enables embedded systems to acquire environmental information, such as images, sound, temperature, and humidity, and make corresponding decisions and responses to adapt to environmental changes and complete specific tasks. For example, in an automotive camera system, environmental perception capability can help vehicles recognize road signs, detect obstacles, and monitor traffic conditions, thereby improving driving safety and the reliability of autonomous driving.
[0067] Artificially designed filters: This is a traditional image enhancement method. The principle is to model noise as a filter and then obtain a corresponding image enhancement filter by mathematically calculating the inverse process of the filter.
[0068] Convolutional Neural Network (CNN) is a deep neural network model used to process data with a grid structure. In image enhancement tasks, CNNs learn the mapping relationship between low-quality input images and corresponding high-quality images, extract features from the images, and perform optimization processing to improve the visual quality and usability of the images.
[0069] Generative Adversarial Networks (GANs) are neural network models consisting of a generator and a discriminator. In image enhancement tasks, GANs train the generator and discriminator adversarially. The generator continuously learns to generate realistic image data to deceive the discriminator, while the discriminator strives to distinguish between real and generated image data, ultimately making it difficult to distinguish between the generated image data and real image data.
[0070] Diffusion models are a class of deep generative models based on Markov chains. In image enhancement tasks, diffusion models gradually degrade image data by adding noise, and then regenerate image data through a reverse denoising process. For example, the Denoising Diffusion Probabilistic Model (DDP) is a generative model based on this diffusion model. It transforms the image data distribution into a Gaussian noise distribution by gradually adding noise, and then generates a high-quality image through a denoising process.
[0071] Multi-Input Multi-Output (MIMO) is a neural network model that can simultaneously process multiple input data sources (such as multiple frames of images) and generate multiple output results (such as multiple frames of images after image enhancement).
[0072] In embedded systems, image enhancement methods effectively improve the system's environmental perception capabilities, enabling the system to acquire more accurate environmental information. However, existing image enhancement methods suffer from a trade-off between performance and resource consumption when improving image quality. For example, related technologies typically employ methods based on manually designed filters for image enhancement; however, this approach relies on human experience and struggles to handle complex noise combinations, resulting in poor image enhancement performance and making it unsuitable for embedded systems with high image enhancement performance requirements. Related technologies also use GAN-based methods for image enhancement; however, this method learns a single-step mapping from low-quality to high-quality images, leading to limited model fitting ability and poor image enhancement performance. Furthermore, related technologies employ diffusion-based methods for image enhancement; however, this method has high computational complexity, making it difficult to apply to embedded systems with limited computing resources. Therefore, this application provides an image enhancement method based on airborne acquisition, aiming to improve image enhancement performance while reducing computational complexity.
[0073] The image enhancement method based on airborne acquisition provided in this application relates to the field of image processing technology. This airborne acquisition-based image enhancement method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the airborne acquisition-based image enhancement method, but is not limited to the above forms.
[0074] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0075] Figure 1 This is an optional flowchart of an image enhancement method based on airborne acquisition provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S102.
[0076] Step S101: Obtain multiple frames of target images to be enhanced;
[0077] Step S102: Perform image enhancement on multiple target images based on the first enhanced image model to obtain multiple enhanced target images.
[0078] In step S101 of some embodiments, the multi-frame target image to be enhanced may refer to a set of continuous image sequences that require image enhancement. For example, the multi-frame target image may refer to a set of continuous video frames in a low-light nighttime scene, where the image quality is poor due to insufficient light and image enhancement is required; or, the multi-frame target image may refer to a set of continuous road image sequences in an adverse weather scene, where the image quality is poor due to adverse weather conditions and image enhancement is required. It is understood that the multi-frame target image to be enhanced can be obtained by extracting a set of continuous video frames from a video acquired by a vehicle-mounted image acquisition system, or it can be obtained by acquiring a set of continuous image sequences in real time using an airborne image acquisition system; the specific method is not limited.
[0079] In step S102 of some embodiments, image enhancement can refer to the process of improving the quality of multiple target images given multiple frames of target images. The first image enhancement model can refer to a neural network model with image enhancement capabilities. For example, the first image enhancement model can be a lightweight GAN model, utilizing adversarial training between the generator and discriminator of a lightweight GAN to perform feature learning and quality enhancement on multiple target images; or, the first image enhancement model can be a lightweight CNN model, using convolutional layers to extract and process features from multiple target images to achieve image enhancement. It is understood that the type of the first image enhancement model is not limited, but the first image enhancement model should have low computational complexity and require fewer computational resources. Multi-frame target image enhancement can refer to the multi-frame enhanced images obtained after inputting multiple target images into the first image enhancement model for image enhancement.
[0080] It is understandable that the first image enhancement model can be pre-trained. The training process of the first image enhancement model is explained below. (Refer to...) Figure 2 The training process of the first enhanced image model may include steps S201 to S204:
[0081] Step S201: Obtain the first multi-frame sample image pair, which includes two sets of multi-frame sample images with different image qualities;
[0082] Step S202: Call the first enhanced image model to perform image enhancement based on the first multi-frame sample image pair to obtain the first enhanced image, and determine the enhancement loss based on the first enhanced image and the first multi-frame sample image pair;
[0083] Step S203: Call the pre-trained second image enhancement model to perform image enhancement based on the first multi-frame sample image pair to obtain the second enhanced image, and determine the distillation loss based on the first enhanced image and the second enhanced image;
[0084] Step S204: Perform enhanced image source detection based on the first enhanced image, the second enhanced image, and multiple sample images to obtain enhanced image source detection results, and update the parameters of the first enhanced image model based on the enhanced image source detection results, enhancement loss, and distillation loss.
[0085] In step S201 of some embodiments, the first multi-frame sample image pair may refer to an image pair comprising two sets of multi-frame sample images with different image qualities. For example, the first multi-frame sample image pair may include a set of low-quality multi-frame sample images and a corresponding set of high-quality multi-frame sample images. The low-quality multi-frame sample images may refer to a continuous image sequence acquired under conditions such as low light at night or in bad weather, which may have problems such as high noise or blurred details; the high-quality multi-frame sample images may refer to the corresponding continuous image sequence acquired under conditions such as normal lighting and good weather, which has clear details and low noise. It is understood that the first multi-frame sample image pair can be formed by extracting a set of high-quality multi-frame sample images from an existing video, and then performing operations such as adding noise or blurring on this set of high-quality multi-frame sample images to obtain the corresponding low-quality multi-frame sample images; the first multi-frame sample image pair can also be formed by using professional image acquisition equipment to capture a set of low-quality continuous image sequences and a corresponding set of high-quality continuous image sequences under different conditions, thereby forming a complete image pair, without specific limitations.
[0086] In step S202 of some embodiments, the first enhanced image may refer to the image obtained after the first enhanced image model enhances the low-quality multi-frame sample images in the first multi-frame sample image pair. The enhancement loss may refer to the difference between the first enhanced image and the corresponding high-quality multi-frame sample image in the first multi-frame sample image pair; that is, the enhancement loss can be calculated using the first enhanced image and the corresponding high-quality multi-frame sample image in the first multi-frame sample image pair. For example, the enhancement loss can be calculated according to the following formula 1.
[0087] L s =L2(y1,g t )(Formula 1)
[0088] Among them, L s Let y1 represent the enhancement loss, and g represent the first enhanced image. t L2 represents the high-quality multi-frame sample image corresponding to the first multi-frame sample image pair, and L2 represents the Euclidean distance. L2 is used to calculate y1 and g. t The differences between them.
[0089] In step S203 of some embodiments, the pre-trained second image enhancement model may refer to a neural network model pre-trained according to a certain image enhancement task. For example, the pre-trained second image enhancement model may be a trained MIMO model or a trained DDP model. It is understood that the type of the second image enhancement model is not limited, but the second image enhancement model should have high image enhancement performance. The second enhanced image may refer to the image obtained after the second image enhancement model enhances the low-quality multi-frame sample images in the second multi-frame sample image pair. Distillation loss may refer to the difference between the first enhanced image and the second enhanced image, that is, the distillation loss can be calculated from the first enhanced image and the second enhanced image. For example, the distillation loss can be calculated according to the following formula 2:
[0090] L kd =L2(y1,y2)(Formula 2)
[0091] Among them, L kd Let y1 represent the distillation loss, y2 represent the first enhanced image, y2 represent the second enhanced image, and L2 represent the Euclidean distance, which is used to calculate the difference between y1 and y2.
[0092] In step S204 of some embodiments, enhanced image source detection can refer to the process of comparing and analyzing a first enhanced image, a second enhanced image, and high-quality multi-frame sample images using a model or algorithm. The enhanced image source detection result can refer to the differences between the first enhanced image, the second enhanced image, and the corresponding high-quality multi-frame sample images in the first multi-frame sample image pair. For example, a lightweight discriminator model (such as a CNN model or a transformer model) can be used to compare the differences between the first and second enhanced images, and another lightweight discriminator model (with the same network structure for both models) can be used to compare the differences between the first enhanced image and the high-quality multi-frame sample images. The enhanced image source detection result is then determined based on the outputs of the two lightweight discriminator models. Updating the parameters of the first enhanced image model can be achieved by calculating the error of the first enhanced image model using the enhanced image source detection result, enhancement loss, and distillation loss, and then adjusting the parameters of the first enhanced image model based on this error. It is understood that, in the embodiments of this application, the first enhanced image model can be trained by multiple sets of first multi-frame sample images to continuously update the model parameters of the first enhanced image model, thereby gradually improving the image enhancement performance of the first enhanced image model.
[0093] In some embodiments, the training process of the first image enhancement model can employ a random early stopping strategy. That is, when the evaluation score of the second enhanced image output by the second image enhancement model is greater than or equal to a preset score, a stopping probability (which can be denoted as "p") is set for the training process of the first image enhancement model to stop the process of calling the pre-trained second image enhancement model for image enhancement in advance, thereby reducing the computational resource requirements of the training process. The evaluation score can refer to the score obtained after evaluating the second enhanced image using feature distribution distance (which can be denoted as "FID") or a no-reference quality assessment model. The preset score can refer to a pre-set threshold used to determine whether the image quality is acceptable; for example, the preset score can be 25 or 26 (out of 30). It is understood that the preset score can be freely set according to actual needs. The stopping probability can refer to the probability of stopping the call to the second image enhancement model in advance when the evaluation score is greater than or equal to the preset score. For example, when the evaluation score is greater than or equal to the preset score, p can be 0.7, meaning there is a 70% probability of stopping the call to the second image enhancement model in advance. It is understood that p can be freely set according to actual needs.
[0094] In some embodiments, this application can perform error analysis on the first enhanced image using a second enhanced image and multiple sample images to obtain error parameters. A third weighting coefficient is then determined based on the error parameters. The enhanced image source detection result, enhancement loss, and distillation loss are then weighted and calculated based on the third weighting coefficient. The parameters of the first enhanced image model are updated based on the calculated weighted loss value. Error analysis can refer to the process of analyzing the differences between the first enhanced image and the second enhanced image, as well as with the multiple sample images. Error parameters can refer to the errors calculated based on the first enhanced image, the second enhanced image, and the multiple sample images. Error parameters include a first sub-error parameter and a second sub-error parameter. The first sub-error parameter can refer to the error calculated based on the first enhanced image and the second enhanced image. The second sub-error parameter can refer to the error calculated based on the first enhanced image and the multiple sample images. The third weighting coefficient can refer to the weighting coefficient calculated based on the error parameters. For example, the third weighting coefficient can be calculated according to the following formula 3.
[0095]
[0096] Where ω1 represents the third weighting coefficient, e1 represents the first sub-error parameter, and e2 represents the second sub-error parameter. The error parameters include the first and second sub-error parameters (the error parameters can be labeled "e", i.e., e = e1 + e2). α and β are used to balance the weighting coefficients of the first and second sub-error parameters. α and β can be freely set according to actual needs, but must satisfy α + β = 1. For example, if e1 is 2.5, e2 is 3.0, α is 0.6, and β is 0.4, then the third weighting coefficient ω2 = (0.6 × 2.5 + 0.4 × 3.0) / (2.5 + 3.0) ≈ 0.50. It can be understood that through formula 3, the third weighting coefficient can be automatically adjusted according to the magnitude of the error parameters. Thus, when updating the parameters of the first enhanced image model, the weighting coefficient corresponding to the weighted loss value of the larger error parameter can be reduced, so that the weighted loss value of the smaller error parameter has a larger weighting coefficient, thereby improving the image enhancement performance of the first enhanced image model. The weighted loss value refers to the loss calculated by weighting the enhanced image source detection result, enhancement loss, and distillation loss according to a third weighting coefficient. For example, if the enhanced image source detection result is 2.5, the enhancement loss is 1.5, the distillation loss is 3.5, and the third weighting coefficient is 0.5, then the weighted loss value is 0.5 × (2.5 + 1.5 + 3.5) = 3.25. Updating the parameters of the first enhanced image model can be achieved by using the backpropagation algorithm to calculate the parameters that need updating based on the weighted loss value, and then adjusting the model parameters of the first enhanced image model accordingly. For example, the stochastic gradient descent method (which can be labeled "SGD method") can be used to calculate the gradient value of the first enhanced image model based on the weighted loss, and the direction and magnitude of the model parameter adjustment can be determined based on the gradient value, thus adjusting the model parameters to update the parameters of the first enhanced image model.
[0097] Please see Figure 3 In some embodiments, step S204 may include, but is not limited to, steps S301 to S303:
[0098] Step S301: Perform enhanced image source detection based on the first enhanced image and the second enhanced image to obtain the first sub-source detection result;
[0099] Step S302: Perform enhanced image source detection based on the first enhanced image and multiple frame sample images to obtain the second sub-source detection result;
[0100] Step S303: Determine the enhanced image source detection result based on the first sub-source detection result and the second sub-source detection result.
[0101] In step S301 of some embodiments, the first sub-source detection result may refer to the difference between the first enhanced image and the second enhanced image; that is, the first sub-source detection result can be calculated from the first enhanced image and the second enhanced image. For example, the first sub-source detection result can be obtained by performing enhanced image source detection on the first enhanced image and the second enhanced image according to a dual adversarial function (which can be labeled "dual adversarial function"); or, the first sub-source detection result can also be obtained by performing enhanced image source detection on the first enhanced image and multiple frame sample images according to a mean squared error function. It is understood that the calculation method of the first sub-source detection result can be freely set according to actual needs.
[0102] In step S302 of some embodiments, the second sub-source detection result may refer to the difference between the first enhanced image and the multi-frame sample images; that is, the second sub-source detection result can be calculated from the first enhanced image and the multi-frame sample images. For example, the second sub-source detection result can be obtained by performing enhanced image source detection on the first enhanced image and the multi-frame sample images according to the dual adversarial function. It is understood that the second sub-source detection result uses the same calculation method as the first sub-source detection result.
[0103] In step S303 of some embodiments, the enhanced image source detection result may refer to the loss value determined by the first sub-source detection result and the second sub-source detection result. For example, the enhanced image source detection result can be calculated according to the following formula 4:
[0104] L d =L d1 +L d2 (Formula 4)
[0105] Among them, L d1 Indicates the first source detection result, L d2 Indicates the result of the second source detection, L d This indicates the result of enhanced image source detection. For example, if L d1 It is 1.5, L d2 If it is 2.0, then L d The result is 1.5 + 2.0 = 3.5. It is understood that this embodiment of the application determines the enhanced image source detection result by combining the first sub-source detection result and the second sub-source detection result, enabling the first enhanced image model to simultaneously focus on the consistency between the first enhanced image and the second enhanced image, as well as the matching degree with high-quality multi-frame sample images, thereby more comprehensively evaluating the quality of the first enhanced image and improving the model's image enhancement performance.
[0106] In some embodiments, this application can set corresponding weight coefficients for the first sub-source detection result and the second sub-source detection result (i.e., the first weight coefficient corresponding to the first sub-source detection result and the second weight coefficient corresponding to the second sub-source detection result), so that the first image enhancement model can dynamically adjust the degree of emphasis on different losses during training. For example, when the first image enhancement model needs to pay more attention to the consistency between the first enhanced image and the second enhanced image, the first weight coefficient can be increased; or, when the first image enhancement model needs to pay more attention to the matching degree between the first enhanced image and the high-quality multi-frame sample image, the second weight coefficient can be increased. In this case, the enhanced image source detection result can refer to the loss value obtained by weighting the first sub-source detection result with the first weight coefficient and the loss value obtained by weighting the second sub-source detection result with the second weight coefficient. For example, if L d1 It is 1.5, L d2 If the first weighting factor is 2.0, the second weighting factor is 0.6, and the third weighting factor is 0.4, then L... d The formula is 0.6 × 2 + 0.4 × 2.5 = 2.2. It is understood that the first and second weighting coefficients can be freely set according to actual needs and are not limited. It is also understood that this embodiment of the application, through this flexible weight adjustment mechanism, can improve the adaptability of the first enhanced image model and better meet the image enhancement needs in different scenarios.
[0107] Please refer to Figure 4 The training process of the second enhanced image model may include, but is not limited to, steps S401 to S403:
[0108] Step S401: Obtain a second multi-frame sample image pair, which includes two sets of multi-frame sample images with different image qualities;
[0109] Step S402: Call the second enhanced image model to perform image enhancement based on the second multi-frame sample image pair to obtain the third enhanced image, and determine the diffusion loss based on the third enhanced image and the second multi-frame sample image pair;
[0110] Step S403: Update the parameters of the second enhanced image model according to the diffusion loss.
[0111] In step S401 of some embodiments, the second multi-frame sample image pair may refer to an image pair comprising two sets of multi-frame sample images with different image qualities. For example, the second multi-frame sample image pair may include a set of low-quality multi-frame sample images and a corresponding set of high-quality multi-frame sample images. It is understood that the second multi-frame sample image pair may include the same content as the first multi-frame sample image pair, and the method of obtaining the second multi-frame sample image pair may also be the same as that of the first multi-frame sample image pair.
[0112] In step S402 of some embodiments, the third enhanced image may refer to the image obtained after enhancing the low-quality multi-frame sample images in the second multi-frame sample image pair according to the second enhanced image model. The diffusion loss may refer to the loss value calculated based on the second enhanced image and the high-quality multi-frame sample images in the second multi-frame sample image pair.
[0113] In step S403 of some embodiments, updating the parameters of the second enhanced image model can be achieved by using the backpropagation algorithm to calculate the parameters that need to be updated in the model parameters based on the diffusion loss, and then adjusting the model parameters of the second enhanced image model according to the parameters to be updated. For example, the gradient value of the second enhanced image model can be calculated based on the diffusion loss using the Adam optimization method (which can be labeled as the "Adam method"), and the adjustment direction and magnitude of the model parameters can be determined based on the gradient value, and then the model parameters can be adjusted accordingly to update the parameters of the second enhanced image model.
[0114] Please see Figure 5 In some embodiments, step S402 may include, but is not limited to, steps S501 to S504:
[0115] Step S501: Calculate the feature consistency loss based on the third enhanced image and the second multi-frame sample image pair to obtain the feature consistency loss;
[0116] Step S502: Calculate the image spatial consistency loss based on the third enhanced image and the second multi-frame sample image pair to obtain the image spatial consistency loss;
[0117] Step S503: Calculate the foreground object consistency loss based on the third enhanced image and the second multi-frame sample image pair to obtain the foreground object consistency loss;
[0118] Step S504: Determine the diffusion loss based on feature consistency loss, image spatial consistency loss, and foreground object consistency loss.
[0119] In step S501 of some embodiments, the feature consistency loss calculation can refer to a method used to measure the similarity between the third enhanced image and the high-quality multi-frame sample images in the second multi-frame sample image pair. The feature consistency loss can refer to calculating a loss value based on the feature consistency loss of the third enhanced image and the high-quality multi-frame sample images. For example, the feature consistency loss can be calculated according to the following formula 5:
[0120] L f =L2(f y3 ,f gt ) (Formula 5)
[0121] Among them, Lf This represents the feature consistency loss. This represents the encoded features corresponding to the third enhanced image. This represents the encoded features corresponding to high-quality multi-frame sample images, where L2 represents the Euclidean distance, and L2 is used to calculate... and The differences between them. Among them, the encoded features corresponding to the third enhanced image can refer to the features obtained by encoding the third enhanced image model through the encoder in the third enhanced image model; the encoded features corresponding to the high-quality multi-frame sample images can refer to the features obtained by encoding the high-quality multi-frame sample images through the encoder in the third enhanced image model.
[0122] In step S502 of some embodiments, the image spatial consistency loss calculation can refer to a method used to measure the overall similarity in pixel space between the third enhanced image and the high-quality multi-frame sample images in the second multi-frame sample image pair. The image spatial consistency loss can refer to calculating the image spatial consistency loss on the third enhanced image and the high-quality multi-frame sample images to obtain a loss value. For example, the image spatial consistency loss can be calculated according to the following formula 6:
[0123] L k =L2(y3,g t ) (Formula 6)
[0124] Among them, L k Let y3 represent the spatial consistency loss of the image, and g represent the third augmented image. t This represents a high-quality multi-frame sample image, where L2 represents the Euclidean distance, and L2 is used to calculate y3 and g. t The differences between them.
[0125] In step S503 of some embodiments, the foreground object consistency loss calculation can refer to a method used to measure the similarity between the foreground object portion in the third enhanced image and the foreground object portion in the high-quality multi-frame sample image. The foreground object consistency loss can refer to calculating the foreground object consistency loss on the third enhanced image and the high-quality multi-frame sample image to obtain a loss value. For example, the foreground object consistency loss can be calculated according to the following formula 7:
[0126] L q =L2(mask·y3,mask·g) t ) (Formula 7)
[0127] Among them, L q y3 represents the foreground object consistency loss, g represents the third augmented image, and g represents the foreground object consistency loss. tThis represents a high-quality multi-frame sample image, where L2 represents the Euclidean distance, mask represents the mask image corresponding to the foreground region of the image (e.g., the foreground region of the image can be marked as 1, and other regions can be marked as 0, thus highlighting the foreground region), mask*y3 represents the mask image corresponding to the foreground region of the third enhanced image, and mask*g t This represents the mask image corresponding to the foreground region of a high-quality multi-frame sample image. L2 is used to calculate mask*y3 and mask*g. t The differences between them.
[0128] In step S504 of some embodiments, the diffusion loss can refer to the loss value calculated based on the feature consistency loss, image spatial consistency loss, and foreground object consistency loss. For example, if the feature consistency loss is 1.5, the image spatial consistency loss is 0.75, and the foreground object consistency loss is 1.0, then the diffusion loss is 1.5 + 0.75 + 1.0 = 3.25. It is understood that embodiments of this application can set corresponding loss weights for the feature consistency loss, image spatial consistency loss, and foreground object consistency loss, thereby enabling the model to flexibly adjust the emphasis on different types of loss during training, thus more effectively improving the image enhancement performance of the third enhanced image model.
[0129] Please see Figure 6 In some embodiments, the second enhanced image model includes an encoding layer, a diffusion layer, a decoding layer, and a gradient embedding layer, and step S402 may also include, but is not limited to, steps S601 to S603:
[0130] Step S601: Call the coding layer to perform feature encoding on the second multi-frame sample image pair to obtain multi-frame encoded features;
[0131] Step S602: Call the gradient embedding layer to determine the gradient embedding features, and call the diffusion layer to perform image enhancement on the gradient embedding features and multi-frame coding features to obtain multi-frame enhanced coding features.
[0132] Step S603: Call the decoding layer to perform feature decoding on the multi-frame enhanced coding features to obtain the third enhanced image.
[0133] In step S601 of some embodiments, the encoding layer may refer to a module used for feature encoding of the second multi-frame sample image pair. For example, the encoding layer may refer to a multi-frame variational autoencoder (VAE) or a convolutional module in a CNN model, without specific limitations. Feature encoding may refer to the process of extracting features from low-quality multi-frame sample images given that low-quality multi-frame sample images are determined. Multi-frame encoded features may refer to the encoded features obtained after feature encoding of the low-quality multi-frame sample images by the encoding layer. For example, if the encoding layer is a VAE encoder, the multi-frame encoded features may be the encoded features obtained after inputting the low-quality multi-frame sample images into the VAE encoder for feature encoding.
[0134] In step S602 of some embodiments, the gradient embedding layer can refer to a module used to process gradient information. For example, the gradient embedding layer can be an Adaptive Prompt Embedding Module or other similar modules, without specific limitations. The gradient embedding feature can refer to a feature containing gradient information extracted by the gradient embedding layer. The diffusion layer can refer to a module used to enhance the image of the gradient embedding feature and the multi-frame encoded feature. For example, the diffusion layer can be a Multi-Modal Diffusion Transformer (MMDIT) or other diffusion modules, without specific limitations. The multi-frame enhanced encoded feature can refer to the encoded feature obtained after enhancing the image of the gradient embedding feature and the multi-frame encoded feature through the diffusion layer. For example, if the diffusion layer is an MMDIT module, the multi-frame enhanced encoded feature can be the encoded feature obtained by inputting the multi-frame sample image features and the gradient information from the Adaptive Prompt Embedding Module into the MMDIT module for image enhancement.
[0135] In step S603 of some embodiments, the decoding layer can refer to a module used to convert multi-frame enhanced coding features into an image. For example, the decoding layer can be a multi-frame variational autoencoder decoder (VAD) or a transposed convolution layer, without specific limitations. Feature decoding can refer to the process of converting multi-frame enhanced coding features into visualized image data through the decoding layer. The third enhanced image can refer to the image obtained after feature decoding of multi-frame enhanced coding features through the decoding layer. For example, if the decoding layer is a VAD decoder, the third enhanced image can be the image obtained after inputting multi-frame enhanced coding features into the VAD decoder for feature decoding.
[0136] Please see Figure 7In some embodiments, step S102 may include, but is not limited to, steps S701 to S704:
[0137] Step S701: Perform quality detection on multiple frames of target images to obtain a first quality score;
[0138] Step S702: When the first quality score is less than a preset threshold, perform image enhancement on multiple target images based on the first enhanced image model to obtain multiple preliminary enhanced images;
[0139] Step S703: Perform quality detection on multiple frames of preliminary enhanced images to obtain a second quality score;
[0140] Step S704: When the second quality score is greater than or equal to a preset threshold, determine the target enhancement image based on the multiple preliminary enhancement images.
[0141] In step S701 of some embodiments, quality detection may refer to the process of evaluating the quality of multiple frames of target images using a specific algorithm or model. The first quality score may refer to the score obtained by performing quality detection on multiple frames of target images.
[0142] In step S702 of some embodiments, the preset threshold can refer to a pre-set value used to determine whether the multi-frame target images need image enhancement. For example, the preset threshold can be set to 0.7 or 0.8 (out of 1), and there is no specific limitation. The multi-frame preliminary enhanced image can refer to the image obtained after image enhancement of the multi-frame target images by the first enhanced image model when the first quality score is less than the preset threshold. It can be understood that when the first quality score is greater than or equal to the preset threshold, the embodiments of this application may not perform image enhancement on the multi-frame target images. At this time, the multi-frame target enhanced image is the multi-frame target image. In this way, image enhancement can be performed only on multi-frame target images with a first quality score lower than the preset threshold, thereby reducing the computational complexity of the first enhanced image model.
[0143] In steps S703 to S704 of some embodiments, the second quality score may refer to the score obtained by quality detection of multiple preliminary enhanced images. The multiple target enhanced images may refer to the images determined based on the multiple preliminary enhanced images when the second quality score is greater than or equal to a preset threshold; in this case, the multiple target enhanced images are the multiple preliminary enhanced images. It is understood that when the second quality score is less than the preset threshold, this embodiment can perform multiple iterations of image enhancement on the multiple preliminary enhanced images to make the quality score of the enhanced image finally output by the first enhanced image model greater than the preset threshold, and determine the multiple target enhanced images based on the enhanced image finally output by the first enhanced image model. Furthermore, to avoid infinite iteration, this embodiment can also preset an iteration number threshold to determine whether iteration needs to be stopped. The iteration number threshold may refer to the maximum number of iterations allowed during image enhancement; for example, the iteration number threshold may be set to 5 or 6 times, without specific limitation. When the iteration number of the first enhanced image model is greater than or equal to the iteration number threshold, iteration will stop, and the multiple target enhanced images will be determined based on the enhanced image finally output by the first enhanced image model.
[0144] In some embodiments, step S701 may include, but is not limited to:
[0145] A pre-trained discriminator model is invoked to perform quality detection on multiple frames of target images to obtain a first quality score. The discriminator model includes a sampling layer and a prediction layer.
[0146] In some embodiments, the pre-trained discriminator model can refer to a model pre-trained for a specific quality detection task. The discriminator model includes a sampling layer and a prediction layer. The sampling layer can refer to a module used to extract features from multiple frames of target images. For example, the sampling layer can be multiple convolutional layers or a multilayer perceptron (MLP), without specific limitations. The prediction layer can refer to a module used for quality detection of the features extracted by the sampling layer. For example, the prediction layer can refer to a CNN module or a residual network (ResNet), without specific limitations. The first quality score can refer to the score obtained after performing quality detection on multiple frames of target images using the pre-trained discriminator model.
[0147] Please see Figure 8 In some embodiments, the discriminator model includes a sampling layer and a prediction layer. Calling the pre-trained discriminator model to perform quality detection on multiple frames of target images to obtain a first quality score may include, but is not limited to, steps S801 to S802:
[0148] Step S801: Call the sampling layer to perform feature sampling on multiple frames of target images to obtain multiple frames of target features;
[0149] Step S802: Call the prediction layer to perform quality detection on the target features of multiple frames and obtain the first quality score.
[0150] In steps S801 to S802 of some embodiments, feature sampling can refer to features extracted from multiple frames of target images that can represent the image content. Multi-frame target features can refer to features obtained after feature sampling of multiple frames of target images through a sampling layer. For example, if the sampling layer is multiple convolutional layers, multi-frame target features can be obtained by inputting multiple frames of target images into multiple convolutional layers respectively, performing average pooling and sigmoid operations on the output of each convolutional layer to calculate a gate value (which can be labeled as a "gate value"), and then performing routing selection based on the gate value (e.g., selecting the top three features) to finally determine the features. The first quality score can refer to the score obtained after quality detection of multi-frame target features through a prediction layer.
[0151] This application also provides a specific application process for an image enhancement method based on airborne acquisition, specifically including: MIMO high-performance diffusion model training, cross-architecture distillation learning, multi-stage enhancement discriminator training, and model deployment and application. The MIMO high-performance diffusion model training (i.e., training of the second enhanced image model) is achieved through pairs of multi-frame image data (i.e., second multi-frame sample image pairs). The MIMO high-performance diffusion model includes a multi-frame VAE encoder (i.e., encoding layer), an MMDIT diffusion model (i.e., diffusion layer), a multi-frame VAE decoder (i.e., decoding layer), and an adaptive prompt embedding module (i.e., gradient embedding layer). During the training of the MIMO high-performance diffusion model, low-quality multi-frame images (i.e., low-quality multi-frame sample images from the first multi-frame sample image pair) are input into the multi-frame VAE encoder for feature extraction. The extracted multi-frame encoded features and the gradient information (i.e., gradient embedding features) from the adaptive prompt embedding module are then input into the MMDIT diffusion model for image enhancement, resulting in multi-frame enhanced encoded features. Finally, the multi-frame enhanced encoded features are input into the multi-frame VAE decoder for feature decoding, resulting in the third enhanced image. Finally, based on the third enhanced image and paired multi-frame image data, feature consistency loss, spatial consistency loss, and foreground object consistency loss are calculated. Then, based on these losses, the model parameters of the MIMO high-performance diffusion model are optimized using the Adam method until the model converges. It is understood that the trained MIMO high-performance diffusion model in this embodiment can process multiple frames of images simultaneously, fully utilizing the relationships between image frames for image enhancement, thereby reducing computational complexity and improving image enhancement performance.
[0152] Cross-architecture distillation learning is achieved by using a pre-trained high-performance MIMO diffusion model (i.e., the pre-trained second-enhancement image model) as the teacher network and a lightweight GAN network (i.e., the first-enhancement image model) as the student network, with the teacher network guiding the training of the student network. During the training of the lightweight GAN network, low-quality multi-frame sample images are input into the teacher network to determine the second-enhancement image, and low-quality multi-frame sample images are input into the student network to determine the first-enhancement image. Then, through a dual adversarial function and two discriminators, the first-enhancement image is subjected to adversarial learning against the high-quality multi-frame sample image and the second-enhancement image to determine the source detection result of the enhanced image. Further, the enhancement loss and distillation loss are determined by the lightweight GAN network based on the first-enhancement image and the high-quality multi-frame sample image, respectively. Finally, the model parameters of the lightweight GAN network are optimized using the Adam method based on the enhanced image source detection result, the enhancement loss, and the distillation loss until the model converges. Understandably, in this application embodiment, a well-trained MIMO high-performance diffusion model is used as a teacher to guide a low-complexity lightweight GAN network to perform distillation learning, enabling the lightweight GAN network to perform image enhancement on multiple frames of images, while making its image enhancement performance close to that of the teacher network and significantly reducing the computational resource requirements.
[0153] The multi-stage enhanced discriminator training (i.e., discriminator model training) involves acquiring multiple frames of sample images and their corresponding sample quality scores as training sample data. Then, a multi-stage enhanced discriminator (i.e., discriminator model) is constructed, comprising a multi-expert sampling module (i.e., sampling layer) and a quality score prediction module (i.e., prediction layer). The multi-expert sampling module contains multiple convolutional layers, and the quality score prediction module predicts the quality scores of the multiple frames of sample images. Further, feature sampling is performed on the multiple frames of sample images through the multiple convolutional layers of the multi-expert sampling module, and average pooling and sigmoid operations are applied to the features output by the multiple convolutional layers to calculate a gating value. Routing is then performed based on the gating value to obtain the features corresponding to the multiple frames of sample images. Further, the quality score prediction module performs quality detection on the features corresponding to the multiple frames of sample images to obtain the predicted quality score. Finally, regression loss is calculated on the predicted quality score and the corresponding sample quality score, and the model parameters of the discriminator model are updated based on the calculated loss value using the Adam method. It is understood that the embodiments of this application can iteratively train the multi-stage enhanced discriminator through multiple training sample data to gradually improve its accuracy in judging image quality.
[0154] The deployment and application of the model involves deploying a trained lightweight GAN network and a trained multi-stage enhancement discriminator into an embedded system. When the embedded system acquires multiple frames of image data (i.e., the multiple target images to be enhanced), the multi-stage enhancement discriminator first evaluates the quality score s (i.e., the first quality score) of the multiple frame image data. If the quality score s is less than a preset threshold T, the lightweight GAN network performs image enhancement on the multiple target images, resulting in enhanced images (i.e., the initial enhanced images). Then, the multi-stage enhancement discriminator evaluates the quality score s1 (i.e., the second quality score) of the enhanced images. If the quality score s1 is greater than or equal to the preset threshold T, the enhanced images are determined based on the final enhanced images output by the lightweight GAN network. If the quality score s1 is less than the preset threshold T, the lightweight GAN network continues to perform further image enhancement on the previously enhanced images, and the multi-stage enhancement discriminator evaluates the quality score of the enhanced images again. This application's embodiments can be understood as follows: by iteratively enhancing the image input to the lightweight GAN network multiple times, the quality score of the image output by the lightweight GAN network eventually exceeds a preset threshold T, thereby gradually improving the image quality. Furthermore, if the number of image enhancements performed through the lightweight GAN network is greater than or equal to a preset iteration threshold N (e.g., 5 times), the target enhanced images for multiple frames are directly determined based on the last image enhancement result, thus preventing the lightweight GAN network from getting stuck in infinite iteration and ensuring efficient utilization of the system's computational resources. It is understood that the image enhancement method based on airborne acquisition in this application not only has the characteristics of high-efficiency operation and can be deployed in embedded acquisition systems, but also possesses strong image enhancement performance. Moreover, this method can also perform quality detection on the enhanced image and determine whether further enhancement is needed based on the quality detection score, thereby ensuring that the quality of the output image meets the system requirements.
[0155] Please see Figure 9 This application also provides an image enhancement device based on airborne acquisition, which can implement the above-described image enhancement method based on airborne acquisition. The device includes:
[0156] Image acquisition module 901 is used to acquire multiple frames of target images to be enhanced;
[0157] Image enhancement module 902 is used to enhance multiple frames of target images based on a first enhanced image model to obtain multiple frames of enhanced target images;
[0158] The training process of the first enhanced image model includes the following steps:
[0159] Obtain the first multi-frame sample image pair, which includes two sets of multi-frame sample images with different image qualities;
[0160] The first image enhancement model is invoked to perform image enhancement based on the first multi-frame sample image pair to obtain the first enhanced image, and the enhancement loss is determined based on the first enhanced image and the first multi-frame sample image pair.
[0161] The pre-trained second image enhancement model is invoked to perform image enhancement based on the first multi-frame sample image pair to obtain the second enhanced image, and the distillation loss is determined based on the first and second enhanced images;
[0162] Enhanced image source detection is performed based on the first enhanced image, the second enhanced image, and multiple sample images to obtain enhanced image source detection results. The parameters of the first enhanced image model are then updated based on the enhanced image source detection results, enhancement loss, and distillation loss.
[0163] The specific implementation of this device is basically the same as the specific embodiment of the image enhancement method based on airborne acquisition described above, and will not be repeated here.
[0164] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described image enhancement method based on airborne acquisition. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0165] Please see Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0166] The processor 1001 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0167] The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 using the image enhancement method based on airborne acquisition according to the embodiments of this application.
[0168] Input / output interface 1003 is used to implement information input and output;
[0169] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0170] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0171] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0172] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image enhancement method based on airborne acquisition.
[0173] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0174] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0175] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0177] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0178] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0179] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0180] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0181] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0182] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0184] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An image enhancement method based on airborne acquisition, characterized in that, The method includes: Acquire multiple frames of the target image to be enhanced; Image enhancement is performed on the multi-frame target images based on the first enhanced image model to obtain multi-frame target enhanced images; The training process of the first enhanced image model includes the following steps: Acquire a first multi-frame sample image pair, which includes a set of low-quality multi-frame sample images and a corresponding set of high-quality multi-frame sample images obtained by using an image acquisition device under different acquisition conditions. The first image enhancement model is invoked to perform image enhancement based on the low-quality multi-frame sample images to obtain the first enhanced image, and the enhancement loss is determined based on the first enhanced image and the high-quality multi-frame sample images. The pre-trained second image enhancement model is invoked to perform image enhancement based on the low-quality multi-frame sample images to obtain a second enhanced image, and the distillation loss is determined based on the first enhanced image and the second enhanced image; A discriminator model is invoked to perform enhanced image source detection based on the first enhanced image and the second enhanced image to obtain the first sub-source detection result; Another discriminator model is invoked to perform enhanced image source detection based on the first enhanced image and the high-quality multi-frame sample images, to obtain a second sub-source detection result; wherein, the other discriminator model and the first discriminator model have the same network structure; The enhanced image source detection result is determined based on the first sub-source detection result and the second sub-source detection result; The parameters of the first enhanced image model are updated based on the enhanced image source detection results, the enhancement loss, and the distillation loss; The training process of the second enhanced image model includes the following steps: A second multi-frame sample image pair is obtained, which includes a set of low-quality multi-frame sample images and a corresponding set of high-quality multi-frame sample images. The second multi-frame sample image pair is obtained in the same way as the first multi-frame sample image pair. The second enhanced image model is invoked to perform image enhancement based on the low-quality multi-frame sample images in the second multi-frame sample image pair to obtain a third enhanced image, and the diffusion loss is determined based on the third enhanced image and the high-quality multi-frame sample images in the second multi-frame sample image pair. The parameters of the second enhanced image model are updated based on the diffusion loss.
2. The method according to claim 1, characterized in that, The determination of diffusion loss based on high-quality multi-frame sample images from the third enhanced image and the second multi-frame sample image pair includes: The feature consistency loss is calculated based on the high-quality multi-frame sample images in the third enhanced image and the second multi-frame sample image pair. Image spatial consistency loss is calculated based on the high-quality multi-frame sample images in the third enhanced image and the second multi-frame sample image pair to obtain the image spatial consistency loss. Based on the high-quality multi-frame sample images in the third enhanced image and the second multi-frame sample image pair, the foreground object consistency loss is calculated to obtain the foreground object consistency loss. The diffusion loss is determined based on the feature consistency loss, the image spatial consistency loss, and the foreground object consistency loss.
3. The method according to claim 1, characterized in that, The second enhanced image model includes an encoding layer, a diffusion layer, a decoding layer, and a gradient embedding layer. The step of calling the second enhanced image model to perform image enhancement based on low-quality multi-frame sample images from the second multi-frame sample image pair to obtain a third enhanced image includes: The coding layer is invoked to perform feature encoding on the low-quality multi-frame sample images in the second multi-frame sample image pair to obtain multi-frame encoded features; The gradient embedding layer is invoked to determine the gradient embedding features, and the diffusion layer is invoked to perform image enhancement on the gradient embedding features and the multi-frame coding features to obtain multi-frame enhanced coding features; The decoding layer is invoked to perform feature decoding on the multi-frame enhanced coding features to obtain the third enhanced image.
4. The method according to claim 1, characterized in that, The step of enhancing the multi-frame target images based on the first enhanced image model to obtain multi-frame enhanced target images includes: The multi-frame target images are subjected to quality detection to obtain a first quality score; When the first quality score is less than a preset threshold, the multi-frame target images are enhanced based on the first enhanced image model to obtain multi-frame preliminary enhanced images. The quality of the multiple pre-enhanced images is tested to obtain a second quality score; When the second quality score is greater than or equal to the preset threshold, the multi-frame target enhancement image is determined based on the multi-frame preliminary enhancement image.
5. The method according to claim 4, characterized in that, The quality of the multi-frame target images is measured to obtain a first quality score, including: The pre-trained discriminator model is invoked to perform quality detection on the multi-frame target images to obtain the first quality score, wherein the discriminator model includes a sampling layer and a prediction layer; The process of calling a pre-trained discriminator model to perform quality detection on the multi-frame target images to obtain the first quality score includes: The sampling layer is invoked to perform feature sampling on the multi-frame target images to obtain multi-frame target features; The prediction layer is invoked to perform quality detection on the target features of the multi-frame dataset to obtain the first quality score.
6. An image enhancement device based on airborne acquisition, characterized in that, The device includes: The image acquisition module is used to acquire multiple frames of target images to be enhanced; The image enhancement module is used to enhance the multi-frame target images based on the first enhanced image model to obtain multi-frame enhanced target images; The training process of the first enhanced image model includes the following steps: Acquire a first multi-frame sample image pair, which includes a set of low-quality multi-frame sample images and a corresponding set of high-quality multi-frame sample images obtained by using an image acquisition device under different acquisition conditions. The first image enhancement model is invoked to perform image enhancement based on the low-quality multi-frame sample images to obtain the first enhanced image, and the enhancement loss is determined based on the first enhanced image and the high-quality multi-frame sample images. The pre-trained second image enhancement model is invoked to perform image enhancement based on the low-quality multi-frame sample images to obtain a second enhanced image, and the distillation loss is determined based on the first enhanced image and the second enhanced image; A discriminator model is invoked to perform enhanced image source detection based on the first enhanced image and the second enhanced image to obtain the first sub-source detection result; Another discriminator model is invoked to perform enhanced image source detection based on the first enhanced image and the high-quality multi-frame sample images, to obtain a second sub-source detection result; wherein, the other discriminator model and the first discriminator model have the same network structure; The enhanced image source detection result is determined based on the first sub-source detection result and the second sub-source detection result; The parameters of the first enhanced image model are updated based on the enhanced image source detection results, the enhancement loss, and the distillation loss; The training process of the second enhanced image model includes the following steps: A second multi-frame sample image pair is obtained, which includes a set of low-quality multi-frame sample images and a corresponding set of high-quality multi-frame sample images. The second multi-frame sample image pair is obtained in the same way as the first multi-frame sample image pair. The second enhanced image model is invoked to perform image enhancement based on the low-quality multi-frame sample images in the second multi-frame sample image pair to obtain a third enhanced image, and the diffusion loss is determined based on the third enhanced image and the high-quality multi-frame sample images in the second multi-frame sample image pair. The parameters of the second enhanced image model are updated based on the diffusion loss.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Unsupervised low-illumination image enhancement method and system, equipment and medium
CN117893456A
Underwater image enhancement method, storage medium and computer program product
CN118674642A