Image enhancement method and device, computer equipment and storage medium

By using a pre-trained image enhancement model and multi-scale feature fusion technology, the problem of image detail enhancement in railway power supply system monitoring was solved, thereby improving image quality and detection rate.

CN120912447APending Publication Date: 2025-11-07SHUOHUANG RAILWAY DEV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510979154.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing image enhancement methods are insufficient to accurately enhance the details of images in railway power supply system monitoring, resulting in poor visual quality, affecting equipment defect detection rates, and increasing safety hazards.

Method used

The image is enhanced at the pixel level by using a pre-trained image enhancement model, and the final enhanced image is generated by combining feature extraction at different scales and multiple fusions, including pixel enhancement parameters and feature fusion.

Benefits of technology

It effectively improves the clarity of local details and overall integrity of images, enhances the visual quality of images in insufficient lighting scenarios, and improves the accuracy of equipment defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912447A_ABST
    Figure CN120912447A_ABST
Patent Text Reader

Abstract

The invention relates to an image enhancement method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a to-be-enhanced original image; enhancing each pixel point based on the pixel enhancement parameter corresponding to each pixel point in the original image through a pre-trained image enhancement model to obtain enhanced pixel points, and determining a first intermediate image according to each enhanced pixel point; performing feature extraction of different scales on the original image to obtain a plurality of feature images with different feature scales, and fusing the feature images to obtain a second intermediate image; performing first fusion based on the original image, the first intermediate image and the second intermediate image to obtain a fused image; and performing second fusion on the fused image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image, thereby effectively enhancing the image in complex scenes such as insufficient light supplement and the like, and considering image details and overall integrity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an image enhancement method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND

[0002] With the continuous advancement of informatization and digitization construction of the railway industry, in order to reduce the work intensity and personal safety risk of manual online inspection, vehicle-mounted image collection equipment has been widely deployed along the railway to capture the running state of facilities and equipment along the line in real time, providing reliable data support for the maintenance of railway facilities.

[0003] In the monitoring of the railway power supply system, the image collection equipment is a key component, but due to the on-site environment, insufficient light compensation occurs from time to time, resulting in poor visual quality of the collected images, thereby affecting the detection rate of equipment defects, weakening the performance of the monitoring equipment, making it difficult to discover and handle potential faults and safety hazards of the equipment in a timely manner, and threatening the safe and stable operation of the railway power supply system.

[0004] In order to improve the visual quality of the collected images, the images are usually enhanced, but the existing enhancement methods have poor enhancement effect when enhancing the images, and it is difficult to accurately enhance the image details. SUMMARY

[0005] Therefore, it is necessary to provide an image enhancement method, device, computer equipment, computer readable storage medium and computer program product capable of accurately identifying and enhancing image details in view of the above technical problems.

[0006] In a first aspect, the present application provides an image enhancement method, which comprises:

[0007] obtaining an original image to be enhanced;

[0008] enhancing each pixel point in the original image based on a pixel enhancement parameter corresponding to each pixel point through a pre-trained image enhancement model, obtaining an enhanced pixel point, and determining a first intermediate image according to each enhanced pixel point;

[0009] performing feature extraction of different scales on the original image to obtain a plurality of feature images with different feature scales, and fusing each feature image to obtain a second intermediate image;

[0010] performing first fusion based on the original image, the first intermediate image and the second intermediate image to obtain a fused image;

[0011] performing second fusion on the fusion image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image.

[0012] In one of the embodiments, the performing second fusion on the fusion image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image comprises:

[0013] performing weighted aggregation on the original image, the first intermediate image and the second intermediate image according to the respective fusion weights corresponding to the original image, the first intermediate image and the second intermediate image to obtain an aggregated image;

[0014] performing second fusion on the fusion image and the aggregated image to obtain an enhanced image for the original image.

[0015] In one of the embodiments, the performing feature extraction on the original image according to at least two feature scales to obtain feature data of the original image under different feature scales comprises:

[0016] performing feature extraction on the original image according to at least two feature scales to obtain feature data of the original image under different feature scales;

[0017] determining a feature sequence according to the respective feature scales corresponding to each of the feature data;

[0018] aligning each of the feature data according to the respective feature scales corresponding to each of the feature data from large to small to obtain aligned feature data;

[0019] performing fusion on the aligned feature data according to the feature scales from large to small in sequence to obtain a second intermediate image.

[0020] In one of the embodiments, the performing feature extraction on the original image according to at least two feature scales to obtain feature data of the original image under different feature scales comprises:

[0021] determining at least two feature scales for performing feature extraction on the original image;

[0022] for each of the feature scales, performing local feature extraction on the original image according to the feature scale to obtain a local feature map of the original image under the feature scale;

[0023] extracting the correlation between each of the pixel points in the local feature map based on a self-attention mechanism, and determining a global correlation parameter corresponding to each of the pixel points according to the correlation.

[0024] aggregate features of the pixel points based on the global correlation parameters corresponding to the pixel points respectively, to obtain feature data of the original image at the feature scale.

[0025] In one of the embodiments, the pixel points are enhanced based on the pixel enhancement parameters corresponding to the pixel points respectively, to obtain enhanced pixel points, and a first intermediate image is determined according to the enhanced pixel points, including:

[0026] extracting features of the pixel points in the original image, to obtain pixel features corresponding to the pixel points respectively;

[0027] determining pixel enhancement parameters for the pixel points, and enhancing the pixel points according to the pixel enhancement parameters, to obtain enhanced pixel points.

[0028] In one of the embodiments, the pixel enhancement parameters for the pixel points are determined, and the pixel points are enhanced according to the pixel enhancement parameters, to obtain enhanced pixel points, including:

[0029] determining initial pixel features of the pixel points and initial enhancement parameters for enhancing the initial pixel features;

[0030] determining intermediate pixel features based on the initial pixel features and the initial enhancement parameters;

[0031] in a case where the intermediate pixel features do not satisfy a preset enhancement condition, determining intermediate enhancement parameters for enhancing the intermediate pixel features, and enhancing the intermediate pixel features based on the intermediate enhancement parameters until the intermediate pixel features satisfy the preset enhancement condition, to obtain enhanced pixel features;

[0032] determining enhanced pixel points based on the enhanced pixel features.

[0033] In one of the embodiments, the image enhancement model is obtained by training based on the following method:

[0034] collecting sample images under different light intensities to form a training sample set; the sample images are unlabeled images;

[0035] inputting the training sample set into a pre-constructed image enhancement model, and obtaining enhanced sample images based on forward transmission of the pre-constructed image enhancement model;

[0036] determine a gray-scale prior loss value of the enhanced sample image based on the enhanced sample image and a prior loss function, train the pre-constructed image enhancement model based on the gray-scale prior loss value until the training is completed, and obtain a pre-trained image enhancement model; wherein the prior loss function is determined based on a statistical parameter of a gray-scale channel in the sample image.

[0037] In a second aspect, the present application further provides an image enhancement device, which comprises:

[0038] an image acquisition module, configured to acquire an original image to be enhanced;

[0039] a pixel enhancement module, configured to enhance each pixel point in the original image based on a pixel enhancement parameter corresponding to the pixel point, obtain an enhanced pixel point, and determine a first intermediate image based on each enhanced pixel point;

[0040] a feature enhancement module, configured to extract features of different scales from the original image, obtain a plurality of feature images with different feature scales, and fuse each feature image to obtain a second intermediate image;

[0041] a first fusion module, configured to perform first fusion based on the original image, the first intermediate image and the second intermediate image, and obtain a fused image;

[0042] a second fusion module, configured to perform second fusion on the fused image, the original image, the first intermediate image and the second intermediate image, and obtain an enhanced image for the original image.

[0043] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method as described above when executing the computer program.

[0044] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the method as described above when executed by a processor.

[0045] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, and the computer program implements the steps of the method as described above when executed by a processor.

[0046] The image enhancement method, device, computer device, computer readable storage medium and computer program product obtain an original image to be enhanced. The pre-trained image enhancement model enhances each pixel point based on a pixel enhancement parameter corresponding to each pixel point in the original image to obtain an enhanced pixel point, and determines a first intermediate image according to each enhanced pixel point. Different scale feature extraction is performed on the original image to obtain a plurality of feature images with different feature scales, and each feature image is fused to obtain a second intermediate image. The first fusion is performed based on the original image, the first intermediate image and the second intermediate image to obtain a fusion image. The second fusion is performed on the fusion image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image. Through the pre-trained image enhancement model, each pixel point can be enhanced based on the enhancement parameter corresponding to each pixel point in the original image, and different scale feature extraction can be performed on the original image. Not only can the pixel enhancement process realize the pixel-level targeted brightness adjustment to improve the clarity of the local details of the image, but also the multi-scale feature extraction process can include local features and global features. Finally, the original image, the first intermediate image and the second intermediate image are fused twice. The first fusion is performed on the original image, the first intermediate image and the second intermediate image to obtain a fusion image containing all feature information in the three images. Then, the original image, the first intermediate image after pixel enhancement and the second intermediate image after feature extraction are combined with the fusion image to determine the final enhanced image. The image in a complex scene such as insufficient light can be effectively enhanced, and the image details and overall integrity are considered. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creative labor.

[0048] Figure 1 An application environment diagram of an image enhancement method in an embodiment;

[0049] Figure 2 A flowchart of an image enhancement method in an embodiment;

[0050] Figure 3 A flowchart of a model training step in an embodiment;

[0051] Figure 4 A flowchart of an image enhancement method in an application example;

[0052] Figure 5 Fig. 1 is a structural block diagram of an image enhancement device in an embodiment;

[0053] Figure 6 Fig. 2 is an internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0054] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application.

[0055] It should be noted that the terms "first", "second", and the like used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" and any variations thereof used in the present application are intended to cover non-exclusive inclusion. The term "multiple" used in the present application refers to two or more. The term "and / or" used in the present application refers to one of the options or any combination of multiple options.

[0056] The image enhancement method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 The terminal 102 communicates with the server 104 through a network. The data storage system can store data required to be processed by the server 104. The data storage system can be integrated on the server 104, or placed on a cloud or other network server. When a user initiates an image enhancement operation on the terminal 102, the terminal 102 can respond to the image enhancement operation and initiate an image enhancement request to the server 104, the server 104 receives and responds to the image enhancement request, and obtains an original image to be enhanced. Then, the server 104 calls a pre-trained image enhancement model, and enhances each pixel point based on the pixel enhancement parameter corresponding to each pixel point in the original image through the pre-trained image enhancement model, to obtain an enhanced pixel point, and determines a first intermediate image according to each enhanced pixel point. The server 104 performs feature extraction of different scales on the original image to obtain a plurality of feature images with different feature scales, and fuses each feature image to obtain a second intermediate image. Finally, the server 104 performs first fusion based on the original image, the first intermediate image and the second intermediate image to obtain a fusion image, performs second fusion on the fusion image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image, and returns the enhanced image to the terminal 102.

[0057] The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, unmanned aerial vehicles, low-altitude aerial vehicles, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle-mounted device, a projection device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0058] In an exemplary embodiment, as shown in Figure 2 , an image enhancement method is provided, which is applied to the server 104 in Figure 1 for example, and includes the following steps 202 to 210. Among them:

[0059] Step 202, obtaining an original image to be enhanced.

[0060] The original image to be enhanced refers to the initial image data that needs to be enhanced. For example, in the monitoring of the railway power supply system, the original image to be enhanced can include the 4C image collected by the 4C device, which is the abbreviation of the High-precision Inspection and Monitoring Device for Catenary Condition. In specific implementation, the original image can have quality problems such as insufficient brightness, low contrast, noise, and blur, and needs to be enhanced so that the enhanced image can meet the image analysis, recognition, and other tasks.

[0061] For example, the server can select the original image to be enhanced from the pre-collected original image according to the image enhancement requirement. In other embodiments, the server can also collect images in real time through an image collection device as the original image to be enhanced according to the image analysis requirement.

[0062] Step 204, enhancing each pixel point in the original image based on the pixel enhancement parameter corresponding to each pixel point through the pre-trained image enhancement model, obtaining an enhanced pixel point, and determining a first intermediate image according to each enhanced pixel point.

[0063] The pre-trained image enhancement model refers to a model that has been trained and can be used to enhance the original image. The image enhancement model can enhance the original image by learning the mapping relationship between the original image and the enhanced image, and obtain an enhanced image. In specific implementation, the image enhancement model can be constructed based on a deep learning algorithm, for example, one or more of a convolutional neural network (CNN), a generative adversarial network (GAN), and self-attention (SA).

[0064] The pixel enhancement parameter refers to a parameter used to adjust the enhancement intensity when each pixel point in the original image is enhanced. The pixel enhancement parameter can include but is not limited to a brightness enhancement parameter, a contrast enhancement parameter, a color saturation enhancement parameter, etc. The brightness enhancement parameter can be used to change the brightness level of the pixel point, the contrast enhancement parameter can be used to control the brightness difference between the pixel point and its surrounding pixel points, and increasing the parameter can enhance the contrast to make the objects in the image clearer. The color saturation enhancement parameter can be used to adjust the vividness of the color of the pixel point, and increasing the parameter can make the color more saturated. By configuring a corresponding pixel enhancement parameter for each pixel point, each pixel point in the original image can be adjusted accordingly to improve the overall quality of the image. For example, increasing the brightness parameter value of a certain pixel point can make the pixel point brighter in the enhanced image.

[0065] The enhanced pixel point refers to a new pixel point obtained after the pixel point in the original image is adjusted based on the pixel enhancement parameter. The enhanced pixel point has different attribute values such as brightness, contrast, color, etc. from the original pixel point. In the image enhancement model, the original pixel point is adjusted according to the pixel enhancement parameter corresponding to each pixel point, thereby generating the enhanced pixel point. The first intermediate image refers to an image composed of the enhanced pixel points according to the arrangement order of the pixel points in the original image, and the first intermediate image is obtained after the original image is enhanced at the pixel level based on the pixel enhancement parameter.

[0066] For example, the server can call the pre-trained image enhancement model and input the original image to be enhanced into the pre-trained image enhancement model. The server can determine each pixel point in the original image and the pixel enhancement parameter corresponding to each pixel point through the pre-trained image enhancement model, and enhance each pixel point according to the corresponding pixel enhancement parameter to obtain the enhanced pixel point. Then, the server can arrange the enhanced pixel points according to the positions of the original pixel points to obtain the first intermediate image.

[0067] At step 206, different scale feature extraction is performed on the original image to obtain a plurality of feature images with different feature scales, and each feature image is fused to obtain a second intermediate image.

[0068] Different scale feature extraction refers to a process of extracting feature representations with different resolutions from the original image. In specific implementation, different scale feature extraction can be achieved by using different size convolution kernels or pooling operations to capture information at different levels of the original image; for example, large scale feature extraction can obtain the overall structure of the original image, such as the shape and position of an object; small scale feature extraction can obtain local details and texture information of the original image, such as edges and textures. A feature image refers to an image with a specific scale of feature representation obtained in the process of different scale feature extraction, and each feature image can correspond to a specific feature scale. In the feature image corresponding to the feature scale, the feature information of the image at the feature scale is included. Different scale feature images can complement each other to jointly describe the information of the original image.

[0069] Fusing the feature images refers to combining a plurality of feature images with different feature scales to obtain a fused feature image that integrates the feature information of each scale. The fusion method can include but is not limited to simple weighted average, splicing, attention mechanism, etc. Through the fusion of feature images, the advantages of different scale feature images can be fully utilized to improve the representation ability of features, so that the fused feature image can more comprehensively describe the content of the original image. The second intermediate image refers to an image generated by the fused plurality of feature images. The second intermediate image is a comprehensive feature representation of the original image after different scale feature extraction and fusion, which can highlight the local details and texture information of the image on the basis of preserving the overall structure of the original image.

[0070] For example, the server can perform feature extraction on the original image according to different scales by using the trained image enhancement model to obtain a plurality of feature images with different feature scales; then, the server can splice each feature image to obtain a second intermediate image. When splicing, since each feature image has a different feature scale, the server can first align the scales of each feature image and then splice them to obtain the final second intermediate image.

[0071] At step 208, a first fusion is performed based on the original image, the first intermediate image, and the second intermediate image to obtain a fused image.

[0072] The first fusion refers to a process of performing element-wise summation on respective corresponding parallel features in the original image, the first intermediate image and the second intermediate image and merging the features obtained by the summation. The parallel features refer to features located in the same area range in the original image, the first intermediate image and the second intermediate image in a certain area range. The certain area range can be an area divided in units of pixels, or an area divided in units of several pixels or a certain range. For example, in units of pixels, the position of each pixel in the corresponding image can be represented by a pixel coordinate. In the original image, the first intermediate image and the second intermediate image, the features corresponding to the pixels with the same pixel coordinate can be regarded as parallel features. The first fusion is a process of sorting and fusing the features corresponding to each pixel after performing summation on the features, to obtain a fused image. In the fused image, image information of different sources is merged to comprehensively utilize the advantages of the original image, the pixel-level enhancement result and the feature-level enhancement result, and to improve the image enhancement effect. For example, the original image can provide rich detail information, the first intermediate image can improve the direct adjustment effect of pixel attributes, and the second intermediate image can enhance the feature representation of the image. The original detail information, the pixel-level feature and the multi-scale feature are merged, so that the fused image can contain multi-source information.

[0073] For example, the server can determine the area for parallel feature merging in the original image, the first intermediate image and the second intermediate image through the pre-trained image enhancement model, for example, taking a single pixel as an area or an element. The server indexes the pixel coordinates, adds the features of the same pixel in the original image, the first intermediate image and the second intermediate image, and then combines the results obtained by the addition in the order of the pixel coordinates to obtain a fused image. For example, the pixel value of pixel A in the original image is 0.2, the pixel value of pixel A in the first intermediate image is 0.3, and the pixel value of pixel A in the second intermediate image is 0.35. Adding the pixel values, the pixel value of pixel A in the fused image is 0.85.

[0074] In step 210, the fused image, the original image, the first intermediate image and the second intermediate image are subjected to second fusion to obtain an enhanced image for the original image.

[0075] The second fusion refers to a process of performing feature fusion again on the fused image, the original image, the first intermediate image and the second intermediate image. In a specific implementation, the original image, the first intermediate image and the second intermediate image can be fused in a certain manner first, and then fused again with the fused image. For example, the original image, the first intermediate image and the second intermediate image can be fused by weighting, splicing or the like, and then fused again with the fused image combined with the original information, the pixel-level information and the multi-scale information to obtain the final enhanced image. Alternatively, the fused image can be fused with the original image, the first intermediate image and the second intermediate image respectively to obtain respective intermediate fused images, and then the intermediate fused images are fused in a certain manner (such as weighting, splicing or the like) to obtain the final enhanced image.

[0076] For example, the server can fuse the original image, the first intermediate image and the second intermediate image in a weighted manner to obtain a weighted image. For example, the server can fuse the pixel values of each pixel point in the original image, the first intermediate image and the second intermediate image in units of pixel points to obtain fused pixel values of each pixel point, and aggregate the pixel values based on the positions of the pixel points to obtain the weighted image. Then, the server can perform second fusion on the weighted image and the fused image to obtain the enhanced image for the original image. In the second fusion, the server can perform second fusion on the pixel values of each pixel point in the weighted image and the fused image in units of pixel points to obtain second fused pixel values of each pixel point, and aggregate the second fused pixel values based on the positions of the pixel points to obtain the enhanced image for the original image.

[0077] In the image enhancement method, an original image to be enhanced is obtained; each pixel point in the original image is enhanced based on a pixel enhancement parameter corresponding to each pixel point, to obtain an enhanced pixel point, and a first intermediate image is determined according to each enhanced pixel point, by using a pre-trained image enhancement model; feature extraction of different scales is performed on the original image to obtain a plurality of feature images with different feature scales, and each feature image is fused to obtain a second intermediate image; first fusion is performed based on the original image, the first intermediate image and the second intermediate image to obtain a fused image; and second fusion is performed on the fused image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image. By using the pre-trained image enhancement model, each pixel point in the original image can be enhanced based on the enhancement parameter corresponding to each pixel point, and feature extraction of different scales can be performed on the original image. In the pixel enhancement process, the brightness of each pixel point can be adjusted to improve the clarity of local details of the image, and in the multi-scale feature extraction process, local features and global features can be included. Finally, the original image, the first intermediate image and the second intermediate image are fused twice. In the first fusion, the original image, the first intermediate image and the second intermediate image are directly fused to obtain a fused image containing all feature information of the three images. Then, the original image, the first intermediate image after pixel enhancement and the second intermediate image after feature extraction are combined with the fused image to determine the final enhanced image. The image in a complex scene such as insufficient light can be effectively enhanced, and the image details and overall integrity are considered.

[0078] In one embodiment, the second fusion of the fused image, the original image, the first intermediate image and the second intermediate image obtains an enhanced image for the original image, including:

[0079] The original image, the first intermediate image and the second intermediate image are weighted and aggregated according to the fusion weights corresponding to the original image, the first intermediate image and the second intermediate image, to obtain an aggregated image; and the second fusion of the fused image and the aggregated image obtains an enhanced image for the original image.

[0080] The fusion weight refers to a weight coefficient assigned to each image when the original image, the first intermediate image and the second intermediate image are weighted and aggregated. The fusion weight is used to represent the proportion of each image in the aggregation result. In specific implementation, the fusion weight can be determined according to the enhancement target. For example, if the enhanced image needs to be closer to the original image, the fusion weight of the original image can be set to be larger. The fusion weight can also be determined according to the effect or accuracy of the image enhancement model in generating the first intermediate image and the second intermediate image. For example, if the image enhancement model has higher accuracy in generating the first intermediate image, it means that the processing effect of the first intermediate image is better, and the fusion weight of the first intermediate image can be set to be higher than that of the second intermediate image.

[0081] The weighted aggregation refers to a process of weighted summation of the original image, the first intermediate image and the second intermediate image according to their respective fusion weights. The weighted aggregation can comprehensively utilize the information of the original image, the first intermediate image and the second intermediate image, and adjust the contribution of each image in the aggregation result according to their respective weights, so as to obtain a more balanced and optimized image result. By reasonably setting the fusion weight, the advantages of certain images can be highlighted, and the shortcomings of other images can be suppressed, thereby improving the overall quality of the image. The aggregated image refers to an image obtained by weighted aggregation of the original image, the first intermediate image and the second intermediate image. The aggregated image fuses the information of the original image, the first intermediate image and the second intermediate image, and under the action of the fusion weight, the features of each image can be reasonably combined and adjusted, so that the aggregated image has more abundant information.

[0082] The fusion image and the aggregated image are secondly fused, that is, the information of the fusion image and the aggregated image is used to further determine the enhanced image. In specific implementation, the fusion image and the aggregated image can be multiplied pixel by pixel to determine the enhanced image. By determining the enhanced image according to the fusion image and the aggregated image, the effect of image enhancement can be further improved. The fusion image and the aggregated image can provide different feature information and enhanced perspectives, so that the enhanced image can better utilize the advantages of the original image, the first intermediate image and the second intermediate image, thereby obtaining a better enhanced image.

[0083] For example, the server can obtain the fusion weights corresponding to the original image, the first intermediate image and the second intermediate image respectively. Then, the server can perform weighted aggregation on the original image, the first intermediate image and the second intermediate image according to the fusion weights corresponding to the original image, the first intermediate image and the second intermediate image respectively, to obtain an aggregated image; for example, the server can multiply the original image, the first intermediate image and the second intermediate image respectively by the fusion weights corresponding to the original image, the first intermediate image and the second intermediate image respectively, and then add them to obtain the aggregated image. Finally, the server can perform the second fusion on the fusion image and the aggregated image to obtain the enhanced image for the original image; for example, the server can multiply the fusion image and the aggregated image pixel by pixel to obtain the final enhanced image.

[0084] In one application example, when the server performs the first fusion and the second fusion on the original image, the first intermediate image and the second intermediate image in sequence, the server can perform fusion based on a nonlinear manner, integrate the features of each image by using the fusion weights corresponding to each of the original image, the first intermediate image and the second intermediate image, to refine the detailed features in different images. In specific implementation, the server can first merge the parallel features in the original image, the first intermediate image and the second intermediate image according to element (pixel) summation to obtain a fusion image, that is:

[0085] (1)

[0086] wherein, is the fusion image, which includes a plurality of first fused pixel points, and the pixel value of each first fused pixel point is obtained by adding the original pixel values of the original image, the first intermediate image and the second intermediate image at the corresponding pixel point; is the original image; is the first intermediate image, is the second intermediate image.

[0087] Then, the server can aggregate the original image, the first intermediate image and the second intermediate image according to the respective corresponding fusion weights, that is, the server can multiply the original image by the corresponding fusion weight to obtain a weighted original image, multiply the first intermediate image by the corresponding fusion weight to obtain a weighted first intermediate image, and multiply the second intermediate image by the corresponding fusion weight to obtain a weighted second intermediate image. Then, the server adds the weighted original image, the weighted first intermediate image and the weighted second intermediate image to obtain an aggregated image. Finally, the server multiplies the aggregated image and the fusion image pixel by pixel to obtain an enhanced image. In specific implementation, when multiplying each image by the corresponding fusion weight and adding the weighted images, the server can weight (multiply) and add the feature values corresponding to each pixel point in units of pixel points, and finally combine the pixel points to obtain the corresponding image. The enhanced image can be represented as:

[0088] (2)

[0089] wherein, is the enhanced image; is the original image, is the fusion weight of the original image; is the first intermediate image, is the fusion weight of the first intermediate image; is the second intermediate image, is the fusion weight of the second intermediate image; is the fusion image; is pixel by pixel multiplication, that is, two images are multiplied pixel by pixel.

[0090] In this embodiment, by using multi-level image fusion, on the one hand, the original image, the first intermediate image and the second intermediate image are weighted and aggregated according to different fusion weights, which can integrate the advantages of multiple image information to generate an aggregated image, and on the other hand, the fusion image and the aggregated image are fused twice to obtain an enhanced image, which can further optimize the image quality, strengthen the key information, effectively improve the image quality, etc., so that the final enhanced image is better in visual effect and information expression, and meets the needs of various application scenarios.

[0091] In one embodiment, different scale feature extraction is performed on the original image to obtain a plurality of feature images with different feature scales, and each feature image is fused to obtain the second intermediate image, including:

[0092] The feature extraction is performed on the original image according to at least two feature scales to obtain feature data of the original image under different feature scales; a feature sequence is determined according to the feature scales corresponding to each feature data; the feature data is aligned according to the feature scales corresponding to each feature data from large to small to obtain aligned feature data; and the aligned feature data is fused in sequence according to the feature scales from large to small to obtain a second intermediate image.

[0093] The feature scale refers to the size of the image feature in the spatial resolution. Different feature scales can reflect the information of the image at different levels, for example, small-scale features usually correspond to local details of the image, such as edges, textures, etc.; large-scale features correspond to the overall structure and semantic information of the image, such as the shape and position of the object. The feature scale can be represented in different ways, such as the size of the convolution kernel, the size of the pooling window, etc. In specific implementation, different convolution kernels can be set, such as 1x1, 3x3, etc., and different feature scales can be achieved by setting different convolution operations on the original image. Feature extraction refers to the process of extracting feature information from the original image, and the feature information can include but is not limited to color features, texture features, shape features, etc. Feature data refers to data representing image features obtained in the feature extraction process; the feature data can exist in the form of a multidimensional array or a vector to include feature values of the image under different feature scales. For example, the feature data can include various types of features, such as color histograms, texture feature vectors, shape descriptors, etc. Each feature data corresponds to the feature representation of the original image under a specific feature scale.

[0094] The feature sequence refers to an ordered sequence or sequence determined according to the feature scales corresponding to each feature data. The feature sequence sorts the feature data of different feature scales according to certain rules, for example, arranges the feature data in order from small to large or from large to small. In an embodiment, if the extracted feature scales are 3x3, 5x5 and 7x7, the feature sequence can be [3x3 feature data, 5x5 feature data, 7x7 feature data] or [7x7 feature data, 5x5 feature data, 3x3 feature data].

[0095] The alignment refers to matching and adjusting the feature data of different feature scales in space, so that each feature data has the same spatial resolution and positional relationship, to correct the different sizes and positional offsets of the feature data extracted under different feature scales, so that each feature data can be effectively fused. In specific implementation, the alignment method can include but is not limited to one or more of interpolation method, affine transformation, etc. For example, for small-scale feature data, it can be up-sampled to the size of large-scale feature data by interpolation method, or the spatial position of the feature data is adjusted by affine transformation. The fusion refers to combining multiple aligned feature data to obtain a fused feature representation that integrates information of each feature scale. The fusion method can include but is not limited to simple weighted average, splicing, attention mechanism, etc. By fusing feature data of different feature scales, the information of the image at different levels can be fully utilized, and the representation ability and robustness of the feature can be improved. The fused feature can better describe the complex structure and semantic information of the image, which is helpful for subsequent image classification, target detection, etc.

[0096] For example, the server can perform feature extraction under different feature scales on the original image respectively according to different feature scales, and obtain feature data corresponding to each feature scale respectively. Then, the server can sort the feature data according to the feature scales corresponding to each feature data respectively in ascending order of the feature scales, to obtain a feature sequence containing each feature data. Next, the server can start from the last feature data in the feature sequence, up-sample the last feature data, so that the last feature data is aligned with the previous feature data in the feature scale, to obtain an aligned feature data (last). Finally, the server fuses the aligned feature data with the previous feature data, and repeats the process until all feature data in the feature sequence are processed, to obtain a second intermediate image.

[0097] In an application example, when aligning and fusing the feature sequence, the server takes a feature sequence including three feature data as an example, which are feature data F1, feature data F2, and feature data F3. The feature data F1 is the data feature obtained by the server performing feature extraction on the original image according to the feature scale D1. The feature data F2 is the data feature obtained by the server performing feature extraction on the original image according to the feature scale D2. The feature data F3 is the data feature obtained by the server performing feature extraction on the original image according to the feature scale D3. First, the server integrates the feature data F1~F3 into a feature sequence [F1, F2, F3] according to the feature scale from small to large. Then, the server up-samples the feature data F3 under the feature scale D3 to align with the feature data F2 by interpolation, that is, the server up-samples the feature scale D3 of the feature data F3 to the feature scale D2. The server then obtains the intermediate feature data M1 by performing weighted averaging on the feature data F3 under the feature scale D2 and the feature data F2. Next, the server again up-samples the intermediate feature data M1 under the feature scale D2 to align with the feature data F1 by interpolation, that is, the server up-samples the feature scale D2 of the intermediate feature data M1 to the feature scale D1. The server then obtains the intermediate feature data M2 by performing weighted averaging on the intermediate feature data M1 under the feature scale D1 and the feature data F1. Finally, the server can determine the second intermediate image based on the intermediate feature data M2, or obtain the second intermediate image by performing convolution on the intermediate feature data M2.

[0098] In this embodiment, by extracting feature data of the original image according to multiple feature scales, aligning each feature data based on the determined feature sequence, and then fusing the aligned feature data to generate the second intermediate image, the details and structural information of different levels of the original image can be comprehensively captured, the advantages of different scale features can be fully utilized, the image quality can be effectively improved, the visual effect and information integrity of the image can be ensured, and the data basis for subsequent image fusion is prepared.

[0099] In one embodiment, feature extraction is performed on the original image according to at least two feature scales to obtain feature data of the original image under different feature scales, including:

[0100] At least two feature scales for performing feature extraction on the original image are determined. For each feature scale, local feature extraction is performed on the original image according to the feature scale to obtain a local feature map of the original image under the feature scale. The correlation between each pixel point in the local feature map is extracted based on a self-attention mechanism, and a global correlation parameter corresponding to each pixel point is determined according to the correlation. Feature aggregation is performed on each pixel point based on the global correlation parameter corresponding to each pixel point to obtain feature data of the original image under the feature scale.

[0101] wherein the local feature extraction refers to a process of extracting features for local regions in the original image. For each determined feature scale, feature extraction is performed on the original image according to the corresponding feature scale, and information describing the features of the local regions can be obtained. In specific implementation, a convolution operation can be used to extract local features in the original image according to different feature scales; for example, for a 3x3 feature scale, a 3x3 convolution kernel can be used to slide on the original image, and a convolution calculation is performed on each 3x3 local region to obtain the feature value of the region, thereby generating a local feature map under the 3x3 feature scale. The local feature map refers to an image representation obtained after local feature extraction. The size of the local feature map can be the same as or different from the size of the original image, and the value of each element in the local feature map can be used to represent the feature information of the corresponding local region in the original image. The local feature map can highlight the features of each local region in the original image, and compared with the original image, the local feature map can focus on the feature expression of the local region; for example, in the local feature map of edge detection, the value of the edge region will be relatively high, and the value of other regions will be low.

[0102] The self-attention mechanism can be used to determine the association relationship between each element in the local feature map. In specific implementation, the self-attention mechanism can regard each pixel point in the local feature map as an element in a sequence, and determine the association degree between each pixel point by calculating the similarity between the pixel points. For each pixel point in the local feature map, the self-attention mechanism can determine the similarity scores between the pixel point and all other pixel points, and then perform weighted summation on the features of the other pixel points according to the similarity scores to obtain a new feature representation of the pixel point. Based on this, in the local feature map, the association relationship refers to the mutual connection and influence degree between each pixel point. The association relationship can be determined by the similarity scores of the pixel points. The global association parameter is a parameter determined based on the association relationship between the pixel points, and is used to reflect the global importance of each pixel point in the entire local feature map or the comprehensive association degree with other pixel points. In specific implementation, the global association parameter can be obtained by normalizing (such as using a softmax function) the similarity scores (i.e., the association relationship) determined by the self-attention mechanism.

[0103] Feature aggregation refers to the process of integrating and merging the features of the pixel points according to the global correlation parameters corresponding to each pixel point. Through feature aggregation, the dispersed pixel point features in the local feature map can be fused into more representative and comprehensive feature representations. For each pixel point, the global correlation parameters are weighted and summed with the features of all pixel points in the local feature map to obtain the new features of the pixel point after feature aggregation. In this way, the new features of each pixel point not only contain its own local information, but also fuse the information of other related pixel points, thereby better reflecting the overall features of the image.

[0104] For example, the server can determine multiple feature scales for feature extraction of the original image according to the effect requirements of image enhancement; for example, taking the size of the convolution kernel as an example, it can be determined that the multiple feature scales include 3x3, 5x5, and 7x7. The server can perform local feature extraction on the original image according to each feature scale to obtain the local feature map of the original image under the corresponding feature scale; for example, the server can perform local feature extraction on the original image according to the feature scales of 3x3, 5x5, and 7x7 to obtain the corresponding local feature maps under each feature scale, such as the local feature map of 3x3, the local feature map of 5x5, and the local feature map of 7x7. The server can extract the correlation between each pixel point in the local feature map based on the self-attention mechanism, and determine the global correlation parameters corresponding to each pixel point according to the correlation; for example, for each pixel point in each local feature map, the server can determine the similarity scores between each pixel point and all other pixel points, and then normalize each similarity score to obtain the global correlation parameters corresponding to each pixel point. The server can perform feature aggregation on each pixel point based on the global correlation parameters corresponding to each pixel point to obtain the feature data of the original image under the corresponding feature scale; for example, for each pixel point, the server can perform weighted summation on the features of other pixel points of the corresponding pixel point according to the global correlation parameters corresponding to the corresponding pixel point to obtain a new feature representation of the corresponding pixel point, and obtain the feature data under each feature scale based on the new feature representation of each pixel point.

[0105] In this embodiment, by determining multiple feature scales and extracting local feature maps according to different feature scales, then determining the correlation between each pixel point using the self-attention mechanism and determining the global correlation parameters of each pixel point, and finally aggregating the features of each pixel point based on the global correlation parameters to obtain the feature data, the local information and global information of different levels of the original image can be comprehensively captured, the local and global features can be effectively integrated, more representative feature data can be obtained, and the accuracy and effectiveness of image feature expression can be improved.

[0106] In an embodiment, the respective pixel points in the original image are enhanced based on respective pixel enhancement parameters corresponding to the respective pixel points, to obtain enhanced pixel points, and a first intermediate image is determined according to the respective enhanced pixel points, including:

[0107] The features of the respective pixel points in the original image are extracted to obtain respective pixel features corresponding to the respective pixel points, and for each pixel point, a pixel enhancement parameter of the pixel point is determined, and the pixel point is enhanced according to the pixel enhancement parameter to obtain an enhanced pixel point.

[0108] In the embodiment, the extraction of the features of the respective pixel points refers to a process of extracting the pixel features of the pixel points from the pixel points of the original image. The pixel features refer to feature vectors or feature values obtained by feature extraction of the pixel points, and the pixel features can be used to describe the attributes of the pixel points, such as color features, texture features, brightness features, etc.

[0109] For example, the server can extract the features of the respective pixel points in the original image to obtain respective pixel features corresponding to the respective pixel points. For example, the server can use a convolutional neural network (CNN) to input the original image into the network and automatically extract the features of the respective pixel points through convolution operation. In other embodiments, the server can also use traditional image processing algorithms, such as a scale invariant feature transform algorithm (SIFT), to analyze each pixel point to obtain a feature vector of each pixel point as a pixel feature. The server can determine a pixel enhancement parameter of each pixel point, and enhance the pixel point according to the pixel enhancement parameter to obtain an enhanced pixel point. For example, the server can analyze the features of each pixel point and its surrounding area to determine respective brightness, contrast, color saturation, and other enhancement parameters, such as a brightness enhancement parameter. Then, the server adjusts the pixel features (brightness) of the respective pixel points according to the brightness enhancement parameter to obtain the enhanced pixel points.

[0110] In the embodiment, by extracting the features of the respective pixel points in the original image and enhancing each pixel point using the corresponding pixel enhancement parameter, individualized pixel optimization can be achieved, and the expressiveness of the key pixels in the image can be effectively improved, and the overall quality of the image can be improved.

[0111] In an embodiment, the respective pixel points in the original image are enhanced based on respective pixel enhancement parameters corresponding to the respective pixel points, to obtain enhanced pixel points, and a first intermediate image is determined according to the respective enhanced pixel points, including:

[0112] determine an initial pixel feature of the pixel point and an initial enhancement parameter for enhancing the initial pixel feature; determine an intermediate pixel feature based on the initial pixel feature and the initial enhancement parameter; in a case where the intermediate pixel feature does not satisfy a preset enhancement condition, determine an intermediate enhancement parameter for enhancing the intermediate pixel feature, and enhance the intermediate pixel feature based on the intermediate enhancement parameter until the intermediate pixel feature satisfies the preset enhancement condition, to obtain an enhanced pixel feature; and determine an enhanced pixel point based on the enhanced pixel feature.

[0113] The initial pixel feature refers to a feature representing original information of the pixel point in the original image. Based on the description of the pixel feature, the initial pixel feature can be a color value, a brightness value, or a feature obtained through a simple feature extraction method (such as calculating the average color of a local region, the maximum and minimum brightness difference, etc.). The initial enhancement parameter refers to a parameter used for preliminary adjustment and enhancement of the initial pixel feature. The initial enhancement parameter can be set according to the target of image processing and experience, or can be automatically calculated based on statistical information of the image (such as average brightness, contrast, etc.).

[0114] The intermediate pixel feature refers to a pixel feature calculated based on the initial pixel feature and the initial enhancement parameter. The intermediate pixel feature is the result of the initial pixel feature after preliminary enhancement, and is used to reflect the pixel feature of the pixel point after the initial enhancement operation. The preset enhancement condition refers to a condition preset for judging whether the intermediate pixel feature satisfies the enhancement target. The preset enhancement condition is used as a condition for judging whether the intermediate pixel feature needs to be further enhanced, and can ensure that the final obtained pixel feature has appropriate intensity (such as brightness, contrast, etc.) to meet the needs of subsequent image processing and analysis. The preset enhancement condition can be determined according to the specific application scenario and the image processing target. For example, in the image enhancement of the overhead contact system, in order to clearly identify the image of the overhead contact system and facilitate feature extraction and identification, a suitable brightness range can be set as the preset enhancement condition. The intermediate enhancement parameter refers to a parameter used for further enhancing the intermediate pixel feature in a case where the intermediate pixel feature does not satisfy the preset enhancement condition. Similar to the initial enhancement parameter, the intermediate enhancement parameter can also be determined according to the state of the current intermediate pixel feature and the preset enhancement condition, with the purpose of further adjusting the brightness, contrast, etc. of the pixel feature so that the adjusted pixel feature meets the requirements.

[0115] The enhanced pixel feature refers to the final pixel feature obtained after multiple enhancement operations (including enhancement based on the initial enhancement parameter and the intermediate enhancement parameter). The enhanced pixel feature satisfies the preset enhancement condition, and the enhanced pixel feature is optimized in brightness, contrast, etc. compared with the initial pixel feature, and is more suitable for subsequent image processing tasks such as image classification, target detection, image segmentation, etc. The enhanced pixel point refers to a pixel point determined based on the enhanced pixel feature. The enhanced pixel point has color and brightness information corresponding to the enhanced pixel feature. The enhanced pixel point can be combined to form an enhanced image.

[0116] For example, the server can determine the initial pixel feature of the pixel point and the initial enhancement parameter for enhancing the initial pixel feature; for example, the server can use a mapping estimation network to determine the optimal mapping relationship of each pixel point in the original image in the gray channel, determine the corresponding enhancement parameter in the case of different pixel features, and in specific implementation, the mapping estimation network can be implemented by using a convolutional neural network (CNN) or a U-Net (U-shaped network) type encoder-decoder structure, and by training the mapping estimation network, the mapping estimation network can learn that when an image with more dark pixels is input, the enhancement parameter corresponding to the darker pixel point is larger, and the enhancement parameter corresponding to the brighter pixel point is smaller. The server can determine the intermediate pixel feature based on the initial pixel feature and the initial enhancement parameter; for example, the server can enhance the initial pixel feature of each pixel point by using the corresponding initial enhancement parameter to obtain the corresponding intermediate pixel feature. Then, the server can compare the intermediate pixel feature with the preset enhancement condition, that is, judge whether the intermediate pixel feature meets the enhancement target, such as whether the brightness feature reaches the target brightness range, if the intermediate pixel feature meets the preset enhancement condition, the iteration is ended, and the intermediate enhancement feature is determined as the enhanced pixel feature. If the intermediate pixel feature does not meet the preset enhancement condition, the intermediate enhancement parameter for enhancing the intermediate pixel feature is determined again, and the intermediate pixel feature is enhanced based on the intermediate enhancement parameter until the intermediate pixel feature meets the preset enhancement condition, and the enhanced pixel feature is obtained; that is, the server inputs the intermediate pixel feature into the mapping estimation network, outputs the corresponding intermediate enhancement parameter through the mapping estimation network, and enhances the intermediate pixel feature of each pixel point by using the corresponding intermediate enhancement parameter until the intermediate pixel feature meets the preset enhancement condition, and the enhanced pixel feature is obtained. Finally, the server can determine the enhanced pixel point based on the enhanced pixel feature.

[0117] In an application example, the server can input the initial pixel features corresponding to each pixel point into the mapping estimation network, output the initial enhancement parameters corresponding to the initial pixel features of each pixel point through the mapping estimation network, and enhance the initial pixel features by using the initial enhancement parameters to obtain intermediate pixel features, denoted as:

[0118] (3)

[0119] wherein, is the intermediate pixel feature after the first enhancement, is the initial pixel feature corresponding to the pixel point at the pixel coordinate in the original image, is the initial enhancement parameter corresponding to the initial pixel feature of the pixel point at the pixel coordinate in the original image.

[0120] Then, if the server judges that the intermediate pixel feature does not satisfy the preset enhancement condition, the server can input the intermediate pixel feature into the mapping estimation network, output the intermediate enhancement parameters corresponding to the intermediate pixel features of each pixel point through the mapping estimation network, and enhance the intermediate pixel features by using the intermediate enhancement parameters to obtain new intermediate pixel features. Repeat this process, and through each iteration, the pixel features of each pixel point can be adaptively enhanced. The iteration process can be represented as:

[0121] (4)

[0122] wherein, is the intermediate pixel feature after the n th enhancement, is the intermediate pixel feature after the (n-1) th enhancement of the pixel point at the pixel coordinate , is the intermediate enhancement parameter corresponding to the intermediate pixel feature of the pixel point at the pixel coordinate after the (n-1) th enhancement, is the iteration number, i.e., the number of enhancements.

[0123] In this embodiment, by determining the initial pixel features and the initial enhancement parameters, the initial pixel features are enhanced based on the initial enhancement parameters to obtain the intermediate pixel features, and then the intermediate pixel features are continuously enhanced. Through multiple rounds of enhancement, and in each round of enhancement process, the enhancement is based on the corresponding enhancement parameter, which can adaptively adjust the enhancement process to avoid over-enhancement or insufficient enhancement, and then generate the enhanced pixel points, which can effectively improve the image pixel quality.

[0124] In one embodiment, as shown in Figure 3 , the image enhancement model is trained based on the following method:

[0125] Step 302, collect sample images under different light intensities to form a training sample set.

[0126] Wherein, the sample image refers to a sample used for training the image enhancement model, and the sample image can be a digital image collected from a real scene. Each sample image can contain visual information in a specific scene, such as visual information in scenes with sufficient / insufficient light during the day, insufficient / insufficient light at night, etc. The sample image can exist in the form of pixel points, each pixel point having different brightness, color, and position attributes, such as some pixel points being brighter and some pixel points being darker. In specific implementation, the sample image is an unlabeled image. The unlabeled image refers to an image without additional artificial annotation information in the sample image. In the unlabeled image, there is original image data without additional semantic or structural annotation for subsequent unsupervised model training. At the same time, training with unlabeled images can reduce the cost and difficulty of data annotation. The training sample set refers to a set composed of a large number of sample images, used for training the image enhancement model. In specific implementation, taking the enhancement of the overhead contact line image as an example, the training sample set can contain a number of sample images without annotation, each sample image being collected under different light intensities for the image enhancement model to identify and process different brightness.

[0127] For example, the server can collect sample images under different light intensities through the vehicle-mounted image collection device, and directly combine each sample image to form a training sample set.

[0128] Step 304, input the training sample set into the pre-constructed image enhancement model, and obtain the enhanced sample image based on the forward propagation of the pre-constructed image enhancement model.

[0129] Wherein, the pre-constructed image enhancement model refers to a model constructed based on deep learning or other machine learning algorithms, used for enhancing the input sample image to improve the quality of the image, such as improving brightness, contrast, and clarity, etc. In specific implementation, the image enhancement model can be constructed using one or more of the convolutional neural network (CNN), generative adversarial network (GAN), and self-attention (SA) algorithms. The forward propagation of the image enhancement model refers to the process of gradually obtaining the enhanced sample image from the input layer to the output layer through the calculation of each layer of the pre-constructed image enhancement model after inputting the sample image in the training sample set into the pre-constructed image enhancement model. In the forward propagation process, each layer will calculate according to the input data and the weight parameters of the layer, and pass the calculation result to the next layer. For example, in the convolution layer, the convolution kernel will perform convolution operation on the input image to extract local features of the image.

[0130] For example, the server can input the training sample set into a pre-constructed image enhancement model, the pre-constructed image enhancement model including two branches, a pixel-level branch and a multi-scale branch, and based on forward transmission of the pre-constructed image enhancement model, each sample image in the training sample set is processed through the pixel-level branch and the multi-scale branch respectively to obtain an enhanced sample image. The processing of the sample image by the pixel-level branch and the multi-scale branch can refer to the description of the above embodiments, and will not be repeated here.

[0131] At step 306, a gray prior loss value of the enhanced sample image is determined based on the enhanced sample image and a prior loss function, and the pre-constructed image enhancement model is trained based on the gray prior loss value until the training is completed, thereby obtaining a pre-trained image enhancement model.

[0132] The prior loss function refers to a function for measuring the difference between the output result of the image enhancement model and the prior knowledge. In specific implementation, the prior loss function can be determined based on the statistical parameters of the gray channel of the sample image, and is used to reflect the deviation between the enhanced sample image and the ideal distribution in the gray distribution. By minimizing the prior loss function, the image enhancement model can be guided to learn an enhancement manner that is more consistent with the prior knowledge, thereby improving the performance and enhancement effect of the model. When determining the prior loss function, the statistical parameters such as the mean and variance of the gray of the sample image under different illumination conditions can be counted, and the prior loss function can be constructed according to the statistical parameters; for example, when the mean gray of the enhanced image needs to be in a preset range, the prior loss function can be designed in the form of the square of the difference between the mean gray of the enhanced image and the center value of the preset range.

[0133] The gray prior loss value refers to a value calculated by inputting the enhanced sample image into the prior loss function, and the gray prior loss value can be used to quantify the consistency of the enhanced sample image with the prior knowledge in terms of gray. The smaller the gray prior loss value, the more consistent the enhanced image is with the prior requirements. In specific implementation, the enhanced sample image can be subjected to gray processing to obtain a gray image; then the gray channel data of the gray image is extracted and input into the prior loss function for calculation to obtain the gray prior loss value.

[0134] The model training refers to adjusting the weight parameters of the pre-constructed image enhancement model through a back propagation algorithm according to the calculated gray prior loss value, so that the output of the model is continuously close to the enhanced sample image of the prior knowledge. In the specific implementation, in each iteration, firstly, the forward transmission is performed to obtain the enhanced sample image and the gray prior loss value; then the gradient of each layer weight parameter is calculated through the back propagation algorithm; and finally, the weight parameters are updated so that the gray prior loss value gradually decreases. The training end condition can be that the preset iteration number is reached, the loss value converges to a certain range, or the performance of the model on the validation set no longer improves, etc. When the training end condition is met, the training is stopped, and the pre-trained image enhancement model is obtained.

[0135] For example, the server can count the statistical parameters such as the mean and variance of the gray value of the sample image under different light conditions, and construct a prior loss function according to the statistical parameters. Then, the server can determine the gray prior loss value of the enhanced sample image according to the enhanced sample image and the prior loss function, for example, the server can perform gray processing on the enhanced sample image to obtain a gray image, and extract the gray channel data of the gray image, and input it into the prior loss function for calculation to obtain the gray prior loss value. The server can train the pre-constructed image enhancement model based on the gray prior loss value until the end, and obtain the pre-trained image enhancement model, for example, the server calculates the gradient of each layer weight parameter through the back propagation algorithm; finally, the weight parameters are updated so that the gray prior loss value gradually decreases, until the gray prior loss value converges, the training is ended, and the pre-trained image enhancement model is obtained.

[0136] In this embodiment, the training sample set is constructed by collecting different light sample images, and the sample images are in the form of unlabeled images, so that the model can learn the image features under various light conditions. At the same time, the enhanced sample is obtained through the forward transmission of the model, and the gray prior loss value is calculated by combining the prior loss function, so as to train the model with this loss value, which can continuously optimize the model parameters and enhance the generalization ability of the model in various real scenes when performing image enhancement, so as to output clear and natural enhanced images in different light scenes, and meet various practical application requirements.

[0137] In one application example, as shown in Figure 4 The image enhancement method includes two branches when performing image enhancement, namely a pixel-level branch and a multi-scale branch. In the specific implementation, the server can input the original image into the image enhancement model, and the image enhancement model processes the original image by using the pixel-level branch and the multi-scale branch respectively.

[0138] In the pixel-level branch, the server can extract pixel features of each pixel in the original image by performing convolution processing on the original image, then separate the original image according to a predetermined region size or pixel, input each pixel after separation into a mapping estimation network, obtain pixel enhancement parameters corresponding to each pixel feature, and then perform iterative enhancement on each pixel feature using the pixel enhancement parameters to obtain enhanced pixels, and combine the enhanced pixels to form a first intermediate image.

[0139] In the multi-scale branch, the server can extract features of the original image through different scale convolution kernels to obtain local feature maps corresponding to each feature scale, and then extract global dependencies of the local feature maps using a self-attention mechanism to obtain feature data corresponding to each feature scale. Then, the server performs upsampling and fusion on each feature data to obtain a second intermediate image.

[0140] Finally, the server can fuse the original image, the first intermediate image and the second intermediate image to obtain an enhanced image.

[0141] It should be understood that, although each step in the flowchart involved in each of the above-described embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each of the above-described embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by the combination are within the scope of protection of the present application.

[0142] Based on the same inventive concept, the embodiments of the present application also provide an image enhancement device for implementing the above-mentioned image enhancement method. The implementation scheme of the problem solving provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more image enhancement device embodiments provided below can refer to the limitations of the image enhancement method in the above text, which will not be repeated here.

[0143] In one exemplary embodiment, as Figure 5As shown, an image enhancement device is provided, comprising: an image acquisition module 502, a pixel enhancement module 504, a feature enhancement module 506, a first fusion module 508 and a second fusion module 510, wherein:

[0144] The image acquisition module 502 is configured to acquire an original image to be enhanced.

[0145] The pixel enhancement module 504 is configured to enhance each pixel point in the original image based on a respective pixel enhancement parameter corresponding to the pixel point by using a pre-trained image enhancement model, to obtain an enhanced pixel point, and to determine a first intermediate image based on each enhanced pixel point.

[0146] The feature enhancement module 506 is configured to perform feature extraction on the original image at different scales to obtain a plurality of feature images having different feature scales, and to fuse each feature image to obtain a second intermediate image.

[0147] The first fusion module 508 is configured to perform first fusion based on the original image, the first intermediate image and the second intermediate image to obtain a fused image.

[0148] The second fusion module 510 is configured to perform second fusion on the fused image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image.

[0149] In an optional embodiment, the pixel enhancement module 504 is further configured to perform feature extraction on the original image at at least two feature scales to obtain feature data of the original image at different feature scales, to determine a feature sequence based on a respective feature scale corresponding to each feature data, to align each feature data in descending order of the respective feature scale corresponding to the feature data to obtain aligned feature data, and to fuse the aligned feature data in descending order of the feature scale to obtain the second intermediate image.

[0150] In an optional embodiment, the pixel enhancement module 504 is further configured to determine at least two feature scales for feature extraction on the original image, to perform local feature extraction on the original image at each feature scale according to the respective feature scale to obtain a local feature map of the original image at the respective feature scale, to extract a correlation between each pixel point in the local feature map based on a self-attention mechanism, and to determine a global correlation parameter corresponding to each pixel point based on the correlation, and to aggregate features of each pixel point based on the global correlation parameter corresponding to the pixel point to obtain feature data of the original image at the respective feature scale.

[0151] In an optional embodiment, the feature enhancement module 506 is further configured to extract features of each pixel in the original image to obtain pixel features corresponding to each pixel respectively; for each pixel, determine a pixel enhancement parameter corresponding to the pixel, and enhance the pixel according to the pixel enhancement parameter to obtain an enhanced pixel.

[0152] In an optional embodiment, the feature enhancement module 506 is further configured to determine an initial pixel feature of the pixel and an initial enhancement parameter for enhancing the initial pixel feature; determine an intermediate pixel feature based on the initial pixel feature and the initial enhancement parameter; in a case where the intermediate pixel feature does not satisfy a preset enhancement condition, determine an intermediate enhancement parameter for enhancing the intermediate pixel feature, and enhance the intermediate pixel feature based on the intermediate enhancement parameter until the intermediate pixel feature satisfies the preset enhancement condition to obtain an enhanced pixel feature; and determine an enhanced pixel based on the enhanced pixel feature.

[0153] In an optional embodiment, the second fusion module 510 is further configured to perform weighted aggregation on the original image, the first intermediate image and the second intermediate image according to respective fusion weights of the original image, the first intermediate image and the second intermediate image to obtain an aggregated image; and perform second fusion on the fusion image and the aggregated image to obtain an enhanced image for the original image.

[0154] In an optional embodiment, the image enhancement apparatus further includes a model training module configured to collect sample images under different light intensities to form a training sample set; the sample images are unlabeled images; input the training sample set into a pre-constructed image enhancement model, obtain enhanced sample images based on forward transmission of the pre-constructed image enhancement model; determine a gray prior loss value of the enhanced sample images based on the enhanced sample images and a prior loss function, train the pre-constructed image enhancement model based on the gray prior loss value until the training ends to obtain a pre-trained image enhancement model; and the prior loss function is determined based on statistical parameters of gray channels in the sample images.

[0155] The above-mentioned modules in the image enhancement apparatus can be realized by software, hardware or a combination thereof. The above-mentioned modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the above-mentioned modules.

[0156] In an exemplary embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store pre-trained image enhancement model, pixel enhancement parameter, fusion weight and other data. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the terminal outside through the network connection. The computer program is executed by the processor to realize an image enhancement method.

[0157] Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0158] In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to realize the image enhancement method of the above embodiment.

[0159] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by the processor to realize the image enhancement method of the above embodiment.

[0160] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by the processor to realize the image enhancement method of the above embodiment.

[0161] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0162] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0163] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0164] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. An image enhancement method characterized by, The method comprises: obtaining an original image to be enhanced; enhancing each pixel point in the original image based on a pixel enhancement parameter corresponding to the pixel point, to obtain an enhanced pixel point, and determining a first intermediate image based on each enhanced pixel point; performing feature extraction on the original image in different scales to obtain a plurality of feature images with different feature scales, and fusing each feature image to obtain a second intermediate image; performing first fusion based on the original image, the first intermediate image and the second intermediate image to obtain a fused image; performing second fusion on the fused image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image.

2. The method of claim 1, wherein, The second fusion of the fused image, the original image, the first intermediate image and the second intermediate image to obtain an enhanced image for the original image comprises: performing weighted aggregation on the original image, the first intermediate image and the second intermediate image according to the fusion weights corresponding to the original image, the first intermediate image and the second intermediate image to obtain an aggregated image; performing second fusion on the fused image and the aggregated image to obtain an enhanced image for the original image.

3. The method of claim 1, wherein, The feature extraction on the original image in different scales to obtain a plurality of feature images with different feature scales, and the fusion of each feature image to obtain a second intermediate image comprises: performing feature extraction on the original image according to at least two feature scales to obtain feature data of the original image under different feature scales; determining a feature sequence according to the feature scales corresponding to each feature data; aligning each feature data according to the feature scales corresponding to each feature data from large to small to obtain aligned feature data; fusing the aligned feature data according to the feature scales from large to small in turn to obtain a second intermediate image.

4. The method of claim 3, wherein, The feature extraction on the original image according to at least two feature scales to obtain feature data of the original image under different feature scales comprises: determining at least two feature scales for feature extraction on the original image; for each feature scale, performing local feature extraction on the original image according to the feature scale to obtain a local feature map of the original image under the feature scale; extracting the association relationship between each pixel point in the local feature map based on a self-attention mechanism, and determining a global association parameter corresponding to each pixel point according to the association relationship; performing feature aggregation on each pixel point based on the global association parameter corresponding to each pixel point to obtain feature data of the original image under the feature scale.

5. The method of claim 1, wherein, The enhancement of each pixel point in the original image based on a pixel enhancement parameter corresponding to the pixel point to obtain an enhanced pixel point, and the determination of a first intermediate image based on each enhanced pixel point, comprises: Features of each pixel point in the original image are extracted to obtain respective pixel features corresponding to each of the pixel points; For each of the pixel points, a pixel enhancement parameter for the pixel point is determined, and the pixel point is enhanced according to the pixel enhancement parameter to obtain an enhanced pixel point.

6. The method of claim 5, wherein, The determination of the pixel enhancement parameter for the pixel point and the enhancement of the pixel point according to the pixel enhancement parameter to obtain an enhanced pixel point include: An initial pixel feature of the pixel point and an initial enhancement parameter for enhancing the initial pixel feature are determined; An intermediate pixel feature is determined based on the initial pixel feature and the initial enhancement parameter; In a case where the intermediate pixel feature does not satisfy a preset enhancement condition, an intermediate enhancement parameter for enhancing the intermediate pixel feature is determined, and the intermediate pixel feature is enhanced based on the intermediate enhancement parameter until the intermediate pixel feature satisfies the preset enhancement condition to obtain an enhanced pixel feature; An enhanced pixel point is determined based on the enhanced pixel feature.

7. The method according to any one of claims 1 to 6, characterized in that, The image enhancement model is obtained by training based on the following method: Sample images under different light intensities are collected to form a training sample set; the sample images are unlabeled images; The training sample set is input into a pre-constructed image enhancement model, and enhanced sample images are obtained based on forward transmission of the pre-constructed image enhancement model; A gray prior loss value of the enhanced sample images is determined based on the enhanced sample images and a prior loss function, the pre-constructed image enhancement model is trained based on the gray prior loss value until the training is completed, and a pre-trained image enhancement model is obtained; the prior loss function is determined based on statistical parameters of a gray channel in the sample images.

8. An image enhancement device, characterized by The apparatus includes: An image acquisition module configured to acquire an original image to be enhanced; A pixel enhancement module configured to enhance each of pixel points in the original image based on respective pixel enhancement parameters corresponding to the pixel points by using a pre-trained image enhancement model to obtain enhanced pixel points, and determine a first intermediate image based on the enhanced pixel points; A feature enhancement module configured to perform feature extraction of different scales on the original image to obtain a plurality of feature images with different feature scales, and fuse the feature images to obtain a second intermediate image; A first fusion module configured to perform first fusion based on the original image, the first intermediate image, and the second intermediate image to obtain a fused image; A second fusion module configured to perform second fusion on the fused image, the original image, the first intermediate image, and the second intermediate image to obtain an enhanced image of the original image. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.