Training method of image processing model, image enhancement method and electronic equipment
By using the focus loss function and the total loss function of the color loss function during the training process of the image processing model, the problem of recovery of medium and high-frequency regions and chromatic aberration of image enhancement is solved, and better image enhancement effect and processing efficiency are achieved.
Patent Information
- Application Number
- CN202510294044.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
Smart Images

Figure CN120219880A_ABST
Abstract
Description
Background Art
[0002] In order to meet storage and transmission requirements, images are compressed in some scenarios. Since problems such as block effect, blurring, and color distortion exist in the compressed images, before using the images subsequently, it is necessary to repair (image enhancement) the images to facilitate subsequent image display, analysis, and other processing.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a training method for an image processing model, an image enhancement method, and an electronic device, which are used to reduce the resource occupation of the image enhancement task, improve the processing efficiency of the image enhancement task, and optimize the image enhancement effect of the image enhancement task.
[0005] According to the first aspect of the embodiments of the present disclosure, a training method for an image processing model is provided, including: obtaining a target image processing model and training data of the target image processing model, where the training data includes an image to be enhanced and an original image corresponding to the image to be enhanced; initializing parameters of the target image processing model; forming a total loss function according to a focal loss function and a color loss function; inputting the image to be enhanced into the target image processing model to obtain an enhanced image; calculating a difference between the original image and the enhanced image through the total loss function, and updating the parameters of the target image processing model according to the difference.
[0006] According to the second aspect of the embodiments of the present disclosure, an image enhancement method is provided, including: sampling an image to be enhanced through an input module of a target image processing model to obtain a first feature map; processing the first feature map through a reparameterization module of the target image processing model to form a second feature map, where the reparameterization module includes at least one reparameterization unit, and the reparameterization unit is trained through a multi-scale feature extraction method; performing deformable convolution processing on the second feature map through a deformable convolution module of the target image processing model to form a third feature map; forming an enhanced image corresponding to the image to be enhanced according to the third feature map and the image to be enhanced through an output module of the target image processing model.
[0007] According to the third aspect of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, where the processor is configured to execute the method as described in any one of the above based on instructions stored in the memory.
[0008] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having a program stored thereon, which when executed by a processor, implements the image enhancement method as described in any one of the above.
[0009] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, characterized in that when the computer program is executed by a processor, it implements the steps of the method described in any one of the above.
[0010] In the embodiments of the present disclosure, by using the total loss function composed of the focal loss function and the color loss function during the training process of the target image processing model, the weights can be adaptively adjusted during the training process, effectively processing high-frequency and low-frequency regions, and improving the detail performance of complex regions; the color loss function can ensure the color consistency and authenticity of the enhanced image, avoiding color differences. Therefore, the embodiments of the present disclosure can solve the problems such as the difficulty in restoring high-frequency regions after image enhancement and obvious color differences in the enhanced image in the related art, and significantly improve the image enhancement effect of the image enhancement model.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0013] Figure 1 is a schematic structural diagram of the target image processing model 100 of the embodiments of the present disclosure.
[0014] Figure 2A and Figure 2B are respectively schematic diagrams of pixel inverse shuffling and pixel shuffling of the exemplary embodiments of the present disclosure.
[0015] Figure 3 is a schematic diagram of the hierarchical setting of the model 100 in the exemplary embodiments of the present disclosure.
[0016] Figure 4 is a schematic structural diagram of the reparameterization unit in the exemplary embodiments of the present disclosure.
[0017] Figure 5 is a flowchart of the image enhancement method in the exemplary embodiments of the present disclosure.
[0018] Figure 6It is a flowchart of a method for training an image processing model in an exemplary embodiment of the present disclosure.
[0019] Figure 7 It is a flowchart of constructing a focal loss function in an exemplary embodiment of the present disclosure.
[0020] Figure 8 It is a flowchart of constructing a color loss function in an exemplary embodiment of the present disclosure.
[0021] Figure 9 It is a flowchart of a training process of a reparameterization module in an exemplary embodiment of the present disclosure.
[0022] Figure 10 It is a block diagram of an electronic device in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0023] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.
[0024] In addition, the accompanying drawings are only schematic illustrations of the present disclosure, and the same reference numerals in the drawings denote the same or similar parts, so repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0025] The following will describe the example embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0026] Existing image enhancement models usually adopt complex network structures and a large number of model parameters to improve feature extraction capabilities and enhancement effects. Therefore, the model has high computational complexity, large video memory occupancy, and slow inference speed. In addition, for the images processed by existing image enhancement models, there are problems such as loss and blurring of image details, and obvious color differences in the enhanced images, which urgently need to be optimized.
[0027] The image enhancement method and the model training method provided by the embodiments of the present disclosure can both be implemented by the target image processing model provided by the embodiments of the present disclosure. The target image processing model can be used to perform image enhancement tasks and output the corresponding enhanced image based on the input image to be enhanced. The image to be enhanced in the embodiments of the present disclosure can be either a compressed image or an image that needs to be enhanced generated in other scenarios, such as an image with blurred details and dull colors caused by low-light shooting, or a damaged image caused by reasons such as device aging and transmission failures. The embodiments of the present disclosure do not make specific limitations on the source and type of the image to be enhanced.
[0028] Next, the target image processing model provided by the embodiments of the present disclosure will be introduced first.
[0029] Figure 1 is a schematic structural diagram of the target image processing model 100 of the embodiments of the present disclosure.
[0030] Refer to Figure 1 , the target image processing model 100 provided by the embodiments of the present disclosure may include:
[0031] An input module 11 for sampling the image to be enhanced to obtain a first feature map;
[0032] A reparameterization module 12 for processing the first feature map to form a second feature map;
[0033] A deformable convolution module 13 for performing deformable convolution processing on the second feature map to form a third feature map;
[0034] An output module 14 for outputting the enhanced image corresponding to the image to be enhanced according to the third feature map and the image to be processed.
[0035] Among them, the reparameterization module 12 includes at least one reparameterization unit, and the reparameterization unit is trained by a multi-scale feature extraction method. Specifically, the multi-scale feature extraction method is implemented through a multi-branch structure, using parallel convolution kernels of different sizes, such as 1×1 convolution kernels, 3×3 convolution kernels, etc., to capture feature information of different scales at the same time, and finally combining the feature information of different sizes extracted by each branch to obtain the feature information of the image to be enhanced. This multi-scale feature extraction method enables the model to understand the image content from multiple perspectives, and has a more comprehensive grasp of both tiny details and overall structural features in complex image scenes, greatly enriching the feature expression.
[0036] The deformable convolution module 13 is used to adaptively adjust the sampling positions according to the image content, improving the accuracy of the model in capturing features of complex shapes and irregular objects. Traditional convolution samples on a fixed grid, making it difficult to adapt to changes in the shape, position, and scale of objects in the image. In contrast, deformable convolution can overcome this problem. When processing images containing objects with various poses and shapes, such as identifying pedestrians and vehicles at different angles and positions in an autonomous driving scenario, the deformable convolution module 13 can dynamically adjust the sampling positions of the convolution kernel according to the pose of the pedestrian and the shape of the vehicle, more accurately extracting key features such as the edges and contours of the target object, and avoiding feature loss or inaccuracy caused by fixed sampling positions.
[0037] Deformable convolution processing means that by introducing learnable offsets, the sampling points of the convolution kernel can "intelligently move" on the image, thereby enhancing the model's perception ability of local details. When processing image regions with rich textures, such as the textures of fabrics and the branches and leaves of trees, deformable convolution can automatically focus the sampling points on the key parts of the texture, obtaining more refined texture information, and more sensitively capturing details such as the orientation and density changes of the texture, making the details of the enhanced image in these regions clearer and more realistic, and significantly improving the visual quality of the image.
[0038] In a compressed image, different regions are affected by compression differently. Some regions may be severely blurred, while others are relatively less affected. Deformable convolution processing can flexibly adjust the sampling strategy according to the distortion situation of each region - for regions with severe distortion, increase the offset range of the sampling points to more comprehensively explore the surrounding potential information and restore the lost details as much as possible; for regions with less distortion, maintain a relatively small offset to avoid introducing noise due to excessive adjustment. Therefore, for an image enhancement model, the deformable convolution module 13 can better balance the enhancement effects of different regions of the image as a whole, enabling the enhanced image to meet high-quality standards in all parts.
[0039] In the model 100 of the embodiments of the present disclosure, a deformable convolution module 13 is set after the reparameterization module 12, enabling the deformable convolution module 13 to work in cooperation with the reparameterization module 12, which can greatly enhance the image enhancement ability of the model 100. Since the deformable convolution module 13 is located after the reparameterization module 12, it can utilize these multi-scale features optimized by the reparameterization module 12 to more accurately perform adaptive sampling and adjustment of the features. For example, the reparameterization module 12 can learn various feature patterns of the image, providing a rich feature information basis for the deformable convolution module 13; based on these multi-scale features, the deformable convolution module 13 performs adaptive sampling and convolution operations, further exploring the correlations and details between features, and further learning the spatial deformation information of the features to achieve more in-depth feature extraction and fusion. The combination of the two can better adapt to various shape and pose changes of objects in the image, enhance the feature learning ability of the model for complex scenes, enabling the model 100 to perform more excellently in complex image enhancement tasks. Whether it is for repairing tiny details or optimizing the overall image structure, it can reach a higher level, achieving more accurate and efficient image processing effects. Moreover, to a certain extent, the reparameterization module 12 can screen and integrate the input features, removing some redundant information and reducing the computational load of the subsequent deformable convolution module 13.
[0040] The image enhancement process in related technologies usually consumes a large amount of computing power and takes a long processing time, with obvious efficiency defects when processing image data such as high-resolution videos. Therefore, in some embodiments of the present disclosure, in order to improve the processing speed of the model and reduce the computational load, an anti-shuffle function and a shuffle function are respectively set in the input module 11 and the output module 14 of the model 100.
[0041] The inventors of the present disclosure analyzed that since there are a large number of convolution operations in the image enhancement process, the image processing speed of the model is mainly affected by the floating-point operation count (FLOPs) of the convolution operations. The higher the floating-point operation count, the slower the image processing speed.
[0042] The formula for calculating the floating-point operation count (FLOPs) of the convolution operation is:
[0043] FLOPs = (C i ×K 2 +C i ×K 2 -1)×H×W×C o (1)
[0044] Where C i represents the number of input channels, C o represents the number of output channels, H and W represent the height and width of the input feature map, and K represents the size of the convolution kernel. It should be noted that this formula is applicable to the bias-free convolution layer. Based on this formula, if the FLOPs are fixed, increasing H and W will result in Ci and C o However, in practical applications, for most computing devices, when the number of channels is below a certain threshold (such as 16 or 32), simply reducing the channel size does not significantly reduce the running time. In addition, the smaller channel size also limits the flexibility of network design and its overall performance.
[0045] To solve this problem, the embodiments of the present disclosure introduce pixel unshuffle operation and pixel shuffle operation: perform pixel shuffle on the input image to reduce the spatial resolution and correspondingly increase the number of channels, then mainly perform calculations in the feature space with a lower resolution, and finally restore the original resolution or even increase the resolution through pixel unshuffle. This operation can effectively reduce the computational overhead. Pixel shuffle and pixel unshuffle are inverse operations of each other. Pixel shuffle rearranges the input elements to generate data with a lower resolution but an increased number of channels. Taking N-fold pixel unshuffle as an example, the number of input channels becomes C×N 2 , and the resolution of the input image drops to According to the above formula (1), under the condition of maintaining the number of output channels unchanged, the number of floating-point operations FLOPs of the first convolutional layer remains unchanged after applying pixel shuffle, but starting from the second convolution, due to the resolution decrease, FLOPs will be reduced by N 2 times.
[0046] Thus, without sacrificing the model enhancement effect, the computational complexity can be effectively reduced, enabling the model to run on resource-constrained devices such as mobile devices or embedded systems. This not only broadens the application scenarios of the model but also provides the possibility for tasks with high requirements for computing resources and processing speed such as real-time image enhancement.
[0047] Correspondingly, when setting the input module 21 to perform pixel unshuffle operation, set the output module 24 to perform pixel shuffle operation to upsample the third feature map to restore or even increase the resolution to form an enhanced image corresponding to the image to be enhanced.
[0048] Figure 2A and Figure 2B are respectively schematic diagrams of pixel unshuffle and pixel shuffle of the exemplary embodiments of the present disclosure.
[0049] Refer to Figure 2A , the pixel unshuffle operation refers to downsampling and channel transformation of the image to be enhanced to reduce the resolution of the image to be enhanced and increase the number of channels of the image to be enhanced.
[0050] Assume that the size of the image to be enhanced is 8×8 (data is simplified for the convenience of example), and it has three channels: R, G, and B. When the pixel deshuffle operation is performed through the input module 21 in step S1, assuming that a 2x pixel deshuffle operation is adopted (i.e., N=2), the resolution of the original 8×8 image will be reduced to (8÷2)×(8÷2)=4×4. In terms of channels, the original 3 channels will become 3×22=12 channels. During this change, the arrangement of image data has changed significantly. Each pixel is no longer a simple combination of RGB colors, but is redistributed to more channels. For example, the RGB value of the pixel originally located at the position (3,5) will be dispersed into the new 12 channels according to specific rules.
[0051] This rearrangement follows the algorithmic logic of pixel shuffling, so that the spatial information of the image is redistributed in the channel dimension. As a result, the subsequent re-parameter module 12 can perform more efficient calculations on feature maps with lower resolution but higher number of channels. As the resolution decreases, the number of pixels involved in each convolution operation decreases, and the amount of calculation decreases accordingly; and the increase in the number of channels allows the model to process more feature information in parallel on different channels. For example, different convolution kernel branches in the re-parameter module 12 can operate on these 12 channels separately to extract features from different angles.
[0052] After being processed by the re-parameter module 12, the obtained second feature map will carry rich multi-scale feature information. When these feature information enters the deformation convolution module 13 again, the deformation convolution module 13 can adaptively adjust the sampling position according to the image content based on the second feature map with a 4×4 resolution and 12 channels. For example, for the edge area in the image that was originally blurred at an 8×8 resolution, on the feature map with a 4×4 resolution and 12 channels, the deformation convolution module 13 can use the learned offset to more accurately sample the pixel position near the edge, thereby more effectively extracting edge features and providing more accurate data support for the subsequent restoration of image details.
[0053] In the embodiments of the present disclosure, by setting the input module 21 to perform pixel inverse shuffling operations, the computational efficiency and feature extraction ability of the image enhancement process can be improved simultaneously. After the pixel inverse shuffling operation reduces the image resolution and increases the number of channels, the subsequent reparameterization module 12 and deformable convolution module 13 can process the feature maps more efficiently. This low-resolution, high-channel-number feature map structure significantly reduces the computational amount of convolution operations, while providing the model with rich channel dimension information, enabling subsequent modules to fully explore the potential connections between features and thus extract more comprehensive and in-depth features. When facing complex image content, such as when processing images containing multiple textures and objects, different channels can focus on different types of features respectively (some channels specifically capture object edge information, while others focus on texture details). These features can be further strengthened and optimized after undergoing multi-scale fusion by the reparameterization module 12 and adaptive sampling processing by the deformable convolution module 13. Good image enhancement effects can be achieved for JPEG-compressed images, images with reduced quality due to insufficient lighting, or images with noise interference.
[0054] In addition, since high-frequency information is usually distributed in the detailed parts of the image, in the high-channel-number feature map, high-frequency information can be represented and processed more meticulously among different channels. This enables the model to better learn the high-frequency information of the image in subsequent processing, more accurately restore and enhance the details of the image, and reduce the problem of detail loss caused by improper processing.
[0055] Reference Figure 2B , when setting the input module 21 to perform pixel inverse shuffling operations to reduce the resolution of the image to be enhanced, a pixel shuffling operation can be set in the output module 24 to perform upsampling and channel transformation on the third feature map, increasing the resolution so that the resolution of the enhanced image output by the model is not lower than that of the image to be enhanced.
[0056] Still introduced according to the above example, when the output module 24 receives the third feature map from the deformable convolution module 13, if the resolution of the enhanced image is set to be equal to that of the image to be enhanced, since the image resolution has been reduced to 4×4 and the number of channels has become 12 after the previous pixel inverse shuffling operation, a 2-fold pixel shuffling operation is performed at this time (corresponding to the previous shuffling multiple).
[0057] During the pixel shuffling operation, the data of 12 channels are recombined according to specific rules to restore the resolution of the image. Specifically, the data of every 4 channels (since it is 2-fold shuffling, 2^2 = 4) are rearranged to generate a pixel value at a new spatial position. For example, specific 4 channels are selected from the 12 channels, and the data at the corresponding positions of these 4 channels are weighted or combined in other ways to obtain a new RGB pixel value (because ultimately it is necessary to restore to a three-channel RGB mode similar to the original image). By performing such an operation on each pixel at each position of the 4×4 resolution image, the resolution of the image is gradually restored to 8×8.
[0058] In this way, after the feature information processed by the reparameterization module 12 and the deformable convolution module 13 is restored in resolution through the pixel shuffling operation of the output module 24, it can be presented at a resolution not lower than that of the original image to be enhanced. Moreover, since the model has fully extracted and optimized the features of the image during the processing of the reparameterization module 12 and the deformable convolution module 13, the enhanced image after restoring the resolution not only has the same size as the original image but also has a significant improvement in quality. For example, the edges that were originally blurred due to JPEG compression become clearer and sharper after being processed by the previous modules and restored in resolution through shuffling; the lost detailed information is also restored to a certain extent, making the overall image look more real and natural. In practical application scenarios, this resolution restoration mechanism can ensure that the output enhanced image meets the user's requirements in terms of resolution whether it is used to process ordinary images or each frame of a video stream.
[0059] Figure 2B Taking the setting that the resolution of the enhanced image is equal to the resolution of the image to be enhanced as an example, in practical applications, it is also possible to set the resolution of the enhanced image to be greater than or equal to the resolution of the image to be enhanced. The upsampling and channel transformation logics are similar and will not be elaborated here.
[0060] Of course, it is also possible not to set the inverse shuffling function in the input module 11. Correspondingly, there is no need to set the shuffling function in the output module 14 to simplify the model structure and avoid potential errors introduced by the inverse shuffling and pixel shuffling operations.
[0061] Figure 3 It is a schematic diagram of the hierarchical setting of the model 100 in an exemplary embodiment of the present disclosure.
[0062] Reference Figure 3, in an exemplary embodiment, the input module 11 may include an inverse shuffling unit 111 and a first dynamic hybrid convolution layer 112, the reparameterization module 12 may include four sequentially connected reparameterization units 121 to 124, the deformable convolution module 13 may include a first deformable convolution unit 131, a second deformable convolution unit 132, and a second dynamic hybrid convolution layer 133, and the output module 14 may include a splicing layer 141, a general convolution layer 142, a shuffling layer 143, and a residual connection unit 144.
[0063] Among them, the inverse shuffling unit 111 is used to perform an inverse shuffling operation on the image to be enhanced of the input model, so as to reduce the resolution of the images processed by most other levels of the model, realize model lightweighting. For example, the resolution and number of channels of the above-mentioned 8×8 3-channel image to be enhanced are respectively converted to 4×4 and 12.
[0064] The first dynamic hybrid convolution layer 112 can be expressed as CON3×C for example, where CON is the abbreviation of "Convolution", C refers to channels, and 3 refers to the convolution kernel size of 3×3. That is, the input image is convolved using a 3×3 convolution kernel to extract the features of the input feature map and output a first feature map with C channels. The first dynamic hybrid convolution layer 112 is mainly used to increase the number of channels of the first feature map and the information of the first feature map, so that subsequent modules can better extract the information of the first feature map.
[0065] In the exemplary embodiment of the present disclosure, the number of intermediate feature channels (i.e., the number of channels of the first feature map and the second feature map) during the processing of the model 100 can be set to 32, that is, the above-mentioned C can be set to 32. This decision is based on the detailed test results of the inference speed for different numbers of channels. The experimental results show that when the number of channels is set to 32, the model 100 can not only retain a large channel capacity but also maintain a high operation speed. Therefore, considering performance optimization and resource occupancy comprehensively, 32 is considered to be a better number of channels. It should be noted that the numbers in the examples of the present disclosure are for convenience of understanding and are not the numbers in actual applications. Those skilled in the art can set the number of intermediate feature channels according to the actual hardware environment, and the present disclosure does not make special restrictions on this.
[0066] The reparameterization module 12 may include four sequentially connected reparameterization units 121 to 124, and the internal structures of each reparameterization unit may be the same. By setting four reparameterization units, a better processing effect can be achieved. Of course, the number of reparameterization units can be adjusted, and the present disclosure does not make special restrictions on this. The specific structure of the reparameterization unit is described in detail in the subsequent embodiments.
[0067] The deformable convolution module 13 may include a first deformable convolution unit 131, a second deformable convolution unit 132, and a second dynamic hybrid convolution layer 133. Among them, the first deformable convolution unit 231 may perform a convolution process on the second feature map for adaptively adjusting the sampling position based on the image content, capturing some relatively basic and local feature changes. The second deformable convolution unit 232, on the basis of the output of the first deformable convolution unit 231, may further perform deeper and more refined feature extraction on the features that have been adaptively adjusted once, capable of mining more complex and higher-level feature information. For example, when processing image data with a multi-scale and multi-level structure, the two deformable convolutions can process features at different scales or levels respectively. By setting two deformable convolution units and performing two deformable convolution operations on the second feature map, the model can have a stronger adaptability to various changes and differences in the image. For example, the first deformable convolution can enable the model to adapt to some common image deformation situations, and the second deformable convolution can further model some more special and complex deformations or feature changes, thereby improving the robustness of the model in different scenarios, such as when processing images with different perspectives, lighting conditions, or large changes in object morphology. In an exemplary embodiment, more deformable convolution units may also be set, and different deformable convolution units can be connected through a non-linear activation function, etc., further increasing the non-linear expression ability of the model and improving the performance of the model in various tasks such as object detection and image segmentation tasks, and more accurately locating and identifying targets.
[0068] The second dynamic hybrid convolution layer 133 is used to further perform convolution on the feature image that has undergone two deformable convolutions, so as to output a third feature map containing more information. In an exemplary embodiment, the second dynamic hybrid convolution layer 133 can also be denoted as CON3×C, and the convolution kernel size is 3×3. After two deformable convolutions, the feature image already contains a large amount of complex shape and structure information dynamically captured based on the image content. However, these features may vary in different scales and dimensions. The second dynamic hybrid convolution layer 133 can fuse these features with different characteristics by using a specific convolution kernel configuration (3×3 and settings related to the number of channels), and adjust and optimize them in terms of feature dimension, number of channels, and feature distribution, etc., to make them more suitable for the processing requirements of subsequent modules. Through this fusion operation, features from different sources and with different natures can complement each other and make up for each other's deficiencies. It should be emphasized that the convolution operation of this layer is not a simple repetition, but is based on the features extracted by the previous deformable convolution for strengthening and optimization, further exploring the potential relationships between features, enhancing the expression intensity of key features, and suppressing possible noise or irrelevant information. For example, when processing an image with compression artifacts, the two deformable convolutions initially locate the features of the artifact area and the normal image area. The second dynamic hybrid convolution layer can adjust these features, making the features of the artifact area more prominent, so that the subsequent model can better identify and remove the artifacts, while making the features of the normal image area more stable and accurate, thereby improving the quality of the entire image.
[0069] The output module 14 may include a splicing layer 141, a general convolution layer 142, a shuffling layer 143, and a residual connection unit 144. The splicing layer 141 is used to splice the feature maps output by different steps in the channel dimension. For example, Figure 3 in the figure, the feature maps output by the first reparameterization unit, the fourth reparameterization unit, the second deformable convolution unit, and the second dynamic hybrid convolution layer (all with the same size) are stacked in the dimension. Assuming that the number of channels of each of these four feature maps is C, then the number of channels of the feature map after channel splicing output by the splicing layer 141 is 4C, so as to better fuse the features extracted in each stage and facilitate subsequent processing. The specific source of the feature maps spliced by the splicing layer 141 can be adjusted by those skilled in the art according to the image enhancement effect. The figure is only an example.
[0070] The general convolutional layer 142 is used to perform convolution on the feature map with a high number of channels formed by channel concatenation (the convolution kernel size is, for example, 3×3), and make the output feature map meet the number of channels required for the subsequent shuffling unit 143. For example, if the number of channels of the first feature map output by the first dynamic hybrid convolutional layer 122 is 32, then the number of channels of the second feature map and the third feature map are both 32. However, the number of channels of the input feature map of the shuffling unit needs to be the same as the number of channels of the output feature map of the inverse shuffling unit, which is 12. Therefore, the general convolutional layer 142 needs to perform a convolution operation on the feature map with 4×32 = 128 channels from the concatenation layer and output a feature map with 12 channels.
[0071] The shuffling unit 143 is used to re - output the 12 - channel feature map as a 3 - channel feature map through channel transformation and resolution transformation, and the resolution is not lower than the original resolution of the image to be enhanced.
[0072] The residual connection unit 144 is used to fuse the feature map output by the shuffling unit 143 with the image to be enhanced to enhance the image to be enhanced.
[0073] The re - parameterization units in the re - parameterization module 12 are introduced in detail below. In the exemplary embodiment set, the re - parameterization unit includes a convolutional layer and an attention layer, and the attention layer is used to process the feature image output by the convolutional layer through an attention mechanism.
[0074] Figure 4 It is a schematic structural diagram of the re - parameterization unit in an exemplary embodiment of the present disclosure.
[0075] Refer to Figure 4 , a re - parameterization unit may include a first convolutional layer 1201, a second convolutional layer 1202, a third dynamic hybrid convolutional layer 1203, a residual connection layer 1204, and an attention layer 1205. Among them, the first convolutional layer 1201 and the second convolutional layer 1202 are trained through a multi - branch and multi - scale feature extraction method, and the training method is described in detail in the subsequent embodiments. After training, the first convolutional layer 1201 and the second convolutional layer 1202 behave as convolutional layers, with a size of, for example, 3×3, and the weights of each position of the convolution kernel are formed through the training process. The third dynamic hybrid convolutional layer 1203 is, for example, CON3×C, and is used to further extract and fuse the output data of the two convolutional layers trained through the multi - scale feature extraction method. The training process is described in detail in the subsequent embodiments. After training, it behaves as a convolutional layer with a fixed size, for example, 3×3. The residual connection layer 1204 is used to perform local residual connection on the feature map processed by the three convolutional layers and the feature map not processed by the three convolutional layers to further fuse the features. The specific process of the local residual connection can be referred to in the subsequent embodiments.
[0076] In the embodiments of the present disclosure, the reparameterization unit integrates the attention mechanism by setting the attention layer 1205 to further enhance the discriminability and relevance of features. The attention mechanism allows the model to dynamically focus on important regions and suppress irrelevant background noise, thereby improving the accuracy of feature representation. Specifically, the attention layer 1205 can weight the importance of different feature maps during the training process, assign weights to different features according to the image content, highlight important features, enabling the network to focus on key information in complex scenarios, and significantly improving the robustness and generalization ability of the model.
[0077] Specifically, the attention layer 1205 is used to find regions in the fused feature map that have unique features and may play a key role in subsequent tasks. To achieve this distinction of focusing on key points, the attention layer 1205 evaluates each feature region, measuring factors such as the uniqueness of each region's features and the relevance to the target task. For those feature regions considered important, higher weights are assigned; while for those less important regions that may belong to background noise, lower weights are assigned. In subsequent processing, the feature regions with high weights will occupy a more important position in the overall feature representation. After the weighted processing of the attention layer, the important features in the feature map will be more prominently shown, and the key features that may have been masked by background noise or not fully emphasized before can now play a greater role in the entire feature system.
[0078] When these features processed by the attention mechanism, i.e., the second feature map, are passed to the subsequent deformable convolution module 13, the deformable convolution module 13 can analyze and judge more accurately based on these prominent key features. Thus, in the image classification task, the model can more precisely identify the object categories in the image; in the object detection task, it can more accurately locate the position and boundary of the target object. In addition, due to the attention mechanism's ability to suppress irrelevant background noise, the model can work more stably when facing complex and variable scenarios. Even when the background in the image is very complex and there are many interference factors, the model can still focus on key information, make accurate decisions, thereby greatly enhancing the robustness and generalization ability of the model. This enables the model to not only perform well on the training data but also have excellent performance when facing unseen new data, broadening the application scope and practicality of the model.
[0079] Exemplarily, the attention layer 1205 may include, for example Figure 4The structure shown, namely, a 1×1 convolutional layer 12051, a 3×3 convolutional layer 12052, a max pooling layer 12053, a 3×3 convolutional layer 12054, an interpolation layer 12055, a residual connection layer 12056, a 1×1 convolutional layer 12057, etc., connected in sequence. The output of the 1×1 convolutional layer 12051 is also transmitted to the residual connection layer 12056 after passing through a 1×1 convolutional layer 12058. The feature map output by the 1×1 convolutional layer 12057 and the input feature map of the attention layer 1205 are multiplied through a multiplication layer 12059 to form the finally output feature map. The settings of each layer of the attention mechanism can be adjusted by those skilled in the art according to the actual situation, which will not be elaborated here. It should be noted that the hierarchical settings of the attention mechanism are the same in both the training stage and the inference stage.
[0080] By setting a local residual connection layer 1204 and an attention layer 1205 in the reparameterization unit, it helps to stabilize the propagation of features, enabling the deformable convolution module 13 to operate on a more stable feature basis, improving the stability and efficiency of the calculation.
[0081] Due to its structural settings, the image processing model 100 provided by the present disclosure itself has extremely strong image enhancement capabilities. For images with a resolution of 1920x1080, testing was carried out on a Tesla V100 graphics card, and the results showed that 30 pictures can be processed per second, that is, high-definition video at 30 frames per second can be processed in real time.
[0082] Next, an image enhancement method implemented through the target image processing model 100 will be introduced. This image enhancement method can be realized through the inference process of the trained target image processing model 100.
[0083] Figure 5 It is a flowchart of the image enhancement method in an exemplary embodiment of the present disclosure.
[0084] Refer to Figure 5 , the image enhancement method 500 may include:
[0085] Step S1, sampling the image to be enhanced through the input module of the target image processing model to obtain a first feature map;
[0086] Step S2, processing the first feature map through the reparameterization module of the target image processing model to form a second feature map. The reparameterization module includes at least one reparameterization unit, and the reparameterization unit is trained through a multi-scale feature extraction method;
[0087] Step S3, performing deformable convolution processing on the second feature map through the deformable convolution module of the target image processing model to form a third feature map;
[0088] Step S4: The output module of the target image processing model forms an enhanced image corresponding to the image to be enhanced according to the third feature map and the image to be enhanced.
[0089] In the embodiments of the present disclosure, by setting a reparameterization module in the target image processing model to process images, the feature extraction ability of the model is enhanced, and the image enhancement effect is greatly optimized; by introducing a deformable convolution module after the reparameterization module, the ability of the target image processing model to handle complex texture shapes and different distortion degrees in different regions is enhanced, and the ability of the model to extract high-frequency information is enhanced, so that the target image processing model can have better image enhancement ability, achieving a better balance between image restoration quality and inference speed, and improving the processing efficiency and processing quality of the image enhancement task at the same time. In addition, the training method for the target image processing model proposed in the embodiments of the present disclosure is also an important guarantee for enabling the target image processing model to have better image enhancement ability.
[0090] In step S1, the input module of the target image processing model samples the image to be enhanced to obtain a first feature map.
[0091] In an exemplary embodiment, in order to improve the image processing speed, the input module 21 can be set to perform a pixel inverse shuffle operation, and in step S1, the image to be enhanced is downsampled and channel-transformed to reduce the resolution of the image to be enhanced and increase the number of channels of the image to be enhanced. Correspondingly, when the input module 21 is set to perform a pixel inverse shuffle operation, a pixel shuffle operation is set in step S4 to upsample the feature map to restore or even increase the resolution to form an enhanced image corresponding to the image to be enhanced. For the specific processes of the inverse shuffle operation and the shuffle operation, please refer to the foregoing embodiments and will not be elaborated herein.
[0092] In some embodiments, it is also possible to set the input module 21 not to perform a pixel inverse shuffle operation on the image to be enhanced in step S1, and only rely on the reparameterization module 12 and the deformable convolution module 13 to process the image with the original resolution. At this time, the input module 21 can directly perform a convolution operation on the original image to be enhanced, and through the convolution operation, the pixel values of each pixel and its surrounding neighborhood pixels in the original image are weighted and summed to extract the basic features of the image. When the input module 21 does not perform a pixel inverse shuffle operation, the output module 24 also does not need to perform a pixel shuffle operation in step S4, and only needs to perform a post-processing operation on the third feature map received from the deformable convolution module 13 to generate the final enhanced image.
[0093] Although not performing a pixel inverse shuffle operation will increase the computational amount of the reparameterization module 12 and the deformable convolution module 13, in certain specific scenarios, this method can give full play to the direct processing ability of the model for the features of the original image and avoid potential errors introduced by the inverse shuffle and pixel shuffle operations.
[0094] In step S2, the first feature map is processed by the reparameterization module of the target image processing model to form a second feature map. The reparameterization module includes at least one reparameterization unit, and the reparameterization unit is trained by a multi-scale feature extraction method.
[0095] Exemplarily, referring to Figure 3 and Figure 4 the model structure shown, in step S2, the first feature map can be sequentially processed by multiple reparameterization units, and the feature map output by the last reparameterization unit is used as the second feature map. Each reparameterization unit can include a convolutional layer, a residual connection layer, and an attention layer trained by a multi-branch, multi-scale feature extraction method, so that the extracted features can fuse feature information of multiple scales. The embodiment of the training method of the reparameterization module is described in detail in the subsequent training method embodiment.
[0096] In step S3, the second feature map is subjected to deformable convolution processing by the deformable convolution module of the target image processing model to form a third feature map.
[0097] In an exemplary embodiment, the deformable convolution operation in the deformable convolution module 13 may specifically include: first predicting an offset amount for the second feature map to obtain offset amount information for each sampling position, and then adjusting the default sampling position based on the offset amount information to obtain an adaptive sampling position, sampling the second feature map based on the adaptive sampling position to obtain sampled feature values, and finally performing convolution on the sampled feature values.
[0098] Among them, the offset prediction parameters are formed through the training process. Exemplarily, the process of deformable convolution can be described as: first obtaining the offset amount for each sampling position according to the trained offset prediction parameters, then sampling at each sampling position according to the offset amount to obtain sampled feature values, and in the final convolution step, performing a convolution operation on the sampled feature values obtained by offset sampling and the convolution kernel. The convolution kernel is a predefined set of weights. By multiplying the corresponding elements of the sampled feature values and the convolution kernel and summing them, new feature values are obtained, and these new feature values are combined to form a feature map after the deformable convolution operation. This feature map contains the image features after adaptive sampling and convolution processing, and can better reflect the complex structure and detailed information of the image.
[0099] In some embodiments, two or more deformable convolutions can be set to improve the deformable convolution effect.
[0100] In an exemplary embodiment, the deformable convolution module 13 may include a first deformable convolution unit 231 and a second deformable convolution unit 232 connected in sequence. Then, step S3 may include:
[0101] Step S31, adaptively sampling the positions of the second feature map through the first deformable convolution unit 231 of the deformable convolution module 13, and performing convolution on the sampling result to form a deformable convolution feature map;
[0102] Step S32, adaptively sampling the positions of the deformable convolution feature map through the second deformable convolution unit of the deformable convolution module 13, and performing convolution on the sampling result to form a third feature map.
[0103] The feature map output after being processed by the two deformable convolution units may further be processed by a dynamic hybrid convolution layer, which will not be elaborated herein.
[0104] Through two deformable convolution processes, various changes and differences in the image can be better recognized, and the target can be more accurately located and recognized.
[0105] In step S4, the output module of the target image processing model forms an enhanced image corresponding to the to-be-enhanced image according to the third feature map and the to-be-enhanced image.
[0106] As described in step S1, when a pixel inverse shuffle operation is set in step S1, a pixel shuffle operation may be set in step S4. Then, post-processing (such as Figure 3 the residual connection in the shown structure) is performed on the image after the pixel shuffle operation to form an enhanced image. If the pixel inverse shuffle operation is not set in step S1, there is no need to set the pixel shuffle operation in step S4, and the third feature map can be normally post-processed to form an enhanced image (for example, removing the inverse shuffle unit 111 and the shuffle unit 143 in Figure 3 ).
[0107] In addition, in an exemplary embodiment, step S4 may further include: obtaining a plurality of feature images from a reparameterization module and a deformable convolution module; performing channel splicing on the plurality of feature images through a splicing layer to obtain spliced image data, and processing the spliced image data through a general convolution layer to form a general convolution map; forming an enhanced image based on the general convolution map and the to-be-enhanced image.
[0108] The detailed process of the image enhancement method may be implemented with reference to the hierarchical settings of the model 100 or Figure 3 the detailed hierarchical settings of the model in, which will not be elaborated in this disclosure.
[0109] In the embodiments of the present disclosure, the extremely strong image enhancement ability of the model 100 is achieved not only through a unique model structure but also further endowed by an improved training method, so as to be able to improve the color deviation and the problem of high-frequency region restoration that occur in the traditional image enhancement process.
[0110] Figure 6 It is a flowchart of a training method for an image processing model in an exemplary embodiment of the present disclosure.
[0111] Figure 6 The training method shown can be used to train an image processing model including Figure 1 the model 100 shown, Figure 1 the model 100 shown can be used to execute the image processing method 500 to achieve image enhancement.
[0112] Referring to Figure 6 , the training method 600 may include:
[0113] Step S61, obtaining a target image processing model and training data of the target image processing model, and initializing the parameters of the target image processing model, where the training data includes an image to be enhanced and an original image corresponding to the image to be enhanced;
[0114] Step S62, forming a total loss function according to a focal loss function and a color loss function;
[0115] Step S63, inputting the image to be enhanced into the target image processing model to obtain an enhanced image corresponding to the image to be enhanced;
[0116] Step S64, calculating the difference between the original image and the enhanced image through the total loss function, and updating the parameters of the target image processing model according to the difference.
[0117] In the embodiments of the present disclosure, the original images and the images to be enhanced used in the training and evaluation processes can both be generated by a Matlab JPEG encoder. The training dataset can select DIV2K and Flickr2K, and randomly extract image patches of size 256×256 from them. In step S61, it is necessary to ensure that all images are standardized to ensure the unity of the input data. To enhance the generalization ability of the model, data augmentation techniques such as random rotation and flipping can also be adopted.
[0118] In an exemplary embodiment, during the implementation process of the training method 600, that is, during the model training process, the Adam optimizer can be used for parameter optimization, the batch size is set to 32, the initial learning rate is 1×10-4, and the learning rate is decayed to 0.5 times the previous one after every 5×104 iterations. The entire training process lasts for 400,000 generations, and during this period, an independent validation set (LIVE1 and BSDS500) is used to regularly evaluate the model performance and perform hyperparameter tuning.
[0119] In the training method 600, the total loss function constructed in step S62 is a key step for the training method 600 to solve technical problems.
[0120] In an exemplary embodiment, the total loss function may be constructed based on a focal loss function, a Fourier frequency domain loss function, and a color loss function, where the focal loss function and the color loss function are unique settings of the embodiments of the present disclosure.
[0121] Figure 7 It is a flowchart for constructing a focal loss function in an exemplary embodiment of the present disclosure.
[0122] Reference Figure 7 , in an exemplary embodiment, step S62 includes constructing a focal loss function, and constructing the focal loss function includes:
[0123] Step S71, constructing a resampled high-frequency prior using the original image, bilinear interpolation downsampling parameters, and bilinear value upsampling parameters;
[0124] Step S72, constructing a fitting error prior using the original image, the image to be enhanced, and the enhanced image;
[0125] Step S73, after normalizing the resampled high-frequency prior and the fitting error prior, converting the normalized resampled high-frequency prior and the fitting error prior into non-zero weights;
[0126] Step S74, forming a focal loss value according to the non-zero weights corresponding to the resampled high-frequency prior and the non-zero weights corresponding to the fitting error prior, and the error term between the original image and the enhanced image.
[0127] To solve the problem of high-frequency region recovery in image enhancement, the embodiments of the present disclosure propose a focal loss function, the core of which is to introduce two priors to measure the membership of pixels as low-frequency or high-frequency, so as to dynamically adjust the loss contribution weights of pixels and optimize the training process of the model.
[0128] In step S71, the resampled high-frequency prior is an index for measuring the high-frequency membership of pixels defined based on resampling. The high-frequency static prior based on resampling (i.e., the resampled high-frequency prior) is defined as shown in formula 2:
[0129] P s =|y - U(D(y))| (2)
[0130] In formula (2), y represents the original image, i.e., the real high-quality image, D represents bilinear interpolation downsampling, and U represents bilinear value upsampling. In P sPixels with smaller median values are more likely to belong to the low - frequency region, while pixels with larger values are more likely to belong to the high - frequency region, making P s a rough label for the high - frequency membership of pixels.
[0131] In step S72, the fitting error prior is an index for inferring the high - frequency membership of pixels defined based on the fitting error. The high - frequency learning prior based on the fitting error (i.e., the fitting error prior) is defined as shown in Equation 3:[[]]
[0132] P l =|y - f(x)| (3)
[0133] In Equation (3), x represents the pre - enhanced image (the image to be enhanced), and f(x) represents the output of the model. Since the model usually performs better in the low - frequency region than in the high - frequency region, P l pixels with smaller median values are pixels that are easily fitted in the low - frequency region, while pixels with larger values are pixels that are difficult to fit in the high - frequency region. Therefore, P l can be used as the high - frequency membership of pixels inferred based on the fitting difficulty.
[0134] In step S73, in order to use these high - frequency priors for model training, a weighting function is introduced. Through the exponential function and hyperparameters α and γ, the high - frequency membership is converted into soft weights to re - balance the loss contribution of pixels. Specifically, first, the two priors are normalized so that they have comparable magnitudes. Then, in step S74, an exponential function is introduced to convert them into non - zero weights. Finally, the resulting focal loss function is shown in Equation (4) (including three parts: 1, 2, and 3):
[0135]
[0136] W(z,α,γ)=α×exp(γ×g(z)) (4 - 2)
[0137] L focal =W(P s ,α s ,γ s )×W(P l ,α l ,γ l )×|f(x)-y| p (4 - 3)
[0138] Therefore, the focal loss function proposed in the embodiments of the present disclosure re - balances the pixel contributions by introducing two priors and a weighting function, enabling the image enhancement model to pay more attention to learning the high - frequency region and other pixels that are difficult to fit. It not only improves the recovery ability of the model in the high - frequency region but also has model - independence and can be widely applied to different enhancement models.
[0139] In addition, to solve the problem of color distortion of images, embodiments of the present disclosure also introduce a color loss function based on the mean absolute error (MAE) to improve the color fidelity of the image enhancement model.
[0140] Figure 8 It is a flowchart for constructing a color loss function in an exemplary embodiment of the present disclosure.
[0141] Refer to Figure 8 , in the exemplary embodiment, step S62 may include constructing a color loss function, and constructing the color loss function includes:
[0142] Step S81, normalizing the pixel values of the original image and the enhanced image respectively to obtain a first image and a second image;
[0143] Step S82, filtering out high-frequency information from the first image and the second image respectively, and only retaining low-frequency information to obtain first low-frequency information corresponding to the first image and second low-frequency information corresponding to the second image;
[0144] Step S83, converting the first low-frequency information and the second low-frequency information into HSV colors;
[0145] Step S84, calculating the differences of the converted first low-frequency information and second low-frequency information respectively according to the hue channel, saturation channel, and brightness channel, and summing the difference values corresponding to each channel to obtain a color loss value.
[0146] Specifically, the implementation steps of this color loss function are as follows:
[0147] First, in step S81, limit the pixel values of the enhanced image and the original image output by the model within the range of [0,1] to ensure the validity of the input values. Then, in step S82, perform transformations on these images to filter out high-frequency texture information and retain their low-frequency components, that is, color information. Next, in step S83, convert the low-frequency color information from RGB to the HSV color space. Since the values of the H channel have the characteristic of wrapping around, the calculation of the hue channel difference is different from that of the saturation and brightness channels. The differences of the saturation and brightness channels are measured by the L1 loss, and the calculation of the hue channel difference is shown in formula (5):
[0148]
[0149] where is the enhanced image, elmin(*,*) is an element-wise minimum operation, and g(*) is the transformation of steps S81 - S83.
[0150] Finally, in step S84, the hue loss, saturation loss, and brightness loss are added according to the channels to obtain the total color loss value.
[0151] In an exemplary embodiment, to further optimize the model performance, a focal loss function can be set for priority training. Exemplarily, the color loss function can be enabled only when the focal loss value of the focal loss function is less than a preset value. For example, it can be set that the color loss function is enabled only when the value of the focal loss function is less than 0.3. Thus, the color loss function is not enabled before the focal loss function completes training, and the color loss function is enabled after the focal loss function is trained qualifiedly, that is, backpropagation is only performed on the color loss near the minimum loss value of the model space, solving the problem that the color loss function is difficult to optimize and ensuring that the model can better balance detail restoration and color fidelity when processing complex images.
[0152] Finally, the total loss function constructed according to the focal loss function, color loss function, and Fourier frequency domain loss function can be as shown in formula (6) to optimize high-frequency detail restoration and color fidelity respectively:
[0153]
[0154] where λ and are adjustable hyperparameters. I(*) is an indicator function. is the Fourier loss proposed in existing research.
[0155] The training involving the total loss function can be carried out on an NVIDIA Tesla V100 GPU.
[0156] The total loss function is used to evaluate the model output during the training process, so as to optimize the model parameters during the backpropagation process.
[0157] In addition to improving the total loss function to optimize the image enhancement ability of the model, when the training method 600 is applied to the training process of the model 100 in the embodiments of the present disclosure, the image enhancement ability of the reparameterization module 12 can also be strengthened by setting a unique training method for the reparameterization module 12.
[0158] Exemplarily, when training the reparameterization module 12 in the embodiments of the present disclosure, parallel convolution kernels of multiple different sizes are used to extract features with different receptive fields, and efficient feature fusion is achieved through local residual connections.
[0159] Figure 9 is a flowchart of the training process of the reparameterization module in an exemplary embodiment of the present disclosure.
[0160] Refer to Figure 9, in an exemplary embodiment, the training process of each reparameterization unit in the reparameterization module 12 may include:
[0161] Step S91, convolve the input feature map with multiple convolutional kernels of different scales in parallel through multiple branches to obtain multiple convolutional feature maps of different scales;
[0162] Step S92, perform local residual connection based on the multiple convolutional feature maps to obtain a fused feature map;
[0163] Step S93, train the attention layer through the fused feature map.
[0164] Refer to Figure 3 Understand the training process of the first convolutional layer and the second convolutional layer inside the reparameterization unit shown. During the training process, first, in step S91, convolve through multiple branches (such as Figure 3 branches 01 to 08 in) using multiple convolutional layers of different scales to obtain multiple convolutional feature maps. Assume that the size of the first feature map is 8×8 and the number of channels is 1 (for the convenience of understanding, take a single-channel example first). It can be regarded as an 8×8 matrix, and each element in the matrix represents the pixel value at that position. Next, convolutional kernels of different scales (such as Figure 3 branches 01 to 08 in) can be set in the convolutional layer of a certain or certain branches: 1×1, 3×3, 3×1, and 1×3. Each convolutional kernel has its own independent weight parameters, and these parameters are continuously optimized during the model training process. The 1×1 convolutional kernel can be regarded as a single weight value. For example, its weight is 0.5. The 1×1 convolutional kernel can perform a linear combination and adjustment of the channels of the feature map, changing the number of channels without changing the spatial size of the feature map. The 3×3 convolutional kernel is a 3×3 matrix, and each bit in the matrix is a weight, and these weights can be updated through the training process later. When the 3×3 convolutional kernel slides on the feature map, it will select the pixel values in the surrounding 3×3 area and perform a weighted calculation with the corresponding weights at that position to obtain the convolutional result. Other convolutional kernels are similar. In Figure 3 the embodiment shown, two or more serial convolutional layers can also be set in one or several branches to implement the training of this branch, such as Figure 3 branches 05 to 08 in. Figure 3 The multi-branch setting shown is only an example, and those skilled in the art can adjust the number of branches and the convolutional layer settings of each branch by themselves.
[0165] During the convolution process, each convolution kernel slides on the first feature map, covering an area of a corresponding size or scale each time, and performing the operation of multiplying weights and pixel values and summing them. After the convolution operations of the above different scales, convolution feature maps of different scales are obtained. These feature maps extract information from the input feature map from different angles and scales respectively.
[0166] The above principle is illustrated by taking a single-channel as an example. If the input feature map is multi-channel, the number of channels of each convolution kernel also needs to be the same as that of the first feature map. During the convolution operation, convolution will be performed on each channel separately, and then the results will be added to obtain the final convolution feature map. In this way, convolution kernels of different scales can simultaneously extract information of different scales and different channels on the multi-channel feature map, further enriching the feature expression. During the training process, the output of the image is enhanced through the joint calculation of multiple branches. The parameters of each convolution layer are optimized during the backpropagation process through the improved total loss function. Finally, after training, the parameters of each convolution layer of each branch are fixed. At this time, normalization can be performed on each branch according to the setting of the convolution kernel of the largest size. For example, convolution kernels of 1×1, 3×3, 3×1, and 1×3 are all expanded to 3×3 size. The positions without weights after expansion are set to 0, and then the weights of each convolution kernel are added according to the positions to obtain the comprehensive weight of this convolution layer (a 3×3 convolution kernel) for implementing the inference process. Since each branch is trained separately and can extract features of different scales, the convolution layer formed by the comprehensive superposition of each branch can achieve the optimized extraction of the features of the input image, extract rich and diverse features from the input feature map, and provide more comprehensive information for subsequent image analysis and processing.
[0167] The above sizes and numbers of convolution kernels are only examples, and those skilled in the art can adjust them by themselves.
[0168] In an exemplary embodiment, the present disclosure also sets in step S91: an edge detection operator is set in at least one branch to obtain a convolution feature map containing edge information.
[0169] Exemplarily, the edge detection operator includes a Sobel branch edge detection operator and a Laplace branch edge detection operator, such as Figure 3 shown in branches 06 - 08. Among them, the Sobel operator is divided into Sobelx and Sobely, which are used to detect the edge information of the image in the horizontal and vertical directions; the Laplace operator calculates the second derivative of the image to highlight the edges and details in the image, thereby obtaining a feature map containing rich edge information.
[0170] Refer to Figure 3 and two branches can be set to use the Sobelx and Sobely edge detection operators respectively for a certain convolution process (such asFigure 3 Edge feature extraction is performed on the convolution feature map of the 1×1 convolution layer) in. Taking one of the branches as an example, the Sobel operator has two different detection methods: Sobelx in the horizontal direction and Sobely in the vertical direction. When detecting in the horizontal direction, the operator moves pixel by pixel on the feature map. Each time it moves, specific calculations are performed on the pixels in the current covered area to obtain the edge intensity values in the horizontal direction. These values are combined to form the convolution feature map in the horizontal direction. Similarly, when detecting in the vertical direction, similar operations are performed to obtain the convolution feature map in the vertical direction. Then, the convolution feature maps in the horizontal and vertical directions are combined together according to certain rules to obtain the convolution feature map processed by the Sobel operator, which highlights the edge information of the original feature map in the horizontal and vertical directions.
[0171] In another branch, the Laplace branch edge detection operator is used to process the convolution feature map that has undergone a certain convolution process (such as Figure 3 the 1×1 convolution layer) in. The Laplace operator slides on the feature map and performs another specific calculation on the pixels in each covered area to approximately calculate the second derivative to form the convolution feature map processed by the Laplace operator, which can more prominently display the edge and detail changes in the image.
[0172] In the residual connection process of step S92, adding the two convolution feature maps obtained by processing with the Sobel operator and the Laplace operator can obtain a feature map containing richer edge information. In addition to the Sobel operator and the Laplace operator mentioned above, the edge detection operator can also be the Canny operator, Prewitt operator, Roberts operator, etc. Those skilled in the art can set it according to the actual situation and will not be elaborated here.
[0173] Next, in step S92, during the training process of a convolution layer (refer to Figure 3) Local residual connections are performed on multiple convolutional feature maps from multiple branches to obtain a fused feature map. For example, local regions corresponding to the same position in each convolutional feature map are selected for pixel value addition. For instance, assume that three convolutional feature maps of different scales are denoted as F1 (obtained by a 1×1 convolutional kernel), F2 (obtained by a 3×3 convolutional kernel), and F3 (obtained by a 3×1 convolutional kernel), all with a size of 8×8. First, local residual connections are performed on F1 and F2. In the local residual connection, local regions corresponding to the same position in F1 and F2 are selected for operation. For example, a sub-region is selected at the upper left corner of the feature map. For each pixel point in this sub-region, the value of the corresponding pixel in F1 is added to the value of the corresponding pixel in F2 to obtain a new pixel value, forming a new 3×3 sub-region. This addition operation enables the model to retain information of different scales while learning features, preventing information loss during the feature fusion process. By performing such an operation on each 3×3 sub-region of F1 and F2, an intermediate fused feature map F12 is obtained, which also has a size of 8×8.
[0174] Next, F12 and F3 are subjected to the next step of local residual connection. Taking the selection of a 3×3 sub-region as an example, the pixel points in the corresponding sub-regions of F12 and F3 are added. In this way, the feature information in F3 is fused with the information in F12 that has been fused with local details and channel adjustment, further enriching the feature representation. After this operation, the final fused feature map F123 is obtained, which can integrate the advantages of convolutional feature maps of different scales.
[0175] In practical applications, local residual connections can effectively avoid the problem of gradient vanishing. Because during the backpropagation process, the gradient can be directly transmitted along the residual connection, making it easier for the model to converge during training. Moreover, through the local connection method, the information in the local regions of the feature map can be better utilized to improve the effect of feature fusion. For example, in an image recognition task, the fused feature map can more accurately represent the features of the objects in the image. Whether it is the tiny detail parts or the overall shape structure, they can all be reflected in the fused feature map, thereby improving the classification and recognition accuracy of the model for the image. In an image segmentation task, the fused feature map can more precisely divide the boundaries of different objects, helping the model to more accurately determine the belonging category of pixels.
[0176] When edge detection operators are set in one or some branches, the final fused feature map not only contains various features extracted by convolution kernels of different scales, but also has rich edge information. In subsequent tasks such as image enhancement and object recognition, this rich information enables the model to better identify and process key contents such as object boundaries and texture details in the image, improving the performance and accuracy of the model. For example, in the image segmentation task, accurate edge information can help the model more accurately divide the regions where different objects are located; in image super-resolution reconstruction, rich edge details can make the reconstructed image clearer and more realistic, enhancing the overall quality of the image.
[0177] The training methods corresponding to step S91 and step S92 can be used to train the first convolutional layer 1201 and the second convolutional layer 1202 in the reparameterization unit, or more convolutional layers can be set.
[0178] When training the third dynamic hybrid convolutional layer 1203, serial stacked convolutional layers such as Figure 3 the shown convolutional layers 1×1, convolutional layer 3×3, and convolutional layer 1×1 can be set to complete the training process, and the parameters of each convolutional layer are optimized through the improved total loss function during the backpropagation process. After the training is completed, the parameters of each convolutional layer are fixed. At this time, normalization can be performed on each branch according to the setting of the largest convolutional kernel. For example, both the 1×1 and 3×3 convolutional kernels are expanded to the 3×3 size, and the positions without weights after expansion are set to 0, and then the weights of each convolutional kernel are added according to the positions to obtain the comprehensive weight of the third dynamic hybrid convolutional layer 1203 (a 3×3 convolutional kernel) for implementing the inference process. Since each convolutional layer is trained separately and can extract features of different scales, the convolutional layer formed by the comprehensive superposition of each branch can achieve optimized extraction of the input image features, extract rich and diverse features from the input feature map, and provide more comprehensive information for subsequent image analysis and processing.
[0179] In step S93, the fused feature map is used to train the attention layer, and the parameters of the attention layer are optimized through the improved total loss function during the backpropagation process. For the detailed settings of the attention layer, please refer to Figure 3 the corresponding embodiment, which will not be elaborated here.
[0180] It should be noted that in some embodiments, such as Figure 3 the shown embodiment, the fused feature map can be processed by the third dynamic hybrid convolutional layer 1203, and the feature map obtained after processing is connected with the input feature map of this reparameterization unit in a residual manner to form a new fused feature map to achieve better feature fusion, and then this new fused feature map is used to train the attention layer in step S93.
[0181] Due to the parallel extraction of features through multiple branches and the combination of complex operations during the training phase, the computational requirements during the inference phase are simplified. The reparameterization design of the embodiments of the present disclosure can not only improve the training efficiency of the model, but also optimize the inference speed of the model, significantly reducing resource consumption while maintaining high performance, and is particularly suitable for deployment on resource-constrained devices.
[0182] In summary, the lightweight, multi-branch reparameterized adaptive image enhancement network designed in the embodiments of the present disclosure can significantly improve the effect and performance of image enhancement by training with a total loss function that combines a focal loss function and a color loss function. Compared with existing image enhancement models (such as FBCNN), the number of parameters of the trained Model 100 is only 0.6% of that of the existing model, greatly reducing the number of parameters of the model. In terms of the output results, the enhanced images output by the improved Model 100 have little difference in objective indicators and no obvious difference in subjective perception compared with other existing models with a large number of parameters. This indicates that Model 100 significantly reduces the consumption of computing resources while maintaining high efficiency, and has high practical value.
[0183] Therefore, through lightweight design, the embodiments of the present disclosure significantly reduce the number of parameters and computational complexity of the target image processing model, enabling it to achieve real-time inference of 1080p resolution images on a platform equipped with an NVIDIA Tesla V100 GPU, reaching a processing speed of at least 30 FPS; by using a multi-branch structure to process features of different scales and frequencies in parallel during the training process, the effect of the model on detail restoration and noise suppression is greatly improved; by using a complex multi-branch structure during training and converting it into a single branch during inference, high performance and high efficiency can be ensured at the same time. In addition, the focal loss function used during the training process can adaptively adjust the weights, effectively processing high-frequency and low-frequency regions and improving the detail performance in complex regions; the color loss function ensures the color consistency and authenticity of the enhanced images, avoiding color differences. Combining the above advantages, training Model 100 using training method 600 can perform outstandingly in terms of computational efficiency, resource consumption, detail restoration, noise suppression, and color fidelity, and is excellent in completing the task of real-time image enhancement.
[0184] Table 1 shows the PSNR and SSIM evaluation results of each method on the LIVE1 and BSDS500 datasets under the conditions of quality factor QF = 10 and QF = 20, and Table 2 compares the number of parameters of models of different methods.
[0185] Table 1 Performance of different methods on compressed JPEG images
[0186]
[0187] Table 2 Comparison of the number of parameters of models of different methods
[0188] Method MWC QGAC FBCNN CRL FFDNet This application Number of parameters (M) 23.78 247.39 68.59 27.76 32.38 0.42
[0189] During the verification process, it is found that when this model is applied to the image enhancement task, it can solve the problem of image quality degradation caused by JPEG compression artifacts, and while ensuring efficient real-time inference, it can significantly improve the effect of image detail restoration and color fidelity.
[0190] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0191] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0192] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0193] The following refers to Figure 10 to describe the electronic device 1000 according to this embodiment of the present invention. Figure 10 The shown electronic device 1000 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0194] As Figure 10 shown, the electronic device 1000 is presented in the form of a general-purpose computing device. The components of the electronic device 1000 may include but are not limited to: at least one of the above-mentioned processing units 1010, at least one of the above-mentioned storage units 1020, and a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010).
[0195] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present invention described in the above "exemplary method" part of this specification. For example, the processing unit 1010 can execute the method as shown in the embodiments of the present disclosure.
[0196] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202, and may further include a read-only storage unit (ROM) 10203.
[0197] The storage unit 1020 may also include a program / utilities 10204 having a set (at least one) of program modules 10205. Such program modules 10205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0198] The bus 1030 may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.
[0199] The electronic device 1000 may also communicate with one or more external devices 1100 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or may communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 1050. Also, the electronic device 1000 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1060. As shown in the figure, the network adapter 1060 communicates with other modules of the electronic device 1000 through the bus 1030. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0200] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0201] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium having a program product stored thereon that can implement the above-described method of this specification. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0202] The program product for implementing the above method according to an embodiment of the present invention may be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0203] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0204] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0205] The program code contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0206] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0207] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, and are not for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously in, for example, multiple modules.
[0208] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A training method for an image processing model, characterized in that: include: Acquire a target image processing model and training data of the target image processing model, and initialize parameters of the target image processing model, wherein the training data includes an image to be enhanced and an original image corresponding to the image to be enhanced; Forming a total loss function based on the focus loss function and the color loss function; Inputting the image to be enhanced into the target image processing model to obtain an enhanced image corresponding to the image to be enhanced; The difference between the original image and the enhanced image is calculated by the total loss function, and the parameters of the target image processing model are updated according to the difference.
2. The training method according to claim 1, characterized in that: The total loss function formed according to the focus loss function and the color loss function includes: It is set that when the focus loss value of the focus loss function is less than a preset value, the color loss function is enabled.
3. The training method according to claim 1 or 2, characterized in that: Forming a total loss function according to the focus loss function and the color loss function includes constructing a focus loss function, and constructing the focus loss function includes: Construct a resampled high-frequency prior using the original image, bilinear interpolation downsampling parameters, and bilinear value upsampling parameters; Use the original image, the image to be enhanced and the enhanced image to construct a priori of the fitting error; After normalizing the resampled high frequency prior and the fitting error prior, the normalized resampled high frequency prior and the fitting error prior are converted into non-zero weights; A focus loss value is formed according to the non-zero weight corresponding to the resampled high-frequency prior and the non-zero weight corresponding to the fitting error prior, as well as the error term between the original image and the enhanced image.
4. The training method according to claim 1 or 2, characterized in that: Forming a total loss function based on the focus loss function and the color loss function includes constructing a color loss function. Constructing the color loss function includes: Normalizing the pixel values of the original image and the enhanced image to obtain a first image and a second image respectively; filtering out high-frequency information of the first image and the second image respectively, and retaining only low-frequency information, so as to obtain first low-frequency information corresponding to the first image and second low-frequency information corresponding to the second image; Converting the first low-frequency information and the second low-frequency information into HSV colors; The differences of the converted first low-frequency information and the second low-frequency information are calculated according to the hue channel, the saturation channel, and the brightness channel, respectively, and the difference values corresponding to the channels are summed to obtain a color loss value.
5. The training method according to claim 1, characterized in that: The target image processing model at least includes an input module, a re-parameter module, a deformation convolution module and an output module. The re-parameter module includes at least one re-parameter unit, and the re-parameter unit is trained by multi-scale feature extraction.
6. The training method according to claim 5, characterized in that: The multi-scale feature extraction method training process of the heavy parameter unit includes: Through multiple branches, multiple convolution kernels of different scales are used in parallel to convolve the input feature map to obtain multiple convolution feature maps of different scales; Performing local residual connection based on the multiple convolutional feature maps to obtain a fused feature map; The attention layer is trained by the fused feature map.
7. The training method according to claim 6, characterized in that: Using multiple convolution kernels of different scales to convolve the input feature map in parallel to obtain multiple convolution feature maps of different scales includes: An edge detection operator is set in at least one branch to obtain a convolution feature map containing edge information, wherein the edge detection operator includes a Sobel branch edge detection operator and a Laplace branch edge detection operator.
8. An image enhancement method, characterized in that: include: The image to be enhanced is sampled through an input module of the target image processing model to obtain a first feature map; Processing the first feature map by a re-parameter module of the target image processing model to form a second feature map, wherein the re-parameter module includes at least one re-parameter unit, and the re-parameter unit is trained by a multi-scale feature extraction method; Performing deformation convolution processing on the second feature map through the deformation convolution module of the target image processing model to form a third feature map; An enhanced image corresponding to the image to be enhanced is formed according to the third feature map and the image to be enhanced through the output module of the target image processing model.
9. The image enhancement method according to claim 8, characterized in that: Sampling the image to be enhanced through the input module of the target image processing model to obtain a first feature map includes: Performing a pixel deshuffle operation through the input module, downsampling and channel transformation on the image to be enhanced, so as to reduce the resolution of the image to be enhanced and increase the number of channels of the image to be enhanced, and obtain the first feature map; Forming an enhanced image corresponding to the image to be enhanced according to the third feature map by the output module of the target image processing model includes: A pixel shuffling operation is performed through the output module to upsample the third feature map to form an enhanced image corresponding to the image to be enhanced.
10. The image enhancement method according to claim 8, characterized in that: Performing deformation convolution processing on the second feature map through the deformation convolution module of the target image processing model to form a third feature map includes: Adaptively sampling the second feature map by the first deformable convolution unit of the deformable convolution module, and convolving the sampling result to form a deformable convolution feature map; The deformed convolution feature map is adaptively sampled at a sampling position by the second deformed convolution unit of the deformed convolution module, and the sampling result is convolved to form the third feature map.
11. The image enhancement method according to claim 8, characterized in that: The forming of an enhanced image corresponding to the image to be enhanced according to the third feature map and the image to be enhanced by the output module of the target image processing model comprises: Acquire multiple feature images from the re-parameter module and the deformation convolution module; Perform channel stitching on the multiple feature images through a stitching layer to obtain stitching image data, and process the stitching image data through a general convolution layer to form a general convolution graph; The enhanced image is formed based on the general convolution map and the image to be enhanced.
12. The image enhancement method according to claim 8, characterized in that: The re-parameter unit includes a convolution layer and an attention layer, and the attention layer is used to process the feature image output by the convolution layer through an attention mechanism.
13. An electronic device, characterized in that: include: Memory; as well as A processor coupled to the memory, the processor being configured to execute the method according to any one of claims 1 to 12 based on instructions stored in the memory.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Cited By
Semi-supervised underwater image enhancement method based on multi-scale context sensing
CN120430966A
Image fusion method and system based on depth estimation and dual-module attention
CN121437293A