Image processing method, computer readable medium, device, and construction method

By combining the U-net architecture and attention mechanism with multi-scale fusion and residual group convolution, this image processing method solves the image enhancement problem of large-size DR images, improves contrast and detail recognition capabilities, and is suitable for ultra-high-definition image processing.

CN120013786BActive Publication Date: 2025-10-28CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510156951.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-10-28
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Existing image enhancement methods struggle to simultaneously improve the contrast, visual quality, and image information entropy of large-size DR images. Furthermore, deep learning networks are ineffective in DR image enhancement tasks, resulting in blurred image edge information, uneven grayscale distribution, and unclear details, making it difficult to identify defects in large-size components.

Method used

An image processing method is adopted, which combines the main and auxiliary encoders and decoders, utilizes the U-net architecture and attention mechanism, hybridizes encoded feature maps, and combines multi-scale fusion and residual group convolution to process the image, thereby enhancing the contrast and detail features of the image.

Benefits of technology

It improves the contrast and detail texture capture capabilities of large-size DR images, enabling better identification of defects in large-size components, while balancing hardware consumption and processing time, and is suitable for ultra-high-definition image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013786B_ABST
    Figure CN120013786B_ABST
Patent Text Reader

Abstract

An image processing method, a computer-readable medium, an electronic device, an image processing device, and a method for constructing the same. Let i = 1, and it includes: S1. Process the image to be enhanced to obtain [description of result of S1] and [description of another result of S1]; S2. Encode [result of S1] into [encoded result 1] and [result of S1] into [encoded result 2]; S3. Mix the encoded [encoded result 1] and [encoded result 2] to obtain [mixed encoded result]; S4. If i < M, then let i = i + 1, and return to S2; if i = M, then let [description of operation for i = M] and proceed to the next step; S5. Decode [mixed encoded result] into [decoded result 1] and [decoded result 2]; S6. Mix the decoded [decoded result 1] and [decoded result 2] to obtain [mixed decoded result]; S7. If i > 1, then let i = i - 1, and return to S5; if i = 1, then process [mixed decoded result] to obtain the enhanced image. Using it can improve the low-light display quality of ultra-high-definition images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically to an image processing method, a computer-readable medium, an electronic device, an image processing apparatus, and a method for constructing the same. Background Technology

[0002] Digital radiography (DR) technology receives X-ray energy through a flat panel detector and directly converts it into electrical signals, which are then converted into digital images by a computer. DR imaging technology has advantages such as high spatial resolution, large information capacity, and wide dynamic range. More importantly, it has a fast imaging speed, enabling rapid and effective detection of surface and internal defects in castings, achieving real-time imaging inspection.

[0003] Traditional image enhancement methods often achieve a balance between improving image contrast, visual effect, and image information entropy, but cannot simultaneously improve all of these indicators. Researchers often use specific algorithms for enhancement based on different needs. Therefore, applying traditional image enhancement methods requires manual adjustment of multiple parameters, and these methods are also time-consuming when dealing with large-sized DR images. Furthermore, current deep learning networks rarely address DR image enhancement tasks, and many are ineffective against large DR images. In addition, factors such as scattering, electrical noise, and uneven thickness of castings in X-ray imaging systems can lead to blurred edge information, uneven grayscale distribution, and unclear details in DR images, which are typically generated as large-sized, high dynamic range DR images.

[0004] Patent document CN118115373A discloses a method for constructing a large-size image enhancement model, an image processing method, a computer-readable medium, and an electronic device. The method for constructing the large-size image enhancement model includes steps of creating training and testing sets, establishing a large-size image enhancement model, training the parameters of the large-size image enhancement model, and testing the parameters of the large-size image enhancement model. The large-size image enhancement model includes an encoder, a decoder, a multi-scale fusion module, and a brightness adjustment module. While this approach can process large-size, high dynamic range images to be enhanced, simultaneously improving contrast, visual effect, and image information entropy, it has limitations in enhancing the local contrast of fine details in large-size components, making it difficult to identify defects in large-size components. Summary of the Invention

[0005] The purpose of this invention is to provide an image processing method, a computer-readable medium, an electronic device, an image processing apparatus, and a method for constructing the same, so as to improve image quality.

[0006] This invention is implemented as follows:

[0007] An image processing method, let the pointer i = 1, includes the following steps:

[0008] Step 1, process the image M0 to be enhanced to obtain a feature map and a feature map

[0009] Step 2, encode the feature map into a feature map Encode the feature map into a feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map The method of encoding the feature map into a feature map is different from the method of encoding the feature map into a feature map ;

[0010] Step 3, mix and encode the feature map and the feature map into a feature map

[0011] Step 4, if i < M, then let i = i + 1, return to Step 2 to execute the loop of Step 2 to Step 4; if i = M, then let the feature map be the feature map The feature map be the feature map for the next step;

[0012] Step 5, decode the feature map into a feature map Decode the feature map into a feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map ;

[0013] Step 6, mix and encode the feature map and the feature map into a feature map

[0014] Step 7, if i > 1, then let i = i - 1, return to Step 5 to execute the loop of Step 5 to Step 7; if i = 1, then process the feature map The enhanced image M0′.

[0015] Preferably, hybrid encoded feature maps and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M i1 ; 1×1 convolution to process feature maps Obtain matrix M i2 M i1 With M i2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. i ;FE i =σE i ×M i2 +(1-σE i M i1 ; feature map Set to q, and set the feature map Set it to k, and set FE i The feature map is obtained after processing with the attention mechanism by setting it to v. Hybrid coding feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M i3 ; 1×1 convolution to process feature maps Obtain matrix M i4 M i3 With M i4 After multiplication, the sigmoid activation function is applied to obtain the feature map σD. i ;FD i =σD i ×M i4 +(1-σD i M i3 ; feature map Set to q, and set the feature map Set it to k, and set FD i The feature map is obtained after processing with the attention mechanism by setting it to v.

[0016] Preferably, the encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Multi-scale fusion processing feature maps Obtain feature map Encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Residual group convolution processing feature map Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Multi-scale fusion processing feature maps Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Residual group convolution processing feature map Obtain feature map

[0017] Further preferably, let the depth of the multi-scale fusion processing feature map be N, and process the image to be enhanced M0 to obtain the feature map. and feature map The method is:

[0018] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0019] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0020] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained For each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map.

[0021] For images Perform 3×3 convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map consistent with the feature map size. With the same dimensions, feature maps are obtained.

[0022] Further preferably, let the depth of the multi-scale fusion processing feature map be N, and process the image to be enhanced M0 to obtain the feature map. and feature map The method is:

[0023] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0024] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0025] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map.

[0026] For images Perform convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map equal to the feature map size. The feature maps are of the same size, and then multi-scale fusion processing is performed on the obtained feature maps to obtain the final feature map.

[0027] Hybrid coding feature map and feature map For feature map

[0028] Furthermore, preferably, if supplementary blocks are used in step 1 during the processing of the image M0 to be enhanced, then in step 7, the feature map is processed. The enhanced image M0′ includes the following steps:

[0029] S711, Down-channel processing feature map Obtain the image MD;

[0030] S712. Cropping the supplementary block in image MD yields the enhanced image M0′.

[0031] If, in step 1, no supplementary blocks are used during the processing of the image M0 to be enhanced, then in step 7, the feature map is processed. The enhanced image M0′ includes the following steps:

[0032] S711, Down-channel processing feature map The enhanced image M0′ is obtained.

[0033] A computer-readable medium storing an image processing program, which, when loaded by a processor, is used to execute the aforementioned image processing method.

[0034] An electronic device includes a computer-readable medium storing an image processing program and a processor, the image processing program being loaded by the processor to perform the aforementioned image processing method.

[0035] An image processing apparatus for processing an image M0 to be enhanced into an enhanced image M0′, comprising:

[0036] The main encoder is used to encode feature maps. For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0037] Auxiliary encoder, used to encode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image, encoding feature map For feature map The method differs from encoding feature maps For feature map Methods;

[0038] Main decoder, used to decode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0039] Auxiliary decoder, used for decoding feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0040] Dual-channel mixer for mixing and encoding feature maps and feature map For feature map and for hybrid encoding feature maps and feature map For feature map

[0041] The primary processor processes the image M0 to be enhanced to obtain feature maps. and feature map and processing feature maps The enhanced image M0′;

[0042] Where a = 1, 2, ..., M-1; b = 1, 2, ..., M-1; c = 1, 2, ..., M; d = 2, 3, ..., M; e = 2, 3, ..., M; f = 1, 2, ..., M; M ≥ 3; feature map For feature map Feature map For feature map

[0043] Preferably, hybrid encoded feature maps and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M c1 ; 1×1 convolution to process feature maps Obtain matrix M c2 M c1 With M c2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. c ;FE c =σE c ×M c2 +(1-σE c M c1 ; feature map Set to q, and set the feature map Set it to k, and set FE c The feature map is obtained after processing with the attention mechanism by setting it to v.

[0044] Hybrid coding feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M f3 ; 1×1 convolution to process feature maps Obtain matrix M f4 M f3 With M f4 After multiplication, the sigmoid activation function is applied to obtain the feature map σD. f ;FD f =σD f ×M f4 +(1-σD f M f3 ; feature map Set to q, and set the feature map Set it to k, and set FD f The feature map is obtained after processing with the attention mechanism by setting it to v.

[0045] Preferably, the encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Multi-scale fusion processing feature maps Obtain feature map Encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Residual group convolution processing feature map Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Multi-scale fusion processing feature maps Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Residual group convolution processing feature map Obtain feature map

[0046] More preferably, let the depth of the multi-scale fusion processing feature map be N, and in the primary processor, process the image to be enhanced M0 to obtain the feature map. and feature map The method is:

[0047] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0048] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0049] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained For each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map.

[0050] For images Perform 3×3 convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map consistent with the feature map size. With the same dimensions, feature maps are obtained.

[0051] Further preferably, let the depth of the multi-scale fusion processing feature map be N, and process the image to be enhanced M0 to obtain the feature map. and feature map The method is:

[0052] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0053] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0054] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map.

[0055] For images Perform convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map equal to the feature map size. The feature maps are of the same size, and then multi-scale fusion processing is performed on the obtained feature maps to obtain the final feature map.

[0056] Hybrid coding feature map and feature map For feature map

[0057] Furthermore, if supplementary blocks are used in processing the image M0 to be enhanced, then in the primary processor, the feature map is processed. The enhanced image M0′ includes the following steps:

[0058] S711, Down-channel processing feature map Obtain the image MD;

[0059] S712. Cropping the supplementary block in image MD yields the enhanced image M0′.

[0060] In step 1, if no supplementary block is used during the processing of the image M0 to be enhanced, then the feature map is processed in the primary processor. The enhanced image M0′ includes the following steps:

[0061] S711, Down-channel processing feature map The enhanced image M0′ is obtained.

[0062] The aforementioned method for constructing an image processing device includes the following steps:

[0063] The steps of creating training and testing sets, wherein each sample in the training and testing sets includes an image to be enhanced M0 and a label image M0″;

[0064] Steps for establishing an image processing simulation device;

[0065] The step of training the parameters of the image processing simulation device involves using the training set to train the image processing simulation device and obtaining the application parameters of the image processing simulation device.

[0066] The steps for testing the parameters of the image processing simulation device are as follows: the parameters of the image processing simulation device are set to the application parameters; the image processing simulation device is tested using the test set; if the error between the enhanced image output by the image processing simulation device and the corresponding label image meets the requirements, then the image processing device is constructed according to the image processing simulation device and the application parameters.

[0067] Preferably, in the step of establishing the image processing simulation device, a loss function is also established, and the loss function is:

[0068]

[0069]

[0070] Among them, I i (i = 1, 2, ..., M) is To perform label image Feature maps obtained by downsampling the size; For I i and The L1 norm loss, for for The image value, y(h,w), is I. i Image values; For I i and The perceived loss, specifically, is to reduce I i Input the feature map extracted by the VGG network and The L2 norm loss between feature maps extracted by the input VGG network is used for... To be Input the feature map extracted by the VGG network, φ j (y) represents I i Input the feature map extracted by the VGG network; Y is For matching The label image; For Y and The L1 norm loss, for for The image value of y is y(h,w), and the image value of Y is y(h,w). For Y and The perceptual loss is specifically calculated by inputting Y into the feature map extracted by the VGG network and then... The L2 norm loss between feature maps extracted by the input VGG network is used for... To be Input the feature map extracted by the VGG network, φ j (y) represents the feature map extracted by inputting Y into the VGG network; λ a , λ p , λ m All are weighting coefficients for the loss; This represents the total loss.

[0071] The beneficial effects of this invention include:

[0072] 1. The image processing method of this invention can process large-sized, high dynamic range images to be enhanced, simultaneously improving their contrast, visual effect, and image information entropy. It can better enhance the local contrast of subtle features, improve the capture of image detail textures, strengthen the contrast of high-frequency detail areas, and highlight important parts of the image. Taking the DR image recognition of railway casting bolsters and side frames as an example, it can locally enhance the contrast of subtle cracks on the side frame, which is beneficial for identifying defects in large-sized components. Both the main and auxiliary paths adopt the U-net architecture, mixing and encoding the feature maps obtained from each layer of the auxiliary path with those obtained from each layer of the main path, using this as the input feature map for the next layer of the main path. This allows the main path to focus on learning the main feature maps from the auxiliary path, thereby improving image quality, such as improving the display effect of low-light images. When processing ultra-high-definition images, this architecture balances hardware consumption and processing time, significantly improving both image processing speed and image quality, and can be applied to processing large-sized ultra-high-definition images.

[0073] 2. The image processing method of the present invention employs an attention mechanism to mix and encode the feature maps obtained from each layer of the auxiliary path with the feature maps obtained from each layer of the main path, and use them as the input feature maps for the next layer of the main path. The main path can focus on learning the main features of the image from the auxiliary path and discard some irrelevant and unimportant information, thereby expanding the applicability of the image processing method.

[0074] 3. The image processing method of the present invention employs residual group convolution in the auxiliary path to process images, which can reduce the amount of computation and the number of parameters, while enhancing the image representation capability by increasing the number of groups. In the main path, multi-scale fusion processing is used to process images, enabling the network to perform inference efficiently even when faced with image inputs of different sizes, without losing detailed features, better extracting the feature parts of the image, and enhancing the display effect of features.

[0075] 4. The computer-readable medium, electronic device, and image processing device of the present invention, which store image processing programs, can process high dynamic range images to be enhanced, such as DR images, and simultaneously enhance their contrast, visual effect, and image information entropy. Furthermore, they are easy to cooperate with automated equipment to achieve automated enhancement processing of high dynamic range images to be enhanced, reducing the need for user experience.

[0076] 5. The method for constructing the image processing device of the present invention can facilitate the design and verification process of the image processing device and reduce the time required to acquire the image processing device. Attached Figure Description

[0077] Figure 1 This is a simplified structural diagram of an image processing method.

[0078] Figure 2 This is a structural diagram of an image processing method.

[0079] Figure 3 This is a method for multi-scale fusion processing of images.

[0080] Figure 4 This is a method for processing images using ResGroup convolution.

[0081] Figure 5 This is a method for hybrid encoding of two-way images.

[0082] Figure 6 This is a DR image that needs to be enhanced.

[0083] Figure 7 Processing the image using the image processing program constructed in Example 4 Figure 6 The processed image is then output.

[0084] Figure 8 This is a DR image that needs to be enhanced.

[0085] Figure 9 Processing the image using the image processing program constructed in Example 4 Figure 8 The processed image is then output. Detailed Implementation

[0086] The present invention will now be described with reference to the accompanying drawings and embodiments to assist those skilled in the art in understanding and implementing the invention. Unless otherwise stated, the following embodiments and the technical terms therein should not be understood without a background of technical knowledge in this field.

[0087] Resolution refers to the number of pixels on a screen or display device, typically composed of horizontal and vertical pixel counts. Early monitors had lower resolutions, such as CRT monitors, which were mostly 1024×768. With technological advancements, resolutions gradually increased, leading to different high-definition standards, the most common being HD (High Definition), FHD (Full High Definition), QHD (Quad High Definition), and UHD (Ultra High Definition).

[0088] A resolution of 1280x720 or higher but lower than 1920x1080 is generally called HD; a resolution of 1920x1080 or higher but lower than 2560x1440 is generally called FHD; a resolution of 2560x1440 or higher but lower than 3840x2160 is generally called QHD; and a resolution of 3840x2160 or higher is generally called UHD.

[0089] In low-light environments, exposure anomalies often result in significant quality degradation in low-light images, including high noise levels, low contrast, low visibility, and extremely low information entropy, making it difficult or impossible for the human eye to perceive information. Traditional Low-Light Image Enhancement (LLIE) methods primarily rely on image priors and distribution mappings, such as methods based on segmentation iteration functions, histogram equalization, homomorphic filtering, and model optimization. Deep learning-based LLIE methods, utilizing large-scale synthetic or real low-light enhancement datasets, have significantly improved image processing speed and image quality, but the applicability of these image processing models is limited. To broaden the applicability of these models, the Transformer architecture has been introduced. However, processing ultra-high-resolution images easily leads to a rapid increase in computational load and GPU memory usage, resulting in reduced model efficiency.

[0090] An image processing method, where pointer i = 1, includes the following steps:

[0091] Step 1: Process the image M0 to be enhanced to obtain the feature map. and feature map

[0092] Step 2: Encode feature maps For feature map Encoded feature map For feature map The feature map The corresponding image size is the feature map. twice the corresponding image size, the feature map the corresponding image size of the feature map twice the corresponding image size, the encoded feature map is the feature map The method of is different from the encoded feature map is the feature map ;

[0093] Step 3, mix the encoded feature map and the feature map is the feature map

[0094] Step 4, if i < M, then set i = i + 1, return to Step 2 to execute the loop of Step 2 to Step 4; if i = M, then set the feature map is the feature map Feature map is the feature map Proceed to the next step;

[0095] Step 5, decode the feature map is the feature map Decode the feature map is the feature map The feature map The corresponding image size is twice the corresponding image size of the feature map ; the feature map The corresponding image size is twice the corresponding image size of the feature map ;

[0096] Step 6, mix the encoded feature map and the feature map is the feature map

[0097] Step 7, if i > 1, then set i = i - 1, return to Step 5 to execute the loop of Step 5 to Step 7; if i = 1, then process the feature map as the enhanced image M0'.

[0098] In the present invention, the method of encoding the feature map is the feature map can be the following steps (see Figure 2 ):

[0099] S211. Shrink and process the feature map to reduce the image feature dimension, obtaining the feature map

[0100] S212. Perform multi-scale fusion processing on the feature map Obtain feature map

[0101] In this invention, the encoded feature map For feature map The method can be the following steps (see Figure 1 ):

[0102] S221, Shrinkage Processing Feature Diagram To reduce the dimensionality of image features, feature maps are obtained.

[0103] S222, ResGroup Processing Feature Map Obtain feature map

[0104] In this invention, the depth of the feature map in S212 is set to N, and the image to be enhanced, M0, is processed to obtain the feature map. and feature map The method can be as follows:

[0105] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0106] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0107] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Sub-image The dimensions are the same, and the cropping order of the sub-images can be determined by following the order from left to right and top to bottom; for each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are concatted and merged at the channel level according to the cropping order to obtain the feature map.

[0108] For images Perform 3×3 convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map consistent with the feature map size. With the same dimensions, feature maps are obtained.

[0109] In actual processing, a kernel size of 3, a stride of 1, and padding of 1 can be used for each sub-image. Perform 3×3 convolution to obtain the feature map. Let the kernel size be 3, the stride be 2, and the padding be 1 for a pair of images. Perform 3×3 convolution to obtain the feature map. This process yields the feature map. and feature map Since the length and width dimensions are the same, the corresponding feature maps obtained from subsequent processing can be used for hybrid encoding.

[0110] In this invention, the depth of the feature map in S212 is set to N, and the image to be enhanced, M0, is processed to obtain the feature map. and feature map The method can be as follows:

[0111] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0112] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0113] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Sub-image The dimensions are the same; each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are concatted and merged at the channel level according to the cropping order to obtain the feature map.

[0114] For images Perform convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map equal to the feature map size. The feature maps are of the same size, and then multi-scale fusion processing is performed on the obtained feature maps to obtain the final feature map.

[0115] Hybrid coding feature map and feature map For feature map

[0116] In actual processing, a kernel size of 3, a stride of 1, and padding of 1 can be used for each sub-image. Perform 3×3 convolution to obtain the feature map. Let the kernel size be 3, the stride be 2, and the padding be 1 for a pair of images. Perform a 3×3 convolution. This process yields the feature map. and feature map Since the length and width dimensions are the same, the corresponding feature maps obtained from subsequent processing can be used for hybrid encoding.

[0117] In this invention, the hybrid coding feature map and feature map For feature map The method can be the following steps (see Figure 2 ):

[0118] S311, 1×1 convolution processing feature map Obtain matrix M i1 ; 1×1 convolution to process feature maps Obtain matrix M i2 M i1 With M i2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. i ;FE i =σE i ×M i2 +(1-σE i M i1 ;

[0119] S312, Feature map Set to q, and set the feature map Set it to k, and set FE i The feature map is obtained after processing with the attention mechanism by setting it to v.

[0120] In this invention, the decoding feature map For feature map The method can be the following steps (see Figure 2 ):

[0121] S511, Extended Processing Feature Map To increase the feature dimension of the image, a feature map is obtained.

[0122] S512, Multi-scale fusion processing feature map Obtain feature map

[0123] In this invention, the decoding feature map For feature map The method can be as follows:

[0124] S521, Extended Processing Feature Map To increase the feature dimension of the image, a feature map is obtained.

[0125] S522 and ResGroup process feature maps Obtain feature map

[0126] In this invention, the hybrid coding feature map and feature map For feature map The method can be the following steps (see Figure 2 ):

[0127] S611, 1×1 convolution processing of feature maps Obtain matrix M i3 ; 1×1 convolution to process feature maps Obtain matrix M i4 M i3 With M i4 After multiplication, the sigmoid activation function is applied to obtain the feature map σD. i ;FD i =σD i ×M i4 +(1-σD i M i3 ;

[0128] S612, Feature map Set to q, and set the feature map Set it to k, and set FD i The feature map is obtained after processing with the attention mechanism by setting it to v.

[0129] In this invention, if supplementary blocks are used in step 1 during the processing of the image M0 to be enhanced, then in step 7, the feature map is processed. The enhanced image M0′ includes the following steps:

[0130] S711, Down-channel processing feature map Obtain the image MD;

[0131] S712. Cropping the supplementary block in image MD yields the enhanced image M0′.

[0132] In this invention, if no supplementary blocks are used in step 1 during the processing of the image M0 to be enhanced, then in step 7, the feature map is processed. The enhanced image M0′ includes the following steps:

[0133] S711, Down-channel processing feature map The enhanced image M0′ is obtained.

[0134] In this invention, the supplementary block is a block of the same color, which can be a pure white block, a pure black block, or other solid color blocks.

[0135] Figure 2 An embodiment of the image processing method of the present invention is shown. In this embodiment, M=3, SAM is a multi-scale fusion image processing method, ResGroup is an image processing method, and DAFM is a method for hybrid encoding of two images.

[0136] In image processing, ResGroup is a specific network module structure commonly used in deep learning models, especially in tasks such as image restoration and super-resolution reconstruction. For example, in the Omni-Kernel Network (OKNet), ResGroup is the basic module of the network, consisting of multiple residual blocks (ResBlocks). Each residual block contains two 3×3 convolutional layers with the non-linear activation function GELU in between.

[0137] ResGroup effectively extracts deep features from images by stacking residual blocks and enhances the expressive power of these features through residual learning, thereby better handling details and structural information in images. In OKNet, ResGroup is used in both the encoder and decoder stages. With proper network structure design, model performance can be improved without significantly increasing computational overhead. ResGroup can be flexibly stacked and combined as needed to adapt to different image processing tasks, such as image dehazing and super-resolution reconstruction.

[0138] Figure 3 An embodiment of a multi-scale fusion image processing method, SAM, is shown. In this embodiment, N = 3. Feature map F r F is obtained by downsampling by 1 / 2. ↓r F was obtained by downsampling by 1 / 4. ↓↓ r For F respectively r F ↓ r F ↓↓ r A ResNet architecture is learned to extract features, resulting in Y0, Y1′, and Y2′. Y1′ is then interpolated to obtain Y1, which has the same image size as Y0; similarly, Y2′ is interpolated to obtain Y2, which has the same image size as Y0. Global average pooling (GAP) is then applied to Y0, Y1, and Y2, followed by multilayer perceptron (MLP) processing to obtain corresponding feature maps. These three feature maps are then fused according to their weights to obtain feature map F. r Feature map F r and feature map F r Adding and fusing them together yields F. out F out A feature map corresponds to a single image. After processing a single feature map using a multi-scale fusion module, N multi-scale feature maps can be obtained.

[0139] Figure 4 This paper presents a ResGroup method for image processing to improve image feature extraction efficiency and reduce hardware requirements. For feature map x, with the decimal pointer i = 1 and feature map x0 as feature map x, the following processing steps are performed: Sa1, perform a 3×3 convolution to extract preliminary features; Sa2, process using the GELU Gaussian error linear unit activation function; Sa3, perform a 3×3 convolution and then merge with feature map x. i-1 Perform residual connections to obtain feature map x i Let i = i + 1; Sa4, if i = N (this N is not the depth of the feature map processed by multi-scale fusion in S212, but a newly defined parameter), then execute step Sa5; if i < N, then return to step Sa1 and execute the loop from step Sa1 to step Sa4; Sa5, for feature map x N 3×3 convolution is performed to extract features; Sa6, GELU Gaussian error linear unit activation function is used for processing; Sa7, the dynamically extracted image features are further refined through the multi-scale fusion module SAM; Sa8, the 3×3 convolution process is then compared with the feature map x. N Perform residual connections to obtain the final feature map of sub-image x.

[0140] For Figure 2 The image processing method shown can split an image into N sub-images in the branch as follows: Figure 1 Feature map shown First, the image is split into N sub-images. Then, each sub-image undergoes multi-layer encoding and multi-layer decoding to obtain the feature map of the corresponding layer. Alternatively, the feature map of each layer can be obtained... Later The image is split into N sub-images. In the next layer, the feature map of each feature map is downsampled and then processed accordingly to obtain the feature map of the corresponding sub-image.

[0141] Figure 5 A hybrid encoding method, DAFM, is presented to reduce the semantic gap from the two paths. This involves hybrid encoding of feature maps. and feature map For feature map For example, it includes the following steps: 1×1 convolution to process the feature map. Obtain matrix M i1 ; 1×1 convolution to process feature maps Obtain matrix M i2 M i1 With M i2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. i At this time σE i Feature maps that represent information closer to reality, FE i =σE i ×M i2 +(1-σE i M i1 Feature map Set to q, and set the feature map Set it to k, and set FE i The feature map is obtained after processing with the attention mechanism by setting it to v.

[0142] In existing technologies, the U-net architecture includes an encoding side, a connector, and a decoding side. The encoding side includes methods for shrinking feature maps to reduce the dimensionality of image features, while the decoding side, in conjunction with the connector, includes methods for expanding feature maps to increase the dimensionality of image features. In this invention, the feature maps are shrunk. Reducing the dimensionality of image features is equivalent to reducing the dimensionality of image features in the U-net architecture. In this invention, the feature map is expanded for processing. This increases the image feature dimension, which is equivalent to increasing the image feature dimension in the U-net architecture. See [link / reference]. Figure 1 In the main path, the feature map of the decoding side is extended for processing. Then, through the connector and the encoding side feature map After connection, the decoding side feature map is obtained.

[0143] The image processing method of the present invention can be used to obtain a computer-readable medium storing an image processing program, an electronic device, and an image processing apparatus.

[0144] Example 1: A computer-readable medium storing an image processing program, which, when loaded by a processor, is used to execute the image processing method of the present invention.

[0145] The image processing program of the present invention is used to process the image to be enhanced M0 into an enhanced image M0′, including:

[0146] The main encoder is used to encode feature maps. For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0147] Auxiliary encoder, used to encode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image, encoding feature map For feature map The method differs from encoding feature maps For feature map Methods;

[0148] Main decoder, used to decode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0149] Auxiliary decoder, used for decoding feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0150] Dual-channel mixer for mixing and encoding feature maps and feature map For feature map and for hybrid encoding feature maps and feature map For feature map

[0151] The primary processor processes the image M0 to be enhanced to obtain feature maps. and feature map and processing feature maps The enhanced image M0′;

[0152] Where a = 1, 2, ..., M-1; b = 1, 2, ..., M-1; c = 1, 2, ..., M; d = 2, 3, ..., M; e = 2, 3, ..., M; f = 1, 2, ..., M; M ≥ 3; feature map For feature map Feature map For feature map

[0153] In this embodiment, the processor refers to the loading and execution hardware such as a microcontroller, CPU, and computing card, while the primary processor refers to the image processing device.

[0154] Preferably, in the main encoder, the encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Multi-scale fusion processing feature maps Obtain feature map

[0155] Preferably, in the auxiliary encoder, the encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. ResGroup processing feature map Obtain feature map

[0156] Preferably, in a hybrid encoder, the hybrid encoded feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M c1 ; 1×1 convolution to process feature maps Obtain matrix M c2 M c1 With M c2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. c ;FE c =σE c ×M c2 +(1-σE c M c1 ; feature map Set to q, and set the feature map Set it to k, and set FE cThe feature map is obtained after processing with the attention mechanism by setting it to v.

[0157] Preferably, in the main decoder, the decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Multi-scale fusion processing feature maps Obtain feature map

[0158] Preferably, in the auxiliary decoder, the decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. ResGroup processing feature map Obtain feature map

[0159] Preferably, in a hybrid encoder, the hybrid encoded feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M f3 ; 1×1 convolution to process feature maps Obtain matrix M f4 M f3 With M f4 After multiplication, the sigmoid activation function is applied to obtain the feature map σD. f ;FD f =σD f ×M f4 +(1-σD f M f3 ; feature map Set to q, and set the feature map Set it to k, and set FD f The feature map is obtained after processing with the attention mechanism by setting it to v.

[0160] More preferably, let the depth of the multi-scale fusion processing feature map be N, and in the primary processor, process the image to be enhanced M0 to obtain the feature map. and feature map The method can be as follows:

[0161] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0162] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0163] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Sub-image The dimensions are the same, and the cropping order of the sub-images can be determined by following the order from left to right and top to bottom; for each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map.

[0164] For images Perform 3×3 convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map consistent with the feature map size. With the same dimensions, feature maps are obtained.

[0165] In actual processing, a kernel size of 3, a stride of 1, and padding of 1 can be used for each sub-image. Perform 3×3 convolution to obtain the feature map. Let the kernel size be 3, the stride be 2, and the padding be 1 for a pair of images. Perform 3×3 convolution to obtain the feature map. This process yields the feature map. and feature map Since the length and width dimensions are the same, the corresponding feature maps obtained from subsequent processing can be used for hybrid encoding.

[0166] More preferably, let the depth of the multi-scale fusion processing feature map be N, and in the primary processor, process the image to be enhanced M0 to obtain the feature map. and feature map The method can be the following steps (see Figure 2 ):

[0167] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0168] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0169] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Sub-image The dimensions are the same; each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are concatted and merged at the channel level according to the cropping order to obtain the feature map.

[0170] For images Perform convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map equal to the feature map size. The feature maps are of the same size, and then multi-scale fusion processing is performed on the obtained feature maps to obtain the final feature map.

[0171] Hybrid coding feature map and feature map For feature map

[0172] In actual processing, a kernel size of 3, a stride of 1, and padding of 1 can be used for each sub-image. Perform 3×3 convolution to obtain the feature map. Let the kernel size be 3, the stride be 2, and the padding be 1 for a pair of images. Perform a 3×3 convolution. This process yields the feature map. and feature map Since the length and width dimensions are the same, the corresponding feature maps obtained from subsequent processing can be used for hybrid encoding.

[0173] Furthermore, if supplementary blocks are used in processing the image M0 to be enhanced, then in the primary processor, the feature map is processed. The enhanced image M0′ includes the following steps:

[0174] S711, Down-channel processing feature map Obtain the image MD;

[0175] S712. Cropping the supplementary block in image MD yields the enhanced image M0′.

[0176] In step 1, if no supplementary block is used during the processing of the image M0 to be enhanced, then the feature map is processed in the primary processor. The enhanced image M0′ includes the following steps:

[0177] S711, Down-channel processing feature map The enhanced image M0′ is obtained.

[0178] In this invention, the supplementary block is a block of the same color, which can be a pure white block, a pure black block, or other solid color blocks.

[0179] Example 2: An electronic device includes a computer-readable medium storing an image processing program and a processor, the image processing program being loaded by the processor to execute the aforementioned image processing method.

[0180] Example 3: An image processing device for processing an image M0 to be enhanced into an enhanced image M0′, comprising:

[0181] The main encoder is used to encode feature maps. For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0182] Auxiliary encoder, used to encode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image, encoding feature map For feature map The method differs from encoding feature maps For feature map Methods;

[0183] Main decoder, used to decode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0184] Auxiliary decoder, used for decoding feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image;

[0185] Dual-channel mixer for mixing and encoding feature maps and feature map For feature map and for hybrid encoding feature maps and feature map For feature map

[0186] The primary processor processes the image M0 to be enhanced to obtain feature maps. and feature map and the feature map of downchannel processing The enhanced image M0′;

[0187] Where a = 1, 2, ..., M-1; b = 1, 2, ..., M-1; c = 1, 2, ..., M; d = 2, 3, ..., M; e = 2, 3, ..., M; f = 1, 2, ..., M; M ≥ 3; feature map For feature map Feature map For feature map

[0188] Preferably, in the main encoder, the encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Multi-scale fusion processing feature maps Obtain feature map

[0189] Preferably, in the auxiliary encoder, the encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. ResGroup processing feature map Obtain feature map

[0190] Preferably, in a hybrid encoder, the hybrid encoded feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M c1 ; 1×1 convolution to process feature maps Obtain matrix M c2 M c1 With M c2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. c ;FE c =σE c ×M c2 +(1-σE c M c1 ; feature map Set to q, and set the feature map Set it to k, and set FE c The feature map is obtained after processing with the attention mechanism by setting it to v.

[0191] Preferably, in the main decoder, the decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Multi-scale fusion processing feature maps Obtain feature map

[0192] Preferably, in the auxiliary decoder, the decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. ResGroup processing feature map Obtain feature map

[0193] Preferably, in a hybrid encoder, the hybrid encoded feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M f3 ; 1×1 convolution to process feature maps Obtain matrix M f4 M f3 With M f4 After multiplication, the sigmoid activation function is applied to obtain the feature map σD. f ;FD f =σD f ×M f4 +(1-σD f Mf3 ; feature map Set to q, and set the feature map Set it to k, and set FD f The feature map is obtained after processing with the attention mechanism by setting it to v.

[0194] More preferably, let the depth of the multi-scale fusion processing feature map be N, and in the primary processor, process the image to be enhanced M0 to obtain the feature map. and feature map The method can be as follows:

[0195] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0196] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0197] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Sub-image The dimensions are the same, and the cropping order of the sub-images can be determined by following the order from left to right and top to bottom; for each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map.

[0198] For images Perform 3×3 convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map consistent with the feature map size. With the same dimensions, feature maps are obtained.

[0199] In actual processing, a kernel size of 3, a stride of 1, and padding of 1 can be used for each sub-image. Perform 3×3 convolution to obtain the feature map. Let the kernel size be 3, the stride be 2, and the padding be 1 for a pair of images. Perform 3×3 convolution to obtain the feature map. This process yields the feature map. and feature map Since the length and width dimensions are the same, the corresponding feature maps obtained from subsequent processing can be used for hybrid encoding.

[0200] More preferably, let the depth of the multi-scale fusion processing feature map be N, and in the primary processor, process the image to be enhanced M0 to obtain the feature map. and feature map The method can be the following steps (see Figure 2 ):

[0201] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0202] If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced;

[0203] Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Sub-image The dimensions are the same; each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are concatted and merged at the channel level according to the cropping order to obtain the feature map.

[0204] For images Perform convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map equal to the feature map size. The feature maps are of the same size, and then multi-scale fusion processing is performed on the obtained feature maps to obtain the final feature map.

[0205] Hybrid coding feature map and feature map For feature map

[0206] In actual processing, a kernel size of 3, a stride of 1, and padding of 1 can be used for each sub-image. Perform 3×3 convolution to obtain the feature map. Let the kernel size be 3, the stride be 2, and the padding be 1 for a pair of images. Perform a 3×3 convolution. This process yields the feature map. and feature map Since the length and width dimensions are the same, the corresponding feature maps obtained from subsequent processing can be used for hybrid encoding.

[0207] Furthermore, if supplementary blocks are used in processing the image M0 to be enhanced, then in the primary processor, the feature map is processed. The enhanced image M0′ includes the following steps:

[0208] S711, Down-channel processing feature map Obtain the image MD;

[0209] S712. Cropping the supplementary block in image MD yields the enhanced image M0′.

[0210] In step 1, if no supplementary block is used during the processing of the image M0 to be enhanced, then the feature map is processed in the primary processor. The enhanced image M0′ includes the following steps:

[0211] S711, Down-channel processing feature map The enhanced image M0′ is obtained.

[0212] In this invention, the supplementary block is a block of the same color, which can be a pure white block, a pure black block, or other solid color blocks.

[0213] The image processing program of Embodiment 1 or the method for constructing the image processing device of Embodiment 3 includes the following steps:

[0214] The steps of creating training and testing sets, wherein each sample in the training and testing sets includes an image to be enhanced M0 and a label image M0″;

[0215] Steps for establishing an image processing simulation device;

[0216] The step of training the parameters of the image processing simulation device involves using the training set to train the image processing simulation device and obtaining the application parameters of the image processing simulation device.

[0217] The steps for testing the parameters of the image processing simulation device include setting the parameters of the image processing simulation device to the application parameters, testing the image processing simulation device using the test set, and if the error between the enhanced image output by the image processing simulation device and the corresponding label image meets the requirements, then an image processing program or image processing device is built according to the image processing simulation device and the application parameters. In the image processing device, the main encoder, auxiliary encoder, main decoder, auxiliary decoder, dual-channel mixer, and primary processor can be constructed using mechanical structures, circuit structures, optical paths, or a combination thereof.

[0218] Preferably, in the steps of creating the training and test sets, high-dynamic-range, large-size original DR images are acquired from the detector. In this embodiment, the DR images of the bolster and side frame of railway castings are scanned, with a resolution of 5732×2333. The DR images are cropped to 556×583, and the excess parts are discarded. Forty images to be enhanced are obtained. Users can change this according to the actual image size. After processing the obtained images to be enhanced using traditional image enhancement methods such as histogram equalization, HDR, and window width and level adjustment, the best enhanced image is selected as the label of this image to be enhanced, i.e., the label image. The image to be enhanced and the label image are paired and randomly cropped to integer multiples of 32 (544×576 in this embodiment). After normalization, they can be used as samples for the training set. The size of the images to be enhanced and the label images in the test set is 5440×2304.

[0219] Preferably, in the step of establishing the image processing simulation device, a loss function is also established, and the loss function is:

[0220]

[0221] Among them, I i (i = 1, 2, ..., M) is To perform label image Feature maps obtained by downsampling the size; For I i and The L1 norm loss, for for The image value, y(h,w), is I. i Image values; For I i and The perceived loss, specifically, is to reduce I i Input the feature map extracted by the VGG network and The L2 norm loss between feature maps extracted by the input VGG network is used for... To be Input the feature map extracted by the VGG network, φ j (y) represents I i Input the feature map extracted by the VGG network; Y is For matching The label image; For Y and The L1 norm loss, for for The image value of y is y(h,w), and the image value of Y is y(h,w). For Y and The perceptual loss is specifically calculated by inputting Y into the feature map extracted by the VGG network and then... The L2 norm loss between feature maps extracted by the input VGG network is used for... To be Input the feature map extracted by the VGG network, φ j (y) represents the feature map extracted by inputting Y into the VGG network; λ a , λ p , λ m All are weighting coefficients for the loss; This represents the total loss.

[0222] In the primary processor, the image M0 to be enhanced is processed to obtain a feature map. and feature map In the method, sub-image This is how it was obtained: Design sub-image The number of pixels in both the length and width directions is 2. M+N-1 Integer multiples, sliding cut feature map with fixed step size and fixed ratio A set of sub-images was obtained Therefore, matching The label images can also be processed in the same way: by sliding and cropping the label images according to the same fixed ratio and fixed step size, a set of sub-label images can be obtained. The sub-label images can then be matched with their corresponding labels in the order they were generated. Matching.

[0223] In this embodiment, λ a =0.5, λ p =λ m =1.

[0224] VGG networks are a common type of network. In Example 4, a pre-trained VGG16 network is specifically used. Its structure can be found in the convolutional neural network mentioned in the paper "Very Deep Convolutional NetWorks for Large-Scale Image Recognition" (authors: Karen Simonyan and Andrew Zisserman).

[0225] Example 4: A method for constructing an image processing program, comprising the following steps:

[0226] Step 1: Acquire high-dynamic, large-size DR original images from the detector. This invention scans DR images of the bolster and side frame of railway castings, with a resolution of 5732×2333. The DR images are cropped to 556×583, resulting in forty smaller images (40 original images). Users can adjust this according to the actual image size.

[0227] Step 2: Apply three traditional image enhancement methods—histogram equalization, HDR, and window width / level adjustment—to the obtained original image, and adjust the parameters accordingly. Select the best enhanced image as the label for the original image. Based on this, construct a neural network training dataset and divide it into training set T. train and test set T test ;

[0228] Step 3: Randomly cut the training set and test set data into multiples of 32, i.e., 544×576, and then perform normalization and image enhancement operations such as image flipping and adding noise to improve generalization. Finally, feed them into the dual-branch multi-scale fusion neural network S(·). Figure 1 A dual-branch multi-scale fusion neural network S(·) is shown.

[0229] Step 4: The dual-branch backbone of the neural network S(·) adopts the classic U-net structure, which consists of an encoder and a decoder. A skip connection is used between the encoder and the decoder. In this invention, an attention fusion module is designed at the fusion point of the two branches to reduce the semantic gap.

[0230] To better enhance the characteristics of high dynamic range (DR) images and achieve large-size inference, the multi-scale fusion network structure of this invention is specifically as follows:

[0231] This module provides high dynamic range features obtained in the encoder and decoder. Figure 1First, two downsampling operations of different sizes are performed to obtain three feature maps of different sizes. These are then fed into a residual neural network structure for learning. Finally, the two smaller sizes are interpolated back to the original size to obtain three feature maps of the same size. Figure 2 Afterwards, global average pooling is performed on each feature and then passed through a multilayer perceptron structure to obtain the feature. Figure 3 Then, the weights are fused and combined with the features. Figure 1 Adding them together yields the final new feature. Figure 4 Finally, the new features Figure 4 It is then fed into the encoder or decoder of the next layer.

[0232] Step 5: In order to obtain better DR enhancement results, this invention creates an attention fusion module, which aims to reduce the semantic gap in feature fusion between the two paths, so as to ensure that the network can better extract features of large 4K images;

[0233] Specifically, this module is placed at the point where the two branches of the network fuse and interact. If a simple fusion were performed at this stage, a semantic gap would result because the two branches learn different features. Therefore, this module effectively reduces the semantic gap and maximizes the interaction and fusion of features from both paths. Specifically, the feature maps Fa and Fm extracted from the two branches are first multiplied by a convolutional kernel and then subjected to a sigmoid operation to fuse pixel-level information, resulting in Fout. We set the fused feature map to v, Fm to q, and Fa to k, and then execute a classic attention mechanism to obtain the final output. Because the initial fusion may still leave some semantic gaps, we want the main branch to focus on learning the main network features from the auxiliary branch and discard some irrelevant and unimportant information.

[0234] Step Six: After the input data passes through the aforementioned neural network S(·), the main branch will obtain three different sizes of network outputs. First, the obtained labels are downsampled once at half size and once at quarter size to obtain three different sizes of comparison labels. This invention designs two main loss terms: a main path and an auxiliary path. The main path's three loss terms are the three sizes of feature maps I output by the network. i (i = 1, 2, ..., M) and corresponding comparison labels for the three sizes The loss term is based on the first norm, while the other three loss terms are calculated by taking the three feature maps of different sizes output by the network. i (i = 1, 2, ..., M) and corresponding comparison labels for the three sizes After feeding all the images into a pre-trained VGG16 network to obtain their feature maps, a norm-1 loss is applied to each. The auxiliary path outputs a set of predicted outputs Y for each sub-image, along with the corresponding contrast label for each sub-image. The norm-1 loss and perception loss are calculated as the losses for the auxiliary path. Here, λ is set... a =0.5, λ p =λ m =1.

[0235] Step 7: We build a neural network in PyTorch. The parameters of the convolutional layers are initialized using a normal distribution with a mean of 0 and a standard deviation of 0.02. The weights of the remaining layers are randomly initialized. The other hyperparameters can be set reasonably within a certain range. For example, we set a batch size of 4 in the 24GB VRAM of an NVIDIA RTX 4090 and trained for 150 iterations. We use the Adam optimizer and a cosine annealing learning rate adjustment strategy, setting each iteration to 50 epochs with an initial learning rate of 0.0002. Finally, we use the backpropagation algorithm to train the DR image enhancement neural network model, thus obtaining the DR image enhancement model.

[0236] Figure 6 A DR image to be enhanced is shown. Figure 7 To process the image using the image processing program obtained in this embodiment. Figure 6 The resulting label image.

[0237] Step 8: When deploying the trained model, since the image width and height required by the network design are multiples of 32, if the actual inference DR image is not a multiple of 32, its height and width can be resized to the nearest integer multiple of 32, and then the final enhanced DR image can be obtained through the trained neural network S(·).

[0238] This invention enhances high dynamic range (DR) images by designing a multi-scale fusion module. This not only effectively extracts feature information from the original DR image but also enables efficient inference when dealing with image inputs of varying sizes. Furthermore, this invention strengthens feature extraction by designing an attention fusion module. Since auxiliary branches capture small-scale details of the image, direct fusion with features captured by the main path can create a semantic gap. Ablation experiments demonstrate the effectiveness of this module.

[0239] The present invention has been described in detail above with reference to the accompanying drawings and embodiments. It should be understood that it is impossible to exhaustively describe all possible implementations in practice; the inventive concept of the present invention is illustrated to the extent possible through examples. Without departing from the inventive concept of the present invention and without any creative effort, any specific embodiments formed by selecting and combining technical features in the above embodiments, experimentally changing specific parameters, or conventionally replacing the disclosed technical means of the present invention using existing technology should be considered as implicit disclosures of the present invention.

Claims

1. An image processing method, wherein pointer i = 1, characterized in that, Includes the following steps: Step 1: Process the image M0 to be enhanced to obtain the feature map. and feature map Step 2: Encode feature maps For feature map Encoded feature map For feature map The feature map The corresponding image size is the feature map. The feature map is twice the size of the corresponding image. The corresponding image size is the feature map. Twice the size of the corresponding image, encoding feature map For feature map The method differs from encoding feature maps For feature map Methods; Step 3: Hybrid Encoding Feature Map and feature map For feature map Step 4. If i < M, then set i = i + 1, and return to Step 2 to execute the loop from Step 2 to Step 4; if i = M, then set the feature map as the feature map Feature map as the feature map to proceed to the next step. Step 5: Decode the feature map For feature map Decoding feature map For feature map The feature map The corresponding image size is the feature map. The feature map is twice the size of the corresponding image. The corresponding image size is the feature map. Twice the size of the corresponding image; Step 6: Hybrid Encoding Feature Map and feature map For feature map Step 7: If i > 1, then let i = i - 1, return to step 5 and execute the loop from step 5 to step 7; if i = 1, then process the feature map. The enhanced image M0′; Here, N is the depth of the multi-scale fusion feature map, and the image M0 to be enhanced is processed to obtain the feature map. and feature map The method is: Query the number of pixels in the length and width directions of the image M0 to be enhanced; If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced; Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained For each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map. For images Perform 3×3 convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map consistent with the feature map size. With the same dimensions, feature maps are obtained. Alternatively, the image M0 to be enhanced can be processed to obtain a feature map. and feature map The method is: Query the number of pixels in the length and width directions of the image M0 to be enhanced; If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are not both 2 M+N-1 If the number is an integer multiple, then a padding block is used at the end in the corresponding direction to padded the pixel count to 2. M+N-1 Integer multiples, let the image The image is after resizing; if the number of pixels in both the length and width directions of the image M0 to be enhanced is 2. M+N-1 If it is an integer multiple, then let the image... Image M0 to be enhanced; Design sub-images The number of pixels in both the length and width directions is 2. M+N-1 Image cropping by sliding in fixed increments at integer multiples. A set of sub-images was obtained Each sub-image Perform convolution processing to obtain feature maps. Based on the feature maps of all sub-images The sub-images are merged at the channel level according to the cropping order to obtain the feature map. For images Perform convolution processing, and adjust the stride of the convolution to make the size of the resulting feature map equal to the feature map size. The feature maps are of the same size, and then multi-scale fusion processing is performed on the obtained feature maps to obtain the final feature map. Hybrid coding feature map and feature map For feature map 2. The image processing method as described in claim 1, characterized in that, Hybrid coding feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M i1 ; 1×1 convolution to process feature maps Obtain matrix M i2 M i1 With M i2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. i ;FE i =σE i ×M i2 +(1-σE i M i1 ; feature map Set to q, and set the feature map Set it to k, and set FE i The feature map is obtained after processing with the attention mechanism by setting it to v. Hybrid coding feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M i3 ; 1×1 convolution to process feature maps Obtain matrix M i4 M i3 With M i4 After multiplication, the sigmoid activation function is applied to obtain the feature map σD. i ;FD i =σD i ×M i4 +(1-σD i M i3 ; feature map Set to q, and set the feature map Set it to k, and set FD i The feature map is obtained after processing with the attention mechanism by setting it to v.

3. The image processing method as described in claim 1, characterized in that, Encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Multi-scale fusion processing feature maps Obtain feature map Encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Residual group convolution processing feature map Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Multi-scale fusion processing feature maps Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Residual group convolution processing feature map Obtain feature map 4. The image processing method as described in claim 1, characterized in that, If supplementary blocks were used in step 1 during the processing of the image M0 to be enhanced, then in step 7, the feature map is processed. The enhanced image M0′ includes the following steps: S711, Down-channel processing feature map Obtain the image MD; S712. Crop the supplementary block in image MD to obtain the enhanced image M0′; If, in step 1, no supplementary blocks are used during the processing of the image M0 to be enhanced, then in step 7, the feature map is processed. The enhanced image M0′ includes the following steps: S711, Down-channel processing feature map The enhanced image M0′ is obtained.

5. A computer-readable medium storing an image processing program, characterized in that, The image processing program is loaded by the processor to execute the image processing method as described in any one of claims 1-4.

6. An electronic device comprising a computer-readable medium storing an image processing program and a processor, characterized in that, The image processing program is loaded by the processor to execute the image processing method as described in any one of claims 1-4.

7. An image processing apparatus constructed using the image processing method as described in claim 1, for processing an image M0 to be enhanced into an enhanced image M0′, characterized in that, include: The main encoder is used to encode feature maps. For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image; Auxiliary encoder, used to encode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image, encoding feature map For feature map The method differs from encoding feature maps For feature map Methods; Main decoder, used to decode feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image; Auxiliary decoder, used for decoding feature maps For feature map The feature map The corresponding image size is the feature map. Twice the size of the corresponding image; Dual-channel mixer for mixing and encoding feature maps and feature map For feature map and for hybrid encoding feature maps and feature map For feature map as well as The primary processor processes the image M0 to be enhanced to obtain feature maps. and feature map and processing feature maps The enhanced image M0′.

8. The image processing apparatus as described in claim 7, characterized in that, Hybrid coding feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M i1 ; 1×1 convolution to process feature maps Obtain matrix M i2 M i1 With M i2 After multiplication, the sigmoid activation function is applied to obtain the feature map σE. i ;FE i =σE i ×M i2 +(1-σE i M i1 ; feature map Set to q, and set the feature map Set it to k, and set FE i The feature map is obtained after processing with the attention mechanism by setting it to v. Hybrid coding feature map and feature map For feature map The method is: 1×1 convolution to process feature maps Obtain matrix M i3 ; 1×1 convolution to process feature maps Obtain matrix M i4 M i3 With M i4 After multiplication, the sigmoid activation function is applied to obtain the feature map σD. i ;FD i =σD i ×M i4 +(1-σD i M i3 ; feature map Set to q, and set the feature map Set it to k, and set FD i The feature map is obtained after processing with the attention mechanism by setting it to v.

9. The image processing apparatus as described in claim 7, characterized in that, Encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Multi-scale fusion processing feature maps Obtain feature map Encoded feature map For feature map The method is: shrink the feature map. To reduce the dimensionality of image features, feature maps are obtained. Residual group convolution processing feature map Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Multi-scale fusion processing feature maps Obtain feature map Decoding feature map For feature map The method is: expand the processing feature map To increase the feature dimension of the image, a feature map is obtained. Residual group convolution processing feature map Obtain feature map 10. The image processing apparatus as claimed in claim 7, characterized in that, If supplementary blocks are used during the processing of the image M0 to be enhanced, then during the processing of the feature map... The enhanced image M0′ includes the following steps: S711, Down-channel processing feature map Obtain the image MD; S712. Crop the supplementary block in image MD to obtain the enhanced image M0′; If no supplementary blocks are used during the processing of the image M0 to be enhanced, then during the processing of the feature map... The enhanced image M0′ includes the following steps: S711, Down-channel processing feature map The enhanced image M0′ is obtained.

11. A method for constructing an image processing apparatus as described in any one of claims 7-10, characterized in that, Includes the following steps: The steps of creating training and testing sets, wherein each sample in the training and testing sets includes an image to be enhanced M0 and a label image M0″; Steps for establishing an image processing simulation device; The step of training the parameters of the image processing simulation device involves using the training set to train the image processing simulation device and obtaining the application parameters of the image processing simulation device. The steps for testing the parameters of the image processing simulation device are as follows: the parameters of the image processing simulation device are set to the application parameters; the image processing simulation device is tested using the test set; if the error between the enhanced image output by the image processing simulation device and the corresponding label image meets the requirements, then the image processing device is constructed according to the image processing simulation device and the application parameters.

12. The method for constructing the image processing device as described in claim 11, characterized in that, In the step of establishing the image processing simulation device, a loss function is also established, which is: Among them, I i (i = 1, 2, ..., M) is To perform label image Feature maps obtained by downsampling the size; For I i and The L1 norm loss, for for The image value, y(h, w), is I. i Image values; For I i and The perceived loss, specifically, is to reduce I i Input the feature map extracted by the VGG network and The L2 norm loss between feature maps extracted by the input VGG network is used for... To be Input the feature map extracted by the VGG network, φ j (y) represents I i Input the feature map extracted by the VGG network; Y is For matching The label image; For Y and The L1 norm loss, for for The image value of y is y(h, w); For Y and The perceptual loss is specifically calculated by inputting Y into the feature map extracted by the VGG network and then... The L2 norm loss between feature maps extracted by the input VGG network is used for... To be Input the feature map extracted by the VGG network, φ j (y) represents the feature map extracted by inputting Y into the VGG network; λ a , λ p , λ m All are weighting coefficients for the loss; This represents the total loss.

Citation Information

Patent Citations

  • Human body detection method in low-illumination environment based on image enhancement, electronic equipment and storage medium

    CN114708615A

  • Construction method of large-size image enhancement model, image processing method, computer readable medium and electronic equipment

    CN118115373A