Image processing method, computer readable medium, device and construction method

By encoding and decoding feature maps, combined with multi-scale fusion and attention mechanism image processing methods, the problem that large-size DR images in the prior art is difficult to improve contrast and visual effects at the same time, achieving better local contrast enhancement and defect recognition effects.

CN120013786AActive Publication Date: 2025-05-16CHONGQING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510156951.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The prior art is difficult to simultaneously improve the contrast, visual effect and image information entropy of large-sized DR images in X-ray imaging systems, especially in terms of local contrast enhancement, which affects defect recognition.

Method used

An image processing method is adopted to enhance the local contrast and detail texture of the image by encoding and decoding feature maps, mixing encoding and multi-scale fusion processing, and combining attention mechanisms.

Benefits of technology

It effectively improves the contrast, visual effect and information entropy of the image, especially in the enhancement of local contrast, which can better identify defects of large-sized components and improve image processing speed and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013786A_ABST
    Figure CN120013786A_ABST
Patent Text Reader

Abstract

An image processing method, a computer readable medium, an electronic device, an image processing device and a construction method thereof, which make i = 1, comprising: S1, processing an image to be enhanced to obtain # imgabs0 # and # imgabs1 # S2, a code # imgabs2 # as # imgabs3 #, a code # imgabs4 # as # imgabs5 # S3, a mixed code # imgabs6 # and a mixed code # imgabs7 # as # imgabs8 # S4, if ilt; if M, making i = i + 1, and returning to S2; if i is equal to M, the # imgabs9 is set as the # imgabs10, the # imgabs11 is set as the # imgabs12, and the next step is carried out; s5, decoding the # imgabs13 # as the # imgabs14 #, decoding the # imgabs15 # as the # imgabs16 # S6, performing hybrid coding on the # imgabs17 # and the # imgabs18 # as the # imgabs19 # S7, and if so, performing hybrid coding on the # imgabs17 # and the # imgabs18 # as the # imgabs19 # S7; if 1, i = i-1, and returning to S5; and if i is equal to 1, processing # imgabs20 # as an enhanced image. The low-light display quality of the super-definition image can be improved by using the super-definition image display device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image processing method, a computer-readable medium, an electronic device, an image processing device and a construction method thereof. Background Art

[0002] Digital X-ray radiography (DR) technology receives X-ray energy through a flat-panel detector and directly converts it into an electrical signal, which is then converted into a digital image by a computer. DR technology imaging has the advantages of high spatial resolution, large amount of information, and large dynamic range. More importantly, it has a fast photography speed and can quickly and effectively detect surface and internal defects of castings, realizing real-time imaging detection.

[0003] Traditional image enhancement methods often achieve a balance in enhancing image contrast, visual effects, and image information entropy, and cannot improve all of the above indicators at the same time. Researchers often use specific algorithms for enhancement according to different enhancement requirements. Therefore, the application of traditional image enhancement methods requires manual multi-parameter adjustment, and traditional enhancement methods are time-consuming when facing large-size DR images. Today's deep learning networks rarely have DR image enhancement tasks, and many deep learning networks today are helpless with large-size DR images. Scattering, electrical noise, and uneven thickness of the casting itself in the X-ray imaging system will cause DR image edge information to be blurred, grayscale distribution to be uneven, and details to be unclear, and the generated DR images are large-size DR images with a high dynamic range.

[0004] Patent document CN118115373A discloses a method for constructing a large-size image enhancement model, an image processing method, a computer-readable medium, and an electronic device. The method for constructing a large-size image enhancement model includes the steps of making a training set and a test set, establishing a large-size image enhancement model, training the parameters of the large-size image enhancement model, and testing the parameters of the large-size image enhancement model. The large-size image enhancement model includes an encoder, a decoder, a multi-scale fusion module, and a brightness adjustment module. Although this solution can process large-size images with a high dynamic range to be enhanced, so that their contrast, visual effects, and image information entropy can be enhanced at the same time, it is insufficient in enhancing the local contrast of the subtle parts of large-size parts, which is not conducive to identifying defects in large-size parts. Summary of the invention

[0005] The object of the present invention is to provide an image processing method, a computer-readable medium, an electronic device, an image processing device and a construction method thereof to improve the quality of an image.

[0006] The present invention is achieved in that:

[0007] An image processing method, let the pointer i = 1, including the following steps:

[0008] Step 1, process the image M0 to be enhanced to obtain a feature map and a feature map

[0009] Step 2, encode the feature map into a feature map Encode the feature map into a feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map The method of encoding the feature map into a feature map is different from the method of encoding the feature map into a feature map ;

[0010] Step 3, mix and encode the feature map and the feature map into a feature map

[0011] Step 4, if i < M, then let i = i + 1, return to step 2 and execute the loop of steps 2 to 4; if i = M, then let the feature map be the feature map Feature map be the feature map for the next step;

[0012] Step 5, decode the feature map into a feature map Decode the feature map into a feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map The corresponding image size of the feature map is twice the corresponding image size of the feature map ;

[0013] Step 6, mix and encode the feature map and the feature map into a feature map

[0014] Step 7, if i > 1, then let i = i - 1, return to step 5 and execute the loop of steps 5 to 7; if i = 1, then process the feature map is the enhanced image M0′.

[0015] Preferably, the hybrid coding feature map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M i1 ; 1×1 convolution processing feature map Get the matrix M i2 ;M i1 With M i2 After multiplication, the sigmoid activation function is used to obtain the feature map σE i ; FE i =σE i ×M i2 +(1-σE i )M i1 ; The feature map Set to q, the feature map Set to k, and FE i Set to v, and get the feature map after processing according to the attention mechanism Hybrid coding feature map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M i3 ; 1×1 convolution processing feature map Get the matrix M i4 ;M i3 With M i4 After multiplication, the sigmoid activation function is used to obtain the feature map σD i FD i =σD i ×M i4 +(1-σD i )M i3 ; The feature map Set to q, the feature map Set to k, and set FD i Set to v, and get the feature map after processing according to the attention mechanism

[0016] Preferably, the encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map Encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Residual group convolution processing feature map Get feature map Decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map Decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Residual group convolution processing feature map Get feature map

[0017] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, the image to be enhanced M0 is processed to obtain the feature map and feature map The method is:

[0018] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0019] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0020] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images For each sub-image Perform convolution processing to obtain the feature map According to the feature maps of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map

[0021] For images Perform 3×3 convolution processing and set the convolution step size to make the size of the feature map equal to the feature map The size is the same as that of

[0022] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, the image to be enhanced M0 is processed to obtain the feature map and feature map The method is:

[0023] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0024] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0025] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Each sub-image Perform convolution processing to obtain the feature map According to the feature maps of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map

[0026] For images Perform convolution processing and set the step size of the convolution to make the size of the feature map equal to the feature map The size of the feature map is the same as that of the feature map, and then the feature map is obtained by multi-scale fusion processing.

[0027] Hybrid Encoding Feature Map and feature map The feature map

[0028] Further preferably, in step 1, if the supplementary block is used in the process of processing the image to be enhanced M0, then in step 7, the feature map is processed The enhanced image M0′ includes the following steps:

[0029] S711, down-channel processing characteristic diagram Get image MD;

[0030] S712, cutting off the supplementary block in the image MD to obtain the enhanced image M0'.

[0031] In step 1, if no supplementary block is used in the process of processing the image to be enhanced M0, then in step 7, the feature map is processed The enhanced image M0′ includes the following steps:

[0032] S711, down-channel processing characteristic diagram The enhanced image M0′ is obtained.

[0033] A computer-readable medium storing an image processing program, wherein the image processing program is used to execute the aforementioned image processing method after being loaded by a processor.

[0034] An electronic device comprises a computer-readable medium storing an image processing program and a processor. The image processing program is loaded by the processor and used to execute the above-mentioned image processing method.

[0035] An image processing device, for processing an image to be enhanced M0 into an enhanced image M0′, comprising:

[0036] Main encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0037] Auxiliary encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice that of the encoded feature map The feature map The method is different from encoding feature maps The feature map Methods;

[0038] Main decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0039] Auxiliary decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0040] Two-way mixer for mixed coding feature maps and feature map The feature map And for mixed encoding feature maps and feature map The feature map

[0041] The primary processor is used to process the image M0 to be enhanced to obtain the feature map and feature map And process feature maps is the enhanced image M0′;

[0042] Among them, a=1,2,…,M-1;b=1,2,…,M-1;c=1,2,…,M;d=2,3,…,M;e=2,3,…,M;f=1,2,…,M;M≥3;characteristic graph The feature map Feature Map The feature map

[0043] Preferably, the hybrid coding feature map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M c1 ; 1×1 convolution processing feature map Get the matrix M c2 ;M c1 With M c2 After multiplication, the sigmoid activation function is used to obtain the feature map σE c ; FE c =σE c ×M c2 +(1-σE c )M c1 ; The feature map Set to q, the feature map Set to k, and FE c Set to v, and get the feature map after processing according to the attention mechanism

[0044] Hybrid Encoding Feature Map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M f3 ; 1×1 convolution processing feature map Get the matrix M f4 ;M f3 With M f4 After multiplication, the sigmoid activation function is used to obtain the feature map σD f FD f =σD f ×M f4 +(1-σD f )M f3 ; The feature map Set to q, the feature map Set to k, and set FD f Set to v, and get the feature map after processing according to the attention mechanism

[0045] Preferably, the encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map Encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Residual group convolution processing feature map Get feature map Decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map Decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Residual group convolution processing feature map Get feature map

[0046] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, in the primary processor, the image to be enhanced M0 is processed to obtain a feature map and feature map The method is:

[0047] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0048] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0049] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images For each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map

[0050] For images Perform 3×3 convolution processing and set the convolution step size to make the size of the feature map equal to the feature map The size of is the same as that of

[0051] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, the image to be enhanced M0 is processed to obtain the feature map and feature map The method is:

[0052] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0053] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0054] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map

[0055] For images Perform convolution processing and set the step size of the convolution to make the size of the feature map equal to the feature map The size of the feature map is the same as that of the feature map, and then the feature map is obtained by multi-scale fusion processing.

[0056] Hybrid coding feature map and feature map The feature map

[0057] Still further preferably, if a supplementary block is used in the process of processing the image to be enhanced M0, then in the primary processor, the feature map is processed The enhanced image M0′ includes the following steps:

[0058] S711, down-channel processing characteristic diagram Get image MD;

[0059] S712, cutting off the supplementary block in the image MD to obtain the enhanced image M0'.

[0060] In step 1, if no supplementary block is used in the process of processing the image to be enhanced M0, then in the preliminary processor, the feature map is processed The enhanced image M0′ includes the following steps:

[0061] S711, down-channel processing characteristic diagram The enhanced image M0′ is obtained.

[0062] The method for constructing the aforementioned image processing device comprises the following steps:

[0063] The step of preparing a training set and a test set, wherein each sample of the training set and the test set includes an image to be enhanced M0 and a label image M0″;

[0064] Steps to build an image processing simulation device;

[0065] a step of training parameters of the image processing simulation device, using the training set to train the image processing simulation device to obtain application parameters of the image processing simulation device;

[0066] The step of testing the parameters of the image processing simulation device, setting the parameters of the image processing simulation device as the application parameters, using the test set to test the image processing simulation device, if the error after comparison between the enhanced image output by the image processing simulation device and the corresponding label image meets the requirements, then building the image processing device according to the image processing simulation device and the application parameters.

[0067] Preferably, in the step of establishing the image processing simulation device, a loss function is also established, and the loss function is:

[0068]

[0069]

[0070] Among them, I i (i=1,2,…,M) is To label the image The feature map obtained by downsampling the size; For I i and The L1 norm loss for for The image value of y(h,w) is I i The image value of For I i and The perceptual loss is specifically to convert I i Input the feature map extracted by the VGG network and The L2 norm loss between the feature maps extracted by the input VGG network is For the general Input the feature map extracted by the VGG network, φ j (y) is I i Input the feature map extracted by the VGG network; Y is To match Label image of Y and The L1 norm loss for for y(h,w) is the image value of Y; Y and The perceptual loss is as follows: Y is input into the feature map extracted by the VGG network and The L2 norm loss between the feature maps extracted by the input VGG network is For the general Input the feature map extracted by the VGG network, φ j (y) is the feature map extracted by inputting Y into the VGG network; λ a , p , m are the weight coefficients of loss; is the total loss.

[0071] The beneficial effects of the present invention include:

[0072] 1. The image processing method of the present invention can process large-scale high dynamic range images to be enhanced, so that the contrast, visual effect and image information entropy can be enhanced at the same time, and the local contrast of subtle features can be better enhanced, the capture of image detail texture can be improved, and the contrast of high-frequency detail areas of the image can be enhanced, so that the important parts of the image can be highlighted. Taking the DR image recognition of the bolster and side frame of railway castings as an example, the contrast of fine cracks on the side frame can be locally enhanced, which is conducive to identifying defects in large-scale parts. The main and auxiliary roads both adopt the U-net architecture, and the feature map obtained from each layer of the auxiliary road is mixed and encoded with the feature map obtained from each layer of the main road as the input feature map of the next layer of the main road. In this way, the main road can focus on learning the main feature map from the auxiliary, thereby improving the quality of the image, such as improving the display effect of low-light images. When this architecture processes ultra-high-definition images, it can take into account both hardware consumption and processing time, and has significant improvements in image processing speed and image quality, and can be applied to the processing of large-size ultra-high-definition images.

[0073] 2. The image processing method of the present invention adopts an attention mechanism to mix and encode the feature map obtained from each layer of the auxiliary path with the feature map obtained from each layer of the main path, and use them as the input feature map of the next layer of the main path. The main path can focus on learning the main features of the image from the auxiliary path and discard some irrelevant and unimportant information, thereby expanding the scope of application of the image processing method.

[0074] 3. The image processing method of the present invention uses residual group convolution to process images in the auxiliary path, which can reduce the amount of calculation and the number of parameters, and enhance the image representation ability by increasing the number of groups. Multi-scale fusion is used in the main path to process images, so that the network can still perform inference efficiently even when facing image inputs of different sizes, without losing detailed features, better extracting the characteristic parts of the image, and enhancing the display effect of the features.

[0075] 4. The computer-readable medium, electronic device, and image processing device storing an image processing program of the present invention can process high dynamic range images to be enhanced, such as DR images, so that their contrast, visual effects, and image information entropy are enhanced at the same time. In addition, they can also be easily coordinated with automated equipment to realize automated enhancement processing of high dynamic range images to be enhanced, thereby reducing the requirements for user experience.

[0076] 5. The method for constructing an image processing device of the present invention can facilitate the design verification process of the image processing device and reduce the time required to obtain the image processing device. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 A simplified structural diagram of an image processing method.

[0078] Figure 2 A structural diagram of an image processing method.

[0079] Figure 3 It is a method for multi-scale fusion processing of images.

[0080] Figure 4 It is a method of processing images using residual group convolution (ResGroup).

[0081] Figure 5 It is a method of hybrid encoding of two-way images.

[0082] Figure 6 is a DR image to be enhanced.

[0083] Figure 7 The image processing program constructed using Example 4 is used to process Figure 6 The processed image is output.

[0084] Figure 8 is a DR image to be enhanced.

[0085] Fig. 9 The image processing program constructed using Example 4 is used to process Figure 8 The processed image is output. DETAILED DESCRIPTION

[0086] The present invention is described in the form of embodiments in conjunction with the accompanying drawings to assist those skilled in the art in understanding and implementing the present invention. Unless otherwise specified, the following embodiments and the technical terms therein should not be understood without departing from the technical knowledge background of the technical field.

[0087] Resolution refers to the number of pixels on a screen or display device, usually composed of horizontal and vertical pixels. Early monitors had lower resolutions, such as CRT monitors with 1024×768. With the development of technology, the resolution has gradually increased, and different high-definition standards have emerged. The more common ones are HD (High Definition), FHD (Full High Definition), QHD (Quad High Definition), and UHD (Ultra High Definition).

[0088] Resolutions above 1280x720 and below 1920x1080 are generally referred to as HD; resolutions above 1920x1080 and below 2560x1440 are generally referred to as FHD; resolutions above 2560x1440 and below 3840x2160 are generally referred to as QHD; resolutions of 3840x2160 and above are generally referred to as UHD.

[0089] When photographing in low-light environments, the resulting low-light images usually show significant quality degradation due to exposure anomalies, including high noise levels, low contrast, low visibility, and extremely low information entropy, making it difficult or impossible for the human eye to perceive information. Traditional low-light image enhancement (LLIE) methods mainly rely on image priors and distribution mapping, such as segmentation iterative function-based methods, histogram equalization-based methods, homomorphic filtering-based methods, and model optimization-based methods. Such as segmentation iterative function-based methods, histogram equalization-based methods, homomorphic filtering-based methods, and model optimization-based methods. Low-light image enhancement methods based on deep learning use large-scale synthetic or real low-light enhancement datasets, which have significantly improved image processing speed and image quality, but the scope of application of image processing models is greatly limited. In order to expand the scope of application of image processing models, after the introduction of the Transformer architecture, processing ultra-high-definition images can easily lead to a rapid increase in computational load and GPU memory usage, and reduce model efficiency.

[0090] An image processing method, assuming that the pointer i=1, comprises the following steps:

[0091] Step 1: Process the image to be enhanced M0 to obtain the feature map and feature map

[0092] Step 2: Encode feature map The feature map Encoding feature map The feature map The feature map The corresponding image size is the feature map twice the corresponding image size, the feature map the corresponding image size is the feature map twice the corresponding image size, the encoded feature map is the feature map The method for is different from the method for encoding the feature map is the feature map ;

[0093] Step 3, mix the encoded feature map and the feature map is the feature map

[0094] Step 4, if i < M, then set i = i + 1, and return to Step 2 to execute the loop of Step 2 to Step 4; if i = M, then set the feature map is the feature map Feature map is the feature map Proceed to the next step;

[0095] Step 5, decode the feature map is the feature map Decode the feature map is the feature map The feature map The corresponding image size is the feature map twice the corresponding image size; the feature map The corresponding image size is the feature map twice the corresponding image size;

[0096] Step 6, mix the encoded feature map and the feature map is the feature map

[0097] Step 7, if i > 1, then set i = i - 1, and return to Step 5 to execute the loop of Step 5 to Step 7; if i = 1, then process the feature map as the enhanced image M0'.

[0098] In the present invention, the method for encoding the feature map is the feature map can be the following steps (see Figure 2 ):

[0099] S211. Shrink and process the feature map to reduce the image feature dimension, obtaining the feature map

[0100] S212. Perform multi-scale fusion processing on the feature map Get feature map

[0101] In the present invention, the encoding feature map The feature map The method may be such steps (see Figure 1 ):

[0102] S221, shrinkage processing feature map To reduce the image feature dimension, we get the feature map

[0103] S222, ResGroup processing feature map Get feature map

[0104] In the present invention, the depth of the multi-scale fusion processing feature map in S212 is assumed to be N, and the image to be enhanced M0 is processed to obtain the feature map and feature map The method can be steps like this:

[0105] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0106] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0107] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Sub-image The sizes are the same, and the order of cropping and obtaining the sub-images can be determined from left to right and from top to bottom; for each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, concat merge them at the channel level to obtain the feature map

[0108] For images Perform 3×3 convolution processing and set the convolution step size to make the size of the feature map equal to the feature map The size is the same as that of

[0109] In actual processing, the kernel size can be set to 3, the stride to 1, and the padding to 1 for each sub-image. Perform 3×3 convolution to obtain the feature map Let kernel size be 3, stride be 2 and padding be 1 for a pair of images Perform 3×3 convolution to obtain the feature map In this way, the feature map obtained and feature map The length and width of the image are the same, and the corresponding feature maps obtained by subsequent processing can be mixed encoded.

[0110] In the present invention, the depth of the multi-scale fusion processing feature map in S212 is assumed to be N, and the image to be enhanced M0 is processed to obtain the feature map and feature map The method can be steps like this:

[0111] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0112] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0113] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Sub-image are of the same size; each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, concat merge them at the channel level to obtain the feature map

[0114] For images Perform convolution processing and set the step size of the convolution to make the size of the feature map equal to the feature map The size of the feature map is the same as that of the feature map, and then the feature map is obtained by multi-scale fusion processing.

[0115] Hybrid coding feature map and feature map The feature map

[0116] In actual processing, the kernel size can be set to 3, the stride to 1, and the padding to 1 for each sub-image. Perform 3×3 convolution to obtain the feature map Let kernel size be 3, stride be 2 and padding be 1 for a pair of images Perform 3×3 convolution processing. In this way, the feature map obtained and feature map The length and width of the image are the same, and the corresponding feature maps obtained by subsequent processing can be mixed encoded.

[0117] In the present invention, the hybrid coding feature map and feature map The feature map The method may be such steps (see Figure 2 ):

[0118] S311, 1×1 convolution processing feature map Get the matrix M i1 ; 1×1 convolution processing feature map Get the matrix M i2 ;M i1 With M i2 After multiplication, the sigmoid activation function is used to obtain the feature map σE i ; FE i =σE i ×M i2 +(1-σE i )M i1 ;

[0119] S312, feature map Set to q, the feature map Set to k, and FE i Set to v, and get the feature map after processing according to the attention mechanism

[0120] In the present invention, the decoding feature map The feature map The method may be such steps (see Figure 2 ):

[0121] S511, extended processing feature map To increase the image feature dimension, we get the feature map

[0122] S512, multi-scale fusion processing feature map Get feature map

[0123] In the present invention, the decoding feature map The feature map The method can be steps like this:

[0124] S521, extended processing feature map To increase the image feature dimension, we get the feature map

[0125] S522, ResGroup processing feature map Get feature map

[0126] In the present invention, the hybrid coding feature map and feature map The feature map The method may be such steps (see Figure 2 ):

[0127] S611, 1×1 convolution processing feature map Get the matrix M i3 ; 1×1 convolution processing feature map Get the matrix M i4 ;M i3 With M i4 After multiplication, the sigmoid activation function is used to obtain the feature map σD i FD i =σD i ×M i4 +(1-σD i )M i3 ;

[0128] S612: feature map Set to q, the feature map Set to k, and set FD i Set to v, and get the feature map after processing according to the attention mechanism

[0129] In the present invention, in step 1, if the supplementary block is used in the process of processing the image to be enhanced M0, then in step 7, the feature map is processed The enhanced image M0′ includes the following steps:

[0130] S711, down-channel processing characteristic diagram Get image MD;

[0131] S712, cutting off the supplementary block in the image MD to obtain the enhanced image M0'.

[0132] In the present invention, in step 1, if the supplementary block is not used in the process of processing the image to be enhanced M0, then in step 7, the feature map is processed The enhanced image M0′ includes the following steps:

[0133] S711, down-channel processing characteristic diagram The enhanced image M0′ is obtained.

[0134] In the present invention, the supplementary block is a block of the same color, which may be a pure white block, a pure black block, or other pure color blocks.

[0135] Figure 2 An embodiment of the image processing method of the present invention is shown. In this embodiment, M=3, SAM is a method for multi-scale fusion processing of images, ResGroup is a method for processing images, and DAFM is a method for hybrid encoding of two-channel images.

[0136] In image processing, ResGroup is a specific network module structure, which is often used in deep learning models, especially in tasks such as image restoration and super-resolution reconstruction. For example, in Omni-Kernel Network (OKNet), ResGroup is the basic module of the network, which consists of multiple residual blocks (ResBlock), each of which contains two 3×3 convolutional layers with a nonlinear activation function GELU in the middle.

[0137] ResGroup can effectively extract deep features of images by stacking residual blocks, and enhance the expressiveness of features by residual learning, so as to better process details and structural information in images. In OKNet, ResGroup is used in the encoder and decoder stages. Through reasonable network structure design, model performance can be improved without significantly increasing computational overhead. ResGroup can be flexibly stacked and combined as needed to adapt to different image processing tasks, such as image dehazing, super-resolution reconstruction, etc.

[0138] Figure 3 An embodiment of a multi-scale fusion image processing method SAM is shown. In this embodiment, N=3. Feature map F r F is obtained by downsampling by 1 / 2 ↓r , and get F by downsampling by 1 / 4 ↓↓ r , respectively for F r 、F ↓ r 、F ↓↓ r Resnet structure learning is performed to extract features, and Y0, Y1′, and Y2′ are obtained. Then Y1′ is interpolated to obtain Y1 with the same image size as Y0; Y2′ is interpolated to obtain Y2 with the same image size as Y0. Y0, Y1, and Y2 are processed by global average pooling (GAP) and then multi-layer perceptron (MLP) to obtain the corresponding feature maps. The three feature maps are fused according to the weights to obtain the feature map F. r ′. Feature map F r and feature map F r ′ and add them together to get F out . F out Corresponding to the feature map of an image. After processing a feature map using the multi-scale fusion module, N multi-scale feature maps can be obtained.

[0139] Figure 4 A method for processing images with ResGroup is shown, which is used to improve the efficiency of image feature extraction and reduce the requirements for hardware specifications. For feature map x, let the times pointer i = 1, let feature map x0 be feature map x, and perform the following processing steps: Sa1, perform 3×3 convolution to extract preliminary features; Sa2, use GELU Gaussian error linear unit activation function for processing; Sa3, perform 3×3 convolution processing and combine with feature map x i-1 Perform residual connection to obtain feature map x i , let i=i+1; Sa4, if i=N (this time N is not the depth of the multi-scale fusion processing feature map in S212, but a newly defined parameter), then execute step Sa5; if i<N, then return to step Sa1 and execute the loop from step Sa1 to step Sa4; Sa5, for the feature map x N Perform 3×3 convolution to extract features; Sa6, use GELU Gaussian error linear unit activation function processing; Sa7, further refine the dynamic extraction of image features through the multi-scale fusion module SAM; Sa8, after 3×3 convolution processing and feature map x N Perform residual connection to obtain the final feature map of sub-image x.

[0140] For Figure 2 In the image processing method shown in FIG. 1 , an image is split into N sub-images in the branch as follows: Figure 1 The feature map shown First split it into N sub-images, and then perform multi-layer encoding and multi-layer decoding operations on each sub-image to obtain the sub-image feature map of the corresponding layer. Later Split into N sub-images, down-sample each feature map in the next layer, and perform subsequent corresponding operations to obtain the sub-image feature map of the corresponding layer.

[0141] Figure 5 A hybrid encoding method DAFM is shown to reduce the semantic gap between the two images. and feature map The feature map For example, the following steps are included: 1×1 convolution processing feature map Get the matrix M i1 ; 1×1 convolution processing feature map Get the matrix M i2 ;M i1 With M i2 After multiplication, the sigmoid activation function is used to obtain the feature map σE i , at this time σE i Represents a feature map that is closer to the real information, FE i =σE i ×M i2 +(1-σE i )M i1 . The feature map Set to q, the feature map Set to k, and FE i Set to v, and get the feature map after processing according to the attention mechanism

[0142] In the prior art, the U-net architecture includes an encoding side, a connector, and a decoding side. The encoding side includes a method for shrinking feature maps to reduce the feature dimension of the image, and the decoding side cooperates with the connector to include a method for expanding feature maps to increase the feature dimension of the image. Reducing the image feature dimension is equivalent to reducing the image feature dimension in the U-net architecture. In the present invention, the extended processing feature map To increase the image feature dimension, which is equivalent to increasing the image feature dimension in the U-net architecture, see Figure 1 , in the main path, the extended processing decoding side feature map Then through the connector and the encoding side feature map After connection, the decoding side feature map is obtained

[0143] By applying the image processing method of the present invention, a computer-readable medium storing an image processing program, an electronic device, and an image processing device can be obtained.

[0144] Embodiment 1: A computer-readable medium storing an image processing program, wherein the image processing program is loaded by a processor and used to execute the image processing method of the present invention.

[0145] The image processing program of the present invention is used to process the image to be enhanced M0 into an enhanced image M0′, comprising:

[0146] Main encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0147] Auxiliary encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice that of the encoded feature map The feature map The method is different from encoding feature maps The feature map Methods;

[0148] Main decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0149] Auxiliary decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0150] Two-way mixer for mixed coding feature maps and feature map The feature map And for mixed encoding feature maps and feature map The feature map

[0151] The primary processor is used to process the image M0 to be enhanced to obtain the feature map and feature map And process feature maps is the enhanced image M0′;

[0152] Among them, a=1,2,…,M-1;b=1,2,…,M-1;c=1,2,…,M;d=2,3,…,M;e=2,3,…,M;f=1,2,…,M;M≥3;characteristic graph The feature map Feature Map The feature map

[0153] In this embodiment, the processor refers to a single-chip microcomputer, a CPU, a computing card, or other loading and executing hardware, and the primary processor refers to an image processing device.

[0154] Preferably, in the main encoder, the encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map

[0155] Preferably, in the auxiliary encoder, the encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map ResGroup processes feature maps Get feature map

[0156] Preferably, in the hybrid encoder, the hybrid encoding feature map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M c1 ; 1×1 convolution processing feature map Get the matrix M c2 ;M c1 With M c2 After multiplication, the sigmoid activation function is used to obtain the feature map σE c ; FE c =σE c ×M c2 +(1-σE c )M c1 ; The feature map Set to q, the feature map Set to k, and FE cSet to v, and get the feature map after processing according to the attention mechanism

[0157] Preferably, in the main decoder, the decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map

[0158] Preferably, in the auxiliary decoder, the decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map ResGroup processes feature maps Get feature map

[0159] Preferably, in the hybrid encoder, the hybrid encoding feature map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M f3 ; 1×1 convolution processing feature map Get the matrix M f4 ;M f3 With M f4 After multiplication, the sigmoid activation function is used to obtain the feature map σD f FD f =σD f ×M f4 +(1-σD f )M f3 ; The feature map Set to q, the feature map Set to k, and set FD f Set to v, and get the feature map after processing according to the attention mechanism

[0160] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, in the primary processor, the image to be enhanced M0 is processed to obtain a feature map and feature map The method can be steps like this:

[0161] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0162] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0163] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Sub-image The sizes are the same, and the order of cropping and obtaining the sub-images can be determined from left to right and from top to bottom; for each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map

[0164] For images Perform 3×3 convolution processing and set the convolution step size to make the size of the feature map equal to the feature map The size is the same as that of

[0165] In actual processing, the kernel size can be set to 3, the stride to 1, and the padding to 1 for each sub-image. Perform 3×3 convolution to obtain the feature map Let kernel size be 3, stride be 2 and padding be 1 for a pair of images Perform 3×3 convolution to obtain the feature map In this way, the feature map obtained and feature map The length and width of the image are the same, and the corresponding feature maps obtained by subsequent processing can be mixed encoded.

[0166] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, in the primary processor, the image to be enhanced M0 is processed to obtain a feature map and feature map The method may be such steps (see Figure 2 ):

[0167] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0168] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0169] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Sub-image are of the same size; each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, concat merge them at the channel level to obtain the feature map

[0170] For images Perform convolution processing and set the step size of the convolution to make the size of the feature map equal to the feature map The size of the feature map is the same as that of the feature map, and then the feature map is obtained by multi-scale fusion processing.

[0171] Hybrid coding feature map and feature map The feature map

[0172] In actual processing, the kernel size can be set to 3, the stride to 1, and the padding to 1 for each sub-image. Perform 3×3 convolution to obtain the feature map Let kernel size be 3, stride be 2 and padding be 1 for a pair of images Perform 3×3 convolution processing. In this way, the feature map obtained and feature map The length and width of the image are the same, and the corresponding feature maps obtained by subsequent processing can be mixed encoded.

[0173] Still further preferably, if a supplementary block is used in the process of processing the image to be enhanced M0, then in the primary processor, the feature map is processed The enhanced image M0′ includes the following steps:

[0174] S711, down-channel processing characteristic diagram Get image MD;

[0175] S712, cutting off the supplementary block in the image MD to obtain the enhanced image M0'.

[0176] In step 1, if no supplementary block is used in the process of processing the image to be enhanced M0, then in the preliminary processor, the feature map is processed The enhanced image M0′ includes the following steps:

[0177] S711, down-channel processing characteristic diagram The enhanced image M0′ is obtained.

[0178] In the present invention, the supplementary block is a block of the same color, which may be a pure white block, a pure black block, or other pure color blocks.

[0179] Embodiment 2: An electronic device comprises a computer-readable medium storing an image processing program and a processor, wherein the image processing program is loaded by the processor and used to execute the aforementioned image processing method.

[0180] Embodiment 3: An image processing device, for processing an image to be enhanced M0 into an enhanced image M0′, comprising:

[0181] Main encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0182] Auxiliary encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice that of the encoded feature map The feature map The method is different from encoding feature maps The feature map Methods;

[0183] Main decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0184] Auxiliary decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice;

[0185] Two-way mixer for mixed coding feature maps and feature map The feature map And for mixed encoding feature maps and feature map The feature map

[0186] The primary processor is used to process the image M0 to be enhanced to obtain the feature map and feature map And the down-channel processing feature map is the enhanced image M0′;

[0187] Among them, a=1,2,…,M-1;b=1,2,…,M-1;c=1,2,…,M;d=2,3,…,M;e=2,3,…,M;f=1,2,…,M;M≥3;characteristic graph The feature map Feature Map The feature map

[0188] Preferably, in the main encoder, the encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map

[0189] Preferably, in the auxiliary encoder, the encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map ResGroup processes feature maps Get feature map

[0190] Preferably, in the hybrid encoder, the hybrid encoding feature map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M c1 ; 1×1 convolution processing feature map Get the matrix M c2 ;M c1 With M c2 After multiplication, the sigmoid activation function is used to obtain the feature map σE c ; FE c =σE c ×M c2 +(1-σE c )M c1 ; The feature map Set to q, the feature map Set to k, and FE c Set to v, and get the feature map after processing according to the attention mechanism

[0191] Preferably, in the main decoder, the decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map

[0192] Preferably, in the auxiliary decoder, the decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map ResGroup processes feature maps Get feature map

[0193] Preferably, in the hybrid encoder, the hybrid encoding feature map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M f3 ; 1×1 convolution processing feature map Get the matrix M f4 ;M f3 With M f4 After multiplication, the sigmoid activation function is used to obtain the feature map σD f FD f =σD f ×M f4 +(1-σD f )Mf3 ; The feature map Set to q, the feature map Set to k, and set FD f Set to v, and get the feature map after processing according to the attention mechanism

[0194] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, in the primary processor, the image to be enhanced M0 is processed to obtain a feature map and feature map The method can be steps like this:

[0195] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0196] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0197] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Sub-image The sizes are the same, and the order of cropping and obtaining the sub-images can be determined from left to right and from top to bottom; for each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map

[0198] For images Perform 3×3 convolution processing and set the convolution step size to make the size of the feature map equal to the feature map The size is the same as that of

[0199] In actual processing, the kernel size can be set to 3, the stride to 1, and the padding to 1 for each sub-image. Perform 3×3 convolution to obtain the feature map Let kernel size be 3, stride be 2 and padding be 1 for a pair of images Perform 3×3 convolution to obtain the feature map In this way, the feature map obtained and feature map The length and width of the image are the same, and the corresponding feature maps obtained by subsequent processing can be mixed encoded.

[0200] Further preferably, assuming that the depth of the multi-scale fusion processing feature map is N, in the primary processor, the image to be enhanced M0 is processed to obtain a feature map and feature map The method may be such steps (see Figure 2 ):

[0201] Query the number of pixels in the length and width directions of the image M0 to be enhanced;

[0202] If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the image after resizing; if the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0;

[0203] Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Sub-image are of the same size; each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, concat merge them at the channel level to obtain the feature map

[0204] For images Perform convolution processing and set the step size of the convolution to make the size of the feature map equal to the feature map The size of the feature map is the same as that of the feature map, and then the feature map is obtained by multi-scale fusion processing.

[0205] Hybrid coding feature map and feature map The feature map

[0206] In actual processing, the kernel size can be set to 3, the stride to 1, and the padding to 1 for each sub-image. Perform 3×3 convolution to obtain the feature map Let kernel size be 3, stride be 2 and padding be 1 for a pair of images Perform 3×3 convolution processing. In this way, the feature map obtained and feature map The length and width of the image are the same, and the corresponding feature maps obtained by subsequent processing can be mixed encoded.

[0207] Still further preferably, if a supplementary block is used in the process of processing the image to be enhanced M0, then in the primary processor, the feature map is processed The enhanced image M0′ includes the following steps:

[0208] S711, down-channel processing characteristic diagram Get image MD;

[0209] S712, cutting off the supplementary block in the image MD to obtain the enhanced image M0'.

[0210] In step 1, if no supplementary block is used in the process of processing the image to be enhanced M0, then in the preliminary processor, the feature map is processed The enhanced image M0′ includes the following steps:

[0211] S711, down-channel processing characteristic diagram The enhanced image M0′ is obtained.

[0212] In the present invention, the supplementary block is a block of the same color, which may be a pure white block, a pure black block, or other pure color blocks.

[0213] The method for constructing the image processing program of embodiment 1 or the image processing device of embodiment 3 comprises the following steps:

[0214] The step of preparing a training set and a test set, wherein each sample of the training set and the test set includes an image to be enhanced M0 and a label image M0″;

[0215] Steps to build an image processing simulation device;

[0216] a step of training parameters of the image processing simulation device, using the training set to train the image processing simulation device to obtain application parameters of the image processing simulation device;

[0217] The step of testing the parameters of the image processing simulation device, setting the parameters of the image processing simulation device to the application parameters, using the test set to test the image processing simulation device, if the error between the enhanced image output by the image processing simulation device and the corresponding label image meets the requirements after comparison, then constructing an image processing program or image processing device according to the image processing simulation device and the application parameters. In the image processing device, the main encoder, auxiliary encoder, main decoder, auxiliary decoder, dual mixer and primary processor can be constructed with mechanical structure, circuit structure, optical structure or a combination thereof.

[0218] Preferably, in the step of making a training set and a test set, a high-dynamic, large-size DR original image is obtained from the detector. This embodiment scans the DR images of the bolster and side frame of the railway casting with a resolution of 5732×2333. The DR image is cropped to a size of 556×583, and the excess part is discarded. Forty images to be enhanced can be obtained. The user here can change it according to the actual image size. After the obtained image to be enhanced is processed by traditional image enhancement methods such as histogram equalization, HDR, and window width and window position adjustment, the best enhanced image is selected as the label of the image to be enhanced, that is, the label image; the image to be enhanced and the label image are paired and randomly cut into integer multiples of 32 (544×576 in this embodiment), and after normalization, they can be used as samples of the training set. The size of the image to be enhanced and the label image of the test set is 5440×2304.

[0219] Preferably, in the step of establishing the image processing simulation device, a loss function is also established, and the loss function is:

[0220]

[0221] Among them, I i (i=1,2,…,M) is To label the image The feature map obtained by downsampling the size; For I i and The L1 norm loss for for The image value of y(h,w) is I i The image value of For I i and The perceptual loss is specifically to convert I i Input the feature map extracted by the VGG network and The L2 norm loss between the feature maps extracted by the input VGG network is For the general Input the feature map extracted by the VGG network, φ j (y) is I i Input the feature map extracted by the VGG network; Y is To match Label image of Y and The L1 norm loss for for y(h,w) is the image value of Y; Y and The perceptual loss is as follows: Y is input into the feature map extracted by the VGG network and The L2 norm loss between the feature maps extracted by the input VGG network is For the general Input the feature map extracted by the VGG network, φ j (y) is the feature map extracted by inputting Y into the VGG network; λ a , p , m are the weight coefficients of loss; is the total loss.

[0222] In the primary processor, the image to be enhanced M0 is processed to obtain a feature map and feature map In the method, the sub-image This is how we get it: Design sub-images The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, sliding and clipping feature map at a fixed ratio and fixed step size Get a set of sub-images So, matching The label image can also be used in the same way: slide and crop the label image at the same fixed ratio and fixed step size to obtain a set of sub-label images, and then the sub-label images can be combined with the corresponding Match.

[0223] In this embodiment, λ a =0.5,λ p =λ m =1.

[0224] The VGG network is a common network. In Example 4, it is specifically a pre-trained VGG16 network, and its structure can be found in the convolutional neural network mentioned in the document "Very Deep Convolutional NetWorks for Large-Scale Image Recognition" (translated as: Ultra-deep convolutional network for large-scale image recognition, author: Karen Simonyan, Andrew Zisserman).

[0225] Embodiment 4: A method for constructing an image processing program, comprising the following steps:

[0226] Step 1: Obtain a high dynamic large-size DR original image from the detector. The present invention scans the DR image of the bolster and side frame of the railway casting with a resolution of 5732×2333. The DR image is cut to 556×583 size to obtain 40 small-size images (original image, 40 images). The user can change the image size according to the actual image size.

[0227] Step 2: Implement three traditional image enhancement methods, namely histogram equalization, HDR, and window width and window position adjustment, on the original image, and adjust the parameters accordingly to select the best enhanced image as the label of the original image; based on this, construct a neural network training data set and divide it into training set T train and the test set T test ;

[0228] Step 3: The training set and test set data are randomly cut into integer multiples of 32, i.e. 544×576, and then normalized, flipped, and enhanced by adding noise to improve generalization. Finally, they are sent to the dual-branch multi-scale fusion neural network S(·); Figure 1 A dual-branch multi-scale fusion neural network S(·) is shown.

[0229] Step 4: The dual-branch trunk of the neural network S(·) adopts the classic U-net structure, which consists of an encoder and a decoder. A jump connection is used between the encoder and the decoder. The present invention designs an attention fusion module at the fusion point of the two branches to reduce the semantic gap.

[0230] In order to better enhance the characteristics of DR high dynamic range images and achieve large-scale reasoning, the multi-scale fusion network structure of the present invention is specifically as follows:

[0231] This module uses the high dynamic range features obtained in the encoder and decoder Figure 1First, two downsampling operations of different sizes are performed to obtain three feature maps of different sizes. After they are respectively sent to the residual neural network structure for learning, the two smaller sizes are interpolated back to the original size to obtain three feature maps of the same size. Figure 2 , and then perform global average pooling operations on them and then pass through the multi-layer perceptron structure to obtain the features Figure 3 , and then fuse the weights and features Figure 1 Add the final new features Figure 4 Finally, the new feature Figure 4 Send to the encoder or decoder of the next layer.

[0232] Step 5: In order to obtain better DR enhancement results, the present invention creates an attention fusion module, which aims to reduce the semantic gap of feature fusion between the two paths to ensure that the network can better extract 4K large-size image features;

[0233] The specific method is to place it at the position where the two branches of the network merge and interact. If it is just a simple fusion at this time, there will be a semantic gap because the two branches learn different features. Because of this, designing this module can effectively reduce the semantic gap and facilitate the interactive fusion of the two features to the greatest extent. Specifically, the featuremaps, Fa, and Fm extracted from the two branches are first passed through a convolution kernel and then multiplied and then sigmoid operation is performed to fuse the pixel-level information to obtain Fout. We set the fused feature map to v, Fm to q, Fa to k, and perform the classic attention mechanism to obtain the final output. Because the initial fusion may still leave some semantic gaps, we hope that the main branch can focus on learning the main network features from the auxiliary branch and discard some irrelevant and unimportant information.

[0234] Step 6: After the input data passes through the aforementioned neural network S(·), the main branch will obtain the output of the network in three different sizes. First, the obtained labels are downsampled once to 1 / 2 size and once to 1 / 4 size to obtain three different size comparison labels. The present invention designs two major losses, the main path and the auxiliary path. The three losses in the main path are the three size feature maps I output by the network. i (i=1,2,…,M) and the corresponding comparison labels of the three sizes A loss term of one norm is performed, and the other three losses are the three-size feature maps I output by the network i (i=1,2,…,M) and the corresponding comparison labels of the three sizes They are all thrown into the pre-trained VGG16 network to obtain their feature maps and then subjected to a norm loss. The auxiliary path outputs a set of sub-image prediction outputs Y, Y and the corresponding contrast labels of the sub-images Calculate the one-norm loss and perceptual loss as the loss of the auxiliary path. a =0.5,λ p =λ m =1.

[0235] Step 7: We build a neural network on PyTorch. The parameters of the convolutional layer in the network are initialized according to a normal distribution with a mean of 0 and a standard deviation of 0.02. The weight parameters of the remaining layers are randomly initialized. The remaining hyperparameters of the network can be reasonably set within a certain range. For example, we set a batch size of 4 in the 24G video memory of NVIDIA RTX 4090 and train 150 times. The Adam optimizer is used, and the learning rate adjustment strategy of cosine annealing is adopted. In this strategy, 50 rounds are set as one round, and 0.0002 is the initial learning rate; finally, the error back propagation algorithm is used to train the DR image enhancement neural network model, thereby obtaining the DR image enhancement model.

[0236] Figure 6 shows a DR image to be enhanced, Figure 7 To process the image using the image processing program obtained in this embodiment Figure 6 The resulting label image.

[0237] Step 8: When deploying the trained model, since the image width and height required by the network design are multiples of 32, if the actual inferenced DR image is not a multiple of 32, its height and width can be resized to the nearest integer multiple of 32, and then the final enhanced DR image can be obtained through the trained neural network S(·).

[0238] The present invention enhances DR images with high dynamic range by designing a multi-scale fusion module, which can not only effectively extract the information features of the original DR image, but also can perform efficient reasoning when dealing with image inputs of different sizes. In addition, the present invention strengthens feature extraction by designing an attention fusion module. Since the auxiliary branch captures small-scale detail information of the image, there will be a semantic gap when it is directly fused with the features captured by the main branch. The effectiveness of this module can be verified by ablation experiments.

[0239] The present invention is described in detail above with reference to the accompanying drawings and embodiments. It should be understood that it is impossible to describe all possible implementation methods in practice, and the inventive concept of the present invention is described as much as possible by way of example. Without departing from the inventive concept of the present invention and without creative work, the technical personnel in this technical field make selections and combinations of the technical features in the above embodiments, make experimental changes to the specific parameters, or use the prior art in this technical field to conventionally replace the disclosed technical means of the present invention to form specific embodiments, which should all belong to the implicit disclosure of the present invention.

Claims

1. An image processing method, wherein the pointer i=1, is characterized in that: The following steps are involved: Step 1: Process the image to be enhanced M0 to obtain the feature map and feature map Step 2: Encode feature map The feature map Encoding feature map The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice that of the feature map The corresponding image size is the feature map The corresponding image size is twice that of the encoded feature map The feature map The method is different from encoding feature maps The feature map Methods; Step 3: Hybrid coding feature map and feature map The feature map Step 4. If i < M, then set i = i + 1, and return to Step 2 to execute the loop from Step 2 to Step 4; if i = M, then set the feature map as the feature map The feature map as the feature map proceed to the next step; Step 5: Decode feature map The feature map Decoding feature map The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice that of the feature map The corresponding image size is the feature map The corresponding image size is twice; Step 6: Hybrid coding feature map and feature map The feature map Step 7: If i>1, set i=i-1 and return to step 5 to execute the loop from step 5 to step 7; if i=1, process the feature map is the enhanced image M0′.

2. An image processing device, used for processing an image to be enhanced M0 into an enhanced image M0′, characterized in that: include: Main encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice; Auxiliary encoder, used to encode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice that of the encoded feature map The feature map The method is different from encoding feature maps The feature map Methods; Main decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice; Auxiliary decoder, used to decode feature maps The feature map The feature map The corresponding image size is the feature map The corresponding image size is twice; Two-way mixer for mixed coding feature maps and feature map The feature map And for mixed encoding feature maps and feature map The feature map as well as The primary processor is used to process the image M0 to be enhanced to obtain the feature map and feature map And process feature maps is the enhanced image M0′; Among them, a=1,2,…,M-1;b=1,2,…,M-1;c=1,2,…,M;d=2,3,…,M;e=2,3,…,M;f=1,2,…,M;M≥3;characteristic graph The feature map Feature Map The feature map 3. The image processing method according to claim 1 or the image processing device according to claim 2, characterized in that: Hybrid Encoding Feature Map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M i1 ; 1×1 convolution processing feature map Get the matrix M i2 ;M i1 With M i2 After multiplication, the sigmoid activation function is used to obtain the feature map σE i ; FE i =σE i ×M i2 +(1-σE i )M i1 ; The feature map Set to q, the feature map Set to k, and FE i Set to v, and get the feature map after processing according to the attention mechanism Hybrid Encoding Feature Map and feature map The feature map The method is: 1×1 convolution to process the feature map Get the matrix M i3 ; 1×1 convolution processing feature map Get the matrix M i4 ;M i3 With M i4 After multiplication, the sigmoid activation function is used to obtain the feature map σD i FD i =σD i ×M i4 +(1-σD i )M i3 ; The feature map Set to q, the feature map Set to k, and set FD i Set to v, and get the feature map after processing according to the attention mechanism 4. The image processing method according to claim 1 or the image processing device according to claim 2, characterized in that: Encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map Encoding feature map The feature map The method is: shrink the feature map To reduce the image feature dimension, we get the feature map Residual group convolution processing feature map Get feature map Decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Multi-scale fusion processing feature map Get feature map Decoding feature map The feature map The method is: expand the processing feature map To increase the image feature dimension, we get the feature map Residual group convolution processing feature map Get feature map 5. The image processing method or image processing device according to claim 4, characterized in that The depth of the multi-scale fusion processing feature map is N, and the image to be enhanced M0 is processed to obtain the feature map and feature map The method is: Query the number of pixels in the length and width directions of the image M0 to be enhanced; If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the resized image; If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0; Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images For each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map For images Perform 3×3 convolution processing and set the convolution step size to make the size of the feature map equal to the feature map The size is the same as that of Alternatively, process the image to be enhanced M0 to obtain the feature map and feature map The method is: Query the number of pixels in the length and width directions of the image M0 to be enhanced; If either the number of pixels in the length direction or the number of pixels in the width direction of the image M0 to be enhanced is not 2 M+N-1 If the pixel number is an integer multiple, a supplementary block is used at the end of the corresponding direction to supplement the number of pixels to 2. M+N-1 Integer multiples, so that the image is the resized image; If the number of pixels in the length direction and the number of pixels in the width direction of the image M0 to be enhanced are both 2 M+N-1 Integer multiples, then let the image is the image to be enhanced M0; Design sub-image The number of pixels in the length direction and the number of pixels in the width direction are both 2 M+N-1 Integer multiple, slide and crop the image at a fixed ratio and fixed step length Get a set of sub-images Each sub-image Perform convolution processing to obtain the feature map According to the feature map of all sub-images According to the order of cropping and acquisition of sub-images, they are merged at the channel level to obtain the feature map For images Perform convolution processing and set the step size of the convolution to make the size of the feature map equal to the feature map The size of the feature map is the same as that of the feature map, and then the feature map is obtained by multi-scale fusion processing. Hybrid Encoding Feature Map and feature map The feature map 6. The image processing method or image processing device according to claim 5, characterized in that: In step 1, if the supplementary block is used in the process of processing the image to be enhanced M0, then in step 7, the feature map is processed The enhanced image M0′ includes the following steps: S711, down-channel processing characteristic diagram Get image MD; S712, cutting off the supplementary block in the image MD to obtain the enhanced image M0'. In step 1, if no supplementary block is used in the process of processing the image to be enhanced M0, then in step 7, the feature map is processed The enhanced image M0′ includes the following steps: S711, down-channel processing characteristic diagram The enhanced image M0′ is obtained.

7. A computer-readable medium storing an image processing program, characterized in that: The image processing program is loaded by the processor and is used to execute the image processing method as claimed in claim 1, 3, 4, 5 or 6.

8. An electronic device comprising a computer-readable medium storing an image processing program and a processor, wherein: The image processing program is loaded by the processor and is used to execute the image processing method as claimed in claim 1, 3, 4, 5 or 6.

9. The method for constructing an image processing device according to any one of claims 2 to 6, characterized in that: The following steps are involved: The step of preparing a training set and a test set, wherein each sample of the training set and the test set includes an image to be enhanced M0 and a label image M0″; Steps to build an image processing simulation device; a step of training parameters of the image processing simulation device, using the training set to train the image processing simulation device to obtain application parameters of the image processing simulation device; The step of testing the parameters of the image processing simulation device, setting the parameters of the image processing simulation device as the application parameters, using the test set to test the image processing simulation device, if the error after comparison between the enhanced image output by the image processing simulation device and the corresponding label image meets the requirements, then building the image processing device according to the image processing simulation device and the application parameters.

10. The method for constructing an image processing device according to claim 9, characterized in that: In the step of establishing the image processing simulation device, a loss function is also established, and the loss function is: Among them, I i (i=1,2,…,M) is To label the image The feature map obtained by downsampling the size; For I i and The L1 norm loss for for The image value of y(h,w) is I i The image value of For I i and The perceptual loss is specifically to convert I i Input the feature map extracted by the VGG network and The L2 norm loss between the feature maps extracted by the input VGG network is For the general Input the feature map extracted by the VGG network, φ j (y) is I i Input the feature map extracted by the VGG network; Y is To match Label image of Y and The L1 norm loss for for y(h,w) is the image value of Y; Y and The perceptual loss is as follows: Y is input into the feature map extracted by the VGG network and The L2 norm loss between the feature maps extracted by the input VGG network is For the general Input the feature map extracted by the VGG network, φ j (y) is the feature map extracted by inputting Y into the VGG network; λ a , p , m are the weight coefficients of loss; is the total loss.

Citation Information

Patent Citations

  • Human body detection method in low-illumination environment based on image enhancement, electronic equipment and storage medium

    CN114708615A

  • Construction method of large-size image enhancement model, image processing method, computer readable medium and electronic equipment

    CN118115373A

  • Medical image segmentation method based on u-net

    US20220309674A1