Image optimization method and system

Through technologies such as deep convolutional neural networks and Gaussian pyramids, images under complex lighting conditions are preprocessed and optimized, which solves the problem of improper image processing in the existing technology and achieves better image quality and visual effects.

CN120219221APending Publication Date: 2025-06-27SHENZHEN XIANYANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304841.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process images under complex lighting conditions, resulting in problems such as overexposure, insufficient exposure, loss of details, contrast imbalance and color distortion, affecting the display quality and visual effects of the image.

Method used

The preprocessed images are optimized by using deep convolutional neural networks. By constructing a Gaussian pyramid to remove noise, detect and mark potential points, fit three-dimensional quadratic functions using Taylor expansion, and combined with image alignment technology, the overall visual effect of the image is improved.

Benefits of technology

In low-light environments or severe reflections, effectively process the highlights and shadowy parts in the image to avoid overexposure or insufficient exposure, improve the overall visual effect of the image, and significantly improve the image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219221A_ABST
    Figure CN120219221A_ABST
Patent Text Reader

Abstract

The invention discloses an image optimization method and system, and the method comprises the following steps: S1, carrying out the preprocessing of an obtained original image, removing the noise of the original image, and obtaining a preprocessed image; and S2, performing image optimization on the preprocessed image by using a deep convolutional neural network, and outputting an optimized image. Through an intelligent image processing algorithm and technology, highlight and shadow parts in the image are better processed in a low-light environment or under the condition of serious reflection, the condition of overexposure or insufficient exposure is avoided, and the overall visual effect of the image is improved. After deep convolutional neural network processing, the image quality is significantly improved. According to the method, the processing precision in other fields can be improved, for example, when tasks such as touch recognition, object positioning and defect detection are carried out, noise interference can be reduced through the optimized image, and the algorithm is more stable and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image optimization, and particularly to a method and system for image optimization. Background Art

[0002] Most of the existing technologies rely on fixed image processing algorithms to process images. This method performs well when processing images under conventional lighting conditions and can meet general visual needs. However, when the image acquisition environment becomes complex and variable, such as when shooting in low-light environments, strong backlighting, or in the presence of severe reflections and glare and other adverse conditions, these fixed algorithms become inadequate and are difficult to cope with various challenges.

[0003] Specifically, under complex lighting conditions, the highlight parts in the image often suffer from severe overexposure problems due to excessive light, resulting in the details in the highlighted areas being completely submerged and unidentifiable. While the shadow parts may be underexposed due to insufficient light, becoming dull and lacking necessary details and sense of hierarchy. This improper processing not only fails to effectively balance the bright and dark areas in the image, making the overall image look too contrasty or too flat, but also often leads to problems such as image detail loss, contrast imbalance, and color distortion. These problems ultimately seriously affect the display quality and visual effect of the image, making it difficult for the image to convey accurate information and the true scene atmosphere.

[0004] Therefore, how to improve the image processing algorithm to adapt to image acquisition under different lighting conditions has become an urgent problem to be solved. Summary of the Invention

[0005] Aiming at the above defects, the purpose of the present invention is to propose a method and system for image optimization to solve the problems of excessive image noise and poor display quality caused by lighting problems.

[0006] To achieve this purpose, the present invention adopts the following technical solutions: A method for image optimization, comprising the following steps:

[0007] Step S1: Preprocess the acquired original image to remove the noise in the original image and obtain a preprocessed image;

[0008] Step S2: Use a deep convolutional neural network to optimize the preprocessed image and output an optimized image.

[0009] Preferably, the preprocessing includes noise removal;

[0010] The noise removal includes the following steps:

[0011] Step A1: Construct a Gaussian pyramid. The original image is convolved with Gaussian kernels of different standard deviations through the Gaussian pyramid to obtain first images with multiple different scales, and all the first images are constructed into a first collection;

[0012] Step A2: Obtain the first first image in the first collection as the detection image;

[0013] Step A3: Obtain any pixel point in the detection image as the detection pixel point;

[0014] Step A4: Use the detection pixel point as the center to construct a 3*3 first detection range in the detection image;

[0015] Use the first images of adjacent scales of the detection image as the second images, and at the position of the same detection pixel, construct a 3*3 second detection range in the second images;

[0016] Step A5: Determine whether the detection pixel point is the maximum or minimum point of the gray value in the first detection range and the second detection range. If so, mark the detection pixel point as a potential point;

[0017] Step A6: Replace the next pixel point as the detection pixel point, and repeat steps A4 to A5 until all pixel points in the detection image are looped, and then execute step A7;

[0018] Step A7: Replace the next first image as the detection image, and repeat steps A3 to A6 until all first images in the first collection are looped, and then execute step A8;

[0019] Step A8: Perform Taylor expansion on all potential points to fit a three-dimensional quadratic function D(X), where X = [x, y, σ] T , (x, y) are the coordinates of the potential point respectively, σ is the standard deviation used in the calculation of the scale of the first image where the potential point is located. Take the derivative of the three-dimensional quadratic function and set the derivative to zero to obtain the coordinates of the key point;

[0020] Step A9: Calculate the eigenvalue of each pixel point in each first image through the Hessian matrix; regard the pixel points with eigenvalues lower than the eigenvalue threshold as edge points;

[0021] Obtain the contrast between the key points and the edge points, and delete the key points and / or edge points with the contrast lower than the contrast threshold to obtain the preprocessed image.

[0022] Preferably, the preprocessing further includes image alignment;

[0023] The steps of the image alignment are as follows:

[0024] Step B1: Use the SIFT algorithm to perform point matching on the original image to obtain a set of matching points;

[0025] Step B2: Randomly obtain N pairs of matching points from the set of matching points as the first matching points;

[0026] Step B3: Use the first matching points to perform a two-dimensional affine transformation to obtain a transformed image;

[0027] Step B4: Find the corresponding second matching points of the first matching points in the transformed image, obtain the Euclidean distance between the first matching points and the second matching points, and determine whether the Euclidean distance is less than a distance threshold. If it is less, mark the first matching points and the second matching points as inliers and record the number of inliers;

[0028] Step B5: Repeat Steps B2 - B4 until the number of repetitions reaches a quantity threshold, obtain the transformed image with the largest number of inliers as the third image, and use the second matching points in the third image and the corresponding first matching points in the original image to update the matrix of the two-dimensional affine transformation to obtain an updated matrix;

[0029] Step B6: Use the updated matrix to perform an alignment operation on the image to update the preprocessed image.

[0030] Preferably, before performing Step S1, the original image also needs to be block - processed.

[0031] Preferably, before performing Step S2, the following steps also need to be performed:

[0032] Step S21: Construct a first convolution kernel and a residual module, use the first convolution kernel to extract features from the preprocessed image, and the extracted features are subjected to a convolution operation through the residual module to output an enhanced image after image enhancement;

[0033] Step S22: Use the preprocessed image and the enhanced image to train a convolutional neural network to obtain a trained convolutional neural network;

[0034] The training process is as follows:

[0035] Step S221: Define a loss function through the preprocessed image and the enhanced image, and the loss function includes color loss, texture loss, content loss, and total variation loss;

[0036] Step S222: Weight - combine all the losses to obtain a final loss, and use the backpropagation algorithm and the final loss to update the convolutional neural network;

[0037] The definition rule of the color loss is as follows:

[0038] The preprocessed image and the enhanced image are respectively subjected to blurring operations using a Gaussian operator, and the difference between the two blurred images is calculated as the color loss;

[0039] The calculation process of the color loss is as follows:

[0040]

[0041] where X b , Y b are respectively the preprocessed image and the enhanced image after blurring;

[0042] where X[i + k, j + l] represents the pixel value at the position (i + k, j + l), and k and l are respectively the indices of the Gaussian kernel G(k, l);

[0043] where A exp is an empirical constant with a set value of 0.1 to 1, u X , u Y represent the central positions of the Gaussian distributions of the preprocessed image and the enhanced image, i.e., the positions where the peaks of the Gaussian functions are located, and σ X , σ Y represent the means of the Gaussian distributions of the preprocessed image and the enhanced image;

[0044] The calculation of the texture loss is as follows:

[0045] L texture = ΣlogD(F w (I s ) l I t );

[0046] where D() is the discriminative network, F w () is the convolutional neural network, I t is the preprocessed image, I s is the enhanced image, and l is the index;

[0047] The calculation of the content loss is as follows:

[0048]

[0049] where C j is the number of channels of the feature map output by the ReLU_5_4 layer of the VGG-19 network, H j and W j respectively represent the height and width of the feature map output by the ReLU_5_4 layer of the VGG-19 network, represents that the image generated by the convolutional neural network passes through the VGG-19 network and extracts the features of its ReLU_5_4 layer, It represents that the preprocessed image passes through the VGG-19 network and extracts the features of its ReLU_5_4 layer;

[0050] The total variation loss is calculated as follows:

[0051]

[0052] where C i , H i , W i are the number of channels, height, and width of the image, and represent the gradient operators along the x-axis and y-axis respectively.

[0053] Preferably, the discriminant network specifically performs the following steps:

[0054] Step C1: Convert the preprocessed image and the enhanced image into grayscale images respectively to obtain a grayscale image set;

[0055] Step C2: Perform convolutional layer feature extraction, multi-layer convolution, and pooling operations on the grayscale image set in sequence to obtain the feature map O L ;

[0056] Step C3: Flatten the feature map O L and convert it into a one-dimensional vector V;

[0057] Input the one-dimensional vector V into the fully connected layer;

[0058] Step C4: Obtain the output Z of the fully connected layer, and finally convert the output Z of the fully connected layer into a probability value P through the Sigmoid function. The probability value P is used to determine whether the image is a preprocessed image or an enhanced image;

[0059] When the probability value P is close to 1, it is determined as a preprocessed image; if the probability value P is close to 0, it is determined as an enhanced image.

[0060] An image optimization system, characterized in that it uses the described image optimization method, including: a preprocessing module and an enhancement module;

[0061] The preprocessing module is used to preprocess the acquired original image, remove the noise of the original image, and obtain a preprocessed image;

[0062] The enhancement module is used to perform image optimization on the preprocessed image using a deep convolutional neural network and output an optimized image.

[0063] Preferably, the preprocessing module includes a noise module;

[0064] The noise module includes a pyramid module, a detection image acquisition module, a detection pixel point acquisition module, a detection range construction module, a potential point marking module, a first loop module, a second loop module, a positioning module, and a deletion module;

[0065] The pyramid module constructs a Gaussian pyramid. The original image undergoes Gaussian kernel convolution with different standard deviations through the Gaussian pyramid to obtain first images with multiple different scales, and all the first images are used to construct a first set;

[0066] The detection image acquisition module is used to obtain the first first image in the first set as the detection image;

[0067] The detection pixel point acquisition module is used to obtain any pixel point in the detection image as the detection pixel point;

[0068] The detection range construction module is used to construct a 3*3 first detection range in the detection image with the detection pixel point as the center;

[0069] Use the first images of adjacent scales in the detection image as the second image, and at the positions where the detection pixels are the same, construct a 3*3 second detection range in the second image;

[0070] The potential point marking module is used to determine whether the detection pixel point is a gray value maximum or minimum point within the first detection range and the second detection range. If so, mark the detection pixel point as a potential point;

[0071] The first loop module is used to replace the next first image as the detection image, and re-call the detection range construction module and the potential point marking module until all the first images in the first set are looped, and then call the second loop module;

[0072] The second loop module is used to replace the next first image as the detection image, and re-call the detection pixel point acquisition module, the detection range construction module, the potential point marking module, and the first loop module until all the first images in the first set are looped, and then call the positioning module;

[0073] The positioning module is used to perform Taylor expansion on all potential points to fit a three-dimensional quadratic function D(X), where X = [x, y, σ] T , (x, y) are the coordinates of the potential point respectively, and σ is the standard deviation used in the calculation of the scale of the first image where the potential point is located. Take the derivative of the three-dimensional quadratic function and set the derivative to zero to obtain the coordinates of the key point;

[0074] The deletion module is used to calculate the eigenvalue of each pixel point of each first image through the Hessian matrix; regard the pixel points with eigenvalues lower than the eigenvalue threshold as edge points;

[0075] Obtain the contrast between key points and edge points, delete the key points and / or edge points whose contrast is lower than the contrast pair threshold, and obtain a preprocessed image.

[0076] Preferably, the preprocessing module further includes an alignment module;

[0077] The alignment module includes a matching unit, a selection unit, an affine unit, an inlier recording unit, an update unit, and an alignment unit;

[0078] The matching unit is used to perform point matching on the original image using the SIFT algorithm to obtain a set of matching points;

[0079] The selection unit is used to randomly obtain N pairs of matching points from the set of matching points as the first matching points;

[0080] The affine unit is used to perform a two-dimensional affine transformation using the first matching points to obtain a transformed image;

[0081] The inlier recording unit is used to find the second matching points corresponding to the first matching points in the transformed image, obtain the Euclidean distance between the first matching points and the second matching points, and determine whether the Euclidean distance is less than the distance threshold. If it is less, the first matching points and the second matching points are marked as inliers, and the number of inliers is recorded;

[0082] The update unit is used to re-invoke the selection unit, the affine unit, and the inlier recording unit until the number of repetitions reaches the number threshold, obtain the transformed image with the largest number of inliers as the third image, and update the two-dimensional affine matrix using the second matching points in the third image and the corresponding first matching points in the original image to obtain an updated matrix;

[0083] The alignment unit is used to perform an alignment operation on the image using the updated matrix to update the preprocessed image.

[0084] Preferably, it further includes an image enhancement unit and a training unit;

[0085] The image enhancement unit is used to construct a first convolutional kernel and a residual module, use the first convolutional kernel to extract features from the preprocessed image, and perform convolutional operations on the extracted features through the residual module to output an enhanced image after image enhancement;

[0086] The training unit is used to train a convolutional neural network using the preprocessed image and the enhanced image to obtain a trained convolutional neural network.

[0087] One of the technical solutions in the above technical solutions has the following advantages or beneficial effects: Through intelligent image processing algorithms and technologies, in low-light environments or situations with severe reflections, the highlight and shadow parts in the image are better processed, avoiding overexposure or underexposure, and improving the overall visual effect of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 is a flowchart of an embodiment of the method of the present invention.

[0089] Figure 2 is a schematic structural diagram of an embodiment of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0090] The following describes in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0091] In the description of the embodiments of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the embodiments of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0092] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, "a plurality" means two or more. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0093] As Figures 1-2 shown, a method for image optimization includes the following steps:

[0094] Step S1: Preprocess the acquired original image to remove the noise of the original image and obtain a preprocessed image;

[0095] Step S2: Use a deep convolutional neural network to optimize the preprocessed image and output an optimized image.

[0096] Through intelligent image processing algorithms and technologies, in low-light environments or situations with severe reflections, the highlight and shadow parts in the image are better processed, avoiding overexposure or underexposure, and enhancing the overall visual effect of the image.

[0097] After being processed by a deep convolutional neural network, the image quality is significantly improved. It can improve the processing accuracy in other fields. For example, when performing tasks such as touch recognition, object localization, and defect detection, the optimized image can reduce noise interference, making the algorithm more stable and reliable.

[0098] Preferably, the preprocessing includes noise removal;

[0099] Among them, the noise removal includes the following steps:

[0100] Step A1: Construct a Gaussian pyramid. The original image is convolved with Gaussian kernels of different standard deviations through the Gaussian pyramid to obtain first images with multiple different scales, and all the first images are constructed into a first collection;

[0101] The formula of the Gaussian kernel is as follows:

[0102] Where (x, y) are the coordinates of the pixel points in the first image respectively, and σ is the standard deviation.

[0103] Step A2: Obtain the first first image in the first collection as the detection image;

[0104] Step A3: Obtain any pixel point in the detection image as the detection pixel point;

[0105] Step A4: Use the detection pixel point as the center to construct a 3*3 first detection range in the detection image;

[0106] Use the first images of adjacent scales of the detection image as the second images, and at the position of the same detection pixel, construct a 3*3 second detection range in the second image;

[0107] Step A5: Determine whether the detection pixel point is the maximum or minimum point of the gray value in the first detection range and the second detection range. If so, mark the detection pixel point as a potential point;

[0108] Step A6: Replace the next pixel point as the detection pixel point, and repeat steps A4 to A5 until all pixel points in the detection image are cycled, and then execute step A7;

[0109] Step A7: Replace the next first image as the detection image, and repeat steps A3 to A6 until all first images in the first collection are cycled, and then execute step A8;

[0110] Step A8: Perform Taylor expansion on all potential points to fit a three-dimensional quadratic function D(X), where X = [x, y, σ] T , (x, y) are the coordinates of the potential points respectively, and σ is the standard deviation used in the calculation of the first image scale where the potential points are located. Take the derivative of the three-dimensional quadratic function and set the derivative to zero to obtain the coordinates of the key points;

[0111] Step A9: Calculate the eigenvalues of each pixel point in each first image through the Hessian matrix; Take the pixel points with eigenvalues lower than the eigenvalue threshold as edge points;

[0112] Obtain the contrast between the key points and the edge points, and delete the key points and / or edge points with a contrast lower than the contrast pair threshold to obtain a preprocessed image.

[0113] Due to the influence of factors such as ambient light and screen brightness, problems such as noise, color difference, and blurring will occur in the original image. In this case, it is impossible to determine at which scale the noise is and how its scale behaves. Therefore, during the preprocessing process, a Gaussian pyramid will be constructed first, and the original image will be converted into a first set with multiple different scales. Analyze the noise through the first images of multiple different scales to enhance the robustness to noise.

[0114] Then, by comparing the maximum or minimum points (i.e., potential points) of the gray values in adjacent scales and adjacent ranges, important features in the image can be effectively screened out, and some local noise points can be prevented from being misidentified as important features. This is because noise points usually do not have extreme value manifestations at multiple scales, so they can be excluded during the detection process. Then use Taylor expansion to fit the three-dimensional quadratic function of the potential points, and the position of the key points can be determined at a higher precision level. This process not only helps to remove noise but also can improve the accuracy of key point positioning and reduce the position deviation caused by noise interference.

[0115] Finally, when screening potential points and edge points, use contrast for further filtering. By setting the contrast threshold, low-contrast noise points can be further removed. Low-contrast points are often caused by noise or other interferences. Therefore, eliminating them can effectively reduce the impact of noise on subsequent image enhancement operations and improve the quality of the image.

[0116] Preferably, the preprocessing further includes image alignment;

[0117] The steps of the image alignment are as follows:

[0118] Step B1: Use the SIFT algorithm to perform point matching on the original image to obtain a set of matching points;

[0119] Step B2: Randomly obtain N pairs of matching points from the set of matching points as the first matching points;

[0120] Step B3: Perform a two-dimensional affine transformation using the first matching points to obtain a transformed image;

[0121] Step B4: Find the second matching points corresponding to the first matching points in the transformed image, obtain the Euclidean distance between the first matching points and the second matching points, and determine whether the Euclidean distance is less than the distance threshold. If it is less, mark the first matching points and the second matching points as inliers and record the number of inliers;

[0122] Step B5: Repeat Steps B2 - B4 until the number of repetitions reaches the quantity threshold. Obtain the transformed image with the largest number of inliers as the third image, and use the second matching points in the third image and the corresponding first matching points in the original image to update the matrix of the two-dimensional affine transformation to obtain an updated matrix;

[0123] Step B6: Use the updated matrix to perform an alignment operation on the image and update the preprocessed image.

[0124] In order to facilitate subsequent image enhancement, the original image needs to be aligned. Although some noise has been removed from the preprocessed image, there is still some noise that has not been completely removed. If the preprocessed image is directly used for alignment, this part of the noise will affect the alignment result, thereby affecting the next step of image enhancement. Therefore, in the present invention, the SIFT algorithm is used to extract local features in the original image, and these features have strong stability with respect to scale, rotation, and illumination changes. This enables the image alignment process to still maintain a good matching effect in the presence of certain noise, perspective changes, or scale changes in the image.

[0125] In addition, in the present invention, the Euclidean distance is used to judge the accuracy of the matching points and mark them as inliers, which can effectively filter out the wrong matching points. Wrong matching points usually appear in the background, noise, or repetitive texture areas. After automatically screening these points, it helps to improve the alignment accuracy and robustness. When the number of inliers is the largest, it means that the alignment effect of the third image is the best. Therefore, using the parameters of the third image with the largest number of inliers to update the matrix of the two-dimensional affine transformation can effectively improve the alignment effect of the matrix of the two-dimensional affine transformation. The specific matrix of the two-dimensional affine transformation is as follows:

[0126]

[0127] where a11, a12, a21, a22, t x 、t yThe matrix parameters in the matrix that are two-dimensional affine, (x′, y′) are the coordinates of the second matching point in the transformed image (the third image), and (x, y) are the coordinates of the first matching point in the preprocessed image. Since there are multiple second matching points in the third image, multiple pairs of matching first and second matching points are input into the above two-dimensional affine matrix to form multiple equations, and finally, methods such as the least squares method are used to solve for a11, a12, a21, a22, and t x and t y , and then the updated matrix can be obtained.

[0128] Preferably, before performing step S1, the original image also needs to be block-processed.

[0129] After block-processing the original image, multiple small-sized images of the same size can be obtained. Multiple small-sized images can be preprocessed together, greatly improving the efficiency of data processing.

[0130] Preferably, before performing step S2, the following steps also need to be performed:

[0131] Step S21: Construct a first convolution kernel and a residual module. Use the first convolution kernel to extract features from the preprocessed image. The extracted features are subjected to a convolution operation through the residual module, and an enhanced image after image enhancement is output;

[0132] In one embodiment, the size of the first convolution kernel is set to a 9x9x64 convolution kernel, and the residual module is provided with two 3x3x64x64 convolution kernels and a 9x9x64x3 convolution kernel to further process and transform the features.

[0133] Step S22: Use the preprocessed image and the enhanced image to train the convolutional neural network to obtain a trained convolutional neural network;

[0134] Among them, the training process is as follows:

[0135] Step S221: Define a loss function through the preprocessed image and the enhanced image. The loss function includes color loss, texture loss, content loss, and total variation loss;

[0136] Among them, step S222: Weight-combine all the losses to obtain a final loss, and use the backpropagation algorithm and the final loss to update the convolutional neural network;

[0137] Among them, the definition rule of the color loss is as follows:

[0138] Perform Gaussian operator blurring operations on the preprocessed image and the enhanced image respectively, and calculate the difference between the two blurred images as the color loss;

[0139] The calculation process of the color loss is as follows:

[0140]

[0141] Where X b , Y b are the preprocessed image after blurring and the enhanced image respectively;

[0142] Where X[i + k, j + l] represents the pixel value at the position (i + k, j + l), and k and l are the indices of the Gaussian kernel G(k, l);

[0143] Where A exp is an empirical constant, with a set value of 0.1 to 1, μ X , μ Y represent the central positions of the Gaussian distributions of the preprocessed image and the enhanced image, that is, the peak positions of the Gaussian functions, and σ X , σ Y represent the means of the Gaussian distributions of the preprocessed image and the enhanced image;

[0144] To measure the color difference between the preprocessed image and the enhanced image, Gaussian blur operations are performed on the preprocessed image and the enhanced image respectively, and then the differences between the two blurred images are calculated. The advantage of doing this is that the human visual system is not particularly sensitive to color changes, and the color is relatively smooth locally. Gaussian blur can eliminate this part of the information and only leave the color information for comparison, while also ensuring the local translational invariance of the image.

[0145] The calculation of the texture loss is as follows:

[0146] L texture = ΣlogD(F w (I s ) l I t );

[0147] Where D() is the discriminator network, and F w () is the convolutional neural network, I t is the preprocessed image, I s is the enhanced image, and l is the index;

[0148] The texture loss is learned through the adversarial network and is used to measure the texture quality. The goal is to determine the authenticity of the image, and the texture loss is defined as the cross-entropy loss of the discriminator. This allows the convolutional neural network to learn texture features closer to the original real image.

[0149] The calculation of the content loss is as follows:

[0150]

[0151] Among them, C j represents the number of channels of the feature map output by the ReLU_5_4 layer of the VGG-19 network, H j and W j respectively represent the height and width of the feature map output by the ReLU_5_4 layer of the VGG-19 network. represents that the image generated by the convolutional neural network passes through the VGG-19 network and extracts the features of its ReLU_5_4 layer. represents that the preprocessed image passes through the VGG-19 network and extracts the features of its ReLU_5_4 layer;

[0152] In the present invention, the feature representation generated by the ReLU_5_4 layer of the VGG-19 network is used to calculate the content loss. It is encouraged that the preprocessed image and the enhanced image have similar feature representations to maintain the semantic information of the picture and avoid only comparing pixel-level differences while ignoring the overall content and perceptual quality of the image.

[0153] The total variation loss is calculated as follows:

[0154]

[0155] Among them, C i , H i , W i are the number of channels, height, and width of the image. and respectively represent the gradient operators along the x-axis and y-axis.

[0156] The total variation loss is used to enhance the spatial domain smoothness of the generated picture and remove salt-and-pepper noise, etc. Even with a low weight in the final loss, it can make the image smoother without affecting the high-frequency part of the image.

[0157] Preferably, by continuously training the discriminative network, it can accurately distinguish between real and fake images, thereby providing feedback for the training of the convolutional neural network to improve the processing effect of the convolutional neural network. The discriminative network specifically performs the following steps:

[0158] Step C1: Convert the preprocessed image and the enhanced image into grayscale images respectively to obtain a grayscale image set;

[0159] Step C2: Perform convolutional layer feature extraction, multi-layer convolution, and pooling operations on the grayscale image set in sequence to obtain the feature map O L ;

[0160] Step C3: Flatten the feature map O L to convert it into a one-dimensional vector V;

[0161] Input the one-dimensional vector V into the fully connected layer;

[0162] Step C4: Obtain the output Z of the fully connected layer, and finally convert the output Z of the fully connected layer into a probability value P through the Sigmoid function. The probability value P is used to determine whether the image is a preprocessed image or an enhanced image;

[0163] When the probability value P is close to 1, it is judged as a preprocessed image; if the probability value P is close to 0, it is judged as an enhanced image.

[0164] Of course, for the probability that only needs to distinguish whether it is a true label, the binary cross-entropy loss function can be used to verify the accuracy of the discriminant network. Only when its accuracy reaches the threshold can it be used. When the accuracy is lower than the threshold, the backpropagation algorithm can be used to update all the parameters (convolution kernel weights, biases, fully connected layer weights and biases, etc.) in the discriminant network according to the calculated binary cross-entropy loss function L, so as to continuously improve the performance of the discriminant network. The specific update formula depends on the optimization algorithm used:

[0165] For example, using the Stochastic Gradient Descent (SGD) algorithm: where W can be any trainable parameter in the network, α is the learning rate, is the standard deviation of the binary cross-entropy loss function, is the standard deviation of the training parameter.

[0166] An image optimization system, characterized in that it uses the described image optimization method, including: a preprocessing module and an enhancement module;

[0167] The preprocessing module is used to preprocess the acquired original image, remove the noise of the original image, and obtain a preprocessed image;

[0168] The enhancement module is used to optimize the preprocessed image using a deep convolutional neural network and output an optimized image.

[0169] Preferably, the preprocessing module includes a noise module;

[0170] The noise module includes a pyramid module, a detection image acquisition module, a detection pixel point acquisition module, a detection range construction module, a potential point marking module, a first loop module, a second loop module, a positioning module, and a deletion module;

[0171] The pyramid module constructs a Gaussian pyramid. The original image is convolved with Gaussian kernels of different standard deviations through the Gaussian pyramid to obtain first images with multiple different scales, and all the first images are constructed into a first set;

[0172] The detection image acquisition module is used to acquire the first first image in the first set as the detection image;

[0173] The detection pixel point acquisition module is used to acquire any pixel point in the detection image as the detection pixel point;

[0174] The detection range construction module is used to construct a 3*3 first detection range in the detection image with the detection pixel point as the center;

[0175] Use the first images at adjacent scales of the detection image as the second images, and construct a 3*3 second detection range in the second images at the positions where the detection pixels are the same;

[0176] The potential point marking module is used to determine whether the detection pixel point is a gray value maximum or minimum point within the first detection range and the second detection range. If so, mark the detection pixel point as a potential point;

[0177] The first loop module is used to replace the next first image as the detection image, and re-invoke the detection range construction module and the potential point marking module until all the first images in the first set are looped, and then call the second loop module;

[0178] The second loop module is used to replace the next first image as the detection image, and re-invoke the detection pixel point acquisition module, the detection range construction module, the potential point marking module and the first loop module until all the first images in the first set are looped, and then call the positioning module;

[0179] The positioning module is used to perform Taylor expansion on all potential points to fit a three-dimensional quadratic function D(X), where X = [x, y, σ] T , (x, y) are the coordinates of the potential point respectively, and σ is the standard deviation used in the calculation of the scale of the first image where the potential point is located. Take the derivative of the three-dimensional quadratic function and set the derivative to zero to obtain the coordinates of the key point;

[0180] The deletion module is used to calculate the eigenvalue of each pixel point of each first image through the Hessian matrix; the pixel points with eigenvalues lower than the eigenvalue threshold are used as edge points;

[0181] Obtain the contrast between the key points and the edge points, and delete the key points and / or edge points with a contrast lower than the contrast threshold to obtain the preprocessed image.

[0182] Preferably, the preprocessing module further includes an alignment module;

[0183] The alignment module includes a matching unit, a selection unit, an affine unit, an inlier recording unit, an update unit and an alignment unit;

[0184] The matching unit is used to perform point matching on the original image using the SIFT algorithm to obtain a set of matching points;

[0185] The selection unit is used to randomly obtain N pairs of matching points from the set of matching points as the first matching points;

[0186] The affine unit is used to perform a two-dimensional affine transformation using the first matching points to obtain a transformed image;

[0187] The inlier recording unit is used to find the second matching points corresponding to the first matching points in the transformed image, obtain the Euclidean distance between the first matching points and the second matching points, and determine whether the Euclidean distance is less than the distance threshold. If it is less, the first matching points and the second matching points are marked as inliers, and the number of inliers is recorded;

[0188] The update unit is used to re-invoke the selection unit, the affine unit, and the inlier recording unit until the number of repetitions reaches the quantity threshold, obtain the transformed image with the largest number of inliers as the third image, and update the two-dimensional affine matrix using the second matching points in the third image and the corresponding first matching points in the original image to obtain an updated matrix;

[0189] The alignment unit is used to perform an alignment operation on the image using the updated matrix to update the preprocessed image.

[0190] Preferably, it further includes an image enhancement unit and a training unit;

[0191] The image enhancement unit is used to construct a first convolutional kernel and a residual module, perform feature extraction on the preprocessed image using the first convolutional kernel, and perform convolutional operations on the extracted features through the residual module to output an enhanced image after image enhancement;

[0192] The training unit is used to train a convolutional neural network using the preprocessed image and the enhanced image to obtain a trained convolutional neural network.

[0193] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0194] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for image optimization, characterized in that: The steps include: Step S1: preprocessing the acquired original image to remove noise from the original image to obtain a preprocessed image; Step S2: Use a deep convolutional neural network to optimize the preprocessed image and output the optimized image.

2. A method for image optimization according to claim 1, characterized in that: The preprocessing includes noise removal; The noise removal includes the following steps: Step A1: construct a Gaussian pyramid. The original image is convolved with Gaussian kernels of different standard deviations through the Gaussian pyramid to obtain first images with multiple different scales. A first collection is constructed from all the first images. Step A2: obtaining the first first image of the first collection as a detection image; Step A3: obtaining any pixel point in the detection image as a detection pixel point; Step A4: Using the detection pixel point as the center, construct a 3*3 first detection range in the detection image; The first image of the adjacent scale of the detection image is used as the second image, and a 3*3 second detection range is constructed in the second image at the same position of the detection pixel; Step A5: determining whether the detected pixel point is a grayscale value maximum or minimum point within the first detection range or the second detection range, and if so, marking the detected pixel point as a potential point; Step A6: Replace the next pixel point as the detection pixel point, and repeat steps A4 to A5 until all the pixels in the detection image are cycled, and then execute step A7; Step A7: Replace the next first image as the detection image, and repeat steps A3 to A6 until all the first images in the first collection are cycled, and then execute step A8; Step A8: Perform Taylor expansion on all potential points and fit a three-dimensional quadratic function D(X), where X = [x, y, σ] T , (x, y) are the coordinates of the potential point respectively, σ is the standard deviation used in the calculation of the first image scale where the potential point is located, the three-dimensional quadratic function is differentiated and the derivative is set to zero to obtain the coordinates of the key point; Step A9: For each first image, the eigenvalue of each pixel is calculated by using the Hessian matrix; The pixels whose eigenvalues ​​are lower than the eigenvalue threshold are regarded as edge points; The contrast between key points and edge points is obtained, and the key points and / or edge points with contrast lower than the contrast threshold are deleted to obtain a preprocessed image.

3. A method for image optimization according to claim 2, characterized in that: Preprocessing also includes image alignment; The steps of image alignment are as follows: Step B1: Use SIFT algorithm to perform point matching on the original image to obtain a set of matching points; Step B2: randomly obtain N pairs of matching points from the matching point set as the first matching points; Step B3: using the first matching point to perform a two-dimensional affine transformation to obtain a transformed image; Step B4: Find the second matching point corresponding to the first matching point in the changed image, obtain the Euclidean distance between the first matching point and the second matching point, and determine whether the Euclidean distance is less than a distance threshold. If so, mark the first matching point and the second matching point as inliers, and record the number of inliers. Step B5: Repeat steps B2 to B4 until the number of repetitions reaches a threshold value, obtain a changed image with the largest number of inliers as a third image, and use the second matching points in the third image and the first matching points corresponding to the original image to update the two-dimensional affine matrix to obtain an updated matrix; Step B6: Use the update matrix to perform an image alignment operation to update the pre-processed image.

4. The image optimization method according to claim 2 or 3, characterized in that: Before executing step S1, the original image needs to be divided into blocks.

5. The method for image optimization according to claim 4, characterized in that: Before executing step S2, the following steps need to be performed: Step S21: constructing a first convolution kernel and a residual module, using the first convolution kernel to extract features from the preprocessed image, performing a convolution operation on the extracted features through the residual module, and outputting an enhanced image after image enhancement; Step S22: training the convolutional neural network using the preprocessed image and the enhanced image to obtain a trained convolutional neural network; The training process for: Step S221: A loss function is defined by preprocessing the image and enhancing the image, wherein the loss function includes color loss, texture loss, content loss and total variation loss; Step S222: all losses are weighted and combined to obtain a final loss, and the convolutional neural network is updated using the back propagation algorithm and the final loss; The definition rules of color loss are as follows: The preprocessed image and the enhanced image are respectively subjected to a blur operation of a Gaussian operator, and the difference between the two blurred images is calculated as the color loss; The calculation process of color loss is as follows: Where X b , Y b They are the pre-processed image and enhanced image after blurring respectively; Where X[i+k,j+l] represents the pixel value at position (i+k,j+l), k and l are the indices of Gaussian kernel G(k,l) respectively; Among them A exp is an empirical constant, and its setting value is 0.1~1, u X ,u Y Represents the center position of the Gaussian distribution of the preprocessed image and the enhanced image, that is, the peak position of the Gaussian function, σ X ,σ Y Represents the mean of the Gaussian distribution of the preprocessed image and the enhanced image; The texture loss is calculated as follows: L texture =ΣlogD(F w (I s ) l I t ); Where D() is the discriminant network, F w () is a convolutional neural network, I t To preprocess the image, I s is the enhanced image, l is the index; The content loss is calculated as follows: Among them C j The number of channels of the feature map output by the ReLU_5_4 layer of the VGG-19 network, H j and W j Respectively represent the height and width of the feature map output by the ReLU_5_4 layer of the VGG-19 network, It means that the image generated by the convolutional neural network passes through the VGG-19 network and extracts the features of its ReLU_5_4 layer. It means that the preprocessed image passes through the VGG-19 network and extracts the features of its ReLU_5_4 layer; The total variation loss is calculated as follows: Among them C i , H i , W i is the number of channels, height and width of the image, and denote the gradient operators along the x-axis and y-axis respectively.

6. The method for image optimization according to claim 5, characterized in that: The discriminant network specifically performs the following steps: Step C1: converting the preprocessed image and the enhanced image into grayscale images respectively to obtain a collection of grayscale images; Step C2: The grayscale image collection is subjected to convolutional layer feature extraction, multi-layer convolution and pooling operations in sequence to obtain a feature map O L ; Step C3: For feature map O L Perform a flattening operation to convert it into a one-dimensional vector V; Input the one-dimensional vector V to the fully connected layer; Step C4: Get the output Z of the fully connected layer, and finally convert the output Z of the fully connected layer into a probability value P through the Sigmoid function. The probability value P is used to determine whether the image is a preprocessed image or an enhanced image. When the probability value P is close to 1, it is judged as a preprocessed image; If the probability value P is close to 0, it is judged as an enhanced image.

7. A system for image optimization, characterized in that: An image optimization method according to any one of claims 1 to 6, comprising: a preprocessing module and an enhancement module; The preprocessing module is used to preprocess the acquired original image to remove noise from the original image to obtain a preprocessed image; The enhancement module is used to optimize the preprocessed image using a deep convolutional neural network and output an optimized image.

8. The image optimization system according to claim 7, characterized in that: The preprocessing module includes a noise module; The noise module includes a pyramid module, a detection image acquisition module, a detection pixel acquisition module, a detection range construction module, a potential point marking module, a first cycle module, a second cycle module, a positioning module and a deletion module; The pyramid module constructs a Gaussian pyramid, and the original image is convolved with Gaussian kernels of different standard deviations through the Gaussian pyramid to obtain first images with multiple different scales, and a first collection is constructed from all the first images; The detection image acquisition module is used to acquire the first first image of the first collection as the detection image; The detection pixel point acquisition module is used to acquire any pixel point in the detection image as a detection pixel point; The detection range construction module is used to construct a 3*3 first detection range in the detection image using the detection pixel point as the center; The first image of the adjacent scale of the detection image is used as the second image, and a 3*3 second detection range is constructed in the second image at the same position of the detection pixel; The potential point marking module is used to determine whether the detected pixel point is a grayscale value maximum or minimum point within the first detection range and the second detection range, and if so, mark the detected pixel point as a potential point; The first loop module is used to replace the next first image as the detection image, and re-call the detection range construction module and the potential point marking module until all the first images in the first collection are cycled, and the second loop module is called; The second loop module is used to replace the next first image as the detection image, re-call the detection pixel acquisition module, the detection range construction module, the potential point marking module and the first loop module until all the first images in the first collection are looped, and the positioning module is called; The positioning module is used to perform Taylor expansion on all potential points and fit a three-dimensional quadratic function D(X), where X = [x, y, σ] T , (x, y) are the coordinates of the potential point respectively, σ is the standard deviation used in the calculation of the first image scale where the potential point is located, the three-dimensional quadratic function is differentiated and the derivative is set to zero to obtain the coordinates of the key point; The deletion module is used to calculate the eigenvalue of each pixel point of each first image through the Hessian matrix; The pixels whose eigenvalues ​​are lower than the eigenvalue threshold are regarded as edge points; The contrast between key points and edge points is obtained, and the key points and / or edge points with contrast lower than the contrast threshold are deleted to obtain a preprocessed image.

9. The image optimization system according to claim 8, characterized in that: The preprocessing module also includes an alignment module; The alignment module includes a matching unit, a selection unit, an affine unit, an interior point recording unit, an updating unit and an alignment unit; The matching unit is used to perform point matching on the original image using the SIFT algorithm to obtain a set of matching points; The selection unit is used to randomly obtain N pairs of matching points from the matching point collection as first matching points; The affine unit is used to perform a two-dimensional affine change using the first matching point to obtain a changed image; The inlier recording unit is used to find the second matching point corresponding to the first matching point in the changed image, obtain the Euclidean distance between the first matching point and the second matching point, and determine whether the Euclidean distance is less than a distance threshold. If so, the first matching point and the second matching point are marked as inliers, and the number of inliers is recorded; The updating unit is used to recall the selection unit, the affine unit and the inlier recording unit until the number of repetitions reaches a number threshold, obtain a changed image with the largest number of inliers as a third image, and use the second matching point in the third image and the first matching point corresponding to the original image to update the two-dimensional affine matrix to obtain an updated matrix; The alignment unit is used to perform an alignment operation on the image using an update matrix to update the pre-processed image.

10. The image optimization system according to claim 7, characterized in that: It also includes an image enhancement unit and a training unit; The image enhancement unit is used to construct a first convolution kernel and a residual module, use the first convolution kernel to extract features from the preprocessed image, perform a convolution operation on the extracted features through the residual module, and output an enhanced image after image enhancement; The training unit is used to train the convolutional neural network using the preprocessed image and the enhanced image to obtain the trained convolutional neural network.