Image correction method and device
By building a pre-trained distorted network, using the fast Fourier convolutional network and the upsampling network to determine the image offset field, iteratively correcting image distortion, solving the problem of expensive and cumbersome use of scanners, and achieving effective image correction and improvement of OCR recognition effect.
Patent Information
- Application Number
- CN202410028850.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the scanner is expensive and cumbersome to use, which limits its use scenarios, resulting in bending and distortion of objects such as photographed documents and documents, affecting image readability and OCR recognition effect.
By constructing a pre-trained distorted network, the offset field of the image to be corrected is determined using the fast Fourier convolutional network and the upsampling network, iteratively corrects, and the offset field of the image to be corrected is obtained and corrected to achieve flattening of the image.
The image distortion can be effectively corrected without using a scanner, improve the readability of the image and OCR recognition effect, and simplify the image correction process.
Smart Images

Figure CN120278927A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to an image correction method and apparatus. Background Art
[0002] Currently, when photographing objects such as documents and certificates, they are often interfered by perspective distortion and geometric distortion, making the objects such as documents and certificates in the photographed image present a curved and distorted state, which will have a negative impact on the readability of the image and the recognition effect of OCR (Optical Character Recognition).
[0003] For this reason, a scanner is used. The scanner uses technologies such as lasers or structured light to obtain the three-dimensional structure information of objects such as documents and certificates in the image, and then calculates the distortion parameters through geometric algorithms for the correction of objects such as documents and certificates in the image. However, scanners are expensive, and their use is relatively cumbersome. The bulky machines also limit the usage scenarios. Summary of the Invention
[0004] In order to solve the technical problems that the above-mentioned scanner is expensive, its use is relatively cumbersome, and the bulky machine also limits the usage scenarios, the embodiments of this application provide an image correction method and apparatus. The specific technical solutions are as follows:
[0005] In the first aspect of the embodiments of this application, first, an image correction method is provided. The method includes:
[0006] Obtain an image to be corrected, where the object to be corrected in the image to be corrected presents a distorted state;
[0007] Determine the offset field of the image to be corrected, where the offset field represents the offset direction and offset amount of the corresponding position in the image to be corrected;
[0008] Use the offset field to correct the image to be corrected to obtain a corrected image, where the corrected object in the corrected image presents a flattened state.
[0009] In an optional implementation manner, the determining the offset field of the image to be corrected includes:
[0010] Input the image to be corrected into a pre-trained distortion network, and obtain the offset field of the image to be corrected output by the pre-trained distortion network.
[0011] In an optional implementation manner, the pre-trained distortion network includes a pre-trained downsampling network, a pre-trained fast Fourier convolution network, and a pre-trained upsampling network;
[0012] Inputting the image to be corrected into a pre-trained warping network to obtain the offset field of the image to be corrected output by the pre-trained warping network includes:
[0013] Performing downsampling on the image to be corrected using the pre-trained downsampling network to obtain the downsampled image to be corrected;
[0014] Processing the downsampled image to be corrected using the pre-trained fast Fourier convolutional network to obtain the feature map of the image to be corrected;
[0015] Performing upsampling on the feature map using the pre-trained upsampling network to obtain the offset field of the image to be corrected.
[0016] In an optional embodiment, the pre-trained fast Fourier convolutional network includes a pre-trained first convolutional layer, a pre-trained second convolutional layer, a pre-trained third convolutional layer, and a pre-trained spectral conversion network;
[0017] The processing the downsampled image to be corrected using the pre-trained fast Fourier convolutional network to obtain the feature map of the image to be corrected includes:
[0018] Performing first convolution on the downsampled image to be corrected using the pre-trained first convolutional layer to obtain a first convolution result;
[0019] Performing second convolution on the downsampled image to be corrected using the pre-trained second convolutional layer to obtain a second convolution result;
[0020] Performing third convolution on the downsampled image to be corrected using the pre-trained third convolutional layer to obtain a third convolution result;
[0021] Performing spectral conversion on the downsampled image to be corrected using the pre-trained spectral conversion network to obtain a spectral conversion result;
[0022] Determining the feature map of the image to be corrected according to the first convolution result, the second convolution result, the third convolution result, and the spectral conversion result.
[0023] In an optional embodiment, the pre-trained fast Fourier convolutional network further includes a pre-trained first batch normalization layer, a pre-trained first activation function layer, a pre-trained second batch normalization layer, and a pre-trained second activation function layer;
[0024] Determining the feature map of the image to be corrected according to the first convolution result, the second convolution result, the third convolution result, and the spectral conversion result includes:
[0025] Adding the first convolution result and the third convolution result to obtain a first addition result;
[0026] Performing a first normalization process on the first addition result using the pre-trained first batch normalization layer to obtain a first normalization result;
[0027] Performing a first activation process on the first normalization result using the pre-trained first activation function layer to obtain local features;
[0028] Adding the second convolution result and the spectral conversion result to obtain a second addition result;
[0029] Performing a second normalization process on the second addition result using the pre-trained second batch normalization layer to obtain a second normalization result;
[0030] Performing a second activation process on the second normalization result using the pre-trained second activation function layer to obtain global features;
[0031] Superimposing the local features and the global features to obtain the feature map of the image to be corrected.
[0032] In an optional implementation manner, the pre-trained spectral conversion network includes a pre-trained fourth convolution layer, a pre-trained third batch normalization layer, a pre-trained third activation function layer, and a pre-trained fifth convolution layer;
[0033] The spectral conversion process of using the pre-trained spectral conversion network to perform spectral conversion on the downsampled image to be corrected to obtain a spectral conversion result includes:
[0034] Performing a fourth convolution process on the downsampled image to be corrected using the pre-trained fourth convolution layer to obtain a fourth convolution result;
[0035] Performing a third normalization process on the fourth convolution result using the pre-trained third batch normalization layer to obtain a third normalization result;
[0036] Performing a third activation process on the third normalization result using the pre-trained third activation function layer to obtain a first activation result;
[0037] Performing a spectral conversion process on the first activation result to obtain an original spectral conversion result;
[0038] Add the first activation result to the original spectral conversion result to obtain a third addition result;
[0039] Perform fifth convolution processing on the third addition result using the pre-trained fifth convolution layer to obtain a spectral conversion result.
[0040] In an optional implementation manner, the pre-trained spectral conversion network further includes a two-dimensional Fourier transform layer, a pre-trained sixth convolution layer, a pre-trained fourth batch normalization layer, a pre-trained fourth activation function layer, and a two-dimensional inverse Fourier transform layer;
[0041] The performing spectral conversion processing on the first activation result to obtain an original spectral conversion result includes:
[0042] Perform two-dimensional Fourier transform processing on the first activation result using the two-dimensional Fourier transform layer to obtain a frequency domain result;
[0043] Perform sixth convolution processing on the frequency domain result using the pre-trained sixth convolution layer to obtain a sixth convolution result;
[0044] Perform fourth normalization processing on the sixth convolution result using the pre-trained fourth batch normalization layer to obtain a fourth normalization result;
[0045] Perform fourth activation processing on the fourth normalization result using the pre-trained fourth activation function layer to obtain a second activation result;
[0046] Perform two-dimensional inverse Fourier transform processing on the second activation result using the two-dimensional inverse Fourier transform layer to obtain an original spectral conversion result.
[0047] In an optional implementation manner, after using the offset field to correct the image to be corrected to obtain a corrected image, the method further includes:
[0048] Iterate according to the following steps until a preset iteration termination condition is satisfied to obtain an iterative offset field in each iteration process:
[0049] Obtain the target image of the current iteration, where the target image of the first iteration is the corrected image, and the target image of a non-first iteration is the iterative image obtained in the previous iteration;
[0050] Input the target image into the pre-trained warping network to obtain the iterative offset field of the target image output by the pre-trained warping network;
[0051] Correct the target image using the iterative offset field to obtain an iterative image, and use the iterative image as the target image for the next iteration;
[0052] Correct the image to be corrected according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image.
[0053] In an optional embodiment, the preset iterative termination condition includes one of the following:
[0054] The number of iterations reaches a preset iteration number threshold;
[0055] The variance of the iterative offset field in the current iteration process is less than a preset variance threshold;
[0056] The variance of the iterative offset field in the current iteration process is greater than the variance of the iterative offset field in the previous iteration process.
[0057] In an optional embodiment, the step of correcting the image to be corrected according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image includes:
[0058] Add the offset field and the iterative offset field in each iteration process to obtain a final offset field;
[0059] Use the final offset field to correct the image to be corrected to obtain a final corrected image.
[0060] In an optional embodiment, before executing the method, it further includes:
[0061] Obtain a sample image set, where the objects in the sample images of the sample image set are in a distorted state;
[0062] Input the sample images into a distortion network to obtain the sample offset fields of the sample images output by the distortion network;
[0063] Use the sample offset fields to correct the sample images to obtain sample corrected images, and train the distortion network according to the sample corrected images;
[0064] Count the number of training times of the distortion network, and stop training when the number of training times reaches a preset number threshold to obtain the pre-trained distortion network.
[0065] In an optional embodiment, the step of obtaining the sample image set includes:
[0066] Obtain a preset data set, where the preset data set includes original images to be corrected and the original offset fields corresponding to the original images to be corrected;
[0067] Detect the object regions in the original images to be corrected, and determine the minimum bounding rectangles of the object regions;
[0068] Crop the original image to be corrected using the minimum bounding rectangle to obtain the object region, and determine the object region as the first sample image;
[0069] Crop the original offset field corresponding to the original image to be corrected using the minimum bounding rectangle to obtain the object offset field;
[0070] Obtain the original image, and perform distortion processing on the original image using the object offset field to obtain a distorted image;
[0071] Determine the distorted image as the second sample image, and store the first sample image and the second sample image in the sample image set.
[0072] In an alternative embodiment, the distortion network includes a downsampling network, a fast Fourier convolutional network, and an upsampling network;
[0073] The step of inputting the sample image into the distortion network to obtain the sample offset field of the sample image output by the distortion network includes:
[0074] Perform downsampling processing on the sample image using the downsampling network to obtain the sample image after downsampling processing;
[0075] Process the sample image after downsampling using the fast Fourier convolutional network to obtain the sample feature map of the sample image;
[0076] Perform upsampling processing on the sample feature map using the upsampling network to obtain the sample offset field of the sample image.
[0077] In an alternative embodiment, the fast Fourier convolutional network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a spectral conversion network;
[0078] The step of processing the sample image after downsampling using the fast Fourier convolutional network to obtain the sample feature map of the sample image includes:
[0079] Perform first convolutional processing on the sample image after downsampling using the first convolutional layer to obtain a first sample convolutional result;
[0080] Perform second convolutional processing on the sample image after downsampling using the second convolutional layer to obtain a second sample convolutional result;
[0081] Perform third convolutional processing on the sample image after downsampling using the third convolutional layer to obtain a third sample convolutional result;
[0082] Perform spectral conversion processing on the downsampled sample image using the spectral conversion network to obtain a sample spectral conversion result;
[0083] Determine the sample feature map of the sample image according to the first sample convolution result, the second sample convolution result, the third sample convolution result, and the sample spectral conversion result.
[0084] In an alternative embodiment, the fast Fourier convolution network further includes a first batch normalization layer, a first activation function layer, a second batch normalization layer, and a second activation function layer;
[0085] The determining the sample feature map of the sample image according to the first sample convolution result, the second sample convolution result, the third sample convolution result, and the sample spectral conversion result includes:
[0086] Add the first sample convolution result and the third sample convolution result to obtain a first sample addition result;
[0087] Perform first normalization processing on the first sample addition result using the first batch normalization layer to obtain a first sample normalization result;
[0088] Perform first activation processing on the first sample normalization result using the first activation function layer to obtain sample local features;
[0089] Add the second sample convolution result and the sample spectral conversion result to obtain a second sample addition result;
[0090] Perform second normalization processing on the second sample addition result using the second batch normalization layer to obtain a second sample normalization result;
[0091] Perform second activation processing on the second sample normalization result using the second activation function layer to obtain sample global features;
[0092] Superimpose the sample local features and the sample global features to obtain the sample feature map of the sample image.
[0093] In an alternative embodiment, the spectral conversion network includes a fourth convolution layer, a third batch normalization layer, a third activation function layer, and a fifth convolution layer;
[0094] The performing spectral conversion processing on the downsampled sample image using the spectral conversion network to obtain a sample spectral conversion result includes:
[0095] Performing a fourth convolution process on the downsampled sample image using the fourth convolution layer to obtain a fourth sample convolution result;
[0096] Performing a third normalization process on the fourth sample convolution result using the third batch normalization layer to obtain a third sample normalization result;
[0097] Performing a third activation process on the third sample normalization result using the third activation function layer to obtain a first sample activation result;
[0098] Performing a spectral conversion process on the first sample activation result to obtain an original sample spectral conversion result;
[0099] Adding the first sample activation result and the original sample spectral conversion result to obtain a third sample addition result;
[0100] Performing a fifth convolution process on the third sample addition result using the fifth convolution layer to obtain a sample spectral conversion result.
[0101] In an optional embodiment, the spectral conversion network further includes a two-dimensional Fourier transform layer, a sixth convolution layer, a fourth batch normalization layer, a fourth activation function layer, and a two-dimensional inverse Fourier transform layer;
[0102] The performing a spectral conversion process on the first sample activation result to obtain an original sample spectral conversion result includes:
[0103] Performing a two-dimensional Fourier transform process on the first sample activation result using the two-dimensional Fourier transform layer to obtain a sample frequency domain result;
[0104] Performing a sixth convolution process on the sample frequency domain result using the sixth convolution layer to obtain a sixth sample convolution result;
[0105] Performing a fourth normalization process on the sixth sample convolution result using the fourth batch normalization layer to obtain a fourth sample normalization result;
[0106] Performing a fourth activation process on the fourth sample normalization result using the fourth activation function layer to obtain a second sample activation result;
[0107] Performing a two-dimensional inverse Fourier transform process on the second sample activation result using the two-dimensional inverse Fourier transform layer to obtain an original sample spectral conversion result.
[0108] In an optional embodiment, the training the distortion network according to the sample correction image includes:
[0109] Iterate according to the following steps until the preset iteration termination condition is met, and obtain the sample iteration offset field in each iteration process:
[0110] Obtain the sample target image of the current iteration. Among them, the sample target image of the first iteration is the sample corrected image, and the sample target image of a non-first iteration is the sample iteration image obtained in the previous iteration;
[0111] Input the sample target image into the distortion network to obtain the sample iteration offset field of the sample target image output by the distortion network;
[0112] Correct the sample target image by using the sample iteration offset field to obtain a sample iteration image, and use the sample iteration image as the sample target image for the next iteration;
[0113] Correct the sample image according to the sample offset field and the sample iteration offset field in each iteration process to obtain the final sample corrected image;
[0114] Train the distortion network according to the final sample corrected image.
[0115] In an optional embodiment, the preset iteration termination condition includes one of the following:
[0116] The number of iterations reaches the preset iteration number threshold;
[0117] The variance of the sample iteration offset field in the current iteration process is less than the preset variance threshold;
[0118] The variance of the sample iteration offset field in the current iteration process is greater than the variance of the sample iteration offset field in the previous iteration process.
[0119] In an optional embodiment, the correcting the sample image according to the sample offset field and the sample iteration offset field in each iteration process to obtain the final sample corrected image includes:
[0120] Add the sample offset field and the sample iteration offset field in each iteration process to obtain the final sample offset field;
[0121] Correct the sample image by using the final sample offset field to obtain the final sample corrected image.
[0122] In the second aspect of the embodiments of the present application, an image correction device is further provided. The device includes:
[0123] An image acquisition module, configured to acquire an image to be corrected, where the object to be corrected in the image to be corrected is in a distorted state;
[0124] An offset field determination module for determining the offset field of the image to be corrected, where the offset field characterizes the offset direction and offset amount at the corresponding position in the image to be corrected;
[0125] An image correction module for correcting the image to be corrected by using the offset field to obtain a corrected image, where the correction object in the corrected image is in a flattened state.
[0126] In an optional implementation manner, the offset field determination module is specifically configured to:
[0127] Input the image to be corrected into a pre-trained warping network, and obtain the offset field of the image to be corrected output by the pre-trained warping network.
[0128] In an optional implementation manner, the pre-trained warping network includes a pre-trained downsampling network, a pre-trained fast Fourier convolutional network, and a pre-trained upsampling network;
[0129] The offset field determination module includes:
[0130] An image downsampling processing sub-module for performing downsampling processing on the image to be corrected by using the pre-trained downsampling network to obtain the image to be corrected after downsampling processing;
[0131] An image processing sub-module for processing the image to be corrected after downsampling processing by using the pre-trained fast Fourier convolutional network to obtain the feature map of the image to be corrected;
[0132] A feature map upsampling processing sub-module for performing upsampling processing on the feature map by using the pre-trained upsampling network to obtain the offset field of the image to be corrected.
[0133] In an optional implementation manner, the pre-trained fast Fourier convolutional network includes a pre-trained first convolutional layer, a pre-trained second convolutional layer, a pre-trained third convolutional layer, and a pre-trained spectral conversion network;
[0134] The image processing sub-module includes:
[0135] A first convolutional processing unit for performing first convolutional processing on the image to be corrected after downsampling processing by using the pre-trained first convolutional layer to obtain a first convolutional result;
[0136] A second convolutional processing unit for performing second convolutional processing on the image to be corrected after downsampling processing by using the pre-trained second convolutional layer to obtain a second convolutional result;
[0137] The fourth convolution processing unit is configured to perform a third convolution process on the to-be-corrected image after downsampling using the pre-trained third convolution layer, to obtain a third convolution result;
[0138] The spectral conversion processing unit is configured to perform a spectral conversion process on the to-be-corrected image after downsampling using the pre-trained spectral conversion network, to obtain a spectral conversion result;
[0139] The feature map determination subunit is configured to determine a feature map of the to-be-corrected image according to the first convolution result, the second convolution result, the third convolution result, and the spectral conversion result.
[0140] In an optional embodiment, the pre-trained fast Fourier convolution network further includes a pre-trained first batch normalization layer, a pre-trained first activation function layer, a pre-trained second batch normalization layer, and a pre-trained second activation function layer;
[0141] The feature map determination subunit is specifically configured to:
[0142] Add the first convolution result and the third convolution result to obtain a first addition result;
[0143] Perform a first normalization process on the first addition result using the pre-trained first batch normalization layer to obtain a first normalization result;
[0144] Perform a first activation process on the first normalization result using the pre-trained first activation function layer to obtain local features;
[0145] Add the second convolution result and the spectral conversion result to obtain a second addition result;
[0146] Perform a second normalization process on the second addition result using the pre-trained second batch normalization layer to obtain a second normalization result;
[0147] Perform a second activation process on the second normalization result using the pre-trained second activation function layer to obtain global features;
[0148] Superimpose the local features and the global features to obtain the feature map of the to-be-corrected image.
[0149] In an optional embodiment, the pre-trained spectral conversion network includes a pre-trained fourth convolution layer, a pre-trained third batch normalization layer, a pre-trained third activation function layer, and a pre-trained fifth convolution layer;
[0150] The spectral conversion processing unit includes:
[0151] The fourth convolution processing subunit is configured to perform fourth convolution processing on the to-be-corrected image after downsampling processing by using the pre-trained fourth convolution layer, so as to obtain a fourth convolution result;
[0152] The third normalization processing subunit is configured to perform third normalization processing on the fourth convolution result by using the pre-trained third batch normalization layer, so as to obtain a third normalization result;
[0153] The third activation processing subunit is configured to perform third activation processing on the third normalization result by using the pre-trained third activation function layer, so as to obtain a first activation result;
[0154] The spectral conversion processing subunit is configured to perform spectral conversion processing on the first activation result, so as to obtain an original spectral conversion result;
[0155] The result addition subunit is configured to add the first activation result and the original spectral conversion result, so as to obtain a third addition result;
[0156] The fifth convolution processing subunit is configured to perform fifth convolution processing on the third addition result by using the pre-trained fifth convolution layer, so as to obtain a spectral conversion result.
[0157] In an optional implementation manner, the pre-trained spectral conversion network further includes a two-dimensional Fourier transform layer, a pre-trained sixth convolution layer, a pre-trained fourth batch normalization layer, a pre-trained fourth activation function layer, and a two-dimensional inverse Fourier transform layer;
[0158] The spectral conversion processing subunit is specifically configured to:
[0159] Perform two-dimensional Fourier transform processing on the first activation result by using the two-dimensional Fourier transform layer, so as to obtain a frequency domain result;
[0160] Perform sixth convolution processing on the frequency domain result by using the pre-trained sixth convolution layer, so as to obtain a sixth convolution result;
[0161] Perform fourth normalization processing on the sixth convolution result by using the pre-trained fourth batch normalization layer, so as to obtain a fourth normalization result;
[0162] Perform fourth activation processing on the fourth normalization result by using the pre-trained fourth activation function layer, so as to obtain a second activation result;
[0163] Perform two-dimensional inverse Fourier transform processing on the second activation result by using the two-dimensional inverse Fourier transform layer, so as to obtain an original spectral conversion result.
[0164] In an optional implementation manner, the device further includes:
[0165] An iterative module for iterating according to the following steps until a preset iteration termination condition is met, and obtaining an iterative offset field in each iteration process:
[0166] A target image acquisition module for acquiring a target image of the current iteration. Among them, the target image of the first iteration is the corrected image, and the target image of a non-first iteration is the iterative image obtained in the previous iteration;
[0167] An iterative offset field determination module for inputting the target image into the pre-trained warping network and obtaining the iterative offset field of the target image output by the pre-trained warping network;
[0168] A target image correction module for correcting the target image by using the iterative offset field to obtain an iterative image, and using the iterative image as the target image for the next iteration;
[0169] A to-be-corrected image correction module for correcting the to-be-corrected image according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image.
[0170] In an optional implementation manner, the preset iteration termination condition includes one of the following:
[0171] The number of iterations reaches a preset iteration number threshold;
[0172] The variance of the iterative offset field in the current iteration process is less than a preset variance threshold;
[0173] The variance of the iterative offset field in the current iteration process is greater than the variance of the iterative offset field in the previous iteration process.
[0174] In an optional implementation manner, the to-be-corrected image correction module is specifically configured to:
[0175] Add the offset field and the iterative offset field in each iteration process to obtain a final offset field;
[0176] Use the final offset field to correct the to-be-corrected image to obtain a final corrected image.
[0177] In an optional implementation manner, the device further includes:
[0178] An image set acquisition module for acquiring a sample image set, where the sample images in the sample image set have the state that the objects are distorted;
[0179] A sample offset field determination module for inputting the sample image into the warping network and obtaining the sample offset field of the sample image output by the warping network;
[0180] A sample image correction module, which is used to correct the sample image by using the sample offset field to obtain a sample corrected image;
[0181] A model training module, which is used to train the distortion network according to the sample corrected image;
[0182] A training times statistics module, which is used to count the training times of the distortion network, and stop training when the training times reach a preset times threshold to obtain the pre-trained distortion network.
[0183] In an optional implementation manner, the image set acquisition module is specifically used for:
[0184] Obtain a preset data set, where the preset data set includes original images to be corrected and the original offset fields corresponding to the original images to be corrected;
[0185] Detect the object region in the original image to be corrected, and determine the minimum bounding rectangle of the object region;
[0186] Crop the original image to be corrected by using the minimum bounding rectangle to obtain the object region, and determine the object region as the first sample image;
[0187] Crop the original offset field corresponding to the original image to be corrected by using the minimum bounding rectangle to obtain an object offset field;
[0188] Obtain an original image, and perform distortion processing on the original image by using the object offset field to obtain a distorted image;
[0189] Determine the distorted image as the second sample image, and store the first sample image and the second sample image in the sample image set.
[0190] In an optional implementation manner, the distortion network includes a downsampling network, a fast Fourier convolutional network, and an upsampling network;
[0191] The sample offset field determination module specifically includes:
[0192] A sample image downsampling processing sub-module, which is used to perform downsampling processing on the sample image by using the downsampling network to obtain the sample image after downsampling processing;
[0193] A sample feature map determination sub-module, which is used to process the sample image after downsampling processing by using the fast Fourier convolutional network to obtain the sample feature map of the sample image;
[0194] The sample feature map upsampling processing sub-module is used to perform upsampling processing on the sample feature map by using the upsampling network to obtain the sample offset field of the sample image.
[0195] In an optional embodiment, the fast Fourier convolutional network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a spectral conversion network;
[0196] The sample feature map determination sub-module specifically includes:
[0197] The first convolutional processing unit is used to perform first convolutional processing on the downsampled sample image by using the first convolutional layer to obtain a first sample convolutional result;
[0198] The second convolutional processing unit is used to perform second convolutional processing on the downsampled sample image by using the second convolutional layer to obtain a second sample convolutional result;
[0199] The third convolutional processing unit is used to perform third convolutional processing on the downsampled sample image by using the third convolutional layer to obtain a third sample convolutional result;
[0200] The spectral conversion processing unit is used to perform spectral conversion processing on the downsampled sample image by using the spectral conversion network to obtain a sample spectral conversion result;
[0201] The sample feature map determination unit is used to determine the sample feature map of the sample image according to the first sample convolutional result, the second sample convolutional result, the third sample convolutional result, and the sample spectral conversion result.
[0202] In an optional embodiment, the fast Fourier convolutional network further includes a first batch normalization layer, a first activation function layer, a second batch normalization layer, and a second activation function layer;
[0203] The sample feature map determination unit is specifically used for:
[0204] Adding the first sample convolutional result and the third sample convolutional result to obtain a first sample addition result;
[0205] Performing first normalization processing on the first sample addition result by using the first batch normalization layer to obtain a first sample normalization result;
[0206] Performing first activation processing on the first sample normalization result by using the first activation function layer to obtain sample local features;
[0207] Adding the second sample convolutional result and the sample spectral conversion result to obtain a second sample addition result;
[0208] Perform a second normalization process on the second sample addition result using the second batch normalization layer to obtain a second sample normalization result;
[0209] Perform a second activation process on the second sample normalization result using the second activation function layer to obtain a sample global feature;
[0210] Superimpose the sample local feature and the sample global feature to obtain a sample feature map of the sample image.
[0211] In an alternative embodiment, the spectral conversion network includes a fourth convolutional layer, a third batch normalization layer, a third activation function layer, and a fifth convolutional layer;
[0212] The spectral conversion processing unit specifically includes:
[0213] A fourth convolutional processing sub-unit for performing a fourth convolutional process on the sample image after downsampling using the fourth convolutional layer to obtain a fourth sample convolutional result;
[0214] A third normalization processing sub-unit for performing a third normalization process on the fourth sample convolutional result using the third batch normalization layer to obtain a third sample normalization result;
[0215] A third activation processing sub-unit for performing a third activation process on the third sample normalization result using the third activation function layer to obtain a first sample activation result;
[0216] A spectral conversion processing sub-unit for performing a spectral conversion process on the first sample activation result to obtain an original sample spectral conversion result;
[0217] A result addition sub-unit for adding the first sample activation result and the original sample spectral conversion result to obtain a third sample addition result;
[0218] A fifth convolutional processing sub-unit for performing a fifth convolutional process on the third sample addition result using the fifth convolutional layer to obtain a sample spectral conversion result.
[0219] In an alternative embodiment, the spectral conversion network further includes a two-dimensional Fourier transform layer, a sixth convolutional layer, a fourth batch normalization layer, a fourth activation function layer, and a two-dimensional inverse Fourier transform layer;
[0220] The spectral conversion processing sub-unit is specifically used for:
[0221] Performing a two-dimensional Fourier transform process on the first sample activation result using the two-dimensional Fourier transform layer to obtain a sample frequency domain result;
[0222] Perform a sixth convolution process on the sample frequency domain result using the sixth convolutional layer to obtain a sixth sample convolution result;
[0223] Perform a fourth normalization process on the sixth sample convolution result using the fourth batch normalization layer to obtain a fourth sample normalization result;
[0224] Perform a fourth activation process on the fourth sample normalization result using the fourth activation function layer to obtain a second sample activation result;
[0225] Perform an inverse two-dimensional Fourier transform process on the second sample activation result using the two-dimensional inverse Fourier transform layer to obtain an original sample spectral conversion result.
[0226] In an alternative embodiment, the model training module specifically includes:
[0227] An iteration sub-module for iterating according to the following steps until a preset iteration termination condition is met, to obtain a sample iteration offset field in each iteration process:
[0228] A sample target image acquisition sub-module for acquiring a sample target image of the current iteration, where the sample target image of the first iteration is the sample correction image, and the sample target image of a non-first iteration is the sample iteration image obtained in the previous iteration;
[0229] A sample iteration offset field determination sub-module for inputting the sample target image into the warping network to obtain a sample iteration offset field of the sample target image output by the warping network;
[0230] A sample target image correction sub-module for correcting the sample target image using the sample iteration offset field to obtain a sample iteration image, and using the sample iteration image as the sample target image for the next iteration;
[0231] A sample image correction sub-module for correcting the sample image according to the sample offset field and the sample iteration offset field in each iteration process to obtain a final sample correction image;
[0232] A model training sub-module for training the warping network according to the final sample correction image.
[0233] In an alternative embodiment, the preset iteration termination condition includes one of the following:
[0234] The number of iterations reaches a preset iteration number threshold;
[0235] The variance of the sample iteration offset field in the current iteration process is less than a preset variance threshold;
[0236] The variance of the sample iteration offset field in the current iteration process is greater than the variance of the sample iteration offset field in the previous iteration process.
[0237] In an optional embodiment, the sample image correction sub-module is specifically configured to:
[0238] Add the sample offset field and the sample iteration offset field in each iteration process to obtain a final sample offset field;
[0239] Use the final sample offset field to correct the sample image to obtain a final sample corrected image.
[0240] In a third aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0241] The memory is used to store a computer program;
[0242] The processor is configured to implement the image correction method described in any one of the above first aspects when executing the program stored on the memory.
[0243] In a fourth aspect of the embodiments of the present application, a storage medium is further provided. Instructions are stored in the storage medium, and when it runs on a computer, the computer is caused to execute the image correction method described in any one of the above first aspects.
[0244] In a fifth aspect of the embodiments of the present application, a computer program product including instructions is further provided. When it runs on a computer, the computer is caused to execute the image correction method described in any one of the above.
[0245] The technical solution provided by the embodiments of the present application obtains an image to be corrected. Among them, the object to be corrected in the image to be corrected is in a distorted state. Determine the offset field of the image to be corrected. The offset field represents the offset direction and offset amount of the corresponding position in the image to be corrected. Use the offset field to correct the image to be corrected to obtain a corrected image. Among them, the corrected object in the corrected image is in a flattened state. By determining the offset field of the image to be corrected and using the offset field to correct the image to be corrected to obtain a corrected image, it is possible to avoid using a scanner and still correct the object to be corrected in the image to be corrected, avoiding negative impacts on the readability of the image and the recognition effect of OCR. Description of the Drawings
[0246] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.
[0247] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0248] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the figures in the drawings do not constitute a proportional limitation.
[0249] Figure 1 It is a schematic diagram of the implementation process of an image correction method shown in the embodiments of the present application;
[0250] Figure 2 It is a schematic diagram of a document image to be corrected shown in the embodiments of the present application;
[0251] Figure 3 It is a schematic diagram of an offset field shown in the embodiments of the present application;
[0252] Figure 4 It is a schematic diagram of edge loss completion in a corrected image shown in the embodiments of the present application;
[0253] Figure 5 It is a schematic diagram of the implementation process of another image correction method shown in the embodiments of the present application;
[0254] Figure 6 It is a schematic diagram of the implementation process of another image correction method shown in the embodiments of the present application;
[0255] Figure 7 It is a schematic diagram of the implementation process of a feature map determination method shown in the embodiments of the present application;
[0256] Figure 8 It is a schematic diagram of the structure of a pre-trained spectral conversion network shown in the embodiments of the present application;
[0257] Figure 9 It is a schematic diagram of the implementation process of another image correction method shown in the embodiments of the present application;
[0258] Figure 10 It is a schematic diagram of the implementation process of another image correction method shown in the embodiments of the present application;
[0259] Figure 11 It is a schematic diagram of the implementation process of a distortion network training method shown in the embodiments of the present application;
[0260] Figure 12 Schematic diagram of the implementation process of another distorted network training method shown in the embodiments of the present application;
[0261] Figure 13 Schematic diagram of the structure of an image correction device shown in the embodiments of the present application;
[0262] Figure 14 Schematic diagram of the structure of an electronic device shown in the embodiments of the present application. Detailed implementation manners
[0263] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0264] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0265] As Figure 1 shown, it is a schematic diagram of the implementation process of an image correction method provided by the embodiments of the present application. This method is applied to an electronic device and may specifically include the following steps:
[0266] S101. Obtain an image to be corrected, where the object to be corrected in the image to be corrected is in a distorted state.
[0267] In the embodiments of the present application, when currently shooting objects such as documents and certificates, they are often interfered by perspective distortion and geometric distortion, making the objects such as documents and certificates in the captured image present a curved and distorted state. Therefore, it is necessary to correct the objects such as documents and certificates in the image so that the objects such as documents and certificates in the image present a flattened state.
[0268] Therefore, the above-mentioned image and the like can be used as the image to be corrected, and thus the image to be corrected is obtained, where the object to be corrected in the image to be corrected is in a curved and distorted state (referring to various deformations of the object to be corrected). For example, obtain an image of a document to be corrected, where the document to be corrected in the image of the document to be corrected is in a curved and distorted state, and the curved and distorted state is as Figure 2as shown
[0269] It should be noted that for the object to be corrected in the image to be corrected, for example, in addition to being a document or a certificate, it can also be other objects such as books, and the embodiments of the present application do not limit this.
[0270] S102. Determine the offset field of the image to be corrected, where the offset field represents the offset direction and offset amount at the corresponding position in the image to be corrected.
[0271] In the embodiments of the present application, for the image to be corrected, the offset field of the image to be corrected can be determined, and the offset field represents the offset direction and offset amount at the corresponding position in the image to be corrected. Among them, based on a deep learning model, the offset field of the image to be corrected can be determined. For the offset field, for example, as Figure 3 as shown
[0272] It should be noted that for the offset field, it is an image with a direction and a magnitude, and it is similar to an image gradient. The shape of the offset field is the same as that of the image to be corrected. Each arrow in the offset field is a vector, containing information such as direction and magnitude, that is, the offset field represents the direction in which the corresponding position in the image to be corrected needs to be offset and the offset amount.
[0273] S103. Use the offset field to correct the image to be corrected to obtain a corrected image, where the corrected object in the corrected image is in a flattened state.
[0274] In the embodiments of the present application, for the offset field, the offset field can be used to correct the image to be corrected to obtain a corrected image, where the corrected object in the corrected image is in a flattened state (referring to that the object to be corrected is in a normal state without various deformations), and the flattened state is as Figure 4 as shown. And during the correction of the image to be corrected, it is necessary to use the remap algorithm for correction, which means using the remap algorithm to correct the image to be corrected with reference to the offset field to obtain a corrected image.
[0275] It should be noted that for the corrected image, the corrected image can be mirror-expanded and morphologically processed to cleverly complete the edge loss compensation, such as the effect of edge loss compensation indicated by the arrows as Figure 4 shown. In addition, the offset field can be smoothed. In this way, using the remap algorithm, the image to be corrected is corrected with reference to the smoothed offset field to obtain a corrected image, and the image details in the corrected image are smoother and more natural.
[0276] Through the description of the technical solution provided in the embodiments of the present application above, a to-be-corrected image is obtained, wherein the object to be corrected in the to-be-corrected image is in a distorted state. An offset field of the to-be-corrected image is determined, and the offset field represents the offset direction and offset amount at the corresponding position in the to-be-corrected image. The to-be-corrected image is corrected by using the offset field to obtain a corrected image, wherein the corrected object in the corrected image is in a flattened state.
[0277] By determining the offset field of the to-be-corrected image and using the offset field to correct the to-be-corrected image to obtain a corrected image, the use of a scanner is avoided, and the object to be corrected in the to-be-corrected image can still be corrected, avoiding negative impacts on the readability of the image and the recognition effect of OCR.
[0278] In addition, in the embodiments of the present application, the correction of an object in an image can be implemented based on deep learning. In order to construct a warping network, the offset field of the image can be obtained through this warping network, so as to correct the object in the image by using the offset field. Based on this, as Figure 5 shown, it is a schematic diagram of the implementation process of another image correction method provided in the embodiments of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0279] S501, obtain a to-be-corrected image, wherein the object to be corrected in the to-be-corrected image is in a distorted state.
[0280] In the embodiments of the present application, this step is similar to step S101 above, and the embodiments of the present application will not elaborate here one by one.
[0281] S502, input the to-be-corrected image into a pre-trained warping network, and obtain the offset field of the to-be-corrected image output by the pre-trained warping network. The offset field represents the offset direction and offset amount at the corresponding position in the to-be-corrected image.
[0282] In the embodiments of the present application, a warping network is constructed. The overall design of the warping network is an Encoder-decoder structure. The warping network is pre-trained to obtain a pre-trained warping network. Specifically, the subsequent training process of the warping network will be described.
[0283] For the to-be-corrected image, the to-be-corrected image can be input into a pre-trained warping network, and the offset field of the to-be-corrected image output by the pre-trained warping network can be obtained. The offset field represents the offset direction and offset amount at the corresponding position in the to-be-corrected image.
[0284] It should be noted that for ordinary Encoder-decoder structures, similar to Unet, traditional convolution is used, and the receptive field will be limited. For scenarios such as image correction, it is necessary to grasp the overall distortion process of the object to complete the correction. Therefore, a fast Fourier convolution network is added to the distortion network design to obtain the global receptive field.
[0285] To this end, the pre-trained twisted network includes a pre-trained downsampling network, a pre-trained fast Fourier convolution network and a pre-trained upsampling network, where the pre-trained downsampling network corresponds to the above-mentioned Encoder, and the pre-trained upsampling network corresponds to the above-mentioned decoder.
[0286] The definition of receptive field is: the size of the area on the feature map output by each layer of the convolutional neural network that is mapped back to the input image. In simple terms, the size of a point on the feature map relative to the original image is also the area of the input image that the convolutional neural network feature can see.
[0287] To this end, for the image to be corrected, a pre-trained downsampling network is used to downsample the image to be corrected to obtain the down-sampled image to be corrected, and a pre-trained fast Fourier convolutional network is used to process the down-sampled image to obtain a feature map of the image to be corrected, and a pre-trained upsampling network is used to upsample the feature map to obtain an offset field of the image to be corrected.
[0288] For example, Figure 6 As shown, the image to be corrected is input into a pre-trained downsampling network to obtain the down-sampled image to be corrected output by the pre-trained downsampling network, the down-sampled image to be corrected is input into a pre-trained fast Fourier convolution network to obtain the feature map of the image to be corrected output by the pre-trained fast Fourier convolution network, the feature map of the image to be corrected is input into a pre-trained upsampling network to obtain the offset field of the image to be corrected.
[0289] The pre-trained fast Fourier convolution network may include a pre-trained first convolution layer, a pre-trained second convolution layer, a pre-trained third convolution layer, and a pre-trained spectral conversion network, based on which a feature map of the image to be corrected may be determined. Figure 7 FIG. 1 is a schematic diagram of an implementation flow of a method for determining a feature map provided in an embodiment of the present application. The method is applied to an electronic device and may specifically include the following steps:
[0290] S701. Perform a first convolution process on the image to be corrected after downsampling using a pre-trained first convolutional layer to obtain a first convolution result.
[0291] In an embodiment of the present application, for the image to be corrected after downsampling, a first convolution process can be performed on the image to be corrected after downsampling using a pre-trained first convolutional layer to obtain a first convolution result.
[0292] Among them, for the pre-trained first convolutional layer, for example, it can be Conv1, and its convolution kernel size is 3*3. Use Conv1 to perform a first convolution process on the image to be corrected after downsampling to obtain a first convolution result.
[0293] S702. Perform a second convolution process on the image to be corrected after downsampling using a pre-trained second convolutional layer to obtain a second convolution result.
[0294] In an embodiment of the present application, for the image to be corrected after downsampling, a second convolution process can be performed on the image to be corrected after downsampling using a pre-trained second convolutional layer to obtain a second convolution result.
[0295] Among them, for the pre-trained second convolutional layer, for example, it can be Conv2, and its convolution kernel size is 3*3. Use Conv2 to perform a second convolution process on the image to be corrected after downsampling to obtain a second convolution result.
[0296] It should be noted that during the processing of the pre-trained fast Fourier convolutional network, the input will be divided into 2 different branches based on channels. One branch is responsible for extracting local features, called the Local branch, and the other branch is responsible for extracting global features, called the Global branch.
[0297] Based on this, the above-mentioned pre-trained first convolutional layer and pre-trained second convolutional layer belong to the Local branch, while the above-mentioned pre-trained third convolutional layer and pre-trained spectral conversion network belong to the Global branch.
[0298] S703. Perform a third convolution process on the image to be corrected after downsampling using a pre-trained third convolutional layer to obtain a third convolution result.
[0299] In an embodiment of the present application, for the image to be corrected after downsampling, a third convolution process can be performed on the image to be corrected after downsampling using a pre-trained third convolutional layer to obtain a third convolution result.
[0300] Among them, for the pre-trained third convolutional layer, for example, it can be Conv3, and its convolution kernel size is 3*3. Use Conv3 to perform a third convolution process on the image to be corrected after downsampling to obtain a third convolution result.
[0301] It should be noted that for the first convolution process, the second convolution process, and the third convolution process, a convolution kernel with a size of, for example, 3×3 can be used, and the parameters in the convolution kernel can be different; for the fourth convolution process and the fifth convolution process, convolution kernels with the same size (for example, different from the 3×3 size) can be used, and the parameters in the convolution kernel can be different; for the sixth convolution process, a convolution kernel with a size of, for example, 1×1 can be used, and the parameters in the convolution kernel can be different from the parameters in the above-mentioned convolution kernels.
[0302] S704, perform spectral conversion processing on the to-be-corrected image after downsampling using a pre-trained spectral conversion network to obtain a spectral conversion result.
[0303] In the embodiment of the present application, for the to-be-corrected image after downsampling, a pre-trained spectral conversion network can be used to perform spectral conversion processing on the to-be-corrected image after downsampling to obtain a spectral conversion result.
[0304] Among them, the pre-trained spectral conversion network includes a pre-trained fourth convolution layer, a pre-trained third batch normalization layer, a pre-trained third activation function layer, and a pre-trained fifth convolution layer. For example, for the pre-trained fourth convolution layer, it can be Conv4, for the pre-trained third batch normalization layer, it can be BN3, for the pre-trained third activation function layer, it can be ReLu3, and for the pre-trained fifth convolution layer, it can be Conv5, and its convolution kernel size is 1×1.
[0305] Based on this, perform fourth convolution processing on the to-be-corrected image after downsampling using the pre-trained fourth convolution layer to obtain a fourth convolution result; perform third normalization processing on the fourth convolution result using the pre-trained third batch normalization layer to obtain a third normalization result; perform third activation processing on the third normalization result using the pre-trained third activation function layer to obtain a first activation result; perform spectral conversion processing on the first activation result to obtain an original spectral conversion result; add the first activation result and the original spectral conversion result to obtain a third addition result; perform fifth convolution processing on the third addition result using the pre-trained fifth convolution layer to obtain a spectral conversion result.
[0306] Specifically, the pre-trained spectral conversion network further includes a two-dimensional Fourier transform layer, a pre-trained sixth convolutional layer, a pre-trained fourth batch normalization layer, a pre-trained fourth activation function layer, and an inverse two-dimensional Fourier transform layer. For example, for the two-dimensional Fourier transform layer, it can be Real FFT2d, for the pre-trained sixth convolutional layer, it can be Conv6, for the pre-trained fourth batch normalization layer, it can be BN4, for the pre-trained fourth activation function layer, it can be ReLu4, and for the inverse two-dimensional Fourier transform layer, it can be Inv Real FFT2d.
[0307] Based on this, specifically, the spectral conversion process is as follows: performing two-dimensional Fourier transform processing on the first activation result using the two-dimensional Fourier transform layer to obtain a frequency domain result; performing sixth convolutional processing on the frequency domain result using the pre-trained sixth convolutional layer to obtain a sixth convolutional result; performing fourth normalization processing on the sixth convolutional result using the pre-trained fourth batch normalization layer to obtain a fourth normalization result; performing fourth activation processing on the fourth normalization result using the pre-trained fourth activation function layer to obtain a second activation result; and performing inverse two-dimensional Fourier transform processing on the second activation result using the inverse two-dimensional Fourier transform layer to obtain an original spectral conversion result.
[0308] For example, as Figure 8 shown, inputting the image to be corrected after downsampling into Conv4 for fourth convolutional processing to obtain a fourth convolutional result, inputting the fourth convolutional result into BN3 for third normalization processing to obtain a third normalization result, inputting the third normalization result into ReLu3 for third activation processing to obtain a first activation result, inputting the first activation result into Real FFT2d for two-dimensional Fourier transform processing to obtain a frequency domain result, inputting the frequency domain result into Conv6 for sixth convolutional processing to obtain a sixth convolutional result, inputting the sixth convolutional result into BN4 for fourth normalization processing to obtain a fourth normalization result, inputting the fourth normalization result into ReLu4 for fourth activation processing to obtain a second activation result, inputting the second activation result into Inv Real FFT2d for inverse two-dimensional Fourier transform processing to obtain an original spectral conversion result, adding the first activation result and the original spectral conversion result to obtain a third addition result, and inputting the third addition result into Conv5 for fifth convolutional processing to obtain a spectral conversion result.
[0309] It should be noted that for the first normalization processing, the second normalization processing, the third normalization processing, and the fourth normalization processing, the same normalization formula can be used, but the parameters in the normalization formula can be different. For the first activation processing, the second activation processing, the third activation processing, and the fourth activation processing, the same activation function, such as ReLu, can be used, but the parameters in the activation function can be different.
[0310] S705. Determine the feature map of the image to be corrected based on the first convolution result, the second convolution result, the third convolution result, and the spectral conversion result.
[0311] In an embodiment of the present application, for the above-obtained first convolution result, second convolution result, third convolution result, and spectral conversion result, the feature map of the image to be corrected can be determined based on the first convolution result, the second convolution result, the third convolution result, and the spectral conversion result.
[0312] Among them, the pre-trained fast Fourier convolution network further includes a pre-trained first batch normalization layer, a pre-trained first activation function layer, a pre-trained second batch normalization layer, and a pre-trained second activation function layer. For example, for the pre-trained first batch normalization layer, it can be BN1, for the pre-trained first activation function layer, it can be ReLu1, for the pre-trained second batch normalization layer, it can be BN2, and for the pre-trained second activation function layer, it can be ReLu2.
[0313] Based on this, add the first convolution result and the third convolution result to obtain a first addition result; perform first normalization processing on the first addition result using the pre-trained first batch normalization layer to obtain a first normalization result; perform first activation processing on the first normalization result using the pre-trained first activation function layer to obtain local features; add the second convolution result and the spectral conversion result to obtain a second addition result; perform second normalization processing on the second addition result using the pre-trained second batch normalization layer to obtain a second normalization result; perform second activation processing on the second normalization result using the pre-trained second activation function layer to obtain global features; superimpose the local features and the global features to obtain the feature map of the image to be corrected.
[0314] For example, as Figure 9 shown, add the first convolution result and the third convolution result to obtain a first addition result, input the first addition result into BN1 for first normalization processing to obtain a first normalization result, input the first normalization result into ReLu1 for first activation processing to obtain local features; add the second convolution result and the spectral conversion result to obtain a second addition result, input the second addition result into BN2 for second normalization processing to obtain a second normalization result, input the second normalization result into ReLu1 for second activation processing to obtain global features; superimpose the local features and the global features to obtain the feature map of the image to be corrected.
[0315] S503. Correct the image to be corrected using the offset field to obtain a corrected image, where the correction object in the corrected image is in a flattened state.
[0316] In the embodiment of the present application, this step is similar to the above step S103, and the embodiments of the present application will not be elaborated herein one by one.
[0317] Determine the offset field of the image to be corrected through a pre-trained warping network, and use the offset field to correct the image to be corrected to obtain a corrected image. In this way, without using a scanner, the object to be corrected in the image to be corrected can still be corrected, avoiding negative impacts on the readability of the image and the recognition effect of OCR. In addition, a pre-trained fast Fourier convolutional network is added to the pre-trained warping network to obtain a global receptive field, which can significantly improve the correction effect of the image to be corrected.
[0318] In addition, in the embodiment of the present application, for the image to be corrected, input it into the pre-trained warping network to obtain its corresponding offset field, and use the offset field to correct the image to be corrected to obtain a corrected image. After the first correction, there will usually be some deformations. Therefore, the correction result can be used as the input and sent into the pre-trained warping network again for iterative correction, so as to obtain a better correction effect. For this reason, as Figure 10 shown, it is a schematic flowchart of the implementation process of another image correction method provided by the embodiment of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0319] S1001, obtain the image to be corrected, where the object to be corrected in the image to be corrected is in a distorted state.
[0320] In the embodiment of the present application, this step is similar to the above step S101, and the embodiments of the present application will not be elaborated herein one by one.
[0321] S1002, input the image to be corrected into the pre-trained warping network to obtain the offset field of the image to be corrected output by the pre-trained warping network. The offset field represents the offset direction and offset amount of the corresponding position in the image to be corrected.
[0322] In the embodiment of the present application, this step is similar to the above step S502, and the embodiments of the present application will not be elaborated herein one by one.
[0323] S1003, use the offset field to correct the image to be corrected to obtain a corrected image, where the corrected object in the corrected image is in a flattened state.
[0324] In the embodiment of the present application, this step is similar to the above step S103, and the embodiments of the present application will not be elaborated herein one by one.
[0325] S1004, perform iterations according to the following steps until a preset iteration termination condition is met, and obtain the iterative offset field in each iteration process.
[0326] In the embodiments of the present application, for the image to be corrected, there will usually be some deformations after the first correction. Therefore, the correction result can be used as the input and fed into the pre-trained distortion network again for iterative correction. For this purpose, after using the offset field to correct the image to be corrected to obtain the corrected image, the following steps (S1005 to S1007) are performed iteratively until the preset iteration termination condition is satisfied, so that the iterative offset field in each iteration process can be obtained.
[0327] Among them, the condition for terminating the iteration is set to one of the following three: (1) The maximum number of iterations is set to, for example, 5, and the iteration is terminated when it exceeds 5 times; (2) The variance of the previous offset field is less than the variance of the current offset field, and the iteration can be terminated; (3) The variance of the current offset field is less than the threshold, and the iteration can be terminated.
[0328] For this purpose, the above preset iteration termination condition includes one of the following: The number of iterations reaches the preset iteration number threshold; the variance of the iterative offset field in the current iteration process is less than the preset variance threshold; the variance of the iterative offset field in the current iteration process is greater than the variance of the iterative offset field in the previous iteration process.
[0329] S1005, obtain the target image of the current iteration. Among them, the target image of the first iteration is the corrected image, and the target image of a non-first iteration is the iterative image obtained in the previous iteration.
[0330] In the embodiments of the present application, obtain the target image of the current iteration. Among them, the target image of the first iteration is the corrected image, that is, the corrected image after the first correction, and the target image of a non-first iteration is the iterative image obtained in the previous iteration.
[0331] S1006, input the target image into the pre-trained distortion network to obtain the iterative offset field of the target image output by the pre-trained distortion network.
[0332] In the embodiments of the present application, for the target image, the target image can be input into the pre-trained distortion network to obtain the iterative offset field of the target image output by the pre-trained distortion network. The iterative offset field represents the offset direction and offset amount of the corresponding position in the target image.
[0333] Among them, the pre-trained distortion network includes a pre-trained downsampling network, a pre-trained fast Fourier convolutional network, and a pre-trained upsampling network. Thus, the pre-trained downsampling network can be used to perform downsampling processing on the target image to obtain the target image after downsampling processing, the pre-trained fast Fourier convolutional network can be used to process the target image after downsampling processing to obtain the iterative feature map of the target image, and the pre-trained upsampling network can be used to perform upsampling processing on the iterative feature map to obtain the iterative offset field of the target image.
[0334] For a pre-trained fast Fourier convolutional network, it may include a pre-trained first convolutional layer, a pre-trained second convolutional layer, a pre-trained third convolutional layer, and a pre-trained spectral conversion network. Based on this, an iterative feature map of the target image can be determined. To this end, the pre-trained first convolutional layer is used to perform a first convolutional process on the downsampled target image to obtain a first iterative convolutional result. The pre-trained second convolutional layer is used to perform a second convolutional process on the downsampled target image to obtain a second iterative convolutional result. The pre-trained third convolutional layer is used to perform a third convolutional process on the downsampled target image to obtain a third iterative convolutional result. The pre-trained spectral conversion network is used to perform a spectral conversion process on the downsampled target image to obtain an iterative spectral conversion result. According to the first iterative convolutional result, the second iterative convolutional result, the third iterative convolutional result, and the iterative spectral conversion result, the iterative feature map of the target image is determined.
[0335] Among them, the pre-trained spectral conversion network includes a pre-trained fourth convolutional layer, a pre-trained third batch normalization layer, a pre-trained third activation function layer, and a pre-trained fifth convolutional layer. Thus, the pre-trained fourth convolutional layer is used to perform a fourth convolutional process on the downsampled target image to obtain a fourth iterative convolutional result. The pre-trained third batch normalization layer is used to perform a third normalization process on the fourth iterative convolutional result to obtain a third iterative normalization result. The pre-trained third activation function layer is used to perform a third activation process on the third iterative normalization result to obtain a first iterative activation result. A spectral conversion process is performed on the first iterative activation result to obtain an iterative original spectral conversion result. The first iterative activation result and the iterative original spectral conversion result are added together to obtain a third iterative addition result. The pre-trained fifth convolutional layer is used to perform a fifth convolutional process on the third iterative addition result to obtain an iterative spectral conversion result.
[0336] Specifically, the pre-trained spectral conversion network further includes a two-dimensional Fourier transform layer, a pre-trained sixth convolutional layer, a pre-trained fourth batch normalization layer, a pre-trained fourth activation function layer, and a two-dimensional inverse Fourier transform layer. Therefore, the two-dimensional Fourier transform layer is used to perform a two-dimensional Fourier transform process on the first iterative activation result to obtain an iterative frequency domain result. The pre-trained sixth convolutional layer is used to perform a sixth convolutional process on the iterative frequency domain result to obtain a sixth iterative convolutional result. The pre-trained fourth batch normalization layer is used to perform a fourth normalization process on the sixth iterative convolutional result to obtain a fourth iterative normalization result. The pre-trained fourth activation function layer is used to perform a fourth activation process on the fourth iterative normalization result to obtain a second iterative activation result. The two-dimensional inverse Fourier transform layer is used to perform a two-dimensional inverse Fourier transform process on the second iterative activation result to obtain an iterative original spectral conversion result.
[0337] In addition, the pre-trained fast Fourier convolutional network further includes a pre-trained first batch normalization layer, a pre-trained first activation function layer, a pre-trained second batch normalization layer, and a pre-trained second activation function layer. Based on this, the first iterative convolution result and the third iterative convolution result are added together to obtain a first iterative addition result; the first iterative addition result is subjected to a first normalization process using the pre-trained first batch normalization layer to obtain a first iterative normalization result; the first iterative normalization result is subjected to a first activation process using the pre-trained first activation function layer to obtain iterative local features; the second iterative convolution result and the iterative spectral conversion result are added together to obtain a second iterative addition result; the second iterative addition result is subjected to a second normalization process using the pre-trained second batch normalization layer to obtain a second iterative normalization result; the second iterative normalization result is subjected to a second activation process using the pre-trained second activation function layer to obtain iterative global features; the iterative local features and the iterative global features are superimposed to obtain an iterative feature map of the target image.
[0338] S1007, use the iterative offset field to correct the target image to obtain an iterative image, and use the iterative image as the target image for the next iteration.
[0339] In the embodiment of the present application, for the iterative offset field, the iterative offset field can be used to correct the target image to obtain an iterative image, and the iterative image is used as the target image for the next iteration. Among them, the remap algorithm is adopted to correct the target image with reference to the iterative offset field to obtain an iterative image.
[0340] S1008, correct the image to be corrected according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image.
[0341] In the embodiment of the present application, for the offset field in the first correction process and the iterative offset field in each iteration process, the image to be corrected can be corrected according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image.
[0342] Among them, the offset field and the iterative offset field in each iteration process can be added together to obtain a final offset field, and the final offset field is used to correct the image to be corrected to obtain a final corrected image.
[0343] In addition, for the final corrected image, the final corrected image can be subjected to mirror expansion and morphological processing to cleverly complete the edge loss compensation. The final offset field can also be smoothed. In this way, the remap algorithm is adopted to correct the image to be corrected with reference to the smoothed final offset field to obtain a final corrected image, and the image details in the final corrected image are smoother and more natural.
[0344] In this way, the correction result is used as the input and sent back to the pre-trained distortion network for iterative correction. By using the iterative inference method to gradually approximate the true offset field and using the final offset field to correct the image to be corrected, a better correction effect can be achieved.
[0345] As Figure 11 shown in the figure, it is a schematic flowchart of an implementation process of a distortion network training method provided by an embodiment of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0346] S1101, obtain a sample image set, where the objects in the sample images of the sample image set are in a distorted state.
[0347] In the embodiment of the present application, a sample image set is obtained, where the sample image set includes multiple sample images, and the objects in the sample images are in a distorted state.
[0348] Among them, a preset data set (such as the Doc3d open source data set) can be obtained. The preset data set includes the original image to be corrected and the original offset field corresponding to the original image to be corrected.
[0349] In view of the size problem of the original image to be corrected in the preset data set and the relatively serious image distortion, the original image to be corrected can be cropped to obtain the object area therein, for example, the document area is cropped.
[0350] Based on this, detect the object area in the original image to be corrected, and determine the minimum bounding rectangle of the object area; use the minimum bounding rectangle to crop the original image to be corrected to obtain the object area, and determine the object area as the first sample image.
[0351] In addition, an inverse transformation algorithm is designed. By using the original offset field in the preset data set for mapping, the corresponding distorted image can be obtained. This algorithm can obtain a more high-definition and practical image distribution, and the combination of different images and different deformation information can greatly improve the richness of the samples.
[0352] Based on this, use the minimum bounding rectangle to crop the original offset field corresponding to the original image to be corrected to obtain the object offset field; obtain the original image, use the object offset field to perform distortion processing on the original image to obtain the distorted image; determine the distorted image as the second sample image, and store the first sample image and the second sample image in the sample image set.
[0353] S1102, input the sample image into the distortion network, and obtain the sample offset field of the sample image output by the distortion network.
[0354] In the embodiment of the present application, for a sample image, the sample image can be input into a warping network to obtain a sample offset field of the sample image output by the warping network, and the sample offset field characterizes the offset direction and offset amount of the corresponding position in the sample image.
[0355] Among them, the warping network includes a downsampling network, a fast Fourier convolutional network, and an upsampling network. The downsampling network is used to perform downsampling processing on the sample image to obtain the sample image after downsampling processing; the fast Fourier convolutional network is used to process the sample image after downsampling processing to obtain a sample feature map of the sample image; the upsampling network is used to perform upsampling processing on the sample feature map to obtain a sample offset field of the sample image.
[0356] Specifically, the fast Fourier convolutional network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a spectral conversion network. Thus, the first convolutional layer is used to perform first convolutional processing on the sample image after downsampling processing to obtain a first sample convolutional result; the second convolutional layer is used to perform second convolutional processing on the sample image after downsampling processing to obtain a second sample convolutional result; the third convolutional layer is used to perform third convolutional processing on the sample image after downsampling processing to obtain a third sample convolutional result; the spectral conversion network is used to perform spectral conversion processing on the sample image after downsampling processing to obtain a sample spectral conversion result; according to the first sample convolutional result, the second sample convolutional result, the third sample convolutional result, and the sample spectral conversion result, a sample feature map of the sample image is determined.
[0357] In addition, the fast Fourier convolutional network further includes a first batch normalization layer, a first activation function layer, a second batch normalization layer, and a second activation function layer. Thus, the first sample convolutional result and the third sample convolutional result are added to obtain a first sample addition result; the first batch normalization layer is used to perform first normalization processing on the first sample addition result to obtain a first sample normalization result; the first activation function layer is used to perform first activation processing on the first sample normalization result to obtain a sample local feature; the second sample convolutional result and the sample spectral conversion result are added to obtain a second sample addition result; the second batch normalization layer is used to perform second normalization processing on the second sample addition result to obtain a second sample normalization result; the second activation function layer is used to perform second activation processing on the second sample normalization result to obtain a sample global feature; the sample local feature and the sample global feature are superimposed to obtain a sample feature map of the sample image.
[0358] Among them, the spectral conversion network includes a fourth convolutional layer, a third batch normalization layer, a third activation function layer, and a fifth convolutional layer. Thus, the fourth convolutional layer is used to perform fourth convolutional processing on the sample image after downsampling to obtain a fourth sample convolutional result; the third batch normalization layer is used to perform third normalization processing on the fourth sample convolutional result to obtain a third sample normalization result; the third activation function layer is used to perform third activation processing on the third sample normalization result to obtain a first sample activation result; the first sample activation result is subjected to spectral conversion processing to obtain an original sample spectral conversion result; the first sample activation result and the original sample spectral conversion result are added together to obtain a third sample addition result; the fifth convolutional layer is used to perform fifth convolutional processing on the third sample addition result to obtain a sample spectral conversion result.
[0359] In addition, the spectral conversion network further includes a two-dimensional Fourier transform layer, a sixth convolutional layer, a fourth batch normalization layer, a fourth activation function layer, and a two-dimensional inverse Fourier transform layer. Thus, the two-dimensional Fourier transform layer is used to perform two-dimensional Fourier transform processing on the first sample activation result to obtain a sample frequency domain result; the sixth convolutional layer is used to perform sixth convolutional processing on the sample frequency domain result to obtain a sixth sample convolutional result; the fourth batch normalization layer is used to perform fourth normalization processing on the sixth sample convolutional result to obtain a fourth sample normalization result; the fourth activation function layer is used to perform fourth activation processing on the fourth sample normalization result to obtain a second sample activation result; the two-dimensional inverse Fourier transform layer is used to perform two-dimensional inverse Fourier transform processing on the second sample activation result to obtain an original sample spectral conversion result.
[0360] S1103, use the sample offset field to correct the sample image to obtain a sample corrected image, and train the distortion network according to the sample corrected image.
[0361] S1104, count the number of training times of the distortion network, and stop training when the number of training times reaches a preset number threshold to obtain a pre-trained distortion network.
[0362] In the embodiment of the present application, the sample image can be corrected by using the sample offset field to obtain a sample corrected image, so that the distortion network can be trained according to the sample corrected image. And count the number of training times of the distortion network, and stop training when the number of training times reaches a preset number threshold to obtain a pre-trained distortion network, as Figure 9 shown by the pre-trained distortion network.
[0363] In addition, during the process of training the distortion network, the sample image is corrected by using the sample offset field to obtain a sample corrected image. Usually, there will still be some deformations after the first correction. Therefore, the correction result can be used as an input and sent into the distortion network again for iterative training, so as to obtain a better correction effect. For this reason, as Figure 12As shown in the figure, it is a schematic flowchart of another implementation process of the distortion network training method provided by the embodiment of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0364] S1201, Obtain a sample image set, where the objects in the sample images of the sample image set are in a distorted state.
[0365] In the embodiment of the present application, this step is similar to step S1101 above, and the embodiment of the present application will not elaborate here one by one.
[0366] S1202, Input the sample image into the distortion network to obtain the sample offset field of the sample image output by the distortion network.
[0367] In the embodiment of the present application, this step is similar to step S1102 above, and the embodiment of the present application will not elaborate here one by one.
[0368] S1203, Use the sample offset field to correct the sample image to obtain a sample corrected image.
[0369] In the embodiment of the present application, this step is similar to step S1101 above, and the embodiment of the present application will not elaborate here one by one.
[0370] S1204, Iterate according to the following steps until a preset iteration termination condition is met, and obtain the sample iteration offset field in each iteration process.
[0371] In the embodiment of the present application, for the sample image, there will usually still be some deformations after the first correction. Therefore, the correction result can be used as the input and sent into the distortion network again for iterative training. For this reason, after using the sample offset field to correct the sample image to obtain the sample corrected image, iterate according to the following steps (S1205 - S1207) until a preset iteration termination condition is met, and obtain the sample iteration offset field in each iteration process.
[0372] Among them, the above preset iteration termination condition includes one of the following: the number of iterations reaches a preset iteration number threshold; the variance of the sample iteration offset field in the current iteration process is less than a preset variance threshold; the variance of the sample iteration offset field in the current iteration process is greater than the variance of the sample iteration offset field in the previous iteration process.
[0373] S1205, Obtain the sample target image of the current iteration, where the sample target image of the first iteration is the sample corrected image, and the sample target image of a non-first iteration is the sample iteration image obtained in the previous iteration.
[0374] In the embodiments of the present application, a sample target image of the current iteration is obtained. Among them, the sample target image of the first iteration is the sample correction image, that is, the sample correction image after the first correction, and the sample target image of a non-first iteration is the sample iteration image obtained in the previous iteration.
[0375] S1206. Input the sample target image into the warping network to obtain the sample iteration offset field of the sample target image output by the warping network.
[0376] In the embodiments of the present application, for the sample target image, the sample target image can be input into the warping network to obtain the sample iteration offset field of the sample target image output by the warping network. Among them, the sample iteration offset field represents the offset direction and offset amount of the corresponding position in the sample target image.
[0377] Among them, the warping network includes a downsampling network, a fast Fourier convolutional network, and an upsampling network. Thus, the sample target image is downsampled by the downsampling network to obtain the downsampled sample target image; the downsampled sample target image is processed by the fast Fourier convolutional network to obtain the iterative sample feature map of the sample target image; the iterative sample feature map is upsampled by the upsampling network to obtain the iterative sample offset field of the sample target image.
[0378] Specifically, the fast Fourier convolutional network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a spectral conversion network. Thus, the first convolutional layer performs a first convolution process on the downsampled sample target image to obtain a first iterative sample convolution result; the second convolutional layer performs a second convolution process on the downsampled sample target image to obtain a second iterative sample convolution result; the third convolutional layer performs a third convolution process on the downsampled sample target image to obtain a third iterative sample convolution result; the spectral conversion network performs a spectral conversion process on the downsampled sample target image to obtain an iterative sample spectral conversion result; according to the first iterative sample convolution result, the second iterative sample convolution result, the third iterative sample convolution result, and the iterative sample spectral conversion result, the iterative sample feature map of the sample target image is determined.
[0379] In addition, the fast Fourier convolution network further includes a first batch normalization layer, a first activation function layer, a second batch normalization layer, and a second activation function layer. Thus, the convolution result of the first iterative sample is added to the convolution result of the third iterative sample to obtain the addition result of the first iterative sample; the first batch normalization layer is used to perform first normalization processing on the addition result of the first iterative sample to obtain the normalized result of the first iterative sample; the first activation function layer is used to perform first activation processing on the normalized result of the first iterative sample to obtain the local features of the iterative sample; the convolution result of the second iterative sample is added to the spectral conversion result of the iterative sample to obtain the addition result of the second iterative sample; the second batch normalization layer is used to perform second normalization processing on the addition result of the second iterative sample to obtain the normalized result of the second iterative sample; the second activation function layer is used to perform second activation processing on the normalized result of the second iterative sample to obtain the global features of the iterative sample; the local features of the iterative sample and the global features of the iterative sample are superimposed to obtain the iterative sample feature map of the sample target image.
[0380] Among them, the spectral conversion network includes a fourth convolution layer, a third batch normalization layer, a third activation function layer, and a fifth convolution layer. Thus, the fourth convolution layer is used to perform fourth convolution processing on the sample target image after downsampling to obtain the convolution result of the fourth iterative sample; the third batch normalization layer is used to perform third normalization processing on the convolution result of the fourth iterative sample to obtain the normalized result of the third iterative sample; the third activation function layer is used to perform third activation processing on the normalized result of the third iterative sample to obtain the activation result of the first iterative sample; spectral conversion processing is performed on the activation result of the first iterative sample to obtain the spectral conversion result of the iterative original sample; the activation result of the first iterative sample is added to the spectral conversion result of the iterative original sample to obtain the addition result of the third iterative sample; the fifth convolution layer is used to perform fifth convolution processing on the addition result of the third iterative sample to obtain the spectral conversion result of the iterative sample.
[0381] In addition, the spectral conversion network further includes a two-dimensional Fourier transform layer, a sixth convolution layer, a fourth batch normalization layer, a fourth activation function layer, and a two-dimensional inverse Fourier transform layer. Thus, the two-dimensional Fourier transform layer is used to perform two-dimensional Fourier transform processing on the activation result of the first iterative sample to obtain the frequency domain result of the iterative sample; the sixth convolution layer is used to perform sixth convolution processing on the frequency domain result of the iterative sample to obtain the convolution result of the sixth iterative sample; the fourth batch normalization layer is used to perform fourth normalization processing on the convolution result of the sixth iterative sample to obtain the normalized result of the fourth iterative sample; the fourth activation function layer is used to perform fourth activation processing on the normalized result of the fourth iterative sample to obtain the activation result of the second iterative sample; the two-dimensional inverse Fourier transform layer is used to perform two-dimensional inverse Fourier transform processing on the activation result of the second iterative sample to obtain the spectral conversion result of the iterative original sample.
[0382] S1207. Use the sample iterative offset field to correct the sample target image to obtain a sample iterative image, and use the sample iterative image as the sample target image for the next iteration.
[0383] In the embodiment of the present application, use the sample iterative offset field to correct the sample target image to obtain a sample iterative image, and use the sample iterative image as the sample target image for the next iteration.
[0384] S1208. Correct the sample image according to the sample offset field and the sample iterative offset field in each iteration process to obtain a final sample corrected image.
[0385] In the embodiment of the present application, for the sample offset field and the sample iterative offset field in each iteration process, the sample image can be corrected according to the sample offset field and the sample iterative offset field in each iteration process to obtain a final sample corrected image.
[0386] Among them, add the sample offset field and the sample iterative offset field in each iteration process to obtain a final sample offset field; use the final sample offset field to correct the sample image to obtain a final sample corrected image.
[0387] S1209. Train the distortion network according to the final sample corrected image.
[0388] S1210. Count the number of training times of the distortion network, and stop training when the number of training times reaches a preset number threshold to obtain a pre-trained distortion network.
[0389] In the embodiment of the present application, the distortion network can be trained according to the final sample corrected image, and the number of training times of the distortion network is counted, and training is stopped when the number of training times reaches a preset number threshold to obtain a pre-trained distortion network. Among them, for the distortion network, inputting an image once is regarded as one training.
[0390] In this way, taking the correction result as the input and feeding it into the distortion network again for iterative training, using the iterative training method to gradually approximate the real offset field, and using the final sample offset field to correct the sample image, better correction effects can be obtained.
[0391] Corresponding to the above method embodiment, the embodiment of the present application also provides an image correction device, as Figure 13 shown. The device may include: an image acquisition module 1310, an offset field determination module 1320, and an image correction module 1330.
[0392] The image acquisition module 1310 is used to acquire an image to be corrected, where the object to be corrected in the image to be corrected is in a distorted state;
[0393] An offset field determination module 1320, configured to determine an offset field of the image to be corrected, where the offset field characterizes an offset direction and an offset amount at a corresponding position in the image to be corrected;
[0394] An image correction module 1330, configured to correct the image to be corrected by using the offset field to obtain a corrected image, where a correction object in the corrected image is in a flattened state.
[0395] In an optional implementation manner, the offset field determination module is specifically configured to:
[0396] Input the image to be corrected into a pre-trained distortion network, and obtain an offset field of the image to be corrected output by the pre-trained distortion network.
[0397] In an optional implementation manner, the pre-trained distortion network includes a pre-trained downsampling network, a pre-trained fast Fourier convolutional network, and a pre-trained upsampling network;
[0398] The offset field determination module includes:
[0399] An image downsampling processing sub-module, configured to perform downsampling processing on the image to be corrected by using the pre-trained downsampling network to obtain the image to be corrected after the downsampling processing;
[0400] An image processing sub-module, configured to process the image to be corrected after the downsampling processing by using the pre-trained fast Fourier convolutional network to obtain a feature map of the image to be corrected;
[0401] A feature map upsampling processing sub-module, configured to perform upsampling processing on the feature map by using the pre-trained upsampling network to obtain an offset field of the image to be corrected.
[0402] In an optional implementation manner, the pre-trained fast Fourier convolutional network includes a pre-trained first convolutional layer, a pre-trained second convolutional layer, a pre-trained third convolutional layer, and a pre-trained spectral conversion network;
[0403] The image processing sub-module includes:
[0404] A first convolutional processing unit, configured to perform first convolutional processing on the image to be corrected after the downsampling processing by using the pre-trained first convolutional layer to obtain a first convolutional result;
[0405] A second convolutional processing unit, configured to perform second convolutional processing on the image to be corrected after the downsampling processing by using the pre-trained second convolutional layer to obtain a second convolutional result;
[0406] A fourth convolutional processing unit, configured to perform a third convolutional process on the to-be-corrected image after downsampling using the pre-trained third convolutional layer to obtain a third convolutional result;
[0407] A spectral conversion processing unit, configured to perform a spectral conversion process on the to-be-corrected image after downsampling using the pre-trained spectral conversion network to obtain a spectral conversion result;
[0408] A feature map determination subunit, configured to determine a feature map of the to-be-corrected image according to the first convolutional result, the second convolutional result, the third convolutional result, and the spectral conversion result.
[0409] In an optional embodiment, the pre-trained fast Fourier convolutional network further includes a pre-trained first batch normalization layer, a pre-trained first activation function layer, a pre-trained second batch normalization layer, and a pre-trained second activation function layer;
[0410] The feature map determination subunit is specifically configured to:
[0411] Add the first convolutional result and the third convolutional result to obtain a first addition result;
[0412] Perform a first normalization process on the first addition result using the pre-trained first batch normalization layer to obtain a first normalization result;
[0413] Perform a second activation process on the first normalization result using the pre-trained first activation function layer to obtain local features;
[0414] Add the second convolutional result and the spectral conversion result to obtain a second addition result;
[0415] Perform a second normalization process on the second addition result using the pre-trained second batch normalization layer to obtain a second normalization result;
[0416] Perform a second activation process on the second normalization result using the pre-trained second activation function layer to obtain global features;
[0417] Superimpose the local features and the global features to obtain the feature map of the to-be-corrected image.
[0418] In an optional embodiment, the pre-trained spectral conversion network includes a pre-trained fourth convolutional layer, a pre-trained third batch normalization layer, a pre-trained third activation function layer, and a pre-trained fifth convolutional layer;
[0419] The spectral conversion processing unit includes:
[0420] The fourth convolution processing subunit is configured to perform fourth convolution processing on the to-be-corrected image after downsampling processing by using the pre-trained fourth convolution layer, so as to obtain a fourth convolution result;
[0421] The third normalization processing subunit is configured to perform third normalization processing on the fourth convolution result by using the pre-trained third batch normalization layer, so as to obtain a third normalization result;
[0422] The third activation processing subunit is configured to perform third activation processing on the third normalization result by using the pre-trained third activation function layer, so as to obtain a first activation result;
[0423] The spectral conversion processing subunit is configured to perform spectral conversion processing on the first activation result, so as to obtain an original spectral conversion result;
[0424] The result addition subunit is configured to add the first activation result and the original spectral conversion result, so as to obtain a third addition result;
[0425] The fifth convolution processing subunit is configured to perform fifth convolution processing on the third addition result by using the pre-trained fifth convolution layer, so as to obtain a spectral conversion result.
[0426] In an optional implementation manner, the pre-trained spectral conversion network further includes a two-dimensional Fourier transform layer, a pre-trained sixth convolution layer, a pre-trained fourth batch normalization layer, a pre-trained fourth activation function layer, and a two-dimensional inverse Fourier transform layer;
[0427] The spectral conversion processing subunit is specifically configured to:
[0428] perform two-dimensional Fourier transform processing on the first activation result by using the two-dimensional Fourier transform layer, so as to obtain a frequency domain result;
[0429] perform sixth convolution processing on the frequency domain result by using the pre-trained sixth convolution layer, so as to obtain a sixth convolution result;
[0430] perform fourth normalization processing on the sixth convolution result by using the pre-trained fourth batch normalization layer, so as to obtain a fourth normalization result;
[0431] perform fourth activation processing on the fourth normalization result by using the pre-trained fourth activation function layer, so as to obtain a second activation result;
[0432] perform two-dimensional inverse Fourier transform processing on the second activation result by using the two-dimensional inverse Fourier transform layer, so as to obtain an original spectral conversion result.
[0433] In an optional implementation manner, the device further includes:
[0434] An iterative module for performing iterations according to the following steps until a preset iteration termination condition is met, to obtain an iterative offset field in each iteration process:
[0435] A target image acquisition module for acquiring a target image of the current iteration, wherein the target image of the first iteration is the corrected image, and the target image of a non-first iteration is the iterative image obtained in the previous iteration;
[0436] An iterative offset field determination module for inputting the target image into the pre-trained warping network to obtain the iterative offset field of the target image output by the pre-trained warping network;
[0437] A target image correction module for correcting the target image by using the iterative offset field to obtain an iterative image, and using the iterative image as the target image for the next iteration;
[0438] An image to be corrected module for correcting the image to be corrected according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image.
[0439] In an optional embodiment, the preset iteration termination condition includes one of the following:
[0440] The number of iterations reaches a preset iteration number threshold;
[0441] The variance of the iterative offset field in the current iteration process is less than a preset variance threshold;
[0442] The variance of the iterative offset field in the current iteration process is greater than the variance of the iterative offset field in the previous iteration process.
[0443] In an optional embodiment, the image to be corrected module is specifically configured to:
[0444] Add the offset field and the iterative offset field in each iteration process to obtain a final offset field;
[0445] Use the final offset field to correct the image to be corrected to obtain a final corrected image.
[0446] In an optional embodiment, the device further includes:
[0447] An image set acquisition module for acquiring a sample image set, wherein the objects in the sample images of the sample image set are in a distorted state;
[0448] A sample offset field determination module for inputting the sample image into the warping network to obtain the sample offset field of the sample image output by the warping network;
[0449] A sample image correction module for correcting the sample image by using the sample offset field to obtain a sample corrected image;
[0450] A model training module for training the distortion network according to the sample corrected image;
[0451] A training times statistics module for counting the training times of the distortion network and stopping the training when the training times reach a preset times threshold to obtain the pre-trained distortion network.
[0452] In an optional embodiment, the image set acquisition module is specifically configured to:
[0453] Obtain a preset data set, where the preset data set includes an original image to be corrected and an original offset field corresponding to the original image to be corrected;
[0454] Detect an object region in the original image to be corrected and determine a minimum bounding rectangle of the object region;
[0455] Crop the original image to be corrected by using the minimum bounding rectangle to obtain the object region, and determine the object region as the first sample image;
[0456] Crop the original offset field corresponding to the original image to be corrected by using the minimum bounding rectangle to obtain an object offset field;
[0457] Obtain an original image, and perform distortion processing on the original image by using the object offset field to obtain a distorted image;
[0458] Determine the distorted image as the second sample image, and store the first sample image and the second sample image in the sample image set.
[0459] In an optional embodiment, the distortion network includes a downsampling network, a fast Fourier convolutional network, and an upsampling network;
[0460] The sample offset field determination module specifically includes:
[0461] A sample image downsampling processing sub-module for performing downsampling processing on the sample image by using the downsampling network to obtain the sample image after downsampling processing;
[0462] A sample feature map determination sub-module for processing the sample image after downsampling processing by using the fast Fourier convolutional network to obtain a sample feature map of the sample image;
[0463] The sample feature map upsampling processing sub-module is used to perform upsampling processing on the sample feature map by using the upsampling network to obtain the sample offset field of the sample image.
[0464] In an optional embodiment, the fast Fourier convolutional network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a spectral conversion network;
[0465] The sample feature map determination sub-module specifically includes:
[0466] The first convolutional processing unit is used to perform first convolutional processing on the downsampled sample image by using the first convolutional layer to obtain a first sample convolutional result;
[0467] The second convolutional processing unit is used to perform second convolutional processing on the downsampled sample image by using the second convolutional layer to obtain a second sample convolutional result;
[0468] The third convolutional processing unit is used to perform third convolutional processing on the downsampled sample image by using the third convolutional layer to obtain a third sample convolutional result;
[0469] The spectral conversion processing unit is used to perform spectral conversion processing on the downsampled sample image by using the spectral conversion network to obtain a sample spectral conversion result;
[0470] The sample feature map determination unit is used to determine the sample feature map of the sample image according to the first sample convolutional result, the second sample convolutional result, the third sample convolutional result, and the sample spectral conversion result.
[0471] In an optional embodiment, the fast Fourier convolutional network further includes a first batch normalization layer, a first activation function layer, a second batch normalization layer, and a second activation function layer;
[0472] The sample feature map determination unit is specifically used for:
[0473] Adding the first sample convolutional result and the third sample convolutional result to obtain a first sample addition result;
[0474] Performing first normalization processing on the first sample addition result by using the first batch normalization layer to obtain a first sample normalization result;
[0475] Performing activation processing on the first sample normalization result first by using the first activation function layer to obtain sample local features;
[0476] Adding the second sample convolutional result and the sample spectral conversion result to obtain a second sample addition result;
[0477] Perform a second normalization process on the second sample addition result using the second batch normalization layer to obtain a second sample normalization result;
[0478] Perform a second activation process on the second sample normalization result using the second activation function layer to obtain a sample global feature;
[0479] Superimpose the sample local feature and the sample global feature to obtain a sample feature map of the sample image.
[0480] In an alternative embodiment, the spectral conversion network includes a fourth convolutional layer, a third batch normalization layer, a third activation function layer, and a fifth convolutional layer;
[0481] The spectral conversion processing unit specifically includes:
[0482] A fourth convolutional processing subunit, configured to perform a fourth convolutional process on the sample image after downsampling using the fourth convolutional layer to obtain a fourth sample convolutional result;
[0483] A third normalization processing subunit, configured to perform a third normalization process on the fourth sample convolutional result using the third batch normalization layer to obtain a third sample normalization result;
[0484] A third activation processing subunit, configured to perform a third activation process on the third sample normalization result using the third activation function layer to obtain a first sample activation result;
[0485] A spectral conversion processing subunit, configured to perform a spectral conversion process on the first sample activation result to obtain an original sample spectral conversion result;
[0486] A result addition subunit, configured to add the first sample activation result and the original sample spectral conversion result to obtain a third sample addition result;
[0487] A fifth convolutional processing subunit, configured to perform a fifth convolutional process on the third sample addition result using the fifth convolutional layer to obtain a sample spectral conversion result.
[0488] In an alternative embodiment, the spectral conversion network further includes a two-dimensional Fourier transform layer, a sixth convolutional layer, a fourth batch normalization layer, a fourth activation function layer, and a two-dimensional inverse Fourier transform layer;
[0489] The spectral conversion processing subunit is specifically configured to:
[0490] Perform a two-dimensional Fourier transform process on the first sample activation result using the two-dimensional Fourier transform layer to obtain a sample frequency domain result;
[0491] Perform a sixth convolution process on the sample frequency domain result using the sixth convolution layer to obtain a sixth sample convolution result;
[0492] Perform a fourth normalization process on the sixth sample convolution result using the fourth batch normalization layer to obtain a fourth sample normalization result;
[0493] Perform a fourth activation process on the fourth sample normalization result using the fourth activation function layer to obtain a second sample activation result;
[0494] Perform an inverse two-dimensional Fourier transform process on the second sample activation result using the two-dimensional inverse Fourier transform layer to obtain an original sample spectrum conversion result.
[0495] In an alternative embodiment, the model training module specifically includes:
[0496] An iteration sub-module for iterating according to the following steps until a preset iteration termination condition is met, to obtain a sample iteration offset field in each iteration process:
[0497] A sample target image acquisition sub-module for acquiring a sample target image of the current iteration, wherein the sample target image of the first iteration is the sample correction image, and the sample target image of a non-first iteration is the sample iteration image obtained in the previous iteration;
[0498] A sample iteration offset field determination sub-module for inputting the sample target image into the warping network to obtain a sample iteration offset field of the sample target image output by the warping network;
[0499] A sample target image correction sub-module for correcting the sample target image using the sample iteration offset field to obtain a sample iteration image, and using the sample iteration image as the sample target image for the next iteration;
[0500] A sample image correction sub-module for correcting the sample image according to the sample offset field and the sample iteration offset field in each iteration process to obtain a final sample correction image;
[0501] A model training sub-module for training the warping network according to the final sample correction image.
[0502] In an alternative embodiment, the preset iteration termination condition includes one of the following:
[0503] The number of iterations reaches a preset iteration number threshold;
[0504] The variance of the sample iteration offset field in the current iteration process is less than a preset variance threshold;
[0505] The variance of the sample iteration offset field in the current iteration process is greater than the variance of the sample iteration offset field in the previous iteration process.
[0506] In an optional implementation, the sample image correction sub-module is specifically configured to:
[0507] Add the sample offset field and the sample iteration offset field in each iteration process to obtain a final sample offset field;
[0508] Use the final sample offset field to correct the sample image to obtain a final sample corrected image.
[0509] The embodiment of the present application also provides an electronic device, as Figure 14 shown, including a processor 141, a communication interface 142, a memory 143, and a communication bus 144. Among them, the processor 141, the communication interface 142, and the memory 143 complete mutual communication through the communication bus 144.
[0510] The memory 143 is used to store a computer program;
[0511] When the processor 141 is used to execute the program stored on the memory 143, the following steps are implemented:
[0512] Obtain an image to be corrected, where the object to be corrected in the image to be corrected is in a distorted state; determine the offset field of the image to be corrected, where the offset field represents the offset direction and offset amount at the corresponding position in the image to be corrected; use the offset field to correct the image to be corrected to obtain a corrected image, where the corrected object in the corrected image is in a flattened state.
[0513] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0514] The communication interface is used for communication between the above electronic device and other devices.
[0515] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0516] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0517] In another embodiment provided by this application, a storage medium is further provided. Instructions are stored in the storage medium. When it runs on a computer, it causes the computer to execute the image correction method described in any one of the foregoing embodiments.
[0518] In another embodiment provided by this application, a computer program product containing instructions is further provided. When it runs on a computer, it causes the computer to execute the image correction method described in any one of the foregoing embodiments.
[0519] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0520] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0521] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0522] The above description is only the preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.
Claims
1. An image correction method, characterized in that, The method includes: Obtaining an image to be corrected, wherein the object to be corrected in the image to be corrected is in a distorted state; Determining an offset field of the image to be corrected, the offset field characterizing the offset direction and offset amount of the corresponding position in the image to be corrected; Correcting the image to be corrected by using the offset field to obtain a corrected image, wherein the corrected object in the corrected image is in a flattened state.
2. The method according to claim 1, wherein The determining the offset field of the image to be corrected includes: Inputting the image to be corrected into a pre-trained distortion network to obtain the offset field of the image to be corrected output by the pre-trained distortion network.
3. The method according to claim 2, wherein The pre-trained distortion network includes a pre-trained downsampling network, a pre-trained fast Fourier convolutional network, and a pre-trained upsampling network; The inputting the image to be corrected into a pre-trained distortion network to obtain the offset field of the image to be corrected output by the pre-trained distortion network includes: Performing downsampling processing on the image to be corrected by using the pre-trained downsampling network to obtain the image to be corrected after downsampling processing; Processing the image to be corrected after downsampling processing by using the pre-trained fast Fourier convolutional network to obtain a feature map of the image to be corrected; Performing upsampling processing on the feature map by using the pre-trained upsampling network to obtain the offset field of the image to be corrected.
4. The method according to claim 3, wherein The pre-trained fast Fourier convolutional network includes a pre-trained first convolutional layer, a pre-trained second convolutional layer, a pre-trained third convolutional layer, and a pre-trained spectral conversion network; The processing the image to be corrected after downsampling processing by using the pre-trained fast Fourier convolutional network to obtain a feature map of the image to be corrected includes: Performing first convolutional processing on the image to be corrected after downsampling processing by using the pre-trained first convolutional layer to obtain a first convolutional result; Performing second convolutional processing on the image to be corrected after downsampling processing by using the pre-trained second convolutional layer to obtain a second convolutional result; Performing third convolutional processing on the image to be corrected after downsampling processing by using the pre-trained third convolutional layer to obtain a third convolutional result; Performing spectral conversion processing on the image to be corrected after downsampling processing by using the pre-trained spectral conversion network to obtain a spectral conversion result; Determining the feature map of the image to be corrected according to the first convolutional result, the second convolutional result, the third convolutional result, and the spectral conversion result.
5. The method according to claim 4, wherein The pre-trained fast Fourier convolutional network further includes a pre-trained first batch normalization layer, a pre-trained first activation function layer, a pre-trained second batch normalization layer, and a pre-trained second activation function layer; The determining the feature map of the image to be corrected according to the first convolutional result, the second convolutional result, the third convolutional result, and the spectral conversion result includes: Adding the first convolutional result and the third convolutional result to obtain a first addition result; Performing first normalization processing on the first addition result by using the pre-trained first batch normalization layer to obtain a first normalization result; Perform a first activation process on the first normalization result using the pre-trained first activation function layer to obtain local features; Add the second convolution result and the spectral conversion result to obtain a second addition result; Perform a second normalization process on the second addition result using the pre-trained second batch normalization layer to obtain a second normalization result; Perform a second activation process on the second normalization result using the pre-trained second activation function layer to obtain global features; Superimpose the local features and the global features to obtain the feature map of the image to be corrected; 6. The method according to claim 4, characterized in that The pre-trained spectral conversion network includes a pre-trained fourth convolution layer, a pre-trained third batch normalization layer, a pre-trained third activation function layer, and a pre-trained fifth convolution layer; The performing spectral conversion processing on the downsampled image to be corrected using the pre-trained spectral conversion network to obtain a spectral conversion result includes: Perform a fourth convolution process on the downsampled image to be corrected using the pre-trained fourth convolution layer to obtain a fourth convolution result; Perform a third normalization process on the fourth convolution result using the pre-trained third batch normalization layer to obtain a third normalization result; Perform a third activation process on the third normalization result using the pre-trained third activation function layer to obtain a first activation result; Perform spectral conversion processing on the first activation result to obtain an original spectral conversion result; Add the first activation result and the original spectral conversion result to obtain a third addition result; Perform a fifth convolution process on the third addition result using the pre-trained fifth convolution layer to obtain a spectral conversion result.
7. The method according to claim 6, characterized in that, The pre-trained spectral conversion network further includes a two-dimensional Fourier transform layer, a pre-trained sixth convolution layer, a pre-trained fourth batch normalization layer, a pre-trained fourth activation function layer, and a two-dimensional inverse Fourier transform layer; The performing spectral conversion processing on the first activation result to obtain an original spectral conversion result includes: Perform two-dimensional Fourier transform processing on the first activation result using the two-dimensional Fourier transform layer to obtain a frequency domain result; Perform a sixth convolution process on the frequency domain result using the pre-trained sixth convolution layer to obtain a sixth convolution result; Perform a fourth normalization process on the sixth convolution result using the pre-trained fourth batch normalization layer to obtain a fourth normalization result; Perform a fourth activation process on the fourth normalization result using the pre-trained fourth activation function layer to obtain a second activation result; Perform two-dimensional inverse Fourier transform processing on the second activation result using the two-dimensional inverse Fourier transform layer to obtain an original spectral conversion result.
8. The method according to claim 2, wherein After the performing correction on the image to be corrected using the offset field to obtain a corrected image, the method further includes: Perform iteration according to the following steps until a preset iteration termination condition is met to obtain the iteration offset field in each iteration process: Obtain the target image of the current iteration, where the target image of the first iteration is the corrected image, and the target image of a non-first iteration is the iteration image obtained in the previous iteration; Input the target image into the pre-trained warping network to obtain the iterative offset field of the target image output by the pre-trained warping network; Use the iterative offset field to correct the target image to obtain an iterative image, and use the iterative image as the target image for the next iteration; Correct the image to be corrected according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image.
9. The method according to claim 8, characterized in that, The preset iteration termination condition includes one of the following: The number of iterations reaches a preset iteration number threshold; The variance of the iterative offset field in the current iteration process is less than a preset variance threshold; The variance of the iterative offset field in the current iteration process is greater than the variance of the iterative offset field in the previous iteration process.
10. The method according to claim 8, wherein The step of correcting the image to be corrected according to the offset field and the iterative offset field in each iteration process to obtain a final corrected image includes: Add the offset field and the iterative offset field in each iteration process to obtain a final offset field; Use the final offset field to correct the image to be corrected to obtain a final corrected image.
11. The method according to claim 2, wherein Before executing the method, it further includes: Obtain a sample image set, where the objects in the sample images of the sample image set are in a distorted state; Input the sample images into the warping network to obtain the sample offset fields of the sample images output by the warping network; Use the sample offset fields to correct the sample images to obtain sample corrected images, and train the warping network according to the sample corrected images; Count the number of training times of the warping network, and stop training when the number of training times reaches a preset number threshold to obtain the pre-trained warping network.
12. The method according to claim 11, wherein The step of obtaining the sample image set includes: Obtain a preset data set, where the preset data set includes the original image to be corrected and the original offset field corresponding to the original image to be corrected; Detect the object region in the original image to be corrected and determine the minimum bounding rectangle of the object region; Crop the original image to be corrected using the minimum bounding rectangle to obtain the object region, and determine the object region as the first sample image; Crop the original offset field corresponding to the original image to be corrected using the minimum bounding rectangle to obtain an object offset field; Obtain an original image, and perform distortion processing on the original image using the object offset field to obtain a distorted image; Determine the distorted image as the second sample image, and store the first sample image and the second sample image in the sample image set.
13. An image correction device, characterized in that, The device includes: An image acquisition module, configured to acquire an image to be corrected, where the object to be corrected in the image to be corrected is in a distorted state; An offset field determination module, configured to determine the offset field of the image to be corrected, where the offset field represents the offset direction and offset amount at the corresponding position in the image to be corrected; An image correction module, configured to correct the image to be corrected using the offset field to obtain a corrected image, where the corrected object in the corrected image is in a flattened state.