Image Rotation Correction Method, System, Electronic Device and Storage Medium
Through the combination of significance detection and feature extraction network, the attention mechanism learning and feature aggregation network are used to achieve rapid rotation correction of images at any angle, solving the problem of image rotation time and poor integrity.
Patent Information
- Application Number
- CN202210542374.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-18
AI Technical Summary
In the prior art, image rotation correction takes a long time and poor image integrity. The traditional method is only applicable to images under specific conditions, and has a long processing time and is not universal.
By acquiring the original image and inputting it to a trained complete significance detection model and feature extraction network, the attention mechanism is used to learn the network, feature aggregation network and the fully connected network to generate image rotation results for correction.
It realizes rapid rotation correction of images at any angle in any scene, and maximizes the integrity of image objects, solving the problems of time-consuming image rotation correction and poor image integrity.
Smart Images

Figure CN114723639B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and particularly to an image rotation correction method, system, electronic device, and storage medium. Background Art
[0002] During the process of image acquisition using hardware devices such as mirrorless cameras and mobile phones, due to the shooting angle of the photographer, hand jitter, usage habits, etc., the captured images are not horizontal and have redundant images. Generally, users manually perform image rotation correction through image processing tools. However, rotating and cropping a large number of images is a cumbersome task, so automatic image rotation correction has always been a concern in the academic and industrial fields.
[0003] Traditional angle detection methods include line detection, discrete Fourier transform, etc. However, traditional angle detection methods are only applicable to images with obvious orientations under specific conditions, and the processing time is long, lacking universality. With the development of deep learning, a large number of researchers are committed to developing image rotation correction methods in a data-driven manner. Some current research is to detect the horizon in the image through a semantic line model and rotate the picture to be parallel to the horizon to assist in composition. However, the general methods have insufficient image integrity.
[0004] Aiming at the problems of long image rotation correction time and poor image integrity in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] In this embodiment, an image rotation correction method, system, electronic device, and storage medium are provided to solve the problems of long image rotation correction time and poor image integrity in the related art.
[0006] In the first aspect, in this embodiment, an image rotation correction method is provided, including:
[0007] Obtain an original image, and input the original image into a trained saliency detection model to obtain a saliency map;
[0008] Input the original image and the saliency map into a trained feature extraction network to obtain an image rotation result;
[0009] Perform rotation correction on the original image according to the image rotation result to obtain a target corrected image.
[0010] In some of the embodiments, the feature extraction network includes an attention mechanism learning network, a feature aggregation network, and a fully connected network; the step of inputting the original image and the saliency map into a trained feature extraction network to obtain an image rotation result includes:
[0011] Generate a stitched image based on the original image and the saliency map;
[0012] Input the stitched image into the attention mechanism learning network to obtain an attention image;
[0013] Input the attention image into the feature aggregation network to extract the spatial information features of the attention image and obtain a first feature map;
[0014] Input the first feature map into the fully connected network to obtain the image rotation result.
[0015] In some embodiments, there are at least two attention mechanism learning networks, and each attention mechanism learning network includes a deep learning network and an attention mechanism model; the step of inputting the stitched image into the attention mechanism learning network to obtain an attention image includes:
[0016] Input the stitched image into the deep learning network in the current attention mechanism learning network to obtain a current second feature map;
[0017] Input the current second feature map into the attention mechanism model in the current attention mechanism learning network to capture the direction perception information and position perception information of the current second feature map and obtain a current attention image;
[0018] Input the current attention image into the deep learning network in the next attention mechanism learning network, and repeat the above steps until all the attention mechanism learning networks are traversed to obtain the attention image.
[0019] In some embodiments, the step of inputting the current second feature map into the attention mechanism model in the current attention mechanism learning network to capture the direction perception information and position perception information of the current second feature map and obtain a current attention image includes:
[0020] Input the current second feature map into the residual block in the attention mechanism model to retain the feature information of the current second feature map and obtain a residual map;
[0021] Input the residual map into the first global pooling layer in the attention mechanism model to obtain a first direction perception map;
[0022] Input the residual map into the second global pooling layer in the attention mechanism model to obtain a second direction perception map;
[0023] Stitch the first direction perception map and the second direction perception map to perceive the global features of the first direction perception map and the second direction perception map and generate a third feature map;
[0024] Split the third feature map into a first tensor and a second tensor;
[0025] Input the first tensor into a first convolutional block and a first activation function in the attention mechanism model to obtain a first weight;
[0026] Input the second tensor into a second convolutional block and a first activation function in the attention mechanism model to obtain a second weight;
[0027] Obtain the attention image according to the residual map, the first weight, and the second weight.
[0028] In some embodiments, the fully connected network includes a third convolutional block and fully connected layers; there are at least two fully connected layers; the inputting the first feature map into the fully connected network to extract the spatial information features of the first feature map to obtain an image rotation result includes:
[0029] Input the first feature map into the third convolutional block to reduce the dimensionality of the channel dimension of the first feature map to obtain a dimensionality-reduced image;
[0030] Input the dimensionality-reduced image into the current fully connected layer to obtain a current image rotation result;
[0031] Input the current image rotation result into the next fully connected layer until all the fully connected layers are traversed to obtain the image rotation result.
[0032] In some embodiments, before obtaining the original image, it further includes:
[0033] Obtain a preset number of iterations, a training image set, and a first preset rotation angle of the training image set;
[0034] Input the training image set into the saliency detection model to be trained to obtain a significant training map;
[0035] Input the training image set and the significant training map into the feature extraction network to be trained to obtain an image rotation training result;
[0036] Obtain a first weight loss result according to the image rotation training result and the first preset rotation angle;
[0037] Input the first weight loss result into the feature extraction network to be trained and the saliency detection model to be trained to obtain a first feature extraction network and a first saliency detection model;
[0038] Obtain the current iteration number. When the current iteration number is greater than or equal to the preset iteration number, use the first feature extraction network as the trained feature extraction network, and use the first saliency detection model as the trained saliency detection model;
[0039] When the current iteration number is less than the preset iteration number, obtain a second weight loss result based on the first feature extraction network and the first saliency detection model, and obtain a second feature extraction network and a second saliency detection model based on the second weight loss result.
[0040] In some embodiments, the step of inputting the original image into the trained saliency detection model to obtain a saliency map includes:
[0041] Input the original image into a downsampling module to obtain a downsampled image;
[0042] Input the downsampled image into the trained encoding-decoding model to obtain a binary image;
[0043] Multiply the binary image by the original image to extract the salient region on the original image and obtain the saliency map.
[0044] In a second aspect, in this embodiment, an image rotation correction system is provided, which includes: a terminal device, a transmission device, and a server device; wherein, the terminal device is connected to the server device through the transmission device;
[0045] The server device is configured to execute the image rotation correction method according to any one of the above first aspects;
[0046] The transmission device is configured to send a target corrected image to the terminal device;
[0047] The terminal device is configured to display the target corrected image.
[0048] In a third aspect, in this embodiment, an electronic device is provided, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the image rotation correction method according to the above first aspect is implemented.
[0049] In a fourth aspect, in this embodiment, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the image rotation correction method according to the above first aspect is implemented.
[0050] Compared with the related art, the image rotation correction method, system, electronic device and storage medium provided in this embodiment obtain the original image, input the original image into a trained saliency detection model to obtain a saliency map; input the original image and the saliency map into a trained feature extraction network to obtain an image rotation result; perform rotation correction on the original image according to the image rotation result to obtain a target corrected image, which solves the problems of long time consumption for image rotation correction and poor image integrity, and realizes fast rotation correction of any image in different scenarios.
[0051] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0053] Figure 1 It is an application scenario diagram of the image rotation correction method in one embodiment;
[0054] Figure 2 It is a flowchart of the image rotation correction method in one embodiment;
[0055] Figure 3 It is a flowchart of the steps of the feature extraction network in one embodiment;
[0056] Figure 4 It is a flowchart of the steps of the feature extraction network in another embodiment;
[0057] Figure 5 It is a flowchart of the steps of the attention mechanism model in one embodiment;
[0058] Figure 6 It is a flowchart of the steps of the fully connected network in one embodiment;
[0059] Figure 7 It is a flowchart of the steps of the saliency detection model in one embodiment;
[0060] Figure 8 It is an internal structure diagram of a computer device in one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To more clearly understand the purpose, technical solution, and advantages of the present application, the present application will be described and illustrated below with reference to the drawings and embodiments.
[0062] Unless otherwise defined, technical terms or scientific terms involved in this application shall have the general meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these", etc. do not indicate a limitation in quantity and can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connect", "be connected", "couple" and other similar words involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " means that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application only distinguish similar objects and do not represent a specific order for the objects.
[0063] The image rotation correction method provided in this application can be applied to an application environment as Figure 1 shown. Among them, the terminal device 102 communicates with the server device 104 through the network. The server device 104 acquires the original image, inputs the original image into a trained saliency detection model to obtain a saliency map; the server device 104 inputs the original image and the saliency map into a trained feature extraction network to obtain an image rotation result; the server device 104 performs rotation correction on the original image according to the image rotation result to obtain a target corrected image. Among them, the terminal device 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices, and the server device 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0064] In this embodiment, an image rotation correction method is provided. Figure 2 is a flowchart of the image rotation correction method in this embodiment, as Figure 2 shown, and the process includes the following steps:
[0065] Step S202, acquire the original image, input the original image into a trained saliency detection model to obtain a saliency map.
[0066] Among them, the saliency detection model is used to perform complete object extraction on salient parts such as human figures and foci in the original image, so as to maximize the integrity of the objects in the original image; the saliency detection model can adopt various models, such as Itti (a visual attention model based on Gaussian pyramid to fuse image color, brightness and orientation features), SR (Spectral Residual), FT (Frequency-tuned), HC (Histogram-based Contrast), CA (Context-Aware), GR (Graph-Regularized) algorithms, etc.
[0067] Step S204: Input the original image and the saliency map into the trained feature extraction network to obtain the image rotation result.
[0068] Among them, the feature extraction network is used to perform feature extraction on the original images at any angle collected in different scenarios based on the saliency map obtained by the saliency detection model, so as to quickly obtain the image rotation result based on the result of the feature extraction, and then realize the rotation of the original image; specifically, the image rotation result can be a value in the range of [-1, 1], and this value is used to indicate at which angle the original image is located after subsequent rotation correction.
[0069] Step S206: Perform rotation correction on the original image according to the image rotation result to obtain the target corrected image.
[0070] Specifically, step S206 needs to correspondingly convert the value of the image rotation result in the range of [-1, 1] into a rotation angle value in the range of [-π, π], so as to rotate the original image around the center point of the image according to this rotation angle value, and at the same time crop the rotated image according to the image range of the saliency map to obtain the target corrected image.
[0071] Through the above steps, by performing saliency detection on the original image and extracting the salient region to obtain the saliency map, the subsequent region position of the salient object in the image in the feature extraction network can be made more sensitive, the integrity of the image object can be maximally retained, and based on the original image and the saliency map for feature extraction, the angle rotation of the original image collected at any angle in any scenario can be realized, and fast cropping can be realized based on the salient object in the detected saliency map, thus solving the problems of long time-consuming for image rotation correction and poor image integrity, and realizing the fast rotation correction of any image in different scenarios.
[0072] In some of these embodiments, the feature extraction network includes an attention mechanism learning network, a feature aggregation network, and a fully connected network; inputting the original image and the saliency map into the trained feature extraction network to obtain an image rotation result includes:
[0073] Generating a spliced image according to the original image and the saliency map;
[0074] Inputting the spliced image into the attention mechanism learning network to obtain an attention image;
[0075] Inputting the attention image into the feature aggregation network to extract the spatial information features of the attention image to obtain a first feature map;
[0076] Inputting the first feature map into the fully connected network to obtain the image rotation result.
[0077] Among them, the spliced image is formed by splicing the original image and the saliency map; the attention mechanism learning network is used to perceive features in aspects such as the direction and position of the image according to the spliced image; the feature aggregation network is used to learn global context information, specifically, it can learn global context information through bilinear upsampling and downsampling respectively; the feature aggregation network includes at least two attention feature aggregation networks and a feature splicing layer, and feature aggregation between different images can be achieved through skip connections between layers; the attention feature aggregation network is the same as the attention mechanism learning network; the spatial information features of the attention image refer to features in aspects such as the direction and position of the attention image.
[0078] Specifically, taking two such attention feature aggregation networks as an example, Figure 3 is a flowchart of the steps of a feature extraction network in this embodiment. As Figure 3 shown, the original image and the saliency map are spliced along the channel dimension to obtain a spliced image; the spliced image is input into the attention mechanism learning network to obtain an attention image; each channel value of the attention image is divided by 2 to obtain a first attention aggregation map; the attention image is input into the first attention feature aggregation network of the feature aggregation network to obtain a second attention aggregation map; the second attention aggregation map is input into the second attention feature aggregation network of the feature aggregation network to obtain a third attention aggregation map; each channel value on the third attention aggregation map is multiplied by 2 to obtain a fourth attention aggregation map; the first attention aggregation map, the second attention aggregation map, and the fourth attention aggregation map are spliced to obtain the first feature map; the first feature map is input into the fully connected network to obtain the image rotation result.
[0079] Through the above steps, the attention mechanism learning network and the feature aggregation network of the feature extraction network are used to convert the original image into a first feature map rich in information that can represent both global context and local context, compensate for the loss of global and multi-scale context during feature extraction by the attention mechanism learning network, and further obtain the image rotation result. The feature extraction network has the characteristics of multi-scale and lightweight, and can realize the angle rotation of the original image collected at any angle in any scene. Based on the significant objects in the detected saliency map, fast cropping can be realized, thus solving the problems of long time-consuming image rotation correction and poor image integrity, and realizing the fast rotation correction of arbitrary images in different scenes.
[0080] In some of these embodiments, there are at least two of the attention mechanism learning networks, and each of the attention mechanism learning networks includes a deep learning network and an attention mechanism model; the step of inputting the spliced image into the attention mechanism learning network to obtain an attention image includes:
[0081] Input the spliced image into the deep learning network in the current attention mechanism learning network to obtain the current second feature map;
[0082] Input the current second feature map into the attention mechanism model in the current attention mechanism learning network to capture the direction perception information and position perception information of the current second feature map, and obtain the current attention image;
[0083] Input the current attention image into the deep learning network in the next attention mechanism learning network, and repeat the above steps until all the attention mechanism learning networks are traversed to obtain the attention image.
[0084] Among them, the deep learning network can be a neural network or a convolutional neural network, etc., for example, it can be VGG, ResNet, etc.; the attention mechanism model is used to capture the direction perception information and position perception information of the current second feature map and calculate the attention weight in the direction.
[0085] Specifically, the attention feature aggregation network is the same as the attention mechanism learning network. Taking three of the attention mechanism learning networks and two of the attention feature aggregation networks as an example, Figure 4 is a flowchart of another step of the feature extraction network in this embodiment, as Figure 4As shown, the original image and the saliency map are concatenated along the channel dimension to obtain a concatenated map; the concatenated map is input into the first deep learning network of the first attention mechanism learning network to obtain a first current second feature map, and the first current second feature map is input into the first attention mechanism model of the first attention mechanism learning network to obtain a first attention image; the first attention image is input into the second deep learning network of the second attention mechanism learning network to obtain a second current second feature map, and the second current second feature map is input into the second attention mechanism model of the second attention mechanism learning network to obtain a second attention image; the second attention image is input into the third deep learning network of the third attention mechanism learning network to obtain a third current second feature map, and the third current second feature map is input into the third attention mechanism model of the third attention mechanism learning network to obtain a third attention image; each channel value of the third attention image is divided by 2 to obtain a first attention aggregation map; the third attention image is input into the fourth deep learning network of the first attention feature aggregation network to obtain a fourth current second feature map, and the fourth current second feature map is input into the fourth attention mechanism model of the first attention feature aggregation network to obtain a second attention aggregation map; the second attention aggregation map is input into the fifth deep learning network of the second attention feature aggregation network to obtain a fifth current second feature map, and the fifth current second feature map is input into the fifth attention mechanism model of the second attention feature aggregation network to obtain a third attention aggregation map; each channel value on the third attention aggregation map is multiplied by 2 to obtain a fourth attention aggregation map; the first attention aggregation map, the second attention aggregation map, and the fourth attention aggregation map are concatenated to obtain the first feature map; the first feature map is input into the fully connected network to obtain the image rotation result.
[0086] Through the above steps, by sequentially inputting the concatenated map into each of the attention mechanism learning networks, the direction perception information and position perception information in the image can be captured, and the original image collected at any angle in any scene can be quickly rotated, thus solving the problems of long time-consuming image rotation correction and poor image integrity, and realizing the fast rotation correction of any image in different scenes.
[0087] In some of the embodiments, inputting the current second feature map into the attention mechanism model in the current attention mechanism learning network to capture the direction perception information and position perception information of the current second feature map to obtain the current attention image includes:
[0088] Inputting the current second feature map into the residual block in the attention mechanism model to retain the feature information of the current second feature map to obtain a residual map;
[0089] Input the residual map into the first global pooling layer in the attention mechanism model to obtain a first direction perception map;
[0090] Input the residual map into the second global pooling layer in the attention mechanism model to obtain a second direction perception map;
[0091] Concatenate the first direction perception map and the second direction perception map, perceive the global features of the first direction perception map and the second direction perception map, and generate a third feature map;
[0092] Slice the third feature map into a first tensor and a second tensor;
[0093] Input the first tensor into the first convolutional block and the first activation function in the attention mechanism model to obtain a first weight;
[0094] Input the second tensor into the second convolutional block and the first activation function in the attention mechanism model to obtain a second weight;
[0095] Obtain the attention image according to the residual map, the first weight, and the second weight.
[0096] Among them, the first global pooling layer can be a horizontal global pooling layer, and the first direction perception map can be a horizontal direction perception map; the second global pooling layer can be a horizontal global pooling layer, and the second direction perception map can be a horizontal direction perception map; the above first direction and second direction can also be direction features that can describe the image coordinates; concatenating the first direction perception map and the second direction perception map can be, after concatenating the first direction perception map and the second direction perception map, inputting them into a global feature perception layer that makes the image effect clearer to generate the third feature map, for example, it can be input into a downsampling convolutional layer, a batch normalization layer, and / or a non-linear layer; the first tensor corresponds to the first direction perception map, and the second tensor corresponds to the second direction perception map; the first convolutional block is a downsampling convolutional block, which can be a 1×1 convolutional block; the second convolutional block is a downsampling convolutional block, which can be a 1×1 convolutional block; the first activation function can be a Sigmoid activation function.
[0097] Figure 5 is a flowchart of the steps of the attention mechanism model in this embodiment, as Figure 5As shown, input the current second feature map into the residual block in the attention mechanism model to retain the feature information of the current second feature map, obtaining a residual map; input the residual map into the first global pooling layer in the attention mechanism model to obtain a first direction perception map; simultaneously, input the residual map into the second global pooling layer in the attention mechanism model to obtain a second direction perception map; after successively splicing, downsampling, batch normalization, and / or non-linearity of the first direction perception map and the second direction perception map, generate a third feature map; respectively divide the third feature map into a first tensor and a second tensor, where the first tensor corresponds to the first direction perception map and the second tensor corresponds to the second direction perception map; input the first tensor into the first convolutional block and the first activation function in the attention mechanism model to obtain a first weight; simultaneously, input the second tensor into the second convolutional block and the first activation function in the attention mechanism model to obtain a second weight; obtain the attention image based on the residual map, the first weight, and the second weight.
[0098] Specifically, taking the size of the current second feature map X as C×H×W as an example, c is the number of channels of the current second feature map, H is the height of the current second feature map, and W is the width of the current second feature map; first, use a horizontal global pooling layer of size C×H×1 and a vertical global pooling layer of size C×1×W to encode each channel, respectively generating a horizontal direction perception map and a vertical direction perception map; for the horizontal global pooling layer to obtain the horizontal direction perception map, for the c-th channel with height h, the output value is as follows:
[0099]
[0100] where x is the value on the channel of the current second feature map X.
[0101] For the vertical global pooling layer to obtain the vertical direction perception map, the output value of the c-th channel with width w is as follows:
[0102]
[0103] The above two global pooling layers, namely the horizontal global pooling layer and the vertical global pooling layer, perform feature aggregation along two directions in the spatial direction, which can capture long-range dependencies along one spatial direction, that is, can establish connections with distant pixels in one spatial direction. At the same time, accurate position information is retained along the other spatial direction.
[0104] Splice the horizontal direction perception map and the vertical direction perception map in the spatial dimension, and successively input them into a downsampling convolutional layer, a batch normalization layer, and / or a non-linear layer after splicing to obtain a third feature map, as shown in the following formula, where the generated f is the third feature map generated by spatial information in the horizontal and vertical directions.
[0105] f = δ(F1([z h , z w ))
[0106] where [z h , z w represents the concatenation of the horizontal direction perception map and the vertical direction perception map; F1 is a channel transformation function of 1x1 convolution; δ is a non-linear activation function.
[0107] The third feature map f is sliced along the spatial dimension into two separate first tensors f h and a second tensor f w ; then two 1×1 convolutions F h and F w are used to transform the feature maps f h and f w through the first activation function σ to the same number of channels as the current second feature map X respectively, obtaining the first weight g h and the second weight g w , as shown in the following formula:
[0108] g h = σ(F h (f h ))
[0109] g w = σ(F w (f w ))
[0110] Finally, the output of the attention mechanism model is as follows:
[0111]
[0112] where y c (i, j) represents the values on each channel of the attention image; x c (i, j) represents the values on each channel of the current second feature map.
[0113] Through the above steps, the attention mechanism model aggregates the input features in two directions into two independent direction perception feature maps through two one-dimensional global pooling layers respectively, so that while capturing long-range dependencies along one spatial direction, precise position information can be retained along the other spatial direction, thereby enhancing the representation ability of the attention image. The attention mechanism model can capture the direction perception information and position perception information in the image, realize the rapid rotation of the original image collected at any angle in any scene, thus solving the problems of long time consumption for image rotation correction and poor image integrity, and achieving the rapid rotation correction of arbitrary images in different scenes.
[0114] In some of these embodiments, the fully connected network includes a third convolutional block and fully connected layers; there are at least two fully connected layers; inputting the first feature map into the fully connected network to extract the spatial information features of the first feature map and obtain an image rotation result includes:
[0115] Input the first feature map into the third convolutional block to reduce the dimension of the channel dimension of the first feature map and obtain a dimension-reduced image;
[0116] Input the dimension-reduced image into the current fully connected layer to obtain the current image rotation result;
[0117] Input the current image rotation result into the next fully connected layer until all the fully connected layers are traversed to obtain the image rotation result.
[0118] Among them, the third convolutional block refers to a convolutional layer with a low channel scale, which can be a 1×1 convolutional block; the fully connected layer adopts a Dropout strategy during calculation to alleviate model overfitting; the second activation function of the fully connected layer is the TanH activation function, and the calculation result is a value in the interval [-1, 1].
[0119] Preferably, when there are two fully connected layers, Figure 6 is a flowchart of the steps of the fully connected network in this embodiment. As Figure 6 shown, input the first feature map into the third convolutional block to reduce the dimension of the channel dimension of the first feature map and obtain a dimension-reduced image; input the dimension-reduced image into the current fully connected layer to obtain the current image rotation result; input the current image rotation result into the next fully connected layer to obtain the image rotation result.
[0120] Through the above steps, by using the third convolutional block and at least two fully connected layers, the problems of long image rotation correction time and poor image integrity are solved, and fast rotation correction of arbitrary images in different scenarios is realized.
[0121] In some of these embodiments, before obtaining the original image, it further includes:
[0122] Obtain a preset number of iterations, a training image set, and a first preset rotation angle of the training image set;
[0123] Input the training image set into the saliency detection model to be trained to obtain a significant training map;
[0124] Input the training image set and the significant training map into the feature extraction network to be trained to obtain an image rotation training result;
[0125] Obtain a first weight loss result according to the image rotation training result and the preset rotation angle;
[0126] Input the first weight loss result into the feature extraction network to be trained and the saliency detection model to be trained, and obtain a first feature extraction network and a first saliency detection model;
[0127] Obtain the current iteration number. When the current iteration number is greater than or equal to the preset iteration number, use the first feature extraction network as the trained feature extraction network, and use the first saliency detection model as the trained saliency detection model;
[0128] When the current iteration number is less than the preset iteration number, obtain a second weight loss result based on the first feature extraction network and the first saliency detection model, and obtain a second feature extraction network and a second saliency detection model based on the second weight loss result.
[0129] It should be noted that in the above training process of the image rotation correction method, a verification process of the image rotation correction method is also included. Obtain a verification image set and a second preset rotation angle of the verification image set; calculate an accuracy result, a recall result, and an F-measure result respectively according to the second preset rotation angle and the image rotation training result; calculate an average accuracy result, an average recall result, and an average F-measure result according to the accuracy result, the recall result, and the F-measure result; select a trained saliency detection model and a trained feature extraction network according to the average accuracy result, the average recall result, and the average F-measure result. Among them, the selection of the trained saliency detection model and the trained feature extraction network is determined according to the actual situation.
[0130] Preferably, collect publicly uploaded images taken by users, and let professional image retouchers perform rotation and cropping on the images to obtain 40,000 groups of image rotation correction datasets AlltuuRotate, and record the angle θ after image rotation; divide the dataset AlltuuRotate into 24,000 groups as the training set T train and 8,000 groups as the test set T test , and a verification set T containing 8,000 images val ; The configuration of the server device 104 can be an Intel i7-9750H processor, Ubuntu operating system, 64GB of memory, NVIDIA RTX3090 graphics card (24G video memory), using the Pycharm2020.1.1 editing tool to build a Pytorch1.8 environment.
[0131] The training image set is T train , and the first preset rotation angle of the training image set is θ l; During training, the size of each batch of samples is set to 32, the preset number of iterations is 200, the learning rate lr is set to 5e-4, and the training process is executed using the error backpropagation algorithm optimized by adam; the weight calculation loss formula corresponding to the first weight loss result is:
[0132] L = (sinθ p - sinθ l ) 2 + (cosθ p - cosθ l ) 2
[0133] where θ p represents the rotation angle training value corresponding to the image rotation training result.
[0134] The validation image set is T val , and the second preset rotation angle of the validation image set is θ2; for a pair of the second preset rotation angle and the rotation angle verification value corresponding to the image rotation verification result (θ^, θ2), the calculation formulas of S(θ^, θ2), accuracy, recall rate, and F-measure are as follows:
[0135] S(θ^, θ) = θ^ - θ
[0136]
[0137]
[0138]
[0139] where p is the set of rotation angle verification values, G is the set of the second preset rotation angles, ‖·‖ represents the number of elements in the set. l(·) represents 1 only when the condition is true; ω is the threshold, and the verification set results are calculated respectively when ω = 0.01, 0.02, 0.03,..., 0.99, and a series of accuracy results, recall rate results, and F-measure results are obtained correspondingly. Then, the average accuracy result, average recall rate result, and average F-measure result are calculated, and finally, the trained complete saliency detection model and the trained complete feature extraction network are selected according to the average accuracy result, the average recall rate result, and the average F-measure result.
[0140] Through the above steps, by training and validating the saliency detection model to be trained and the feature extraction network to be trained, a trained complete saliency detection model and a trained complete feature extraction network can be obtained, thus solving the problems of long time-consuming image rotation correction and poor image integrity, and realizing the rapid rotation correction of arbitrary images in different scenarios.
[0141] In some of these embodiments, inputting the original image into a trained saliency detection model to obtain a saliency map includes:
[0142] Inputting the original image into a downsampling module to obtain a downsampled image;
[0143] Inputting the downsampled image into a trained encoder-decoder model to obtain a binary map;
[0144] Multiplying the binary map by the original image to extract the salient regions on the original image to obtain the saliency map.
[0145] Among them, the downsampling module downsamples the size of the original image to 512×512; the encoder-decoder model is a Unet-like saliency detection model, such as UNet+++, U2Net, and U2Net+, etc.; in this embodiment, preferably, a trained U2Net+ model is used to perform deep learning on the downsampled image to obtain a binary map, and the original image is processed according to the binary map to obtain the saliency map; the U2Net+ model is trained using the publicly available dataset DUTS-TR dataset; the U2Net+ model includes multiple RSU-L(Cin,M,Cout) encoder-decoder blocks, where L represents the number of layers of the encoder, Cin and Cout are the number of input and output channels respectively, and M is the internal channel number of the RSU; the RSU is composed as follows:
[0146] RSU = F1(x) + U(F1(x))
[0147] Where x is the downsampled image, F1(x) is the intermediate feature map, and U is a Unet-like structure.
[0148] Preferably, taking the number of layers L of the encoder-decoder model as 6 as an example, Figure 7 is a flowchart of the steps of the saliency detection model in this embodiment, as Figure 7As shown in the figure, the original image is input into the downsampling module to obtain a downsampled image; the downsampled image is input into the first encoder of the first encoding and decoding block in the encoding and decoding model to obtain a first encoding result; the first encoding result is input into the second encoder of the second encoding and decoding block to obtain a second encoding result; the second encoding result is input into the third encoder of the third encoding and decoding block to obtain a third encoding result; the third encoding result is input into the fourth encoder of the fourth encoding and decoding block to obtain a fourth encoding result; the fourth encoding result is input into the fifth encoder of the fifth encoding and decoding block to obtain a fifth encoding result; the fifth encoding result is input into the sixth encoder of the sixth encoding and decoding block to obtain a sixth encoding result; the sixth encoding result and the fifth encoding result are input into the fifth decoder to obtain a fifth decoding result; the fifth decoding result and the fourth encoding result are input into the fourth decoder to obtain a fourth decoding result; the fourth decoding result and the third encoding result are input into the third decoder to obtain a third decoding result; the third decoding result and the second encoding result are input into the second decoder to obtain a second decoding result; the second decoding result and the first encoding result are input into the first decoder to obtain the binary image; the binary image is multiplied by the original image to extract the significant region on the original image to obtain the significant map.
[0149] Through the above steps, by performing saliency detection on the original image and extracting the significant region to obtain the significant map, the subsequent region position of the significant object in the image in the feature extraction network is made more sensitive, and the integrity of the image object can be maximally retained, thereby solving the problems of long time-consuming image rotation correction and poor image integrity, and realizing fast rotation correction of any image in different scenarios.
[0150] It should be understood that although Figure 2-7 the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 2-7 at least a part of the steps in
[0151] In this embodiment, an image rotation correction system is further provided, and the system includes: a terminal device 102, a transmission device, and a server device 104; wherein, the terminal device 102 is connected to the server device 104 through the transmission device;
[0152] The server device 104 is used to execute the image rotation correction method in any one of the above method embodiments;
[0153] The transmission device is used to send the target corrected image to the terminal device;
[0154] The terminal device 102 is used to display the target corrected image.
[0155] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0156] Optionally, the above electronic device may further include a transmission device and input / output devices. Among them, the transmission device is connected to the above processor, and the input / output devices are connected to the above processor.
[0157] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0158] S1. Obtain the original image, input the original image into a trained saliency detection model, and obtain a saliency map.
[0159] S2. Input the original image and the saliency map into a trained feature extraction network to obtain an image rotation result.
[0160] S3. Rotate and correct the original image according to the image rotation result to obtain a target corrected image.
[0161] It should be noted that specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated in this embodiment.
[0162] In addition, in combination with the image rotation correction method provided in the above embodiments, a storage medium may also be provided to implement in this embodiment. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the image rotation correction methods in the above embodiments is implemented.
[0163] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 8As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements an image rotation correction method.
[0164] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0165] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0166] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.
[0167] Obviously, the accompanying drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative work. Additionally, it can be understood that although the work done during this development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only routine technical means and should not be regarded as insufficient disclosure of the present application.
[0168] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean independence or alternative to other embodiments that are mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.
[0169] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several variations and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An image rotation correction method, characterized in that, Including: Obtain an original image, input the original image into a trained saliency detection model to obtain a saliency map; Input the original image and the saliency map into a trained feature extraction network to obtain an image rotation result, including: Generate a spliced image according to the original image and the saliency map; Input the spliced image into an attention mechanism learning network to obtain an attention image, including: Input the spliced image into a deep learning network in the current attention mechanism learning network to obtain a current second feature map; Input the current second feature map into an attention mechanism model in the current attention mechanism learning network to capture the direction perception information and position perception information of the current second feature map to obtain a current attention image, including: Input the current second feature map into a residual block in the attention mechanism model to retain the feature information of the current second feature map to obtain a residual map; Input the residual map into a first global pooling layer in the attention mechanism model to obtain a first direction perception map; Input the residual map into a second global pooling layer in the attention mechanism model to obtain a second direction perception map; Splice the first direction perception map and the second direction perception map to perceive the global features of the first direction perception map and the second direction perception map and generate a third feature map; Slice the third feature map into a first tensor and a second tensor; Input the first tensor into a first convolutional block and a first activation function in the attention mechanism model to obtain a first weight; Input the second tensor into a second convolutional block and a first activation function in the attention mechanism model to obtain a second weight; Obtain the attention image according to the residual map, the first weight and the second weight; Input the current attention image into a deep learning network in the next attention mechanism learning network until all the attention mechanism learning networks are traversed to obtain the attention image; there are at least two attention mechanism learning networks, and each attention mechanism learning network includes a deep learning network and an attention mechanism model; Input the attention image into a feature aggregation network to extract the spatial information features of the attention image to obtain a first feature map; Input the first feature map into a fully connected network to obtain the image rotation result; the feature extraction network includes the attention mechanism learning network, the feature aggregation network and the fully connected network; Perform rotation correction on the original image according to the image rotation result to obtain a target corrected image.
2. The image rotation correction method according to claim 1, wherein The fully connected network includes a third convolutional block and a fully connected layer; there are at least two fully connected layers; the step of inputting the first feature map into the fully connected network to obtain an image rotation result includes: Input the first feature map into the third convolutional block to reduce the channel dimension of the first feature map to obtain a dimension-reduced image; Input the dimension-reduced image into the current fully connected layer to obtain a current image rotation result; Input the current image rotation result into the next fully connected layer until all the fully connected layers are traversed to obtain the image rotation result.
3. The image rotation correction method according to claim 1, characterized in that Before obtaining the original image, it further includes: Obtain a preset number of iterations, a training image set, and a first preset rotation angle of the training image set; Input the training image set into the saliency detection model to be trained to obtain a significant training map; Input the training image set and the significant training map into the feature extraction network to be trained to obtain an image rotation training result; Obtain a first weight loss result according to the image rotation training result and the first preset rotation angle; Input the first weight loss result into the feature extraction network to be trained and the saliency detection model to be trained to obtain a first feature extraction network and a first saliency detection model; Obtain the current number of iterations. When the current number of iterations is greater than or equal to the preset number of iterations, use the first feature extraction network as the trained feature extraction network, and use the first saliency detection model as the trained saliency detection model; When the current number of iterations is less than the preset number of iterations, obtain a second weight loss result according to the first feature extraction network and the first saliency detection model, and obtain a second feature extraction network and a second saliency detection model according to the second weight loss result.
4. The image rotation correction method according to claim 1, wherein The step of inputting the original image into the trained saliency detection model to obtain a significant map includes: Input the original image into a downsampling module to obtain a downsampled image; Input the downsampled image into the trained encoding-decoding model to obtain a binary map; Multiply the binary map by the original image to extract the significant region on the original image to obtain the significant map.
5. An image rotation correction system, characterized in that, It includes: A terminal device, a transmission device, and a server device; wherein, the terminal device is connected to the server device through the transmission device; The server device is used to execute the image rotation correction method according to any one of claims 1 to 4; The transmission device is used to send a target corrected image to the terminal device; The terminal device is used to display the target corrected image.
6. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is set to run the computer program to execute the image rotation correction method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps of the image rotation correction method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method, apparatus, medium and apparatus for image recognition based on fine-grained image
CN109409384A
Attention mechanism-embedded iterative aggregation neural network high-resolution remote sensing scene classification method
CN112232151A
Image saliency region detection method and device
CN112861883A
Methods and systems for automatically correcting image rotation
US20210407045A1