Picture correction method and device
By using a multi-dilation factor dilated convolutional layer and a multimodal supervised correction model, multi-scale feature maps are extracted for image correction, which solves the problems of high cost and low robustness in existing technologies and achieves high-precision and fast image correction results.
Patent Information
- Application Number
- CN202511161964.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, image correction methods often rely on expensive optical scanning equipment or image algorithms that require a large number of hyperparameters, resulting in high cost, bulkiness, low robustness, and limited applicability.
A correction model based on multi-dilation factor dilated convolutional layers and multimodal supervision is adopted. Multi-scale feature maps are extracted through depthwise separable convolution, pyramid pooling units and attention units to generate an accurate offset matrix for image correction.
It achieves high-precision and fast image correction, and is suitable for various scenarios, especially on devices with limited computing power, improving robustness and correction accuracy.
Smart Images

Figure CN120876333A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an image correction method and apparatus. Background Technology
[0002] When photographing documents, perspective distortion and geometric distortion often occur, causing the photographed document images to appear distorted and warped. This results in both the images and text within the document images being distorted and warped, negatively impacting the readability of the document images.
[0003] In related correction methods, expensive optical scanning equipment is often used to obtain the three-dimensional structural information of the document through technologies such as lasers or structured light. Then, the distortion parameters are calculated through geometric algorithms, and the document's curvature is flattened accordingly. However, this optical scanning equipment is expensive and bulky, which not only results in excessive costs, but also limits the applicable scenarios.
[0004] In other cases, image algorithms can be used to estimate and model based on parameters such as the curvature layout information of text lines in the document, thereby correcting the curved document. However, this method requires too many hyperparameters to be set, which often cannot be adapted to various scenarios. Too many parameters involved in the calculation result in insufficient robustness of most image algorithms and a long time consumption, which also limits the applicable scenarios.
[0005] Therefore, there is a lack of image correction methods that are applicable to a wide range of scenarios and have good robustness. Summary of the Invention
[0006] This application provides an image correction method that addresses the limitations of related technologies in image correction due to low computational robustness and efficiency, which restricts its application scenarios. In this method, a trained correction model is subjected to depthwise separable convolution on the image to be corrected. This extracts channel spatial feature maps in both channel and spatial dimensions. The pyramid pooling unit in the correction model then extracts features from these channel spatial feature maps, resulting in multi-scale feature maps of different scales. Furthermore, the attention unit in the correction model applies attention weighting to these multi-scale feature maps, enhancing useful information and suppressing useless information. This allows the offset matrix predicted based on the weighted feature maps to more accurately represent the offset of the image to be corrected. After mapping using the offset matrix, an accurate corrected image can be obtained.
[0007] In a first aspect, embodiments of this application provide an image correction method, the method comprising:
[0008] Obtain the image to be corrected;
[0009] The image to be corrected is input into the correction model, which processes the image to obtain the corrected image. The correction model includes at least one dilated convolutional layer, and different dilated convolutional layers are set with different dilation factors.
[0010] Furthermore, the image to be corrected is processed using a correction model to obtain a corrected image, including:
[0011] The image to be corrected is offset by a correction model to obtain the offset matrix of the image to be corrected.
[0012] The offset matrix is used to map each pixel in the image to be corrected to obtain the corrected image.
[0013] Furthermore, the correction model was trained in the following way:
[0014] Obtain the training dataset, which includes multiple sets of metadata. Each set of metadata includes multiple distorted training images and their corresponding actual corrected images.
[0015] Input the metadata of a predetermined number of groups into the correction model to be trained, and output the prediction offset matrix corresponding to each distorted training image;
[0016] The corresponding corrected prediction image is obtained by mapping each pixel in each distorted training image according to the predicted offset matrix.
[0017] The multimodal loss value corresponding to each corrected prediction image is determined based on the multimodal loss function. Each multimodal loss value represents the difference between the corresponding corrected prediction image and the corresponding actual corrected image in multiple dimensions.
[0018] The corrected model to be trained is trained based on the multimodal loss value to obtain the trained corrected model.
[0019] Furthermore, the corrective model to be trained is trained based on the multimodal loss value to obtain the trained corrective model, including:
[0020] If the multimodal loss value is greater than the preset loss value threshold and / or the training dataset has not completed a predetermined number of training iterations, the parameters in the correction model to be trained are adjusted, and different metadata of the same number of groups are input into the adjusted correction model to generate correction prediction images corresponding to the distorted training images of each group of metadata.
[0021] If the multimodal loss value is less than or equal to the loss value threshold and / or the training dataset has completed a predetermined number of training iterations, the current correction model to be trained is determined as the completed correction model.
[0022] The correction model to be trained includes sequentially connected deep separable convolutional layers, pyramid pooling units, and attention units.
[0023] Furthermore, a predetermined number of metadata sets are input into the correction model to be trained, and the predicted offset matrix corresponding to the distorted training image for each metadata set is output, including:
[0024] Perform the following processing on the warped training images for each set of metadata:
[0025] The distorted training image is subjected to depthwise convolution and pointwise convolution by depthwise separable convolutional layers to obtain the predicted channel spatial feature map;
[0026] The predicted channel spatial feature map is obtained by performing regular convolution and dilated convolution on the pyramid pooling unit;
[0027] Attention units perform attention weighting on multi-scale feature maps to generate predicted weighted feature maps.
[0028] Generate a prediction offset matrix based on the predicted weighted feature map.
[0029] Each set of metadata also includes the actual offset matrix corresponding to the warped training image; the multimodal loss function includes an offset loss sub-function and an image loss sub-function;
[0030] Furthermore, the multimodal loss value corresponding to each corrected and predicted image is determined based on the multimodal loss function, including:
[0031] Each predicted offset matrix and its corresponding actual offset matrix are input into the offset loss sub-function to obtain the corresponding offset loss;
[0032] Determine the total variational denoising value corresponding to each prediction offset matrix;
[0033] The image loss between each corrected prediction image and the corresponding actual corrected image is determined based on the image loss sub-function;
[0034] For each corrected prediction image, the corresponding offset loss, the corresponding total variation denoising value, and the corresponding image loss are weighted to obtain the corresponding multimodal loss.
[0035] Each distorted training image includes a first text region and a first straight line region; the corresponding actual corrected image includes a second text region corresponding to the first text region and a second straight line region corresponding to the first straight line region; the image loss sub-function includes a text region loss function and a straight line region loss function.
[0036] Furthermore, the image loss between each corrected predicted image and the corresponding actual corrected image is determined based on the image loss sub-function, including:
[0037] The text region loss between each first text region and its corresponding second text region is determined based on the text region loss function.
[0038] Based on each straight line mask, determine the first straight line region in the corresponding correction prediction image and the second straight line region in the corresponding actual correction image;
[0039] The image loss is obtained by weighting the loss of each text region and the corresponding line region.
[0040] Accordingly, each set of metadata also includes a text edge mask corresponding to each second text region and a line mask corresponding to each second line region;
[0041] Furthermore, before determining the image loss between each corrected predicted image and the corresponding actual corrected image based on the image loss subfunction, the method also includes:
[0042] The first text region in the corresponding corrected prediction image is determined based on the edge mask of each text.
[0043] The first straight-line region in the corresponding corrected prediction image is determined based on the straight-line region loss function.
[0044] Among them, depthwise separable convolutional layers include depthwise convolutional layers and pointwise convolutional layers;
[0045] Furthermore, depthwise separable convolutional layers are used to perform depthwise convolution and pointwise convolution on the distorted training image to obtain the predicted channel space feature map, including:
[0046] Perform the following processing on the warped training images for each set of metadata:
[0047] Spatial information is obtained by performing spatial depth convolution on the distorted training image using a deep convolutional layer.
[0048] The spatial information of each channel is convolved and mixed pointwise by a pointwise convolutional layer to generate a channel spatial feature map.
[0049] The pyramid pooling unit includes a single regular convolutional layer and at least one dilated convolutional layer. The convolutional kernels of the regular convolutional layer and each dilated convolutional layer are the same and connected in parallel. Each dilation factor is used to control the receptive field of the corresponding dilated convolutional layer.
[0050] Furthermore, the pyramid pooling unit performs regular convolution and dilated convolution on the corresponding predicted channel spatial feature map to obtain the corresponding multi-scale feature map, including:
[0051] Each prediction channel spatial feature map is convolved by a regular convolutional layer to generate the corresponding first-scale feature map.
[0052] Each dilated convolutional layer performs dilated convolution on the spatial feature map of each prediction channel with different receptive fields, thereby generating multiple second-scale feature maps with different receptive fields.
[0053] The first-scale feature map and the corresponding second-scale feature maps of each predicted channel spatial feature map are concatenated to obtain the corresponding multi-scale feature map.
[0054] Furthermore, the offset matrix represents the positional difference between each pixel in the image to be corrected and the corresponding pixel in the preset distortion-free image;
[0055] The offset matrix is used to map each pixel in the image to be corrected, resulting in the corrected image, including:
[0056] Based on the positional differences of each pixel in the offset matrix, the positions of each pixel in the image to be corrected are adjusted to obtain a corrected image with the same pixel positions as the undistorted image.
[0057] Secondly, embodiments of this application also provide an image correction device, which includes: an acquisition module, a prediction module, and a mapping module;
[0058] The acquisition module is configured to acquire the image to be corrected.
[0059] The prediction module is configured to input the image to be corrected into the correction model, and process the image to be corrected through the correction model to obtain the corrected image. The correction model includes at least one dilated convolutional layer, and different dilated convolutional layers are set with different dilation factors.
[0060] Thirdly, embodiments of this application also provide an image correction device, the device comprising:
[0061] One or more processors;
[0062] Storage device, configured to store one or more programs,
[0063] When one or more programs are executed by one or more processors, the one or more processors implement the image correction method of the embodiments of this application.
[0064] Fourthly, embodiments of this application also provide a non-volatile storage medium for storing computer-executable instructions, which are configured to perform the image correction method of embodiments of this application when executed by a computer processor.
[0065] Fifthly, embodiments of this application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads from the computer-readable storage medium and executes the computer program, causing the device to perform the image correction method of embodiments of this application.
[0066] As can be seen from the above, the image correction method and apparatus provided in this application achieve high-precision document image correction through multi-dilation factor dilated convolutional layers. Dilated convolutional layers with different dilation factors can effectively capture multi-scale features from local text edges to global page curvature through different receptive fields, ensuring the accuracy of correction. At the same time, when acquiring multi-scale features with different receptive fields, the convolution kernels of each dilated convolutional layer are not changed. Therefore, compared with the method of expanding the receptive field through different convolution kernels, the method of expanding the receptive field through dilation factors in this application results in a smaller computational load for the trained correction model, which can achieve fast correction of the image to be corrected without introducing too many parameters. Furthermore, since no excessive parameters need to be introduced, the trained correction model can also be applied to more scenarios with limited computing power. Attached Figure Description
[0067] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0068] Figure 1 A flowchart of an image correction method provided in an embodiment of this application;
[0069] Figure 2 A flowchart illustrating a correction model processing method provided in this application embodiment;
[0070] Figure 3 A flowchart illustrating a correction model training method provided in this application embodiment;
[0071] Figure 4 A flowchart illustrating an offset matrix prediction method provided in this application embodiment;
[0072] Figure 5 A structural diagram of an attention unit provided in an embodiment of this application;
[0073] Figure 6 A flowchart illustrating a method for calculating multimodal loss values provided in this application embodiment;
[0074] Figure 7A structural diagram of a depth-separable convolutional layer provided in an embodiment of this application;
[0075] Figure 8 A flowchart illustrating a depthwise separable convolution method provided in this application embodiment;
[0076] Figure 9 A structural diagram of a FasterASPP structure provided in an embodiment of this application;
[0077] Figure 10 A flowchart illustrating a multi-scale feature map determination method provided in this application embodiment;
[0078] Figure 11 A framework diagram of another correction model training method provided in the embodiments of this application;
[0079] Figure 12 A structural block diagram of an image correction device provided in an embodiment of this application;
[0080] Figure 13 This is a schematic diagram of the structure of an image correction device provided in an embodiment of this application. Detailed Implementation
[0081] The embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of this application and are not intended to limit the scope of the embodiments. Furthermore, it should be noted that, for ease of description, only the parts relevant to the embodiments of this application are shown in the accompanying drawings, not the entire structure.
[0082] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and are not limited in number; for example, a first object can be one or more. Furthermore, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. Words such as "comprising" or "including" mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect. "Above," "below," "left," "right," etc., are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0083] As described in the background section, the existing image correction methods are still insufficient to meet the needs of practical use.
[0084] In the process of implementing this application, the applicant discovered that the main problem with the relevant image correction methods is that when correcting images, expensive optical scanning equipment is often used to obtain the three-dimensional structural information of the document through technologies such as lasers or structured light. Then, the distortion parameters are calculated through geometric algorithms, and the document's curvature is flattened accordingly. However, such optical scanning equipment is expensive and bulky, which not only results in excessive costs but also limits the applicable scenarios.
[0085] In other cases, when correcting images, image algorithms are used to estimate and model based on parameters such as the curvature layout information of text lines in the document, thereby correcting the curved document. However, this method requires setting too many hyperparameters, which often cannot adapt to various scenarios. Too many parameters involved in the calculation result in insufficient robustness of most image algorithms and long processing time, which also limits the applicable scenarios.
[0086] Based on this, one or more embodiments of this application provide an image correction method that uses dilated convolution with different dilation factors in the correction model to acquire multi-scale feature maps under different receptive fields. This enables accurate and fast prediction of the offset matrix without involving a large number of parameters, and then the image to be corrected can be accurately corrected by mapping the offset matrix.
[0087] Because the correction model adjusts the receptive field by using different dilation factors, it avoids a large number of parameter calculations and achieves lightweight computation, thus adapting to image correction in different scenarios. Relevant application scenarios include image correction using electronic devices with insufficient computing power, or image correction scenarios where optical scanning equipment is unsuitable, etc., which are not limited to these in this application embodiment.
[0088] The image correction method provided in this application embodiment can be executed by a computer device. The computer device refers to any electronic device with data computing, processing and storage capabilities, such as mobile phones, PCs (Personal Computers), tablet computers and other terminal devices. This application embodiment does not limit this.
[0089] The embodiments of this application are described in detail below with reference to the accompanying drawings.
[0090] Figure 1 This is a flowchart illustrating an image correction method provided in an embodiment of this application. Figure 1 As shown, it includes the following steps:
[0091] Step S101: Obtain the image to be corrected.
[0092] Among them, the image to be corrected is an image whose content is distorted, such as images, graphics, photographs and / or text in the image being displayed as curved or other non-realistic distortions.
[0093] Step S102: Input the image to be corrected into the correction model. The correction model processes the image to be corrected to obtain the corrected image. The correction model includes at least one dilated convolutional layer, and different dilated convolutional layers are set with different dilation factors.
[0094] In the trained correction model, one or more dilated convolutional layers have the same convolutional kernel but different dilation factors. The dilation factor is used to control the receptive field of each dilated convolution.
[0095] Specifically, when performing dilated convolution in a dilated convolution layer, holes can be inserted into the convolution kernel. The corresponding dilation factor can determine the number of inserted holes, thereby creating an effect of expanding the receptive field.
[0096] Based on different dilation factors, each dilated convolutional layer can form different receptive fields, thus enabling each dilated convolution to collect features at different scales of the receptive fields. Furthermore, after fusing the collection results of each dilated convolution, the resulting feature map is a multi-scale feature map that incorporates multiple receptive fields.
[0097] The correction model in this step is trained based on multimodal supervision. Multimodal supervision means that during the training of the correction model, the loss function used in each training round includes differences in multiple dimensions. For example, the difference between the image predicted in this round and the actual distortion-free image in the image itself, and the difference between the offset field predicted using the training image and the actual offset field in the offset field dimension.
[0098] The differences in the dimensions of the image itself can be further included in the differences between the predicted image and the actual distortion-free image in the text region dimension, and the differences between the predicted image and the actual distortion-free image in the line region dimension.
[0099] Based on this, the correction model trained with multi-dimensional multimodal supervision is designed so that the parameters of the trained correction model are set based on multiple dimensions. This allows for accurate prediction of the corrected image when the trained correction model is used for prediction.
[0100] As can be seen, the image correction method of the above embodiments of this application achieves high-precision document image correction through multi-dilation factor dilated convolutional layers and a correction model trained under multimodal supervision. Dilated convolutional layers with different dilation factors can effectively capture multi-scale features from local text edges to global page curvature, ensuring the accuracy of correction. Combined with multimodal supervision, the correction model is trained by jointly training multiple dimensions. Thus, when correcting the image to be corrected, the trained correction model can correct distortions in the image from multiple dimensions, resulting in a more accurate corrected image. Furthermore, the process of correcting from multiple dimensions also ensures higher stability and better robustness of the trained correction model. In addition, compared with expanding the receptive field through different convolutional kernels, the method of expanding the receptive field through dilation factors in this application results in a smaller computational load for the trained correction model, enabling fast correction of the image to be corrected without introducing too many parameters. Moreover, since no excessive parameters are required, the trained correction model can be applied to more scenarios with limited computing power.
[0101] Figure 2 A flowchart illustrating a correction model processing method provided in an embodiment of this application. Figure 2 As shown, it includes the following steps:
[0102] Step S201: The image to be corrected is offset using the correction model to obtain the offset matrix of the image to be corrected.
[0103] The offset matrix represents the distortion of the image to be corrected. Each pixel in the offset matrix can be, for example, the positional difference between each pixel in the distorted image to be corrected and the corresponding pixel in the undistorted image.
[0104] In this step, after the image to be corrected is input into the correction model, the correction model can accurately predict the offset matrix of the image to be corrected.
[0105] Step S202: Map each pixel in the image to be corrected according to the offset matrix to obtain the corrected image.
[0106] Based on the offset matrix predicted by the correction model in S201, each element in the offset matrix can be used to map each pixel in the image to be corrected. Thus, the position of each element in the image to be corrected can be adjusted by the aforementioned positional difference to obtain the corrected image.
[0107] Specifically, based on the obtained offset matrix, which represents the positional difference between each pixel in the image to be corrected and the corresponding pixel in the expected distortion-free image, the correction model can adjust the position of each pixel in the image to be corrected according to this offset matrix. After adjusting the position of each pixel, the corrected image is obtained. Ideally, the position of each pixel in the corrected image should be the same as the position of the corresponding pixel in the distortion-free image.
[0108] In some alternative implementations, the mapping operation of the image to be corrected based on the offset matrix can also be performed by another mapping model. That is, after the correction model predicts the offset matrix, the predicted offset matrix and the image to be corrected are input into the mapping model, and the mapping process is performed through the mapping model to obtain the corresponding corrected image.
[0109] Figure 3 This is a flowchart illustrating a correction model training method provided in an embodiment of this application. Figure 2 As shown, it includes the following steps:
[0110] Step S301: Obtain the training dataset. The training dataset includes multiple sets of metadata. Each set of metadata includes multiple distorted training images and their corresponding actual corrected images.
[0111] In the process of training based on multimodal supervision, it is necessary to explicitly obtain the training dataset for training.
[0112] The training dataset contains multiple sets of metadata. Each set of original data contains distorted training images used as input to the correction model to be trained. These distorted training images are images with distorted content. Each set of metadata also includes the actual correction images corresponding to the distorted training images. These actual correction images are images that are expected to be free of distortion.
[0113] In this step, during the acquisition of each distortion training image, text, images, or graphics can be captured by means of photography or scanning. During the acquisition, the carrier of the text, image, or graphics is bent to obtain a distortion training image with distorted content. If the carrier of the text, image, or graphics is not bent or the bending is negligible, an actual corrected image with no distortion or negligible distortion is acquired.
[0114] Specifically, the carrier of text, images, or graphics can be, for example, a printed piece of paper with text, images, or graphics. Different bending methods of the same paper can produce multiple different distortion training images. That is, paper with the same text, image, or graphic content can be combined with the same actual correction image to form multiple sets of metadata through different distortion methods.
[0115] Among them, unbent sheets of paper taken from different angles can be used as distorted training images with perspective distortion.
[0116] Step S302: Input the metadata of a predetermined number of groups into the correction model to be trained, and output the prediction offset matrix corresponding to each distorted training image.
[0117] Based on the training dataset determined in S301 above, all metadata can be divided into multiple parts. That is, each part contains multiple sets of metadata, and each part constitutes all the metadata. In each round of training, the metadata of each set in one part is input into the correction model to be trained.
[0118] In the correction model to be trained, the number of metadata groups processed in each round of training can be set, and in the process of dividing all metadata in the training dataset into multiple parts, the division can be carried out according to the set number of groups, so that the number of metadata groups in each part is the same as the set number of groups.
[0119] Based on this, after inputting the metadata of any part into the correction model to be trained, the correction model to be trained can predict the prediction offset matrix corresponding to the distorted training image for each set of metadata.
[0120] Step S303: Map each pixel in each distorted training image according to the prediction offset matrix to obtain the corrected prediction image corresponding to each distorted training image.
[0121] Based on the prediction offset matrices determined in S302 above, the corresponding distorted training images can be corrected according to the prediction offset matrices to obtain corrected prediction images.
[0122] Specifically, since each pixel in the prediction offset matrix represents the positional difference between each pixel in the distorted training image and the corresponding pixel in the actual corrected image without distortion, the Remap function can be pre-set, and the aforementioned prediction offset matrix can be input into the Remap function to perform mapping processing on each pixel in the corresponding distorted training image, thereby obtaining the corrected prediction image corresponding to the distorted training image after the mapping processing.
[0123] The Remap function can be a pre-defined function in the database, such as a function in the OpenCV (Open Source Computer Vision) library. The Remap function can adjust the position of each pixel in the distorted training image based on the input prediction offset matrix, so that after adjusting the position of each pixel, a corrected prediction image is obtained. Ideally, the position of each pixel in the corrected prediction image should be the same as the position of the corresponding pixel in the actual corrected image without distortion.
[0124] Step S304: Determine the multimodal loss value corresponding to each corrected prediction image based on the multimodal loss function. Each multimodal loss value represents the difference between the corresponding corrected prediction image and the corresponding actual corrected image in multiple dimensions.
[0125] The multimodal loss function includes an evaluation of the corrected prediction image from multiple dimensions. These dimensions can be, for example, the difference between the corrected prediction image and the corresponding actual corrected image in the image's own dimensions, and the difference between the offset field predicted using the corrected prediction image and the offset field of the actual corrected image in the offset field dimension.
[0126] The differences in the dimensions of the image itself can be further included in the differences between the predicted image and the actual corrected image in the dimension of the text region, and the differences between the predicted image and the actual corrected image in the dimension of the line region.
[0127] Based on the pre-set multimodal loss function, the corrected prediction image determined in S303 can be input into the multimodal loss function for calculation to obtain the multimodal loss value.
[0128] As can be seen, the multimodal loss value is the evaluation result of the corrected prediction image through multiple dimensions. Therefore, the multimodal loss value can represent the differences between the corrected prediction image and the corresponding distorted training image in multiple dimensions.
[0129] Step S305: Train the correction model to be trained based on the multimodal loss value to obtain the trained correction model.
[0130] Based on the multimodal loss value determined in S304 above, the completion of training can be determined by comparing the magnitude of the multimodal loss value with the pre-set loss threshold.
[0131] In cases where the multimodal loss value is greater than the preset loss value threshold and / or the training dataset has not completed a predetermined number of training iterations, the parameters in the correction model to be trained are adjusted, and different metadata of the same number of groups are input into the adjusted correction model to generate correction prediction images corresponding to the distorted training images of each group of metadata.
[0132] Specifically, if the multimodal loss value is greater than the loss threshold, it can be considered that the generated corrected prediction image still has a large error compared with the expected actual corrected image, and the correction model to be trained still needs to be continuously trained.
[0133] Based on this, the parameters in the correction model to be trained can be adjusted according to the difference between the multimodal loss value and the loss threshold. After adjustment, based on the multiple parts of the training dataset mentioned above, the metadata of another part is input into the correction model to be trained for the next round of training. In the next round of training, the prediction offset matrix corresponding to each distorted training image is predicted again, and then Remap mapping is performed again based on each prediction offset matrix to generate each correction prediction image, and the multimodal loss value is calculated again.
[0134] In other cases, the number of training rounds can be preset. Before the training dataset completes the required number of rounds, the parameters in the correction model to be trained can be adjusted based on the difference between the multimodal loss value and the loss threshold. After adjustment, based on the multiple parts of the training dataset mentioned above, the metadata of another part is input into the correction model to be trained for the next round of training. In the next round of training, the prediction offset matrix corresponding to each distorted training image is predicted again, and Remap mapping is performed again based on each prediction offset matrix to generate each corrected prediction image, and the multimodal loss value is calculated again.
[0135] In other cases, the multimodal loss value and the number of training rounds can be combined for judgment. For example, if the multimodal loss value is greater than the loss threshold, it is then determined whether the training dataset has completed the predetermined number of rounds. If the number of rounds has not been completed, the parameters in the correction model to be trained are adjusted based on the difference between the multimodal loss value and the loss threshold. After adjustment, based on the multiple parts of the training dataset mentioned above, the metadata of another part is input into the correction model to be trained for the next round of training. In the next round of training, the prediction offset matrix corresponding to each distorted training image is predicted again, and Remap mapping is performed again based on each prediction offset matrix to generate each corrected prediction image, and the multimodal loss value is calculated again.
[0136] Based on the multimodal loss value determined in S304 above, if the multimodal loss value is less than or equal to the loss value threshold and / or the training dataset has completed a predetermined number of training iterations, the current correction model to be trained is determined as the completed correction model.
[0137] Specifically, if the multimodal loss value is less than or equal to the loss threshold, the generated corrected prediction image can be considered to be the same as or close to the expected actual corrected image, with a small error between them, and the corrected model can be considered to have been trained.
[0138] In other cases, based on a pre-set number of training rounds, it can be determined that a trained correction model has been obtained when the training dataset has completed that number of rounds.
[0139] In other cases, the multimodal loss value and the number of training epochs can be combined for judgment. For example, if the multimodal loss value is less than or equal to the loss threshold, it can be determined whether the training dataset has completed the predetermined number of times. If the training dataset has completed the number of times, it can be determined that the trained correction model has been obtained.
[0140] It can be seen that training based on the multimodal loss function brings significant technical improvements to the correction model. By constructing a training dataset containing distorted training images, actual corrected images, and corresponding labeled data, the model can comprehensively learn the complex features when content distortion occurs. Dynamically adjusting the parameters of the correction model during training, coupled with a predetermined number of iterations, ensures the stability and reliability of model convergence. This results in a significant improvement in both accuracy and robustness of the trained correction model, enabling it to adapt to image correction needs in various complex scenarios.
[0141] Figure 4 This is a flowchart illustrating an offset matrix prediction method provided in an embodiment of this application. Figure 4 As shown, the following steps are performed on the warped training images for each set of metadata:
[0142] Step S401: Perform depthwise convolution and pointwise convolution on the distorted training image using a depthwise separable convolutional layer to obtain a predicted channel spatial feature map.
[0143] The correction model to be trained in this application includes sequentially connected depthwise separable convolutional layers, pyramid pooling units, and attention units.
[0144] In the process of generating the prediction offset matrix using the correction model to be trained, after inputting each distorted training image into the correction model to be trained, it can first be subjected to depthwise convolution and pointwise convolution by a depthwise separable convolutional layer to obtain the prediction channel space feature map.
[0145] Specifically, for each distorted training image, a depthwise separable convolutional layer is used to extract spatial features from multiple channels and fuse the spatial features of each channel to obtain a predicted channel spatial feature map. The data of each spatial feature can be, for example, a local feature within a single channel. Local features can be, for example, features of dimensions such as the edges, corners, and textures of the image. Multiple channels can be, for example, multiple different pixel values or different grayscale values.
[0146] In this step, after processing by depthwise separable convolutional layers, the width and height of the predicted channel space feature map are both half that of the distorted training image, thus significantly reducing the computational load of subsequent pyramid pooling units and attention units.
[0147] Step S402: Perform regular convolution and dilated convolution on the predicted channel spatial feature map using pyramid pooling units to obtain a multi-scale feature map.
[0148] Based on the predicted channel spatial feature map determined by the aforementioned S402, feature extraction can be performed on it using pyramid pooling units to obtain a multi-scale feature map.
[0149] Specifically, the pyramid pooling unit can extract features from the feature map of the prediction channel space according to multiple different receptive field scales, thereby obtaining feature maps at different receptive field scales, and then fuse the feature maps at different receptive field scales to obtain a multi-scale feature map.
[0150] As can be seen, the features in this multi-scale feature map represent the fusion results of features extracted at different receptive field scales. Therefore, the features in the distorted training image can be represented from multiple scales.
[0151] Step S403: The attention unit performs attention weighting on the multi-scale feature map to generate a predicted weighted feature map.
[0152] Based on the multi-scale feature map determined in S402 above, attention units can be set to weight the features in the multi-scale feature map according to the importance of each region or feature. When more important regions or features are given higher weights, the resulting predicted weighted feature map can enhance the performance of important regions or features.
[0153] The dimension of the offset matrix can be twice the product of the length and width of the distorted training image.
[0154] Figure 5 This is a structural diagram of an attention unit provided in an embodiment of this application.
[0155] like Figure 5 As shown, the attention unit can be, for example, SAM (Spatial Attention Unit), where, Figure 5 The input features are the features in the multi-scale feature map obtained in S502 above. By assigning corresponding weights to each feature, attention can be increased to important features. After attention weighting of each feature according to the weights, the optimized feature is the prediction weighted feature map.
[0156] It should be noted that when using the sequentially connected deep separable convolutional layer, pyramid pooling unit, and attention unit in the correction model for feature extraction, after the attention unit outputs the predicted weighted feature map for the first time, the predicted weighted feature map output for the first time can be returned to the deep separable convolutional layer, and the feature extraction operation can be performed again through the deep separable convolutional layer, pyramid pooling unit, and attention unit to obtain a more accurate offset matrix.
[0157] Step S404: Generate a prediction offset matrix based on the prediction weighted feature map.
[0158] Based on the predicted weighted feature map determined in S403 above, the correction model to be trained can generate a predicted offset matrix according to the predicted weighted feature map.
[0159] The prediction offset matrix is a matrix that represents the distortion of the distorted training image. Each pixel in the prediction offset matrix can be, for example, the positional difference between each pixel in the distorted training image and the corresponding pixel in the undistorted image.
[0160] It should be noted that for the trained correction model, the offset matrix of the image to be corrected can also be generated in the manner described in S401-S404 above.
[0161] As described above, through the design of a multimodal loss function, efficient end-to-end model optimization is achieved. This not only automatically generates text edge masks and line masks using paired distorted training images and actual corrected images, significantly reducing manual annotation costs, but also effectively avoids jagged edges or breaks in the corrected images by constraining the smoothness of the offset field through TV Loss. The collaborative design of depthwise separable convolution and the FasterASPP module enables the model to converge quickly during training, significantly improving training efficiency.
[0162] In the embodiments of this application, in the process of determining the multimodal loss using the multimodal loss function, the multimodal loss function can specifically evaluate the corrected and predicted image from the dimensions of the image and the offset matrix. Therefore, the multimodal loss function can specifically include an image loss sub-function and an offset loss sub-function.
[0163] Since the distorted training image includes text regions and straight line regions, the image loss sub-function can further evaluate the corrected prediction image from the dimensions of the text regions and straight line regions of the distorted training image. Therefore, the image loss sub-function can include text region loss function and straight line region loss function.
[0164] Based on this, in the process of evaluating the corrected prediction image from the dimension of the offset matrix, it is necessary to distort the actual offset matrix of the training image; and in the process of evaluating the corrected prediction image from the dimensions of the text region and the straight line region, it is necessary to use the text edge mask corresponding to the text region and the straight line mask corresponding to the straight line region.
[0165] Therefore, the metadata for each group in the training dataset also includes the actual offset matrix corresponding to the distorted training image, the text edge mask corresponding to the text region, and the line mask corresponding to the line region.
[0166] Based on this, a multimodal loss function can be constructed, and a multimodal loss value can be determined to characterize the difference between the corrected prediction image and the corresponding distorted training image from multiple dimensions.
[0167] In a specific example, the text edge mask and the line mask can be determined by text edge detection and line detection.
[0168] Specifically, for the actual corrected image, text edge detection can be performed. After detecting the text edges, the text region and / or edge region are binarized, and the binarized actual corrected image is dilated to obtain a text edge mask.
[0169] Furthermore, since there are both visible and hidden lines in the actual corrected image, different line detection methods can be used to perform two line detections.
[0170] The displayed straight line can be, for example, a straight line that exists and is specifically displayed in the actual corrected image; the hidden straight line can be, for example, a straight line formed by text lines or text columns in the actual corrected image, such as the horizontal straight line represented by each text line, and the vertical straight line displayed between each text line due to formatting settings such as left alignment or right alignment.
[0171] When performing line detection on actual corrected images, explicit lines can be detected first using methods such as Hough line detection, and implicit lines can be detected using methods such as DBNet (Differentiable Binarization Network).
[0172] Based on this, a line mask can be obtained after line detection.
[0173] Figure 6 A flowchart illustrating a method for calculating multimodal loss values provided in an embodiment of this application. Figure 6 As shown, it includes the following steps:
[0174] Step S601: Input each predicted offset matrix and the corresponding actual offset matrix into the offset loss sub-function to obtain the corresponding offset loss.
[0175] Based on the offset loss sub-function set in the multimodal loss function mentioned above, the offset loss can be calculated using this offset loss sub-function.
[0176] Specifically, the offset loss represents the difference between the predicted offset matrix output by the correction model to be trained and the actual offset matrix, which can be represented by, for example, L2 (mean squared error loss).
[0177] Step S602: Determine the total variation denoising value corresponding to each prediction offset matrix.
[0178] In the multimodal loss function, a total variation denoising function can be further set. This total variation denoising function can ensure the smoothness of the predicted offset matrix and avoid sudden changes in pixel offset.
[0179] Based on this, the TVLOSS (total variation denoising value) can be calculated on the predicted offset matrix using the total variation denoising function.
[0180] Step S603: Determine the image loss between each corrected prediction image and the corresponding actual corrected image based on the image loss sub-function.
[0181] Based on the image loss sub-function set in the multimodal loss function mentioned above, the image loss between the corrected prediction image and the corresponding actual corrected image can be calculated using this image loss sub-function.
[0182] Specifically, based on the text region loss function and the line region loss function set in the image loss subfunction, the text region loss and the line region loss can be determined respectively.
[0183] In the process of determining the text region loss using the text region loss function, the text edge mask can be used to determine the first text region in the corrected prediction image and the second text region in the actual corrected image.
[0184] Based on this, the difference between the first and second text regions can be determined using the text region loss function, which is the text region loss. The text region loss function can be expressed as: L11(x1,y1)).
[0185] In the text region loss function, L11 represents the text region loss calculated using the mean absolute error loss equation, x1 represents the region determined in the actual corrected image by the text edge mask, and y1 represents the region determined in the corrected prediction image by the text edge mask. The text region loss can also be called the high-frequency information reconstruction loss, which is used to represent the loss of high-frequency information in the reconstruction process.
[0186] In some examples, L11 can be calculated using formula (1) as shown below:
[0187]
[0188] Where n represents the actual number of corrected images. This represents the i-th x1. Let y1 represent the i-th y1.
[0189] In the embodiments of this application, the text portions in the correction prediction image and the actual correction image, as well as the edge regions of the text portions, can all be regarded as text regions.
[0190] Furthermore, in the process of determining the straight-line region loss using the straight-line region loss function, a straight-line mask can be used to determine the first straight-line region in the corrected prediction image and the second straight-line region in the actual corrected image.
[0191] Based on this, the difference between the first and second straight-line regions can be determined using the straight-line region loss function, which can be expressed as: L12(x2,y2).
[0192] In the linear region loss function, L12 represents the linear region loss calculated using the mean absolute error loss equation, x2 represents the region determined in the actual corrected image by the linear mask, and y2 represents the region determined in the corrected prediction image by the linear mask. The linear region loss can also be called the linear reconstruction loss, which is used to represent the loss of linear data during the reconstruction process.
[0193] In some examples, L12 can be calculated using formula (2) as shown below:
[0194]
[0195] in, This represents the j-th x2, Let y2 represent the j-th y2. After determining the text region loss and the line region loss, the image loss can be obtained by weighting the text region loss and the line region loss.
[0196] The image loss can be expressed as formula (3) as shown below:
[0197] L T =δ×L W +θ×L Z (3)
[0198] Among them, L T L represents image loss. W L represents the text region loss. Z The loss represents the linear region loss, δ represents the weight of the text region loss, which ranges from 0 to 1, and can be, for example, 0.1, 0.2 or 0.3; θ represents the weight of the linear region loss, which ranges from 0 to 1, and can be, for example, 0.9, 0.2 or 0.7. The sum of the weights of the text region loss and the linear region loss is 1.
[0199] In other cases, where there are images or graphics in the corrected prediction image, a graphics loss function can be constructed using the mean absolute error loss, and the difference between the image or graphics region in the corrected prediction image and the corresponding image or graphics region in the actual corrected image can be calculated and used as the image reconstruction loss.
[0200] Specifically, the graph loss function can be represented as: L13(x3,y3).
[0201] In the graph loss function, L13 represents the calculation of image reconstruction loss using the mean absolute error loss calculation equation, x3 represents the image or graph region in the actual corrected image, and y3 represents the corresponding image or graph region in the corrected prediction image.
[0202] In some examples, L13 can be calculated using formula (4) as shown below:
[0203]
[0204] in, This represents the k-th x3. This represents the k-th y3.
[0205] Based on this, the image loss can be obtained by weighting the image reconstruction loss, text region loss, and line region loss.
[0206] In this case, the image loss can be expressed as shown in formula (5) below.
[0207] L T =γ×L C +δ×L W +θ×L Z (5)
[0208] Among them, L C L represents the image reconstruction loss. W L represents the text region loss. Z The image reconstruction loss is represented by γ, which is the weight of the image reconstruction loss, ranging from 0 to 1, and can be, for example, 0.1, 0.2, or 0.3; the text region loss is represented by δ, which is the weight of the text region loss, ranging from 0 to 1, and can be, for example, 0.3, 0.4, or 0.5; the line region loss is represented by θ, which is the weight of the line region loss, ranging from 0 to 1, and can be, for example, 0.6, 0.4, or 0.2; the sum of the weights of the image reconstruction loss, the text region loss, and the line region loss is 1.
[0209] Step S604: For each corrected prediction image, the corresponding offset loss, the corresponding total variation denoising value, and the corresponding image loss are weighted to obtain the corresponding multimodal loss.
[0210] Based on the offset loss determined in S601, the total variation denoising value determined in S602, and the image loss determined in S603, the losses of multiple dimensions can be weighted to obtain a multimodal loss function representing multiple dimensions.
[0211] The multimodal loss function with multiple dimensions can be expressed as shown in the following formula (6):
[0212] Loss=α×L2+β×TVLoss+γ×L C +δ×L W +θ×L Z (6)
[0213] Where Loss represents the multimodal loss across multiple dimensions, α represents the weight of the offset loss, which ranges from 0 to 1, and can be, for example, 0.2, 0.3, or 0.4; β represents the weight of the total variation denoising value, which ranges from 0 to 1, and can be, for example, 0.3, 0.4, or 0.5. The sum of the weights of the offset loss, the total variation denoising value, the image reconstruction loss, the text region loss, and the line region loss is 1.
[0214] As described above, the hierarchical supervision significantly improved the training effect and correction accuracy of the correction model. Based on a joint optimization strategy of the offset loss function and the image loss function, multi-dimensional supervision of the offset matrix and image content was achieved. The offset loss ensures the geometric consistency between the predicted offset matrix and the actual distortion, while TVLoss effectively maintains the spatial smoothness of the offset matrix, avoiding unnatural abrupt changes in the corrected image. By further refining the image loss function into text region loss and line region loss, and by specifically strengthening the supervision of text and line regions, the correction model can maintain good performance in complex distortion correction tasks.
[0215] In the embodiments of this application, the depth-separable convolutional layer includes a depthwise convolutional layer and a pointwise convolutional layer.
[0216] Figure 7 This is a structural diagram of a depth-separable convolutional layer provided in an embodiment of this application.
[0217] like Figure 7 As shown, depth-separable convolutional layers include cascaded depthwise convolutional layers and pointwise convolutional layers.
[0218] based on Figure 7 The structure of the depth-separable convolutional layer is shown. Figure 8 A flowchart illustrating a depthwise separable convolution method provided in an embodiment of this application. Figure 8 As shown, the following steps are performed on the warped training images for each set of metadata:
[0219] Step S801: Perform spatial depth convolution on the distorted training image using a depth convolution layer to obtain spatial information.
[0220] After inputting each distorted training image into the correction model to be trained, the depthwise separable convolutional layer can first extract features from each distorted training image in both channel and spatial dimensions.
[0221] based on Figure 7 The structure shown can first perform spatial depth convolution on each twisted training image by a depth convolutional layer, and then perform spatial convolution on each channel of the twisted training image separately to obtain the corresponding spatial information.
[0222] Step S802: The spatial information of each channel is convolved and mixed by the pointwise convolutional layer to generate a channel spatial feature map.
[0223] Based on the spatial information of each distorted training image determined by the aforementioned S801, it can be input into a series of pointwise convolutional layers.
[0224] The kernel of the pointwise convolutional layer can be 1x1.
[0225] Based on this, pointwise convolutional layers can perform cross-channel fusion on each pixel in the distorted training image, thereby fusing the spatial information of each channel of each pixel to obtain the channel spatial feature map of the distorted training image.
[0226] Step S803: Perform regular convolution and dilated convolution on the predicted channel spatial feature map using pyramid pooling units to obtain a multi-scale feature map.
[0227] Step S804: The attention unit performs attention weighting on the multi-scale feature map to generate a predicted weighted feature map.
[0228] Step S805: Generate a prediction offset matrix based on the prediction weighted feature map.
[0229] As mentioned above, depthwise separable convolutional layers have a significantly reduced number of parameters compared to standard convolutions, enabling lightweight correction models.
[0230] In the embodiments of this application, the structure of the pyramid pooling unit can be, for example, a FasterASPP (Faster Atrous Spatial Pyramid Pooling) structure, which includes one or more parallel dilated convolutional layers, and each dilated convolutional layer can be further connected in parallel with a single conventional convolutional layer.
[0231] Figure 9 This is a structural diagram of a FasterASPP structure provided in an embodiment of this application.
[0232] like Figure 9 As shown, the FasterASPP structure includes a regular convolutional layer with a 3x3 kernel and two dilated convolutional layers connected in parallel with the regular convolutional layer. The kernel of each dilated convolutional layer is also the same 3x3, but the two layers have different dilation factors.
[0233] based on Figure 9 The FasterASPP structure shown is... Figure 10 This is a flowchart illustrating a multi-scale feature map determination method provided in an embodiment of this application. Figure 10 As shown, it includes the following steps:
[0234] Step S1001: Perform depthwise convolution and pointwise convolution on the distorted training image using a depthwise separable convolutional layer to obtain the predicted channel spatial feature map.
[0235] Step S1002: Perform regular convolution on the spatial feature map of each prediction channel by a regular convolutional layer to generate the corresponding first-scale feature map.
[0236] Combination Figure 9 As shown, the data output from the previous layer is simultaneously input into a regular convolutional layer and two dilated convolutional layers.
[0237] In this step, the channel spatial feature map determined in S1001 is as follows: Figure 9 The data output from the previous layer can be convolved using a regular convolutional layer to generate a first-scale feature map.
[0238] Among them, such as Figure 9 As shown, the receptive field of a regular convolutional layer when performing regular convolution is 3x3. The receptive field of the convolution kernel is 1.
[0239] It can be seen that the first-scale feature map represents the features collected when the convolution kernel is 3x3.
[0240] In some cases, such as Figure 9 As shown, after convolution in a regular convolutional layer, an activation function is set. This activation function is used to perform non-linear processing on the convolution result of the regular convolutional layer, thereby enhancing the convergence ability of the correction model. It can be, for example, the ReLU function.
[0241] Furthermore, the result after processing with the ReLU function can be used as the first-scale feature map.
[0242] Step S1003: Each dilated convolutional layer performs dilated convolution on the spatial feature map of each prediction channel with different receptive fields, thereby generating multiple second-scale feature maps with different receptive fields.
[0243] In this step, based on the channel spatial feature map determined in S1001, the channel spatial feature map can be dilated by each dilated convolutional layer, and a second-scale feature map can be generated by each dilated convolutional layer.
[0244] Among them, such as Figure 9 As shown, the two dilated convolutional layers, each with a 3x3 kernel, have different dilation factors: one dilated convolutional layer has a dilation factor of 2, and the other dilated convolutional layer has a dilation factor of 4.
[0245] Based on this, when performing dilated convolution, dilated convolution with an inflation factor of 2 will be based on a 3x3 convolution kernel, with dilated rows and columns inserted between the rows and columns of its acquisition window. The number of dilated rows and columns corresponds to the inflation factor, thereby expanding the area actually covered by the acquisition window, i.e., the receptive field, to be equivalent to the receptive field corresponding to a 5x5 convolution kernel, achieving the effect of expanding the receptive field while keeping the convolution kernel unchanged.
[0246] Meanwhile, the dilated convolution with a dilation factor of 4 is also based on a 3x3 convolution kernel. Dilated rows and columns are inserted between the rows and columns of its acquisition window. The number of dilated rows and columns corresponds to the dilation factor, thereby expanding the area actually covered by the acquisition window, i.e., the receptive field, to be equivalent to the receptive field corresponding to a 9x9 convolution kernel. This achieves the effect of expanding the receptive field while keeping the convolution kernel unchanged.
[0247] It can be seen that each second-scale feature map represents the features acquired under the corresponding receptive field conditions.
[0248] In some cases, such as Figure 9 As shown, after each dilated convolutional layer is convolved, a corresponding activation function is set for each dilated convolutional layer. This activation function is used to perform nonlinear processing on the convolution result of the corresponding dilated convolutional layer, thereby enhancing the convergence ability of the correction model. It can be, for example, the ReLU function.
[0249] Furthermore, the result of processing each corresponding ReLU function can be used as a second-scale feature map.
[0250] Step S1004: Concatenate the first-scale feature map and the corresponding second-scale feature maps of each predicted channel spatial feature map to obtain the corresponding multi-scale feature map.
[0251] Based on the first-scale feature map determined in S1002 and the various second-scale feature maps determined in S1003, they can be stitched together to obtain a multi-scale feature map.
[0252] In this step, such as Figure 9 As shown, after the regular convolution and each dilated convolution, a concatenation layer is set in series. Since the first-scale feature map and each second-scale feature map focus on different features, the first-scale feature map and each second-scale feature map can be simultaneously input into the concatenation layer. The concatenation layer is used to concatenate the first-scale feature map and each second-scale feature map to obtain a multi-scale feature map.
[0253] In some cases, such as Figure 9As shown, a convolutional layer with a 1x1 kernel can be set after the stitching layer to compress the high-dimensional multi-scale features, thereby reducing the amount of computation and performing non-linear processing between channels. This can enhance the representation of key information and reduce noise generated during the stitching process.
[0254] Step S1005: The attention unit performs attention weighting on the multi-scale feature map to generate a predicted weighted feature map.
[0255] Step S1006: Generate a prediction offset matrix based on the prediction weighted feature map.
[0256] As described above, the multi-scale feature fusion mechanism significantly improves the correction model's ability to represent distortions. Based on an architecture design that combines conventional convolutions and dilated convolutions with varying dilation rates, the conventional convolutional layers focus on capturing detailed features within the 3×3 local receptive field, while the dilated convolutional layers with dilation factors of 2 and 4 extract mid-to-long-range geometric features with equivalent receptive fields of 5×5 and 9×9, respectively. This achieves full-scale feature coverage from local detailed distortions to macroscopic global distortions. By organically fusing feature maps of different scales through a stitching layer and performing channel compression and feature reorganization, the correction model can significantly improve the geometric accuracy of offset field prediction while maintaining lightweight computation.
[0257] Figure 11 This is a framework diagram of another correction model training method provided in an embodiment of this application. Figure 11 As shown, the correction model to be trained includes not only the aforementioned depthwise separable convolutional layers, pyramid pooling units, and attention units, but also multiple convolutional layers with 3x3 kernels.
[0258] After inputting the correction training images into the correction model to be trained, the correction training images are first convolved using a 3x3 convolutional layer. The convolution result is then input into a depthwise separable convolutional layer for processing. The processing result is then input into a pyramid pooling unit of a FasterASPP structure for further processing. Finally, the processing result is input into an attention unit for weighted summation to obtain a weighted feature map.
[0259] like Figure 11 As shown, the depthwise separable convolutional layer, the pyramid pooling unit of the FasterASPP structure, and the attention unit are connected in series. After the attention unit obtains the weighted feature map, if this attention weighting is the first attention weighting process, the weighted feature map can be input into the depthwise separable convolutional layer again. That is, the processing of the depthwise separable convolutional layer, the pyramid pooling unit of the FasterASPP structure, and the SAM unit is performed twice, so as to obtain a more accurate weighted feature map.
[0260] Furthermore, if this attention weighting is the second attention weighting process, the weighted feature map obtained after processing the depth-separable convolutional layer, the pyramid pooling unit of the FasterASPP structure, and the SAM unit twice can be input into another convolutional layer with a 3x3 kernel. The output of this convolutional layer with a 3x3 kernel is concatenated with a third convolutional layer with a 3x3 kernel. After the convolution operation of two convolutional layers with a 3x3 kernel, the prediction offset matrix can be obtained.
[0261] like Figure 11 As shown, the corrected prediction image is obtained by inputting the predicted offset matrix into the pre-set remap function to map the corrected training image.
[0262] Furthermore, the difference between the predicted offset matrix and the actual offset matrix can be determined using the offset loss function, that is... Figure 11 The offset loss in the image; using the text region loss function, the difference between the text regions in the predicted image and the actual text regions in the corrected image can be determined, that is... Figure 11 The text region loss function is used to determine the difference between the line region in the predicted image and the actual line region in the corrected image. Figure 11 The loss in the straight-line region.
[0263] Furthermore, by weighting the offset loss function, the text region loss function, and the line region loss function, a multimodal loss function can be used between the predicted image and the actual corrected image, and the multimodal loss can be determined.
[0264] Furthermore, the parameters of the correction model to be trained can be adjusted based on this multimodal loss.
[0265] Based on the same inventive concept, and corresponding to the methods of any of the above embodiments, the embodiments of this application also provide an image correction device.
[0266] Figure 12 This is a structural block diagram of an image correction device provided in an embodiment of this application. The device is configured to execute the image correction method provided in the above embodiment, and has corresponding functional modules and beneficial effects for executing the method. For example... Figure 12 As shown, the device includes: an acquisition module 1201 and a prediction module 1202;
[0267] Module 1201 is configured to acquire the image to be corrected.
[0268] The prediction module 1202 is configured to input the image to be corrected into the correction model, which processes the image to obtain the corrected image. The correction model includes at least one dilated convolutional layer, with different dilation factors set for different dilation factors. The correction model trained with multi-dilation factor dilated convolutional layers and multimodal supervision achieves high-precision document image correction. Dilated convolutional layers with different dilation factors can effectively capture multi-scale features from local text edges to global page curvature, ensuring the accuracy of correction. Combined with multimodal supervision, the correction model is trained by jointly training multiple dimensions, resulting in a more accurate and stable correction model.
[0269] Accordingly, the prediction module 1202 is also specifically configured to perform offset processing on the image to be corrected through the correction model to obtain the offset matrix of the image to be corrected;
[0270] The offset matrix is used to map each pixel in the image to be corrected to obtain the corrected image.
[0271] Furthermore, the prediction module 1202 is specifically configured to train the correction model in the following manner:
[0272] Obtain the training dataset, which includes multiple sets of metadata. Each set of metadata includes multiple distorted training images and their corresponding actual corrected images.
[0273] Input the metadata of a predetermined number of groups into the correction model to be trained, and output the prediction offset matrix corresponding to each distorted training image;
[0274] The corresponding corrected prediction image is obtained by mapping each pixel in each distorted training image according to the predicted offset matrix.
[0275] The multimodal loss value corresponding to each corrected prediction image is determined based on the multimodal loss function. Each multimodal loss value represents the difference between the corresponding corrected prediction image and the corresponding actual corrected image in multiple dimensions.
[0276] The correction model to be trained is trained based on the multimodal loss value to obtain the trained correction model.
[0277] The process of training the correction model to be trained based on the multimodal loss value to obtain the trained correction model includes:
[0278] If the multimodal loss value is greater than the preset loss value threshold and / or the training dataset has not completed a predetermined number of training iterations, the parameters in the correction model to be trained are adjusted, and different metadata of the same number of groups are input into the adjusted correction model to generate correction prediction images corresponding to the distorted training images of each group of metadata.
[0279] If the multimodal loss value is less than or equal to the loss value threshold and / or the training dataset has completed a predetermined number of training iterations, the current correction model to be trained is determined as the completed correction model.
[0280] The correction model to be trained includes sequentially connected deep separable convolutional layers, pyramid pooling units, and attention units.
[0281] Input a predetermined number of metadata sets into the correction model to be trained, and output the prediction offset matrix corresponding to the distorted training image for each metadata set, including:
[0282] Perform the following processing on the warped training images for each set of metadata:
[0283] The depthwise separable convolutional layer performs depthwise convolution and pointwise convolution on the distorted training image to obtain a predicted channel spatial feature map.
[0284] The pyramid pooling unit performs regular convolution and dilated convolution on the predicted channel spatial feature map to obtain a multi-scale feature map.
[0285] The attention unit performs attention weighting on the multi-scale feature map to generate a predicted weighted feature map.
[0286] A prediction offset matrix is generated based on the predicted weighted feature map.
[0287] Accordingly, each set of metadata also includes the actual offset matrix corresponding to the warped training image; the multimodal loss function includes an offset loss subfunction and an image loss subfunction;
[0288] The multimodal loss value corresponding to each corrected and predicted image is determined based on the multimodal loss function, including:
[0289] Each predicted offset matrix and its corresponding actual offset matrix are input into the offset loss sub-function to obtain the corresponding offset loss;
[0290] Determine the total variational denoising value corresponding to each prediction offset matrix;
[0291] The image loss between each corrected prediction image and the corresponding actual corrected image is determined based on the image loss sub-function;
[0292] For each corrected prediction image, the corresponding offset loss, the corresponding total variation denoising value, and the corresponding image loss are weighted to obtain the corresponding multimodal loss.
[0293] Each distorted training image includes a first text region and a first straight line region; the corresponding actual corrected image includes a second text region corresponding to the first text region and a second straight line region corresponding to the first straight line region; the image loss sub-function includes a text region loss function and a straight line region loss function.
[0294] The image loss between each corrected predicted image and its corresponding actual corrected image is determined based on the image loss sub-function, including:
[0295] The text region loss between each first text region and its corresponding second text region is determined based on the text region loss function.
[0296] The straight-line region loss between each first straight-line region and the corresponding second straight-line region is determined based on the straight-line region loss function.
[0297] The image loss is obtained by weighting the loss of each text region and the corresponding line region.
[0298] Accordingly, each set of metadata also includes a text edge mask corresponding to each second text region and a line mask corresponding to each second line region;
[0299] Before determining the image loss between each corrected predicted image and the corresponding actual corrected image based on the image loss subfunction, the method also includes:
[0300] The first text region in the corresponding corrected prediction image is determined based on the edge mask of each text.
[0301] The first straight-line region in the corresponding corrected prediction image is determined based on the straight-line region loss function.
[0302] Accordingly, depthwise separable convolutional layers include depthwise convolutional layers and pointwise convolutional layers;
[0303] The depthwise separable convolutional layer performs depthwise convolution and pointwise convolution on the distorted training image to obtain a predicted channel space feature map, including:
[0304] Perform the following processing on the warped training images for each set of metadata:
[0305] Spatial information is obtained by performing spatial depth convolution on the distorted training image using the deep convolutional layer;
[0306] The pointwise convolutional layer performs pointwise convolution on the spatial information of each channel and mixes them to generate a channel spatial feature map.
[0307] The pyramid pooling unit includes a single regular convolutional layer and at least one dilated convolutional layer. The regular convolutional layer and each dilated convolutional layer have the same convolutional kernel and are connected in parallel. Each dilated convolutional layer is set with a different dilation factor, and each dilation factor is used to control the receptive field of the corresponding dilated convolutional layer.
[0308] Furthermore, the pyramid pooling unit performs regular convolution and dilated convolution on the corresponding predicted channel spatial feature map to obtain the corresponding multi-scale feature map, including:
[0309] Each prediction channel spatial feature map is convolved by a regular convolutional layer to generate the corresponding first-scale feature map.
[0310] Each dilated convolutional layer performs dilated convolution on the spatial feature map of each prediction channel with different receptive fields, thereby generating multiple second-scale feature maps with different receptive fields.
[0311] The first-scale feature map and the corresponding second-scale feature maps of each predicted channel spatial feature map are concatenated to obtain the corresponding multi-scale feature map.
[0312] Accordingly, the mapping module 1203 is further configured such that the offset matrix represents the positional difference between each pixel in the image to be corrected and the corresponding pixel in the preset distortion-free image;
[0313] The step of mapping each pixel in the image to be corrected according to the offset matrix to obtain the corrected image includes:
[0314] Based on the positional differences of each pixel in the offset matrix, the positions of each pixel in the image to be corrected are adjusted to obtain a corrected image with the same pixel positions as the undistorted image.
[0315] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.
[0316] The apparatus of the above embodiments is used to implement the corresponding image correction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0317] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the embodiments of this application also provide an image correction device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the image correction method of any of the above embodiments.
[0318] Figure 13This is a schematic diagram of the structure of an image correction device provided in an embodiment of this application, as shown below. Figure 13 As shown, the device includes a processor 1301, a memory 1302, an input device 1303, and an output device 1304; the number of processors 1301 in the device can be one or more. Figure 13 Taking a processor 1301 as an example; the processor 1301, memory 1302, input device 1303, and output device 1304 in the device can be connected via a bus or other means. Figure 13 Taking a bus connection as an example, the memory 1302, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as program instructions / modules for implementing the image correction method in the embodiments of this application. The processor 1301 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 1302, thereby implementing the aforementioned image correction method. The input device 1303 can be configured to receive input digital or character information and generate key signal inputs related to user settings and function control of the device. The output device 1304 may include a display screen or other display device.
[0319] The apparatus of the above embodiments is used to implement the corresponding image correction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0320] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-volatile storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are configured to perform an image correction method described in the above embodiments. This method includes: a correction model trained using dilated convolutional layers with multiple dilation factors and multimodal supervision, achieving high-precision document image correction. Dilated convolutional layers with different dilation factors can effectively capture multi-scale features from local text edges to global page curvature, ensuring the accuracy of the correction. Combined with multimodal supervision, the correction model is trained by jointly training multiple dimensions, thereby making the trained correction model more accurate and stable.
[0321] It is worth noting that in the above-described embodiments of the image correction device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not configured to limit the protection scope of the embodiments of this application.
[0322] In some possible implementations, various aspects of the methods provided in this application can also be implemented as a program product, which includes program code. When the program product is run on a computer device, the program code is configured to cause the computer device to perform the steps of the methods according to the various exemplary embodiments of this application described above. For example, the computer device can perform the image correction method described in the embodiments of this application. The program product can be implemented using any combination of one or more readable media, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated further here.
Claims
1. An image correction method, characterized in that, include: Obtain the image to be corrected; The image to be corrected is input into the correction model, and the correction model processes the image to obtain the corrected image. The correction model includes at least one dilated convolutional layer, and different dilated convolutional layers are set with different dilation factors.
2. The image correction method according to claim 1, characterized in that, The image to be corrected is processed using the correction model to obtain a corrected image, including: The image to be corrected is offset using the correction model to obtain the offset matrix of the image to be corrected. The offset matrix is used to map each pixel in the image to be corrected to obtain the corrected image.
3. The image correction method according to claim 1, characterized in that, The correction model was trained in the following manner: Obtain the training dataset, which includes multiple sets of metadata, each set of metadata including multiple distorted training images and their corresponding actual corrected images; Input the metadata of a predetermined number of groups into the correction model to be trained, and output the prediction offset matrix corresponding to each distorted training image; The corresponding corrected prediction image is obtained by mapping each pixel in each distorted training image according to the predicted offset matrix. The multimodal loss value corresponding to each corrected prediction image is determined based on the multimodal loss function. Each multimodal loss value represents the difference between the corresponding corrected prediction image and the corresponding actual corrected image in multiple dimensions. The correction model to be trained is trained based on the multimodal loss value to obtain the trained correction model.
4. The image correction method according to claim 3, characterized in that, The step of training the correction model to be trained based on the multimodal loss value to obtain the trained correction model includes: If the multimodal loss value is greater than the preset loss value threshold and / or the training dataset has not completed a predetermined number of training iterations, the parameters in the correction model to be trained are adjusted, and different metadata of the same number of groups are input into the adjusted correction model to generate correction prediction images corresponding to the distorted training images of each group of metadata. If the multimodal loss value is less than or equal to the loss value threshold and / or the training dataset has completed a predetermined number of training iterations, the current correction model to be trained is determined as the completed correction model.
5. The image correction method according to claim 3, characterized in that, The correction model to be trained includes sequentially connected deep separable convolutional layers, pyramid pooling units, and attention units; The step of inputting a predetermined number of metadata sets into the correction model to be trained and outputting a prediction offset matrix corresponding to the distorted training image for each set of metadata sets includes: Perform the following processing on the warped training images for each set of metadata: The depthwise separable convolutional layer performs depthwise convolution and pointwise convolution on the distorted training image to obtain a predicted channel spatial feature map. The pyramid pooling unit performs regular convolution and dilated convolution on the predicted channel spatial feature map to obtain a multi-scale feature map. The attention unit performs attention weighting on the multi-scale feature map to generate a predicted weighted feature map. A prediction offset matrix is generated based on the predicted weighted feature map.
6. The image correction method according to claim 3, characterized in that, Each set of metadata also includes the actual offset matrix corresponding to the warped training image; the multimodal loss function includes an offset loss sub-function and an image loss sub-function; The step of determining the multimodal loss value corresponding to each corrected prediction image based on the multimodal loss function includes: Each predicted offset matrix and its corresponding actual offset matrix are input into the offset loss sub-function to obtain the corresponding offset loss. Determine the total variational denoising value corresponding to each prediction offset matrix; The image loss between each corrected prediction image and the corresponding actual corrected image is determined based on the image loss sub-function. For each corrected prediction image, the corresponding offset loss, the corresponding total variation denoising value, and the corresponding image loss are weighted to obtain the corresponding multimodal loss.
7. The image correction method according to claim 6, characterized in that, Each distortion training image includes a first text region and a first straight line region; the corresponding actual correction image includes a second text region corresponding to the first text region and a second straight line region corresponding to the first straight line region; the image loss sub-function includes a text region loss function and a straight line region loss function; The step of determining the image loss between each corrected predicted image and the corresponding actual corrected image based on the image loss sub-function includes: The text region loss between each first text region and its corresponding second text region is determined based on the text region loss function. Based on each straight line mask, determine the first straight line region in the corresponding correction prediction image and the second straight line region in the corresponding actual correction image; The image loss is obtained by weighting the loss of each text region and the corresponding line region.
8. The image correction method according to claim 7, characterized in that, Each set of metadata also includes a text edge mask corresponding to each second text region and a line mask corresponding to each second line region; Before determining the image loss between each corrected predicted image and the corresponding actual corrected image based on the image loss subfunction, the method further includes: The first text region in the corresponding corrected prediction image is determined based on the edge mask of each text. The first straight-line region in the corresponding corrected prediction image is determined based on the straight-line region loss function.
9. The image correction method according to claim 5, characterized in that, The depth-separable convolutional layer includes a depthwise convolutional layer and a pointwise convolutional layer; The step of performing depthwise convolution and pointwise convolution on the distorted training image by the depthwise separable convolutional layer to obtain the predicted channel space feature map includes: Perform the following processing on the warped training images for each set of metadata: Spatial information is obtained by performing spatial depth convolution on the distorted training image using the deep convolutional layer; The pointwise convolutional layer performs pointwise convolution on the spatial information of each channel and mixes them to generate a channel spatial feature map.
10. The image correction method according to claim 5, characterized in that, The pyramid pooling unit includes a single regular convolutional layer and at least one dilated convolutional layer. The regular convolutional layer and each dilated convolutional layer have the same convolutional kernel and are connected in parallel. Each dilated convolutional layer is set with a different dilation factor, and each dilation factor is used to control the receptive field of the corresponding dilated convolutional layer. The process of performing regular convolution and dilated convolution on the corresponding predicted channel spatial feature map by the pyramid pooling unit to obtain the corresponding multi-scale feature map includes: The conventional convolutional layer performs conventional convolution on the spatial feature map of each prediction channel to generate the corresponding first-scale feature map; Each dilated convolutional layer performs dilated convolution on the spatial feature map of each prediction channel with different receptive fields, thereby generating multiple second-scale feature maps with different receptive fields. The first-scale feature map and the corresponding second-scale feature maps of each predicted channel spatial feature map are concatenated to obtain the corresponding multi-scale feature map.
11. The image correction method according to claim 2, characterized in that, The offset matrix represents the positional difference between each pixel in the image to be corrected and the corresponding pixel in the preset distortion-free image; The step of mapping each pixel in the image to be corrected according to the offset matrix to obtain the corrected image includes: Based on the positional differences of each pixel in the offset matrix, the positions of each pixel in the image to be corrected are adjusted to obtain a corrected image with the same pixel positions as the undistorted image.
12. An image correction device, characterized in that, include: The module consists of an acquisition module, a prediction module, and a mapping module. The acquisition module is configured to acquire the image to be corrected; The prediction module is configured to input the image to be corrected into the correction model, and process the image to be corrected through the correction model to obtain a corrected image. The correction model includes at least one dilated convolutional layer, and different dilated convolutional layers are set with different dilation factors.
Citation Information
Cited By
Frequency domain and scale common sensing network applied to pancreatic tumor segmentation
CN121190771A
Image correction method and related device
CN122156021A