Image Correction Method and Device, Computer Storage Medium, and Electronic Device

By augmenting the initial sample set and training the machine learning model, the automated tilt detection and correction of images are achieved, and the problems of high cost and low efficiency of manual review in the existing technology are solved, the operation process efficiency of merchants entering the platform is improved and the correctness of text recognition is ensured.

CN113780330BActive Publication Date: 2025-06-17BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110395349.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-13
Publication Date
2025-06-17
Estimated Expiration
2041-04-13

AI Technical Summary

Technical Problem

The prior art cannot automatically correct tilted images, resulting in high cost and low efficiency of manual review, affecting the merchant's entry process.

Method used

By augmenting the initial sample set, a preset machine learning model is trained to obtain a classification prediction model, which is used to automatically detect the tilt angle of the image and correct the tilt image.

Benefits of technology

Automatic image tilt detection and correction is realized, the cost of manual review is reduced, the operation process efficiency of merchants entering the platform is improved, and the accuracy of subsequent text recognition is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113780330B_ABST
    Figure CN113780330B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of image processing technology, and provides an image correction method, an image correction device, a computer storage medium, and an electronic device. Among them, the image correction method includes: performing data augmentation on an initial sample set to obtain an augmented sample set; training a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is used to perform tilt detection on a file image to be processed; performing tilt detection on the file image to be processed according to the classification prediction model to obtain the tilt angle range of the file image to be processed; if the tilt angle range is inconsistent with the target angle range, tilt correction is performed on the file image to be processed. The present disclosure can automatically detect the tilt angle of an image and perform tilt correction, solves the technical problem in the prior art that manual review and repeated re-uploading of file images are required, simplifies the image uploading process, and improves the image review efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the continuous development of multimedia technology, multimedia devices such as digital cameras and high-definition camera phones have occupied an increasingly important position in people's lives. By adopting image processing technology, information such as text and pictures collected by digital devices can be converted into other information forms for output, for example, converted into audio output to meet the vision needs of visually impaired patients. However, due to input devices or certain other factors, the text images collected will inevitably be tilted to some extent more or less. Therefore, skew image correction is a very important topic in the current field of text image research.

[0003] Currently, generally, pictures uploaded by merchants are manually reviewed. However, manual review requires a large amount of manpower and material resources, with a high cost, and cannot meet the time requirements of business needs. Also, if it is found during the review process that the image is tilted, the merchant needs to re-upload it. Therefore, this process takes a lot of time and seriously affects the merchant's settlement process.

[0004] In view of this, there is an urgent need in this field to develop a new image correction method and device.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Summary of the Invention

[0006] The purpose of the present disclosure is to provide an image correction method, an image correction device, a computer storage medium, and an electronic device, thereby at least to a certain extent overcoming the defect in the prior art that skew images cannot be automatically corrected.

[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.

[0008] According to a first aspect of the present disclosure, there is provided an image correction method, including: performing data augmentation on an initial sample set to obtain an augmented sample set; training a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is used to perform skew detection on a file image to be processed; performing skew detection on the file image to be processed according to the classification prediction model to obtain a skew angle range of the file image to be processed; if the skew angle range is inconsistent with a target angle range, performing skew correction on the file image to be processed.

[0009] In an exemplary embodiment of the present disclosure, before data augmentation is performed on the initial sample set to obtain an augmented sample set, the method further includes: collecting an original document image and obtaining a label corresponding to the original document image; performing a rotation process on the original document image to obtain a rotated image; and determining the initial sample set according to the original document image and its corresponding label, and the rotated image and its corresponding label.

[0010] In an exemplary embodiment of the present disclosure, the performing a rotation process on the original document image to obtain a rotated image includes: dividing a circumference into N angular ranges at a preset interval; N is an integer greater than 1; randomly selecting an angular value from each angular range; and performing a rotation process on the original document image according to the N selected angular values to obtain N rotated images.

[0011] In an exemplary embodiment of the present disclosure, the label corresponding to the rotated image is determined by the following method: determining the label corresponding to the rotated image according to the angular range to which the angular value belongs.

[0012] In an exemplary embodiment of the present disclosure, after obtaining N rotated images, the method further includes: detecting whether the rotated image exceeds an image border; if so, performing a size correction on the image border according to the size of the original document image and the selected angular value.

[0013] In an exemplary embodiment of the present disclosure, after performing a size correction on the image border, the method further includes: filling a color in a blank area in the image border; cropping the rotated image from the image after the color filling; randomly selecting a background image from a pre-stored background image set and adding random noise to the background image; and pasting the rotated image onto the background image after adding the random noise.

[0014] In an exemplary embodiment of the present disclosure, the data augmentation of the initial sample set to obtain an augmented sample set includes: using the N rotated images corresponding to each original document image as a basic image set, randomly shuffling the image numbers in the basic image set to obtain a target image set; randomly selecting a first image from the basic image set and cropping a first sub-image from the first image; randomly selecting a second image from the target image set and cropping a second sub-image from the second image; the second image having the same number as the first image; performing image mixing on the first sub-image and the second sub-image to obtain a mixed image; determining the label of the mixed image according to the label of the first sub-image and the label of the second sub-image; and obtaining the augmented sample set according to the mixed image and its corresponding label.

[0015] In an exemplary embodiment of the present disclosure, the image mixing of the first sub-image and the second sub-image includes: randomly sampling from a preset numerical range based on a beta distribution to obtain a sampling value; and performing image mixing on the first sub-image and the second sub-image based on the sampling value.

[0016] In an exemplary embodiment of the present disclosure, if the tilt angle range is inconsistent with the target angle range, the tilt correction of the to-be-processed document image includes: obtaining boundary values of the tilt angle range; the boundary values include an upper limit value and a lower limit value; and performing tilt correction on the to-be-processed document image according to an average value of the upper limit value and the lower limit value.

[0017] In an exemplary embodiment of the present disclosure, if the tilt angle range is inconsistent with the target angle range, the tilt correction of the to-be-processed document image further includes: converting the to-be-processed document image into a grayscale image; performing Gaussian blur processing on the grayscale image to obtain a blurred image; performing edge detection on the blurred image to obtain an edge image; performing line detection on the edge image by using a Hough transform method based on polar coordinate space transformation to obtain characteristic lines; obtaining included angle values between each of the characteristic lines and a horizontal line, and selecting target lines whose included angle values are within the tilt angle range; and performing tilt correction on the to-be-processed document image according to an average value of the included angle values corresponding to the target lines.

[0018] According to a second aspect of the present disclosure, there is provided an image correction device, including: a data enhancement module configured to perform data enhancement on an initial sample set to obtain an augmented sample set; a model training module configured to train a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is configured to perform tilt detection on a to-be-processed document image; the classification prediction model is configured to obtain a tilt angle range of the to-be-processed document image according to the classification prediction model; and a tilt correction module configured to perform tilt correction on the to-be-processed document image if the tilt angle range is inconsistent with a target angle range.

[0019] According to a third aspect of the present disclosure, there is provided a computer storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image correction method described in the first aspect is implemented.

[0020] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the image correction method described in the first aspect by executing the executable instructions.

[0021] As can be seen from the above technical solutions, the image correction method, image correction device, computer storage medium, and electronic device in the exemplary embodiments of the present disclosure at least have the following advantages and positive effects:

[0022] In the technical solutions provided in some embodiments of the present disclosure, on the one hand, data augmentation is performed on the initial sample set to obtain an augmented sample set, which can solve the technical problem that the amount of data is small because the sample images are in a confidential state and cannot be obtained through the Internet, enrich the number of images in the training set, and thus ensure a sufficient proportion of samples to guarantee the accuracy of the subsequent model. Further, a classification prediction model is obtained by training a preset machine learning model according to the augmented sample set, and the tilt angle range of the file image to be processed is obtained by detecting the tilt of the file image to be processed according to the classification prediction model, which can automatically detect the tilt angle of the uploaded image and solve the technical problem of high human and material costs caused by manual review in the prior art, reducing the review cost. On the other hand, if the tilt angle range is inconsistent with the target angle range, tilt correction is performed on the file image to be processed, which can simplify the operation process when merchants enter the platform, avoid the technical problems of low efficiency and affecting the entry process caused by the need for merchants to upload file images multiple times when the review fails due to file tilt, reduce human and material costs, and can also solve the technical problems of unrecognizability and affecting the review progress caused by excessive file tilt, improve the file review efficiency, and ensure the accuracy of subsequent text recognition.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0025] Figure 1 A flowchart showing the image correction method in this exemplary embodiment;

[0026] Figure 2 A flowchart showing the determination of the initial sample set in this exemplary embodiment;

[0027] Figures 3A - 3B A schematic diagram showing the original file image collected in this exemplary embodiment;

[0028] Figure 4Schematic diagram showing the process of rotating the original document image to obtain a rotated image in this exemplary embodiment;

[0029] Figure 5 Schematic diagram showing the size correction of the image border according to the size of the original document image and the selected angle value in this exemplary embodiment;

[0030] Figure 6 Schematic diagram showing the process of processing the image after size correction in this exemplary embodiment;

[0031] Figure 7 Schematic diagram showing the image after color filling in this exemplary embodiment;

[0032] Figure 8 Schematic diagram showing the resulting image obtained by pasting the rotated image onto the denoised background image in this exemplary embodiment;

[0033] Figure 9 Schematic diagram showing the process of data augmentation on the initial sample set to obtain an augmented sample set in this exemplary embodiment;

[0034] Figure 10 Schematic diagram showing the mixed image in this exemplary embodiment;

[0035] Figure 11A Schematic diagram showing the architecture of the ResneSt model in this exemplary embodiment;

[0036] Figure 11B Schematic diagram showing the architecture of the Split-Attention module of the ResneSt model in this exemplary embodiment;

[0037] Figure 12 Schematic diagram showing the process of skew correction for the document image to be processed if the skew angle range is inconsistent with the target angle range in this exemplary embodiment;

[0038] Figure 13 Schematic diagram showing the overall process of the image correction method in this exemplary embodiment;

[0039] Figure 14 Schematic diagram showing the structure of the image correction device in the exemplary embodiment of the present disclosure;

[0040] Figure 15 Schematic diagram showing the structure of the electronic device in the exemplary embodiment of the present disclosure. Detailed implementation manners

[0041] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0042] As used in this specification, the terms "a", "an", "the", and "said" are used to denote the presence of one or more elements / components / etc.; the terms "comprising" and "having" are used to mean an open inclusion and mean that there may be additional elements / components / etc. in addition to the listed elements / components / etc.; the terms "first", "second", etc. are used only as labels and are not a limitation on the quantity of their objects.

[0043] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0044] Nowadays, online shopping has become an essential way of life. This is inseparable from successful e-commerce platforms and a large number of suppliers. An e-commerce platform is a platform with a very large number of merchants, users, products, and transaction volumes. For users to find a large number of good products, a large number of merchants need to settle on the platform, which requires rapid review of the document images uploaded by merchants during settlement. However, since many merchants are not familiar with computer and shooting operations, or due to operational errors, the uploaded document photos are excessively tilted, upside down, sideways, etc., and the uploaded document images cannot be recognized, seriously affecting the merchant settlement process.

[0045] Currently, the following difficulties exist in identifying the images uploaded by merchants:

[0046] First, there is a lack of data. Real merchant document images have a certain degree of confidentiality and cannot be obtained through the Internet, so the amount of data obtained is usually small. At the same time, the number of file types to be detected is very large, currently 31 types, and the required amount of data is large. This constitutes a contradiction and increases the difficulty of recognition;

[0047] Second, there is a wide variety of document types, significant differences in document content, and large variations in document backgrounds. The document types include more than 30 types such as business licenses, authorization letters, power of attorney, product lists, etc. There are also significant differences in the sizes of the images. In some images, the part with content takes up only a very small portion of the image. Additionally, when merchants take pictures of document or certificate images, the noise and impacts caused by operations are particularly large, such as reflections, color differences, too dim light, etc. Some merchants' captured document images only have the text part in the middle. These influencing factors greatly increase the difficulty of the project;

[0048] Third, there is no labeled data, and the differences in the skew data are very large. The skew angle of the image can be any angle from 0 to 360 degrees, and such data is lacking in reality. Additionally, when merchants take pictures of images, there are infinitely many image backgrounds, which can easily cause learning biases in the model. Moreover, the data with such noise is very scarce and not labeled either.

[0049] Fourth, the content of the document image is complex, and traditional methods such as the Hough transform and text recognition methods cannot be used. In actual work, for some similar scenarios, some methods of image processing can be used to identify the skew angle of the main body in the image. For example, the Hough transform is used to detect the straight lines in the image, and then the skew angle of the document is judged based on these straight lines. However, the aspect ratio of the image varies within a very large range. Some aspect ratios are 2:1, or even 3:1, some are 1:3. For some photos, there are no straight lines in the image due to shooting reasons, and the edges of the document cannot be seen. Even if there are borders and straight lines in the image, the lengths of the broken lines and straight lines are different, with vertical and horizontal ones. In some images, the horizontal straight lines are longer, and in some images, the vertical straight lines are longer. Some straight lines are in tables, and some are for filling in blanks in the text. Therefore, traditional methods cannot identify the skew angle of the document image. Additionally, although the deep recognition method can find the text boxes where the text is located, for inverted text, it cannot determine whether the image is in the correct position, and it has a large computational amount, slow calculation speed, and requires a large amount of labeled data. Moreover, for the deep recognition algorithm, all the text parts in the image need to be labeled, which requires a very large amount of work. Therefore, the deep recognition method also cannot identify the skew angle of the document image.

[0050] In the embodiments of the present disclosure, first, an image correction method is provided, which at least to some extent overcomes the defect in the prior art that the skew document image cannot be automatically corrected.

[0051] Figure 1 The flowchart of the image correction method in this exemplary embodiment is shown. The execution subject of this image correction method can be a server that corrects the image.

[0052] Refer to Figure 1, an image correction method according to an embodiment of the present disclosure includes the following steps:

[0053] Step S110, performing data augmentation on an initial sample set to obtain an augmented sample set;

[0054] Step S120, training a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is used to perform skew detection on a file image to be processed;

[0055] Step S130, performing skew detection on the file image to be processed according to the classification prediction model to obtain a skew angle range of the file image to be processed;

[0056] Step S140, if the skew angle range is inconsistent with the target angle range, performing skew correction on the file image to be processed.

[0057] In Figure 1 In the technical solution provided by the embodiment shown, on the one hand, performing data augmentation on the initial sample set to obtain an augmented sample set can solve the technical problem that the data volume is small due to the sample images being in a confidential state and unable to be obtained through the Internet, enrich the number of images in the training set, and thus ensure a sufficient proportion of samples and the accuracy of the subsequent model. Further, training a preset machine learning model according to the augmented sample set to obtain a classification prediction model, and performing skew detection on the file image to be processed according to the classification prediction model to obtain the skew angle range of the file image to be processed can automatically detect the skew angle of the uploaded image and solve the technical problem of high labor and material costs caused by manual review in the prior art, reducing the review cost. On the other hand, if the skew angle range is inconsistent with the target angle range, performing skew correction on the file image to be processed can simplify the operation process when merchants enter the platform, avoid the technical problems of low efficiency and affecting the entry process caused by merchants having to upload file images multiple times when the review fails due to file skew, reduce labor and material costs, and also solve the technical problems of unrecognizability and affecting the review progress caused by excessive file skew, improve the file review efficiency, and ensure the correct rate of subsequent text recognition.

[0058] The following Figure 1 elaborates in detail the specific implementation processes of the respective steps:

[0059] It should be noted that the method in the present disclosure can also be used in the following application scenarios. For example, improving the OCR (Optical Character Recognition) recognition rate to improve the efficiency of document automation processing, automatic license plate number recognition and traffic monitoring, automatic recognition of handwritten characters, automatic classification of business cards, etc. It can be set according to the actual situation and all belong to the protection scope of the present disclosure.

[0060] In the present disclosure, an initial sample set can be determined first. Specifically, reference can be made to Figure 2 , Figure 2 which shows a schematic flow chart for determining the initial sample set in an embodiment of the present disclosure, including steps S201 - S203. The following is an explanation in conjunction with Figure 2 :

[0061] In step S201, the original document image is collected and the label corresponding to the original document image is obtained.

[0062] In this step, the original document image can be collected, and the label corresponding to the original document image can be obtained. The label indicates the inclination degree of the original document image. Among them, with reference to Figures 3A - 3B , the original document image can be a document image that has been uploaded by relevant merchants to the Internet platform and is in a positive position (i.e., the inclination angle is 0, or the inclination angle is within the business - tolerable range, that is, the inclination angle is small). Exemplarily, when the inclination angle of the original document image is within the range of (-14.9, 14.9) degrees, it can be determined that the label corresponding to the original document image is 0.

[0063] In step S202, the original document image is rotationally processed to obtain a rotated image.

[0064] In this step, reference can be made to Figure 4 , Figure 4 which shows a schematic flow chart for rotationally processing the original document image to obtain a rotated image, including steps S401 - S403. The following is an explanation in conjunction with Figure 4 for step S202:

[0065] In step S401, the circumference is divided into N angular ranges at a preset interval.

[0066] In this step, the circumference (360 degrees) can be divided into N angular ranges, where N is an integer greater than 1. Taking N = 12 as an example, the circumference can be divided into 12 equal parts to obtain 12 angular ranges, each angular range being 30 degrees. For example: (-15, 15), (15, 45), (45, 75), etc. Each angular range corresponds to a label value, that is, 12 angular ranges correspond to 12 labels (i.e., 0 - 11). Exemplarily, the label value corresponding to (-15, 15) is 0, the label value corresponding to (15, 45) is 1..., thus, there are a total of 12 labels from 0 - 11.

[0067] It should be noted that a boundary distance margin can be introduced in this disclosure. The margin can be set to a relatively small value by oneself (for example, any value between 0 and 2 degrees. The smaller this value is, the smaller the impact on the processing result). The purpose is to reduce the influence of the boundary position and reduce the training difficulty. Exemplarily, the margin can be set to 0.1.

[0068] Furthermore, the boundary values of each angle range can be shrunk inward by 0.1. Thus, (-15, 15) becomes (-14.9, 14.9), and (15, 45) becomes (15.1, 44.9). This avoids the situation where the image angle is equal to or too close to the interval boundary value and it is impossible to determine which angle range it belongs to. For example, if the tilt angle of the image is exactly 15 degrees, it is impossible to determine whether it belongs to (-15, 15) or (15, 45).

[0069] In step S402, an angle value is randomly selected from each angle range respectively.

[0070] In this step, after dividing to generate the above N angle ranges, an angle value can be randomly selected from each angle range. For example, 12 angle values can be selected from the above 12 angle ranges.

[0071] In step S403, according to the N angle values selected, the original file image collected is rotated to obtain N rotated images.

[0072] In this step, the above original file image can be rotated according to the N angle values selected to obtain N rotated images. For example, according to the 12 angle values selected, each original file image is rotated to obtain 12 rotated images. Thus, the angle range to which each angle value belongs determines the label corresponding to the rotated image (12 angle values belong to 12 angle ranges, and 12 angle ranges correspond to labels 0 - 11). Exemplarily, when the rotated images obtained are image A - image L, it can be determined that the label corresponding to image A is 0, the label corresponding to image B is 1... the label corresponding to image L is 11.

[0073] Thus, this disclosure can solve the technical problem of the lack of labeled tags for training samples and ensure the learning accuracy of the machine learning model. Also, it can solve the technical problem in the prior art that the rotated 180 - degree image cannot be recognized when using a deep recognition model to detect the tilt angle of the file image.

[0074] After obtaining N rotated images, it can be detected whether the rotated image exceeds the image border. If so, the size of the image border is corrected according to the size of the original file image and the angle values selected.

[0075] Reference Figure 5 , when the size of the original document image XYZW is w*h (the length is w and the width is h), and the size of the original image border is also w*h, then when the selected angle value is θ, the length of the adjusted image border ABCD is w cosθ + h sinθ, and the width is wsinθ + h cosθ. Thus, it is possible to avoid the situation of partial loss of the image caused by the rotated image exceeding the image border, and ensure the integrity of the original document image.

[0076] After the size correction of the image border, there are some blank areas inside the image border. Thus, reference can be made to Figure 6 , Figure 6 shows a schematic flow chart of processing the image after size correction in the embodiments of the present disclosure, including steps S601 - step S604. The following will be combined with Figure 6 explain the specific implementation manners:

[0077] In step S601, the blank areas in the image border are filled with color.

[0078] In this step, after the size of the image border is adjusted, the blank areas in Figure 5 can be filled with color. Exemplarily, the blank areas (XCY, YBZ, ZAW, WDX) can be filled with black. The image obtained after color filling is as shown in Figure 7 .

[0079] In step S602, the rotated image is cropped from the image after color filling.

[0080] In this step, the above-mentioned rotated image can be cropped from the image after color filling based on the threshold method.

[0081] In step S603, a background image is randomly selected from the pre-stored background image set, and random noise is added to the background image.

[0082] In this step, Gaussian noise, salt-and-pepper noise, etc. can be added to the above-mentioned background image to make up for the problem caused by insufficient background images. The specific type of noise can be set according to the actual situation and belongs to the protection scope of the present disclosure.

[0083] In step S604, the rotated image is pasted onto the background image after adding random noise.

[0084] In this step, the rotated image obtained by the above-mentioned cropping can be pasted onto the background image after adding random noise to obtain a result image (the label corresponding to the result image is the same as the label corresponding to the rotated image). Exemplarily, the result image obtained after this processing step is as shown inFigure 8 As shown. Thus, the present disclosure can solve the problem caused by the inconsistency between the directly rotated image and the real image in the prior art. When directly rotating the image, the blank part will be automatically filled with black or white, and these features are not real features and are easily fitted by the model, resulting in a sharp decline in the generalization ability of the model.

[0085] Continue to refer to Figure 2 , in step S203, according to the original document image and its corresponding label, the rotated image and its corresponding label, an initial sample set is determined.

[0086] In this step, the above-mentioned original document image and its corresponding label, the rotated image and its corresponding label can be determined as the above-mentioned initial sample set. Or, the above-mentioned original document image and its corresponding label, the image obtained in step S604 and its corresponding label can also be determined as the above-mentioned initial sample set.

[0087] In step S110, data augmentation is performed on the initial sample set to obtain an augmented sample set.

[0088] In this step, data augmentation can be performed on the initial sample set to obtain an augmented sample set. Among them, there are many ways of data augmentation, such as geometric transformation of the image (such as rotation, random cropping, deformation, scaling, etc.), color transformation (including noise, blurring, color transformation, random erasing, etc.). Thus, the present disclosure can solve the technical problem that the data volume is small because the sample image is in a confidential state and cannot be obtained through the Internet, enrich the number of images in the training set, and thus ensure a sufficient proportion of samples to ensure the accuracy of the subsequent model.

[0089] Exemplarily, reference can be made to Figure 9 , Figure 9 shows a schematic flowchart of data augmentation on the initial sample set to obtain an augmented sample set in this exemplary embodiment, including steps S901 - S905. The following combines Figure 9 to explain step S110:

[0090] In step S901, the N rotated images corresponding to each original document image are used as a basic image set, and the image numbers in the basic image set are randomly shuffled to obtain a target image set.

[0091] In this step, for example, it is assumed that the basic image set contains 12 images, and their image numbers are 1 - 12 respectively. After shuffling the order of the above 12 images, they can be re - numbered to obtain a target image set.

[0092] In step S902, a first image is randomly selected from the basic image set, and a first sub - image is intercepted from the first image.

[0093] In this step, the first image can be randomly selected from the basic image set (for example, the image numbered 1 is selected). Furthermore, the first sub-image can be intercepted from this first image. Specifically, a ratio m of the area occupancy can be randomly sampled from ([ (1)), and a ratio n of the length-width ratio can be randomly sampled from . Furthermore, the first sub-image with an area occupancy ratio of m and a length-width ratio of n to the total area is randomly cropped from the first image.

[0094] In step S903, the second image is randomly selected from the target image set, and the second sub-image is intercepted from the second image; the serial number of the second image is the same as that of the first image.

[0095] In this step, the second image can be selected from the target image set (with the same serial number as the above-mentioned first image, for example: 1). Furthermore, the second sub-image can be captured from this second image. Specifically, a ratio u of the area occupancy can be randomly sampled from ([ (1)), and a ratio v of the length-width ratio can be randomly sampled from . Furthermore, the second sub-image with an area occupancy ratio of u and a length-width ratio of v to the total area is randomly cropped from the second image.

[0096] In step S904, the first sub-image and the second sub-image are mixed to obtain a mixed image.

[0097] In this step, a sampling value λ can be randomly sampled from a preset numerical interval (for example: (0, 1)) based on the beta distribution. Based on the above sampling value, the first sub-image and the second sub-image are mixed (MixUp). Exemplarily, the obtained mixed image is as Figure 10 shown.

[0098] Among them, the beta distribution is a density function that serves as the conjugate prior distribution of the Bernoulli distribution and the binomial distribution, and has important applications in machine learning and mathematical statistics. In probability theory, the beta distribution is also called the β distribution, which refers to a group of continuous probability distributions defined in the interval (0, 1).

[0099] MixUp is an algorithm for mixing and enhancing images in computer vision. It can mix images between different classes to expand the training data set.

[0100] In step S905, an augmented sample set is obtained according to the mixed image and its corresponding label.

[0101] In this step, the obtained mixed image and its corresponding label can be determined as the augmented sample set. Thus, the present disclosure can perform data augmentation under the condition of limited samples, solve the technical problem of lack of training samples in the prior art, enrich the number of samples, and ensure the training accuracy of the subsequent model.

[0102] It should be noted that after obtaining the above-mentioned first sub-image and second sub-image through the processing in steps S901 - S903, the above-mentioned first sub-image can also be scaled, for example, scaled to 256 * 256 pixels to obtain a first scaled image (the label of the first scaled image is the same as that of the above-mentioned first sub-image). And, the second sub-image is scaled, for example, scaled to 256 * 256 pixels to obtain a second scaled image (the label of the second scaled image is the same as that of the above-mentioned second sub-image). Furthermore, with reference to the relevant explanations in step S904 above, the first scaled image and the second scaled image are mixed to obtain a mixed image. Then, according to the labels of the first scaled image and the second scaled image, the label of the mixed image is determined, and based on the mixed image and its corresponding label, the above-mentioned augmented sample set is obtained.

[0103] Continue to refer to Figure 1 , in step S120, a classification prediction model is obtained by training a preset machine learning model based on the augmented sample set.

[0104] In this step, after obtaining the above-mentioned augmented sample set, based on the TensorFlow 2.0 (the second-generation artificial intelligence learning system developed based on DistBelief, which can be used in multiple machine learning and deep learning fields such as speech recognition or image recognition) machine learning platform, a preset machine learning model is trained according to the above-mentioned augmented sample set to obtain a classification prediction model. Specifically, the above-mentioned augmented sample set can be input into the preset machine learning model. Then, the loss value of the machine learning model can be adjusted to train the machine learning model so that the loss function of the machine learning model tends to converge to obtain a classification prediction model. Exemplarily, the loss value of the model can be corrected based on the following formula 2:

[0105] loss = λ × f loss (y_a, y _pred ) + (1 - λ) × f loss (y _ b, y _pred ) Formula 2

[0106] Among them, loss represents the corrected loss value, λ represents the above-mentioned sampling value, f loss represents the function for calculating the loss value, y_a represents the label of the basic image set, y _predRefers to the label predicted by the model, and y_b represents the label of the target image set.

[0107] Exemplarily, the above-mentioned preset machine learning model can be a ResneSt model. Compared with the network structures of similar structures such as Resnet before, the ResneSt model has higher accuracy. Compared with the Efficientnet series of network structures, it has higher accuracy and higher computational efficiency.

[0108] This disclosure uses a self-supervised method. Self-supervision, as opposed to supervised learning, unsupervised learning, and semi-supervised learning, means that common data generates data labels through manual annotation, and then the model is trained to fit these labels, which is the training process of the model. However, when there is no labeled data and the workload of labeling data is very large, algorithms need to be used to generate labels and define model-customized training tasks, usually called pretext tasks. After completing this training task, then perform the training task for the business scenario, which is called the downstream task. Through the training of the pretext task, the model can learn some basic modalities of the image, such as the color, lines, shape, etc. of the image. Then perform the second-stage task according to the requirements of the business scenario. Using this method, the tilt angle of the image can be judged in the absence of labeled data.

[0109] Reference Figure 11A , Figure 11A shows the architecture schematic diagram of the ResneSt model in the embodiments of this disclosure. The following is combined with Figure 11A for explanation:

[0110] step1, divide all the input feature maps into different cardinality groups;

[0111] step2, divide each cardinality group into different splits;

[0112] step3, calculate the weights of each split using split-attention (attention splitting module), and then fuse them as the output of each cardinality group;

[0113] step4, splice the feature maps of all cardinality groups together in the channel dimension;

[0114] Step 5: Perform conv (changing the number of channels) again and fuse the original input features of the ResNeSt Block using skip connection.

[0115] The following combines Figure 11B to explain the architecture diagram of the Split-Attention module in step 3 above:

[0116] Step 1: Divide the input of this cardinality group into r splits. After some transformations for each split, it enters split-attention. First, fuse the feature maps together using element-wise summation (sum the array elements in sequence) (output dimension: H×W times C). This step can be represented by the following formula 3:

[0117]

[0118] Step 2: Point the fused feature map to global pooling, that is, compress the spatial dimension of the image (output dimension: C). This step can be represented by the following formula 4:

[0119]

[0120] Step 3: Calculate the weight of each split in combination with softmax (normalization layer). The dense c implementation in the figure uses two fully connected layers;

[0121] Step 4: Multiply the feature map of each split input to each split-attention module by the calculated weight of each split to obtain a weighted fusion of a cardinality group (output dimension: H×WtimesC). This step can be represented by the following formula 5:

[0122]

[0123] Among them, represents the weight of each split:

[0124]

[0125] It can be seen that split-attention is actually calculating the corresponding weight for the feature map of each group of splits, and then fusing according to the weight.

[0126] Continue to refer to Figure 1 In step S130, the tilt detection is performed on the file image to be processed according to the classification prediction model, and the tilt angle range of the file image to be processed is obtained.

[0127] In this step, after training the above classification prediction model, the file image to be processed can be input into the above classification prediction model, so that the above classification prediction model performs tilt detection on the file image to be processed. Furthermore, according to the output of the above classification prediction model, the tilt angle range of the above file image to be processed is obtained. Exemplarily, when the output of the classification prediction model is 0, it can be determined that the tilt angle range of the file image to be processed is (-14.9, 14.9), and when the output of the classification prediction model is 1, it can be determined that the tilt angle range of the file image to be processed is (15.1, 44.9).

[0128] After obtaining the above tilt angle range, if the tilt angle range is consistent with the target angle range, it is determined that the file image to be processed is a non-tilted image. For example, when the classification prediction result output by the model is 0 (i.e., the tilt angle range is (-14.9, 14.9)) and the preset target angle range is (-14.9, 14.9), it can be determined that the above file image to be processed is a non-tilted image or an image with a tilt angle within the business acceptable range and does not require tilt correction.

[0129] If the tilt angle range is inconsistent with the target angle range, for example, when the classification prediction result output by the model is 1 (i.e., the tilt angle range is (15.1, 44.9)), it can be determined that the above file image to be processed is a tilted image, and tilt correction is performed on the above file image to be processed.

[0130] Specifically, the boundary values (including the upper limit value and the lower limit value) of the above-obtained tilt angle range can be obtained, and the input image, the file image to be processed, is tilt-corrected according to the average value of the upper limit value and the lower limit value. For example, when the classification prediction result output by the model is 1, the corresponding angle range is (15.1, 44.9), then the lower limit value can be determined to be 15.1, the lower limit value is 44.9, and thus, the average value is Furthermore, the file image to be processed can be rotated by -30 degrees to perform tilt correction on the above file image to be processed.

[0131] Exemplarily, it can also refer to Figure 12 , Figure 12 shows a schematic flow chart of performing tilt correction on the file image to be processed if the tilt angle range is inconsistent with the target angle range, including steps S1201 - S1206. The following combines Figure 12 to explain the specific implementation manners:

[0132] In step S1201, the file image to be processed is converted into a grayscale image.

[0133] In this step, the above-mentioned file image to be processed can be converted into a grayscale image, which is an image with equal RGB three-channel values. Exemplarily, the color image can be converted into a grayscale image based on the following several algorithms:

[0134] ① Floating-point method: Gray = R * 0.3 + G * 0.59 + B * 0.11;

[0135] ② Integer method: Gray = (R * 30 + G * 59 + B * 11) / 100;

[0136] ③ Shift method: Gray = (R * 77 + G * 151 + B * 28) >> 8;

[0137] ④ Average value method: Gray = (R + G + B) / 3;

[0138] ⑤ Only take green: Gray = G;

[0139] ⑥ Gamma correction algorithm:

[0140] In step S1202, the grayscale image is subjected to Gaussian blur processing to obtain a blurred image.

[0141] In this step, the above grayscale image can be subjected to Gaussian blur processing to obtain a blurred image. From a mathematical perspective, the Gaussian blur process of an image is the convolution of the image with a normal distribution. Since the normal distribution is also called the Gaussian distribution, this technology is called Gaussian blur. Convolving the image with a circular box blur will generate a more accurate out-of-focus imaging effect.

[0142] In step S1203, edge detection is performed on the blurred image to obtain an edge image.

[0143] In this step, the Canny edge detection algorithm can be used to perform edge detection on the above blurred image to obtain an edge image. Exemplarily, the Sobel edge detection algorithm, Laplacian operator (second-order differential operator), Roberts operator, Prewitt operator, Kirsch operator, etc. can also be used to perform edge detection on the above blurred image, which can be set according to the actual situation and belongs to the protection scope of the present disclosure.

[0144] Image edge detection greatly reduces the amount of data, eliminates irrelevant information, and retains the important structural attributes of the image.

[0145] In step S1204, a Hough transform method based on polar coordinate space transformation is used to detect straight lines in the edge image, and characteristic straight lines are obtained.

[0146] In this step, a Hough transform method based on polar coordinate space transformation can be used to detect straight lines in the above-mentioned edge image, and characteristic straight lines are obtained. Thus, the technical problem that the Hough transform cannot be used to detect the tilt angle caused by the complex image content in the prior art (the aspect ratio varies greatly, there are no straight lines or edges in the image, the straight lines in the image are of different lengths, horizontal and vertical) can be solved.

[0147] In step S1205, the included angle values between each characteristic straight line and the horizontal line are obtained, and target straight lines with included angle values within the tilt angle range are selected.

[0148] In this step, after obtaining the above-mentioned characteristic straight lines, the included angle values between each characteristic straight line and the horizontal line can be obtained. Furthermore, target straight lines with included angle values within the tilt angle range output by the above model can be determined.

[0149] It should be noted that the included angle obtained through the above steps is the included angle value in polar coordinates. Thus, the above tilt angle range can be first converted into an angle range in the polar coordinate system, and then target straight lines with included angle values within the angle range in the above polar coordinate system can be determined.

[0150] In step S1206, the file image to be processed is corrected for tilt according to the average value of the included angle values corresponding to the target straight lines.

[0151] In this step, the average value of the included angle values corresponding to the above target straight lines can be calculated, and then the file image to be processed is corrected for tilt according to the above average value. For example, when the above average value is , the file image to be processed can be rotated degrees to achieve tilt correction of the file image to be processed.

[0152] Thus, the present disclosure can not only automatically detect the tilt angle of the file image, but also correct the tilted image according to the detection result, thereby improving the operation process when merchants enter the platform, avoiding the technical problems of low efficiency and affecting the entry process caused by the need for merchants to upload the file image multiple times when the file is tilted and the audit fails, and reducing the human and material costs. Moreover, it can solve the technical problems of unrecognizability and affecting the audit progress caused by excessive tilt of the file, improve the file audit efficiency, and ensure the accuracy of subsequent text recognition.

[0153] Refer to Figure 13 , Figure 13 which shows the overall flowchart of the image correction method in the embodiment of the present disclosure, including steps S1301 - S1305. The following is combined withFigure 13 Explanation:

[0154] In step S1301, obtain the original file image and its corresponding rotated image;

[0155] In step S1302, the data augmentation module;

[0156] In step S1303, train the ResneSt model;

[0157] In step S1304, use the trained model to identify the tilt angle of the image and perform correction processing on the tilted image;

[0158] In step S1305, output the image detection result.

[0159] Based on the above technical solutions, on the one hand, the present disclosure can solve the technical problem that the data volume is small because the sample images are in a confidential state and cannot be obtained through the Internet, enrich the number of pictures in the training set, so as to ensure a sufficient proportion of samples and ensure the accuracy of the subsequent model. Further, the present disclosure can also solve the technical problem of the lack of labeled data in the training samples and ensure the learning accuracy of the machine learning model. On the other hand, the present disclosure automatically detects pictures, solves the technical problem of high human and material costs caused by manually detecting the tilt angle of pictures in the prior art, improves the detection efficiency, and, by performing tilt correction on the detected tilted pictures, solves the technical problem of seriously affecting the merchant settlement process caused by the need to contact the merchant to re-upload relevant files in the prior art, and speeds up the review process.

[0160] The present disclosure also provides an image correction device, Figure 14 showing the structural schematic diagram of the image correction device in an exemplary embodiment of the present disclosure; as Figure 14 shown, the image correction device 1400 may include a data augmentation module 1401, a model training module 1402, a classification prediction module 1403, and a tilt correction module 1404. Among them:

[0161] The data augmentation module 1401 is used to perform data augmentation on the initial sample set to obtain an expanded sample set.

[0162] In an exemplary embodiment of the present disclosure, the data augmentation module is used to collect the original file image and obtain the label corresponding to the original file image; perform rotation processing on the original file image to obtain a rotated image; and determine the initial sample set according to the original file image and its corresponding label, the rotated image and its corresponding label.

[0163] In an exemplary embodiment of the present disclosure, the data augmentation module is configured to divide the circumference into N angular ranges at a preset interval; N is an integer greater than 1; randomly select an angular value from each angular range; and perform a rotation process on the original document image according to the N selected angular values to obtain N rotated images.

[0164] In an exemplary embodiment of the present disclosure, the label corresponding to the rotated image is determined by the following method: determining the label corresponding to the rotated image according to the angular range to which the angular value belongs.

[0165] In an exemplary embodiment of the present disclosure, after obtaining the N rotated images, the data augmentation module is configured to detect whether the rotated images exceed the image border; if so, perform a size correction on the image border according to the size of the original document image and the selected angular values.

[0166] In an exemplary embodiment of the present disclosure, after performing the size correction on the image border, the data augmentation module is configured to fill the blank area in the image border with a color; crop the rotated image from the image after color filling; randomly select a background image from a pre-stored set of background images, and add random noise to the background image; and paste the rotated image onto the background image after adding random noise.

[0167] In an exemplary embodiment of the present disclosure, the data augmentation module is configured to use the N rotated images corresponding to each original document image as a basic image set, randomly shuffle the image numbers in the basic image set to obtain a target image set; randomly select a first image from the basic image set, and crop a first sub-image from the first image; randomly select a second image from the target image set, and crop a second sub-image from the second image; the second image has the same number as the first image; perform image mixing on the first sub-image and the second sub-image to obtain a mixed image; determine the label of the mixed image according to the label of the first sub-image and the label of the second sub-image; and obtain an augmented sample set according to the mixed image and its corresponding label.

[0168] In an exemplary embodiment of the present disclosure, the data augmentation module is configured to randomly sample from a preset numerical interval based on a beta distribution to obtain a sampling value; and perform image mixing on the first sub-image and the second sub-image based on the sampling value.

[0169] The model training module 1402 is configured to train a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is used to perform tilt detection on a file image to be processed.

[0170] The classification prediction model 1403 is configured to perform tilt detection on a file image to be processed according to the classification prediction model to obtain the tilt angle range of the file image to be processed.

[0171] The skew correction module 1404 is configured to perform skew correction on the file image to be processed if the skew angle range is inconsistent with the target angle range.

[0172] In an exemplary embodiment of the present disclosure, the skew correction module is configured to obtain boundary values of the skew angle range; the boundary values include an upper limit value and a lower limit value; and perform skew correction on the file image to be processed according to the average value of the upper limit value and the lower limit value.

[0173] In an exemplary embodiment of the present disclosure, the skew correction module is configured to convert the file image to be processed into a grayscale image; perform Gaussian blur processing on the grayscale image to obtain a blurred image; perform edge detection on the blurred image to obtain an edge image; perform line detection on the edge image by using a Hough transform method based on polar coordinate space transformation to obtain characteristic lines; obtain the included angle values between the characteristic lines and the horizontal line, and select target lines whose included angle values are within the skew angle range; and perform skew correction on the file image to be processed according to the average value of the included angle values corresponding to the target lines.

[0174] The specific details of each module in the above image correction device have been described in detail in the corresponding image correction method, and thus will not be elaborated herein.

[0175] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0176] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0177] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the methods according to the embodiments of the present disclosure.

[0178] The present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device.

[0179] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0180] The computer-readable storage medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0181] The computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.

[0182] In addition, in the embodiments of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0183] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0184] The following refers to Figure 15 to describe the electronic device 1500 according to this embodiment of the present disclosure. Figure 15 The shown electronic device 1500 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0185] As Figure 15As shown, the electronic device 1500 is presented in the form of a general-purpose computing device. The components of the electronic device 1500 may include, but are not limited to: at least one of the above-mentioned processing units 1510, at least one of the above-mentioned storage units 1520, a bus 1530 connecting different system components (including the storage unit 1520 and the processing unit 1510), and a display unit 1540.

[0186] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1510, so that the processing unit 1510 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of the present specification above. For example, the processing unit 1510 can execute as Figure 1 shown in: Step S110, performing data augmentation on the initial sample set to obtain an augmented sample set; Step S120, training a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is used to perform skew detection on the file image to be processed; Step S130, performing skew detection on the file image to be processed according to the classification prediction model to obtain the skew angle range of the file image to be processed; Step S140, if the skew angle range is inconsistent with the target angle range, performing skew correction on the file image to be processed.

[0187] The storage unit 1520 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 15201 and / or a cache storage unit 15202, and may further include a read-only storage unit (ROM) 15203.

[0188] The storage unit 1520 may further include a program / utilities 15204 having a set (at least one) of program modules 15205. Such program modules 15205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0189] The bus 1530 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the various bus structures.

[0190] The electronic device 1500 can also communicate with one or more external devices 1600 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1500, and / or communicate with any device that enables the electronic device 1500 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 1550. Moreover, the electronic device 1500 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1560. As shown in the figure, the network adapter 1560 communicates with other modules of the electronic device 1500 through the bus 1530. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0191] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. An image correction method, characterized in that, Including: Performing data augmentation on an initial sample set to obtain an augmented sample set; The performing data augmentation on the initial sample set to obtain the augmented sample set includes: using N rotated images corresponding to each original document image as a basic image set, randomly shuffling the image numbers in the basic image set to obtain a target image set; randomly selecting a first image from the basic image set and cropping a first sub-image from the first image; randomly selecting a second image from the target image set and cropping a second sub-image from the second image; the second image having the same number as the first image; performing image mixing on the first sub-image and the second sub-image to obtain a mixed image; determining the label of the mixed image according to the label of the first sub-image and the label of the second sub-image; obtaining the augmented sample set according to the mixed image and its corresponding label; Training a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is used for performing skew detection on a file image to be processed; Performing skew detection on the file image to be processed according to the classification prediction model to obtain the skew angle range of the file image to be processed; If the skew angle range is inconsistent with the target angle range, performing skew correction on the file image to be processed; The if the skew angle range is inconsistent with the target angle range, performing skew correction on the file image to be processed includes: performing line detection on the file image to be processed by using a Hough transform method based on polar coordinate space transformation to obtain feature lines; obtaining the included angle values between each feature line and the horizontal line, and selecting target lines whose included angle values are within the skew angle range; performing skew correction on the file image to be processed according to the average value of the included angle values corresponding to the target lines.

2. The method according to claim 1, characterized in that, Before performing data augmentation on the initial sample set to obtain the augmented sample set, the method further includes: Collecting original document images and obtaining the labels corresponding to the original document images; Performing rotation processing on the original document images to obtain rotated images; Determining the initial sample set according to the original document images and their corresponding labels, the rotated images and their corresponding labels.

3. The method according to claim 2, characterized in that, The performing rotation processing on the original document images to obtain rotated images includes: Dividing the circumference into N angular ranges at a preset interval; N is an integer greater than 1; Randomly selecting an angular value from each angular range; Performing rotation processing on the original document image according to the N selected angular values to obtain N rotated images.

4. The method according to claim 3, characterized in that, The label corresponding to the rotated image is determined by the following method: Determining the label corresponding to the rotated image according to the angular range to which the angular value belongs.

5. The method according to claim 3, characterized in that, After obtaining the N rotated images, the method further includes: Detecting whether the rotated images exceed the image border; If so, performing size correction on the image border according to the size of the original document image and the selected angular values.

6. The method according to claim 5, characterized in that, After performing size correction on the image border, the method further includes: Filling the blank area in the image border with a color; Cropping the rotated image from the image after the color filling; Randomly selecting a background image from a pre-stored set of background images and adding random noise to the background image; Pasting the rotated image onto the background image after adding the random noise.

7. The method according to claim 1, characterized in that, The image mixing of the first sub-image and the second sub-image includes: Randomly sampling from a preset numerical interval based on a beta distribution to obtain a sampling value; Based on the sampling value, performing image mixing on the first sub-image and the second sub-image.

8. The method according to any one of claims 1 to 7, characterized in that, If the tilt angle range is inconsistent with the target angle range, the tilt correction of the file image to be processed includes: Obtaining the boundary values of the tilt angle range; the boundary values include an upper limit value and a lower limit value; Performing tilt correction on the file image to be processed according to the average value of the upper limit value and the lower limit value.

9. The method according to claim 8, characterized in that, The method further includes: Converting the file image to be processed into a grayscale image; Performing Gaussian blur processing on the grayscale image to obtain a blurred image; Performing edge detection on the blurred image to obtain an edge image; Performing line detection on the edge image by a Hough transform method based on polar coordinate space transformation to obtain characteristic lines; Obtaining the included angle values between each of the characteristic lines and the horizontal line, and selecting target lines whose included angle values are within the tilt angle range; Performing tilt correction on the file image to be processed according to the average value of the included angle values corresponding to the target lines.

10. An image correction device, characterized in that, Including: A data augmentation module for augmenting an initial sample set to obtain an augmented sample set; The data augmentation module uses the N rotated images corresponding to each original file image as a basic image set, randomly shuffles the image numbers in the basic image set to obtain a target image set; randomly selecting a first image from the basic image set, and cropping a first sub-image from the first image; Randomly selecting a second image from the target image set, and cropping a second sub-image from the second image; the second image has the same number as the first image; performing image mixing on the first sub-image and the second sub-image to obtain a mixed image; determining the label of the mixed image according to the label of the first sub-image and the label of the second sub-image; Obtaining the augmented sample set according to the mixed image and its corresponding label; A model training module for training a preset machine learning model according to the augmented sample set to obtain a classification prediction model; the classification prediction model is used for tilt detection of a file image to be processed; The classification prediction model is used for performing tilt detection on the file image to be processed according to the classification prediction model to obtain the tilt angle range of the file image to be processed; A tilt correction module for performing tilt correction on the file image to be processed if the tilt angle range is inconsistent with the target angle range; An inclination correction module is configured to perform straight line detection on the to-be-processed file image by using a Hough transform method based on polar coordinate space transformation to obtain feature straight lines; acquire the included angle values between each of the feature straight lines and the horizontal line, and select target straight lines whose included angle values are within the inclination angle range; and perform inclination correction on the to-be-processed file image according to the average value of the included angle values corresponding to the target straight lines.

11. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the image correction method according to any one of claims 1 to 9.

12. An electronic device, characterized in that, Comprising: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the image correction method according to any one of claims 1 to 9 by executing the executable instructions.

Citation Information

Patent Citations

  • Character image correction processing method and device, equipment and storage medium

    CN109583445A

  • Image mask filtering method and device and storage medium

    CN110717060A

  • Fine-grained image classification method and device for deep learning

    CN112487227A

  • Performance test method for medical image recognition system

    CN112506797A