Document image distortion correction model training method and document image distortion correction method and system using same
By segmenting and training a document image distortion correction model to generate an inverse map, the problem of image distortion when taking pictures of documents with a smartphone is solved, achieving high-quality distortion correction and background information preservation, and supporting higher-quality document processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-07
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for photographing documents using smartphone cameras suffer from the problem that image distortion correction requires the entire document area to be in a quadrilateral shape and background information to be lost.
By receiving an image of a document region, a first grid is segmented into multiple units, a second grid is generated to display the distortion correction state, and a document image distortion correction model is trained to generate an inverse mapping to correct image distortion.
It achieves high-quality distortion correction in non-overall document areas, avoids the disappearance of background information, and supports high-quality processing for subsequent OCR and translation services.
Smart Images

Figure CN121753062A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The disclosure relates to a method of training a document image distortion correction model, and a method and system of correcting a document image distortion using the same, and more particularly, to a method and system of training a document image distortion correction model to generate a reverse map for correcting a distortion of a document image, and correcting a distortion of a document image using the reverse map generated by the trained model. BACKGROUND
[0002] Recently, as the performance of a camera included in a smart phone is improved, it is possible to photograph a document using a camera of a smart phone instead of an optical scanner, and to perform optical character recognition (OCR) on the photographed document or to receive a translation service. However, a document within an image photographed by a camera can be distorted according to a position and a pose of the camera. Therefore, in order to remove such a document distortion, a rectification technique is developed.
[0003] However, in the case of using such a conventional technique, there is a constraint that, in order to remove a document distortion within a photographed image, it is necessary to secure a whole document region within the image, and only one quadrangle-shaped object can exist within the image. In addition, a document image whose distortion is removed according to the conventional technique has a problem that background information existing in an originally photographed image disappears. SUMMARY
[0004] The disclosure provides a method of training a document image distortion correction model and a method and apparatus (system) of correcting a document image distortion for solving the above-described problems.
[0005] The disclosure can be implemented in various ways including a method, an apparatus (system), or a computer program stored in a readable storage medium.
[0006] According to one embodiment of the disclosure, the method of training a document image distortion correction model can include the steps of receiving a first grid that divides a document region on a first image including the document region into a plurality of cells, generating a second grid that displays a state in which a distortion of the first image is corrected, based on the first grid, generating a reverse map that calculates a pixel within the first grid corresponding to each pixel within the second grid, and training a document image distortion correction model in a manner that the reverse map is generated from the first image.
[0007] According to one embodiment of the present disclosure, a document image distortion correction method includes the steps of: receiving a first image including a document region; generating a reverse map for correcting distortion of the first image using a document image distortion correction model; and generating a second image in which distortion of the first image is corrected based on the reverse map.
[0008] A computer program can be provided, which is stored in a computer-readable recording medium, for executing a method according to one embodiment of the present disclosure on a computer.
[0009] An information processing system according to one embodiment of the present disclosure includes: a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one program stored in the memory, the at least one program including a command language for training a document image distortion correction model to perform the steps of: receiving a first grid that divides a document region on a first image including the document region into a plurality of cells; generating a second grid that displays a state in which distortion of the first image is corrected based on the first grid; generating a reverse map that calculates pixels in the first grid corresponding to each pixel in the second grid; and generating the reverse map from the first image.
[0010] According to some embodiments of the present disclosure, a document image distortion correction model can be trained to output a reverse map that can be used to correct distortion of a corresponding image from an original image. In this case, instead of using a forward map that defines a transformation relationship from an input image to an output image in which distortion is corrected, a reverse map that defines a transformation relationship opposite thereto is used, thereby preventing a hole phenomenon in which a portion of pixels is not filled in an image in which distortion is corrected and enabling high-quality image distortion correction to be performed.
[0011] According to some embodiments of the present disclosure, distortion of a document can be corrected regardless of a form of a document and / or text within an original image. In addition, even if a document region as a whole cannot be ensured from an input image, distortion present in a corresponding document region can be corrected, and background information around the document region does not disappear after the distortion is corrected, and thus a high-quality image in which distortion is corrected can be generated.
[0012] According to some embodiments of the present disclosure, by post-processing an image in which distortion is corrected, a higher-quality image can be generated. In addition, by extracting text from an image in which distortion is corrected using a reverse map and translating the text, a higher-quality translation service can be provided.
[0013] Effects of the present disclosure are not limited to the above-mentioned effects, and other effects not mentioned can be clearly understood by persons having ordinary knowledge in the technical field to which the present disclosure pertains (referred to as "those skilled in the art") from the written description in the claims. BRIEF DESCRIPTION OF DRAWINGS
[0014] Embodiments of the present disclosure are explained with reference to the accompanying drawings described below, in which like reference numerals denote like elements, but are not limited thereto.
[0015] Figure 1 is an example of a method of training a document image distortion correction model according to one embodiment of the present disclosure.
[0016] Figure 2 is a block diagram showing an internal structure of an information processing system according to one embodiment of the present disclosure.
[0017] Figure 3 is a diagram showing an internal structure of a processor of an information processing system according to one embodiment of the present disclosure.
[0018] Figure 4 is a diagram showing an example of an internal structure of a document image distortion correction model according to one embodiment of the present disclosure.
[0019] Figure 5 is a diagram showing an example of generating a grid in an original image according to one embodiment of the present disclosure.
[0020] Figure 6 is a diagram showing an example of displaying a plurality of blocks according to one embodiment of the present disclosure.
[0021] Figure 7 is a diagram showing an example of generating a second grid based on a first grid according to one embodiment of the present disclosure.
[0022] Figure 8 is a diagram showing an example of displaying a plurality of blocks according to one embodiment of the present disclosure.
[0023] Figure 9 is a diagram showing an example of calculating an inverse mapping according to one embodiment of the present disclosure.
[0024] Figure 10 is a diagram showing an example of an image whose distortion is corrected according to one embodiment of the present disclosure.
[0025] Figure 11 is a diagram showing an example of post-processing a generated image according to one embodiment of the present disclosure.
[0026] Figure 12is a flowchart representing an example of a training method of a document image distortion correction model according to one embodiment of the present disclosure.
[0027] Figure 13 is a flowchart representing an example of a document image distortion correction method according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] Hereinafter, specific contents for implementing the present disclosure will be explained in detail with reference to the accompanying drawings. However, in the following explanation, specific explanation on well-known functions or structures will be omitted in case that it unnecessarily confuses the gist of the present disclosure.
[0029] In the drawings, the same reference numerals are applied to the same or corresponding constituent elements. Also, in the following explanation of the embodiments, repeated description on the same or corresponding constituent elements can be omitted. However, even if the description on the constituent elements is omitted, it is not intended to exclude the constituent elements in some embodiments.
[0030] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The following detailed description is provided to assist in a comprehensive understanding of one or more embodiments of the present disclosure. Accordingly, those skilled in the art will recognize that other embodiments can be practiced with the same or equivalent functionality, components, and / or structures without departing from the present disclosure. Figure 1 The advantages and features of the disclosed embodiments and the method for achieving them will be apparent from the following embodiments described below. However, the present disclosure is not limited to the embodiments disclosed below, and can be implemented in various forms different from each other, but the embodiments are provided to make the present disclosure complete and to inform the scope of the invention completely to those skilled in the art.
[0031] The terms used in the present specification are simply explained, and the disclosed embodiments are specifically explained. The terms used in the present specification are selected as common terms widely used at present in consideration of the functions in the present disclosure, but this can be different according to the intention of those skilled in the art related to the field or precedents, appearance of new technology, etc. Also, in a specific case, there are terms arbitrarily selected by the applicant, and in this case, the meaning thereof will be detailed in the explanation part of the corresponding invention. Therefore, the terms used in the present disclosure should not be defined based on the name of the simple term, but should be defined based on the meaning that the term has and the overall content of the present disclosure.
[0032] Unless the context clearly indicates otherwise, the expression of the singular in the present specification includes the expression of the plural. Also, unless the context clearly indicates otherwise, the expression of the plural in the present specification includes the expression of the singular. Throughout the specification, when a certain part includes a certain constituent element, unless otherwise noted in the contrary, it means that it can also include other constituent elements, not excluding other constituent elements.
[0033] Also, the term "module" or "part" used in the specification means a software or hardware constituent element, and the "module" or "part" performs certain functions. However, the "module" or "part" is not limited to the software or hardware. The "module" or "part" can be configured to exist in an addressable storage medium, and can be configured to reproduce one or more processors. Accordingly, as one example, the "module" or "part" can include at least one of software constituents, object-oriented software constituents, class constituents, and task constituents, and processes, functions, attributes, procedures, sub-routines, program code segments, drivers, firmware, micro-codes, circuits, data, databases, data structures, tables, arrays, or variables. Constituent elements and functions provided within the "module" or "part" can be combined with a smaller number of constituent elements and "modules" or "parts" or further divided into additional constituent elements and "modules" or "parts".
[0034] According to one embodiment of the disclosure, the "module" or "part" can be implemented as a processor and a memory. The "processor" should be broadly interpreted to include a general processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some circumstances, the "processor" can also designate a semiconductor (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. The "processor" can also refer to a combination of processing devices, such as a combination of a DSP and a microprocessor, a combination of a plurality of microprocessors, a combination of one or more microprocessors with one or more DSP cores, or a combination of any other such structures. In addition, the "memory" should be broadly interpreted to include any electronic components that can store electronic information. The "memory" can also refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage devices, registers, etc. The memory is said to be in electronic communication with the processor as long as the processor is capable of reading information from and / or writing information to the memory. The memory integrated in the processor is in electronic communication with the processor.
[0035] In the disclosure, the "system" can include at least one of a server device and a cloud device, but is not limited thereto. For example, the system can be constituted by one or more server devices. As another example, the system can be constituted by one or more cloud devices. As still another example, the system can be constituted by a server device and a cloud device together and operate.
[0036] In the present disclosure, the "display" can refer to any display device related to the computing device, for example, can refer to any display device capable of displaying any information / data controlled by or provided from the computing device.
[0037] In the present disclosure, "each of a plurality of A" or "each of a plurality of A" can refer to each of all constituent elements included in the plurality of A, or can refer to each of a part of the constituent elements included in the plurality of A.
[0038] In the present disclosure, the "machine learning model" can include any model for inferring an answer for a given input. According to an embodiment, the machine learning model can include an artificial neural network model including an input layer (layer), a plurality of hidden layers, and an output layer. Here, each layer can include a plurality of nodes. In the present disclosure, each of a plurality of machine learning models is described as a different machine learning model, but is not limited thereto, and a part or the entirety of the plurality of machine learning models can be implemented in one machine learning model. In addition, one machine learning model can include a plurality of machine learning models. In the present disclosure, the terms of machine learning model and artificial neural network model can be used interchangeably, thereby being able to represent the same or similar model.
[0039] In the present disclosure, the "document image" can include an image generated by photographing or scanning a document in a printed form, and an image related to an electronic document (for example, an electronic receipt, an electronic business card, an electronic ID, etc.) generated by a program, etc.
[0040] In the present disclosure, the "document region" can mean a region including character recognition of text or becoming a translation object within the document image. In addition, in the present disclosure, the "background region" can mean a remaining region other than the document region within the document image.
[0041] Figure 1An example illustrating a training method of a document image distortion correction model according to one embodiment of the present disclosure is shown. In one embodiment, a grid (120) can be generated on an original image 110 containing a document region. Specifically, a first grid 122 can be generated that partitions the document region of the original image 110 into a plurality of cells. Here, the slope of the horizontal lines and the slope of the vertical lines of the first grid 122 can be determined based on the orientation of the text contained in the document region. The first grid 122 thus generated can roughly show the distortion of the document region contained in the original image 110, such as rotation, bending, wrinkling, etc., at the patch level. For example, the first grid 122 can contain cells partitioned into a 10x10 grid, but is not limited thereto. The generation (120) of the first grid 122 can be performed using an image editing application or using a suitable grid generation algorithm.
[0042] In addition, the generated first grid 122 can be transformed to generate a second grid 132 (130). Here, the second grid 132 can show the state of the distortion of the document region corrected by the first grid 122 at the patch level. The second grid 132 can be generated in a plurality of cells including a horizontal right angle shape based on the average of the coordinate values of a plurality of edge points of the first grid 122 and the average of a plurality of cell intervals.
[0043] In one embodiment, a plurality of patches (140) can be generated from each of the first grid 122 and the second grid 132. Specifically, a plurality of first patches representing the partitioned plurality of cells of the first grid 122 can be generated. Here, the plurality of first patches 142 can contain a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points contained in the first grid 122 and 4 vertices of the original image 110. Similarly, a plurality of second patches 144 representing the partitioned plurality of cells of the second grid 132 can be generated. Here, the plurality of second patches 144 can contain a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points contained in the second grid 132 and 4 vertices of the image after correcting the distortion of the original image 110.
[0044] In one embodiment, a backward map (150) can be generated based on the plurality of first patches 142 and the plurality of second patches 144. Specifically, a transformation matrix defining the transformation relationship from each of the plurality of second patches to each of the plurality of first patches corresponding thereto can be calculated. The plurality of transformation matrices thus calculated can be combined to generate a backward map for the entire image. Thus, the backward map can calculate the pixels within the first grid 122 of the image before the distortion is corrected corresponding to each pixel within the second grid 132 of the image after the distortion is corrected.
[0045] In one embodiment, the document image distortion correction model (160) can be trained in a manner that the inverse mapping 160 is generated from the original image 110. That is, the document image distortion correction model is trained using pairs of the original image 110 and the inverse mapping 160 associated therewith as training data to calculate the inverse mapping capable of correcting distortion on the corresponding image in correspondence with the input image. Here, the document image distortion correction model can be a deep learning model based on an encoder-decoder capable of learning to output the inverse mapping for distortion correction of the corresponding image based on the input image, but is not limited thereto.
[0046] The document image distortion correction model trained by such a configuration can generate the inverse mapping for correcting distortion of the document region from the input image. In addition, the output image in which distortion is corrected can also be generated by applying the inverse mapping to the input image. The output image in which distortion is corrected can be used for subsequent processing such as OCR (optical character recognition) or translation.
[0047] Figure 2 is a block diagram showing an internal structure of an information processing system 200 according to one embodiment of the present disclosure. The information processing system 200 can include a memory 210, a processor 220, a communication module 230, and an input / output interface 240. The information processing system 200 is configured to be able to communicate information and / or data with an external system through a network using the communication module 230.
[0048] The memory 210 can include any computer-readable recording medium that is non-transitory. According to one embodiment, the memory 210 can include a ROM (read only memory), a disk drive, an SSD (solid state drive), a flash memory, and the like non-depleting mass storage device. As other examples, the ROM, the SSD, the flash memory, the disk drive, and the like non-depleting mass storage device can be included in the information processing system 200 as another permanent storage device distinguished from the memory. In addition, the memory 210 can store an operating system and at least one program code (for example, a code for performing document distortion correction, and the like installed in the information processing system 200 and driven).
[0049] Such software components can be loaded from a computer-readable recording medium that is independent of the memory 210. Such independent computer-readable recording medium can include recording media that can be directly connected to such information processing system 200, such as a floppy disk drive, a magnetic disk, a magnetic tape, a DVD / CD-ROM drive, a memory card, and the like computer-readable recording medium. As other examples, software components can also be loaded into the memory 210 through the communication module 230, instead of being loaded from a computer-readable recording medium. For example, at least one program can be loaded into the memory 210 based on a computer program (e.g., a program for training a document distortion correction model, a program for performing document distortion correction, and the like) installed by a file distribution system that distributes files of installation files of application programs developed by a developer or the like provided by the communication module 230.
[0050] The processor 220 can be configured to process commands of a computer program by performing basic arithmetic, logic, and input / output calculations. The commands can be provided to a user terminal (not shown) or other external systems through the memory 210 or the communication module 230. For example, the processor 220 can train a document image distortion correction model in such a way that a second mesh that displays a state in which distortion of a first image is corrected is generated based on a first mesh, a reverse mapping that calculates a pixel within the first mesh corresponding to each pixel within the second mesh is generated from the first image, and the reverse mapping is generated from the first image. In addition, the processor 220 can generate a second image in which distortion of the first image is corrected using the reverse mapping generated from the document image distortion correction model based on the first image. Furthermore, the processor 220 of the information processing system 200 can be configured to manage, process, and / or store information and / or data received from a plurality of user terminals and / or a plurality of external systems.
[0051] The communication module 230 can provide a configuration or a function for a user terminal (not shown) and the information processing system 200 to communicate with each other through a network, and the information processing system 200 can provide a configuration or a function for communication with an external system (as one example, another cloud system, and the like). As one example, a control signal, a command, data, and the like provided by the control of the processor 220 of the information processing system 200 can be transmitted to a user terminal and / or an external system through a communication module of the user terminal and / or the external system via the communication module 230 and a network. For example, the information processing system 200 can transmit an image in which a document distortion is corrected to a user terminal through the communication module 230.
[0052] In addition, the input / output interface 240 of the information processing system 200 can be a unit that interacts with a device (not shown) for input or output connected to or can be included in the information processing system 200. Figure 2The input / output interface 240 is illustrated as an element constituted independently of the processor 220, but is not limited thereto, and the input / output interface 240 can be constituted to be included in the processor 220. The information processing system 200 can include more constituent elements than those of Figure 3 illustrated. However, most of the prior technical constituent elements do not need to be explicitly illustrated.
[0053] Figure 3 is a diagram showing an internal structure of the processor 220 of the information processing system according to one embodiment of the present disclosure. According to one embodiment, the processor 220 can include a learning section 310, a reverse mapping generation section 320, and an image generation section 330. Figure 4 The internal structure of the processor 220 of the information processing system shown in the above is only one example, and can be implemented differently therefrom. For example, at least a part of the constitution of the processor 220 can be omitted, or other constitution can be added, and at least a part of the actions or processes performed by the processor 220 can be performed by a processor of a user terminal connected in a manner capable of communicating with the information processing system.
[0054] The learning section 310 can train a document image distortion correction model using training data including one or more document images and one or more reverse mappings corresponding thereto. Specifically, the processor 220 can receive a first grid that divides a document region on a first image including the document region into a plurality of cells. In this case, the learning section 310 can generate a second grid showing a state in which distortion of the first image is corrected, based on the first grid. Specifically, the learning section 310 can generate the second grid including a plurality of cells in a horizontal right angle shape, based on an average of coordinate values of a plurality of edge points of the first grid and an average of a plurality of cell intervals.
[0055] In one embodiment, the learning section 310 can generate a plurality of first patches showing a plurality of cells divided by the first grid. Here, the plurality of first patches can include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the first grid and 4 vertices of the first image. Similarly, the learning section 310 can generate a plurality of second patches showing a plurality of cells divided by the second grid. The plurality of second patches can include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the second grid and 4 vertices of the second image.
[0056] In an embodiment, the learning part 310 can generate a reverse mapping that calculates pixels in the first grid corresponding to each pixel in the second grid. Specifically, the learning part 310 can calculate a transformation matrix that transforms each of the plurality of second patches to each of the plurality of first patches. In addition, the learning part 310 can generate the reverse mapping based on the transformation matrix. Here, the transformation matrix can include an affine transformation matrix, but is not limited thereto.
[0057] In an embodiment, the learning part 310 can train a document image distortion correction model based on the first image and the reverse mapping generated therefrom. Specifically, the learning part 310 can train the document image distortion correction model in a manner that the reverse mapping is generated from the first image. Here, the document image distortion correction model can be an encoder-decoder based deep learning model that can learn to output the reverse mapping based on the first image.
[0058] The reverse mapping generation part 320 can generate a reverse mapping based on the input image. Here, the reverse mapping generation part 320 can include the document image distortion correction model trained by the learning part 310. Alternatively, the reverse mapping generation part 320 can transmit the input image to a document image distortion correction model existing outside the processor 220 and receive a reverse mapping generated by the document image distortion correction model. The reverse mapping generation part 320 can transmit the generated reverse mapping to the image generation part 330.
[0059] The image generation part 330 can generate an image in which distortion of the input image is corrected using the reverse mapping received from the reverse mapping generation part 320. In addition, the image generation part 330 can remove a blank area occurring due to correction of distortion from the image in which distortion is corrected. Furthermore, the image generation part 330 can extract text from the image in which distortion is corrected using a suitable OCR algorithm. In addition, the image generation part 330 can translate the extracted text into text in a language different from the language of the text using a suitable translation algorithm.
[0060] Figure 5 FIG. 4 is a diagram illustrating an example of an internal structure of a document image distortion correction model 420 according to an embodiment of the disclosure. In an embodiment, the document image distortion correction model 420 can receive a first image 410 including a background region and a document region. Here, the document image distortion correction model 420 can include an encoder 422, a classifier 424, and a decoder 426. For example, the document image distortion correction model 420 can be an encoder-decoder based deep learning model based on U-Net, but is not limited thereto.
[0061] In an embodiment, the document image distortion correction model 420 can be trained to output a reverse mapping as an input and an original image as an output. In this case, the document image distortion correction model 420 can encode the first image 410 using the encoder 422. Here, the encoder 422 can be composed of a transformer-based model, but is not limited thereto. In addition, the document image distortion correction model 420 can output the reverse mapping 430 for the first image 410 using the decoder 426. The generated reverse mapping 430 can be used to generate the second image 440 in which distortion of the first image 410 is corrected.
[0062] In an embodiment, the document image distortion correction model 420 can determine whether the first image 410 has distortion using the classifier 424. Specifically, the document image distortion correction model 420 can calculate a distortion score of the first image 410 using the classifier 424. Here, the distortion score can be normalized to a value of 0 to 1, and the closer the distortion score is to 1, the more it indicates that distortion exists in the image.
[0063] In an embodiment, the document image distortion correction model 420 can be trained in advance to output either a reverse mapping or an input image based on the calculated distortion score. Specifically, in the case where the classifier 424 calculates a distortion score lower than a threshold value (for example, 0.5) set in advance, the document image distortion correction model 420 can be trained to output the first image 410. In contrast, in the case where the classifier 424 calculates a distortion score of 0.5 or more set in advance, the document image distortion correction model 420 can be trained to output the reverse mapping 430 for the first image 410 using the decoder 426.
[0064] In an embodiment, the document image distortion correction model 420 can use pairs of an original image and a reverse mapping as ground truth data pairs or training data pairs for training. Here, the original image of the training data pair can include not only an image in which document distortion exists, but also an image in which no document distortion exists. Thus, the document image distortion correction model 420 can infer a reverse mapping even for an image in which no document distortion exists, and the classifier 424 can be trained using the image in which no document distortion exists. In this case, for an image in which no document distortion exists, an identity backward map using a position of a pixel value of the original image can be used. Further, in order to augment training data, the original image and an image in which a grid is displayed can be rotated and flipped to increase the amount of training data.
[0065] In one embodiment, in order to train the decoder 426 that outputs the inverse mapping, an L1 loss (MAE loss) function can be used. In addition, in order to train the classifier 424, a BCE (Binary Cross Entropy) loss function can be used.
[0066] Figure 5 is a diagram illustrating an example of generating a grid 522 in an original image according to one embodiment of the present disclosure. A first example 510 illustrates an example of an original image. Here, the original image can include a document region 512 and a background region 514. In a case where the original image includes a plurality of document regions, only any one of the plurality of document regions can be determined as the document region 512 that is a distortion correction target, and the remaining document regions are determined as the background region 514. In addition, a second example 520 illustrates an example in which the grid 522 is displayed in the original image.
[0067] In one embodiment, the grid 522 can be generated to cover the document region 512. Specifically, a main subject can be determined within the original image. For example, in the first example 510, a white poster can be determined as the main subject. In this case, a region of the maximum size of text that is not cut out among the text present inside the main subject can be determined as a region of the grid 522.
[0068] In one embodiment, the grid 522 can include a plurality of cells. For example, the grid 522 can include 10 x 10 cells by pointing out 11 x 11 points and connecting them. In this case, in order to display the shape of the document distortion, the slope of the horizontal line and the slope of the vertical line of the grid 522 can be determined to correspond to the direction of the text, the direction of the line, and the like included in the document region 512. Thereby, by the grid 522, it is possible to express the distortion of the rotation, the bending, the wrinkle, and the like of the document at the patch level.
[0069] In Figure 6 the grid 522 is illustrated in a size of 10 x 10, but is not limited thereto. For example, the number of cells of the grid can be determined based on the time and the cost of generating the grid. By such a grid 522, the document region 512 is simplified, and thus it is possible to construct a document distortion correction training data set at a reasonable cost.
[0070] Figure 7is a diagram illustrating an example of displaying a plurality of blocks 630 according to one embodiment of the present disclosure. In one embodiment, the plurality of blocks 630 can be generated based on a plurality of points included in the grid 620 displayed in the original image 610. Specifically, the plurality of blocks 630 can include a plurality of triangular blocks generated by applying Delaunay triangulation to the plurality of points included in the grid 620 and 4 vertices of the original image 610. For example, in a case where the grid 620 includes 11 x 11 points, Delaunay triangulation can be performed using a total of 125 points including the 4 vertices of the original image 610. Here, by using the 4 vertices of the original image 610, distortion correction can be performed on the entire original image 610 including the inside of the grid 620 and an outside region. In addition, by the triangular blocks, distortion can be more gently corrected.
[0071] Figure 5 is a diagram illustrating an example of generating a second grid 730 based on a first grid 712 according to one embodiment of the present disclosure. As shown in the first example 710, an example in which the first grid 712 is marked on a first image (or an original image) is illustrated. In addition, the second example 720 illustrates an example in which the second grid 730 is marked on a second image (or a target image). Figure 8
[0072] In one embodiment, the second grid 730 displaying a state in which distortion of the first image is corrected can be generated based on the first grid 712. Specifically, the second grid 730 can be generated in a horizontal right angle shape based on an average of coordinate values of a plurality of edge points of the first grid 712 and an average of intervals of a plurality of cells.
[0073] For example, a width w of the second grid 730 can be an average of a distance between an x-coordinate of an upper left end vertex and an x-coordinate of an upper right end vertex of the first grid 712 and a distance between an x-coordinate of a lower left end vertex and an x-coordinate of a lower right end vertex of the first grid 712. Similarly, a height h of the second grid 730 can be an average of a distance between a y-coordinate of the upper left end vertex and a y-coordinate of the lower left end vertex of the first grid 712 and a distance between a y-coordinate of the upper right end vertex and a y-coordinate of the lower right end vertex of the first grid 712.
[0074] In addition, the coordinates of the vertices 732, 734, 736, 738 of the second grid 730 can be determined based on the first grid 712. Specifically, the x-coordinate of the top-left vertex 732 of the second grid 730 can be the average of the x-coordinates of the left-side vertices of the first grid 712. In addition, the y-coordinate of the top-left vertex 732 of the second grid 730 can be the average of the y-coordinates of the top-side vertices of the first grid 712. After the coordinates of the top-left vertex 732 of the second grid 730 are calculated, the coordinates of the remaining vertices 734, 736, 738 can be calculated using the previously calculated width w and height h of the second grid 730.
[0075] In addition, the cell intervals of the second grid 730 can be determined based on the cell intervals of the first grid 712. Specifically, the intervals of the horizontal lines of the second grid 730 can be calculated according to the interval ratio of the horizontal lines of the first grid 712. Similarly, the intervals of the vertical lines of the second grid 730 can be calculated according to the interval ratio of the vertical lines of the first grid 712.
[0076] Figure 9 is an example diagram illustrating displaying a plurality of blocks 830 according to one embodiment of the present disclosure. In one embodiment, the plurality of blocks 830 can be generated based on the plurality of points included in the grid 820 displayed in the target image 810. Specifically, the plurality of blocks 830 can include a plurality of triangular blocks generated by applying Delaunay triangulation to the plurality of points included in the grid 820 and the 4 vertices of the target image 810. For example, in the case where the grid 820 includes 11x11 points, Delaunay triangulation can be performed using a total of 125 points including the 4 vertices of the target image 810. Here, by using the 4 vertices of the target image 810, distortion correction can be performed on the entire target image 810 including the inside of the grid 820 and the outside region. In addition, by the triangular blocks, distortion can be corrected more gently.
[0077] Figure 10 is an example diagram illustrating calculating inverse mapping according to one embodiment of the present disclosure, Figure 6 is an example diagram illustrating an image whose distortion is corrected according to one embodiment of the present disclosure. The first example 910 illustrates generating a plurality of first blocks 912 for the first grid as Figure 8 illustrated. In addition, the second example 920 illustrates generating a plurality of second blocks 922 for the second grid as Figure 10 illustrated.
[0078] In one embodiment, a reverse mapping of pixels in the first grid corresponding to each pixel in the second grid can be generated. Specifically, a transformation matrix that transforms each of the plurality of second tiles 922 to each of the plurality of first tiles 912 can be calculated. For example, a first transformation matrix that transforms a first pixel 924 in the plurality of second tiles 922 to a first pixel 914 in the plurality of first tiles 912 can be calculated. Similarly, a second transformation matrix and a third transformation matrix that transform a second pixel 926 and a third pixel 928 in the plurality of second tiles 922 to a second pixel 916 and a third pixel 918 in the plurality of first tiles 912, respectively, can be calculated.
[0079] In one embodiment, the reverse mapping can be generated by combining the plurality of transformation matrices generated by the above-described method. That is, by synthesizing the reverse mappings that transform each of the plurality of second tiles 922 to each of the plurality of first tiles 912, a reverse mapping for the entire target image can be generated. In this way, by using the reverse mapping to fill in the values of the pixels of the target image, it is possible to avoid creating holes in all of the pixels in the target image, thereby generating a high-quality image.
[0080] Referring to Figure 10 , a third example 1010 represents an example of generating an image of the values of all of the pixels in the plurality of second tiles 1012 using the reverse mapping for the entire target image. In addition, a fourth example 1020 is an example of an output image. An image in which the distortion of the original image as shown in Figure 11 is corrected can be output. In this case, the image in which the distortion is corrected can not only maintain the document area of the original image, but also maintain the background area. In addition, not only the distortion of the document area can be corrected, but also the distortion of the background area can be corrected.
[0081] With such a configuration, the distortion of the document can be corrected regardless of the form of the document and / or text in the original image. In addition, even if the entire document area cannot be ensured, the distortion can be corrected, and the background information does not disappear, so a high-quality image in which the distortion is corrected can be generated.
[0082] Figure 2 is a diagram illustrating an example of post-processing a generated image according to one embodiment of the present disclosure. In one embodiment, a second image 1120 in which the distortion is corrected can be generated from a first image 1110 input by a user. Specifically, a processor (e.g., a CPU) of an image processing device can calculate a reverse mapping of the first image 1110, and can generate the second image 1120 in which the distortion is corrected using the reverse mapping. Figure 12The processor 220 can receive a first image including a document region. Here, the document region can refer to a text region included in a main subject. In addition, the processor can generate a reverse mapping for correcting distortion of the first image 1110 using a document image distortion correction model trained in advance. Then, the processor can generate a second image 1120 in which distortion of the first image 1110 is corrected based on the generated reverse mapping.
[0083] In an embodiment, the second image 1120 can include an empty region 1122. Such an empty region 1122 can be generated at an edge of the image in the process of correcting distortion of the first image 1120. In this case, the processor can generate a third image 1130 by cropping the empty region 1122. The process of cropping the empty region 1122 can be performed within a limit of not damaging background information or the like of the original image 1110.
[0084] In an embodiment, the processor can extract text included in the second image 1120 or the third image 1130 using an appropriate OCR method. In addition, the processor can translate the extracted text into text in a language different from the language of the text using an appropriate translation method. For example, the processor can generate an image in which Japanese text included in the second image 1120 or the third image 1130 is translated into Korean text.
[0085] Through such a configuration, post-processing is performed on an image in which distortion is corrected, and thus a higher-quality image can be generated. In addition, by extracting text from an image in which distortion is corrected and translating the text, a higher-quality translation service can be provided.
[0086] Figure 13 FIG. 12 is a flowchart illustrating an example of a training method 1200 of a document image distortion correction model according to an embodiment of the disclosure. In an embodiment, the training method 1200 of the document image distortion correction model can be performed using at least one processor. The method 1200 can begin with the processor receiving a first grid that divides a document region on a first image including the document region into a plurality of cells (S1210). Here, a slope of a horizontal line and a slope of a vertical line of the first grid can be determined based on a direction of text included in the document region.
[0087] The processor can then generate a second mesh displaying a state in which distortion of the first image is corrected based on the first mesh (S1220). Specifically, the processor can generate a plurality of first blocks displaying a plurality of cells into which the first mesh is divided. In addition, the processor can generate a second mesh including a plurality of cells in a horizontal right angle shape based on an average of coordinate values of a plurality of edge points of the first mesh and an average of intervals of the plurality of cells. Here, the plurality of first blocks can include a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points included in the first mesh and 4 vertices of the first image.
[0088] The processor can then generate inverse mapping of pixels in the first mesh corresponding to each pixel within the second mesh (S1230). Specifically, a transformation matrix transforming from each of the plurality of second blocks to each of the plurality of first blocks can be calculated. In addition, the processor can generate the inverse mapping based on the transformation matrix.
[0089] The processor can then train a document image distortion correction model in a manner that the inverse mapping is generated from the first image (S1240). Here, the document image distortion correction model can be an encoder-decoder based deep learning model capable of learning to output the inverse mapping based on the first image.
[0090] In an embodiment, the processor can generate a plurality of second blocks displaying a plurality of cells into which the second mesh is divided. Here, the plurality of second blocks can include a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points included in the second mesh and 4 vertices of the second image.
[0091] is an example of a flowchart illustrating a document image distortion correction method 1300 according to an embodiment of the disclosure. In an embodiment, the document image distortion correction method 1300 can be performed by at least one processor. The method 1300 can begin with the processor receiving a first image including a document region (S1310). In this case, the processor can determine any one of a plurality of document regions included in the first image as a document region that is a distortion correction target, in the case where the first image includes a plurality of document regions. In addition, the processor can determine the remaining document regions of the first image as background regions.
[0092] Then, the processor can generate a backward map for correcting distortion of the first image using a document image distortion correction model (S1320). Here, the document image distortion correction model can be a machine learning model trained based on a first mesh dividing a document region within an input image into a plurality of cells, a second mesh showing a state in which distortion of the first image is corrected based on the first mesh, and a pixel within the first mesh corresponding to each pixel within the second mesh.
[0093] Then, the processor can generate a second image in which distortion of the first image is corrected based on the backward map (S1330). In addition, the processor can remove a blank region of the second image. Further, the processor can extract text from a document region of the generated second image. Then, the processor can translate the extracted text into text in a language different from a language of the text.
[0094] In an embodiment, the backward map can be a transformation matrix varying from each of a plurality of second blocks generated by the second mesh to each of a plurality of first blocks generated by the first mesh. In this case, the plurality of first blocks can include a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points included in the first mesh and 4 vertices of the first image. In addition, the plurality of second blocks can include a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points included in the second mesh and 4 vertices of the second image.
[0095] In an embodiment, the document image distortion correction model can include a classifier for determining whether distortion exists in an input image. Here, the classifier can calculate a distortion score of the first image, and the document image distortion correction model can be trained in advance to output any one of the backward map and the first image based on the calculated distortion score.
[0096] The above-described method can be provided in the form of a computer program stored in a computer-readable recording medium in order to be executed in a computer. The medium can continuously store a computer executable program, or temporarily store for execution or download. In addition, the medium can be a plurality of recording mechanisms or storage mechanisms combined in a single or multiple hardware, and is not limited to a medium directly connected to a certain computer system, but can be a medium dispersed on a network. As an example of the medium, there can be a magnetic medium including a hard disk, a floppy disk, and a magnetic tape; an optical recording medium such as a CD-ROM and a DVD; a magneto-optical medium such as a floptical disk; and a medium configured in a manner capable of storing a program command language, such as a ROM, a RAM, a flash memory, and the like. In addition, as an example of other media, a recording medium or a storage medium managed in an application store that distributes an application program, or a website, a server, and the like that supply or distribute other various software can be cited.
[0097] The methods, acts or blocks of the disclosure can be implemented in a variety of ways. For example, such methods can be implemented in hardware, firmware, software, or combinations thereof. As would be apparent to those skilled in the art, the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the present disclosure can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0098] In terms of hardware implementation, the processing units used to implement the techniques can include one or more ASICs, DSPs, GPUs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, electronic devices, other electronic units designed to perform the functions described in the present disclosure, a computer, or a combination thereof.
[0099] Accordingly, various illustrative logical blocks, modules, and circuits described in connection with the disclosure can be implemented or performed with a general purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. Also, the processor can be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0100] In the case of firmware and / or software, the techniques can be embodied as instructions stored in a computer readable medium, which can be read and executed by one or more processors to perform the functions described herein. A computer readable medium can include any memory storage device, including a floppy disk, a flexible disk, a hard disk, a magnetic tape, a compact disk (CD), a digital versatile disk (DVD), a Blu-ray disk, a memory stick, a ROM, a RAM, a PROM, an EPROM, an EEPROM, a flash memory, a solid state drive, a computer readable card, a magnetic or optical data storage device, or any other suitable memory storage device. The computer readable medium can be a computer readable storage medium or a computer readable signal medium. The computer readable storage medium can include any physical medium, including a floppy disk, a flexible disk, a hard disk, a magnetic tape, a compact disk (CD), a digital versatile disk (DVD), a Blu-ray disk, a memory stick, a ROM, a RAM, a PROM, an EPROM, an EEPROM, a flash memory, a solid state drive, a computer readable card, a magnetic or optical data storage device, or any other suitable memory storage device. The computer readable signal medium can include a computer readable medium that stores data for communication to a computer readable storage medium or that stores a computer readable command, data structure, or program code that is accessible by a processor for execution. A computer readable signal medium can include a computer readable medium that is part of a data signal that can be transmitted to a computer readable storage medium, which can be a computer readable storage medium or a computer readable signal medium. The computer readable storage medium and the computer readable signal medium can be the same or different.
[0101] In the case of software, the techniques described herein can be implemented by one or more computer programs executing on one or more computers. The computer programs can be written in any of a number of high level programming languages such as C, C++, Java, Visual Basic, or other suitable programming languages. The computer programs can be stored in any computer readable medium, such as a hard disk, a floppy disk, a compact disk (CD), a digital versatile disk (DVD), a Blu-ray disk, a memory stick, a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), a flash memory, a solid state drive, a computer readable card, a magnetic or optical data storage device, or any other suitable computer readable medium. The computer programs can be distributed in any form over any computer readable medium, including a computer readable storage medium or a computer readable signal medium.
[0102] For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. Disk and disc, as used in this disclosure, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer readable media.
[0103] The software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. The exemplary storage medium can be coupled to the processor, such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.
[0104] The embodiments described above are described as applying the presently disclosed subject matter in one or more standalone computer systems, but the present disclosure is not so limited and can be implemented in connection with any computing environment, such as a network or distributed computing environment. Further, in the present disclosure, the subject matter can be implemented in multiple processing chips or devices, and storage can be similarly affected across multiple devices. Such devices can include PCs, network servers, and portable devices.
[0105] In this specification, the present disclosure is explained based on some embodiments, but various modifications and changes can be made within the scope of the present disclosure that would be understood by those skilled in the art to which the present disclosure pertains without departing from the scope of the present disclosure. In addition, it should be considered that such modifications and changes also belong to the scope of protection claimed in the present specification.
Claims
1. A training method for a document image distortion correction model, wherein the training method for the document image distortion correction model is executed by at least one processor, wherein, Includes the following steps: Receive a first grid that divides the document region on a first image containing the document region into multiple units; A second grid is generated based on the first grid to display the state where the distortion of the first image has been corrected; Generate and calculate the inverse mapping of pixels in the first grid corresponding to each pixel in the second grid; as well as The document image distortion correction model is trained in a manner that generates the reverse mapping from the first image.
2. The training method for the document image distortion correction model according to claim 1, wherein, The slopes of the horizontal and vertical lines of the first grid are determined based on the direction of the text contained in the document area.
3. The training method for the document image distortion correction model according to claim 1, wherein, The step of generating the second grid includes the step of generating a plurality of first blocks that display the first grid as a plurality of subdivided cells.
4. The training method for the document image distortion correction model according to claim 3, wherein, The step of generating the second mesh also includes: The step of generating a second grid comprising multiple cells with horizontal right-angled shapes based on the average of the coordinate values of multiple edge points of the first grid and the average of the intervals between the multiple cells; and The step of generating multiple second blocks that display the divided multiple cells of the second grid.
5. The training method for the document image distortion correction model according to claim 3, wherein, The plurality of first blocks comprise a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points contained in the first grid and four vertices of the first image.
6. The training method for the document image distortion correction model according to claim 4, wherein, The plurality of second blocks comprise a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points contained in the second grid and four vertices of the second image.
7. The training method for the document image distortion correction model according to claim 4, wherein, The steps for generating the reverse mapping include: The steps of calculating the transformation matrix from each of the plurality of second blocks to each of the plurality of first blocks; and The step of generating the reverse mapping based on the transformation matrix.
8. The training method for the document image distortion correction model according to claim 1, wherein, The document image distortion correction model is an encoder-decoder based deep learning model that can learn by outputting the reverse mapping based on the first image.
9. The training method for the document image distortion correction model according to claim 1, wherein, Also includes: In the case where the first image includes multiple document regions, the step of determining any one of the multiple document regions as the document region to be corrected for distortion; and The step of determining the remaining document area of the first image as the background area.
10. A document image distortion correction method, wherein the document image distortion correction method is executed by at least one processor, wherein, include: The step of receiving the first image containing the document region; The step of generating an inverse mapping for correcting the distortion of the first image using a document image distortion correction model; as well as The step of generating a second image based on the reverse mapping, which corrects the distortion of the first image.
11. The document image distortion correction method according to claim 10, wherein, The document image distortion correction model is a machine learning model trained on a first grid that divides the document region in the input image into multiple units, a second grid that generates a state showing the distortion of the first image that has been corrected based on the first grid, and a reverse mapping of the pixels in the first grid corresponding to each pixel in the second grid.
12. The document image distortion correction method according to claim 11, wherein, The reverse mapping is a transformation matrix that transforms each of the plurality of second blocks generated from the second grid to each of the plurality of first blocks generated from the first grid.
13. The document image distortion correction method according to claim 12, wherein, The plurality of first blocks comprise a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points contained in the first grid and four vertices of the first image.
14. The document image distortion correction method according to claim 12, wherein, The plurality of second blocks comprise a plurality of triangular blocks generated by applying Delaunay triangulation to a plurality of points contained in the second grid and four vertices of the second image.
15. The document image distortion correction method according to claim 10, wherein, The document image distortion correction model includes a classifier that determines whether distortion exists in the input image.
16. The document image distortion correction method according to claim 15, wherein, The classifier calculates the distortion score of the first image. The document image distortion correction model is pre-trained to output the inverse mapping and either the first image based on the calculated distortion score.
17. The document image distortion correction method according to claim 10, wherein, It also includes the step of removing blank areas from the second image.
18. The document image distortion correction method according to claim 10, wherein, Also includes: The steps of extracting text from the document region of the generated second image; and The step of translating the extracted text into a language different from the language of the original text.
19. A computer program stored in a computer-readable recording medium for performing the method of any one of claims 1 to 18 on a computer.
20. An information processing system, in, include: Communication module; Memory; and At least one processor is connected to the memory and configured to execute at least one computer-readable program contained in the memory. The at least one program contains a command language for training a document image distortion correction model to perform the following steps: Receive a first grid that divides the document region on a first image containing the document region into multiple units; A second grid is generated based on the first grid to show the state where the distortion of the first image has been corrected; Generate and calculate the inverse mapping of pixels in the first grid corresponding to each pixel in the second grid; and The reverse mapping is generated from the first image.