Method for training document image distortion correction model, and document image distortion correction method and system using same
The document image distortion correction model addresses the limitations of conventional techniques by generating a backward map to correct distortion without requiring the entire document area, ensuring high-quality image correction and background preservation, enhancing OCR and translation services.
Patent Information
- Application Number
- PCT/KR2024/017471
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2024-11-07
- Publication Date
- 2025-07-10
AI Technical Summary
Conventional document image distortion correction techniques require the entire document area to be secured in the image and can lose background information, limiting their effectiveness and applicability.
A method and system for learning a document image distortion correction model that generates a backward map to correct distortion without relying on the document's shape or text, allowing for partial document capture and preserving background information.
Enables high-quality distortion correction and preservation of background information, facilitating improved optical character recognition and translation services.
Smart Images

Figure KR2024017471_10072025_PF_FP_ABST
Abstract
Description
Learning method for document image distortion correction model and document image distortion correction method and system using the same
[0001] The present disclosure relates to a method for learning a document image distortion correction model and a method and system for correcting document image distortion using the same, and more particularly, to a method and system for learning a document image distortion correction model to generate a backward map used for correcting distortion of a document image, and correcting distortion of a document image using the backward map generated by the learned model.
[0002] Recently, with the advancement of smartphone camera performance, users can now use their smartphone cameras instead of optical scanners to capture documents, perform optical character recognition (OCR) on them, or receive translation services. However, depending on the camera's position and pose, distortion can occur in the document within the captured image. Therefore, rectification techniques have been developed to remove this distortion.
[0003] However, using these conventional techniques poses limitations: To remove document distortion within a captured image, the entire document area must be secured, and only a single rectangular subject must be present within the image. Furthermore, document images from which distortion has been removed using these conventional techniques suffer from the loss of background information present in the original captured image.
[0004] The present disclosure provides a learning method for a document image distortion correction model and a document image distortion correction method and device (system) for solving the above-described problems.
[0005] The present disclosure can be implemented in various ways, including as a method, a device (system), or a computer program stored on a readable storage medium.
[0006] According to one embodiment of the present disclosure, a method for learning a document image distortion correction model may include the steps of receiving a first grid that divides a document area into a plurality of cells on a first image including a document area, generating a second grid that represents a state in which distortion of the first image is corrected based on the first grid, generating a backward map that calculates a pixel in the first grid corresponding to each pixel in the second grid, and training the document image distortion correction model to generate the backward map from the first image.
[0007] According to one embodiment of the present disclosure, a method for correcting distortion of a document image may include the steps of receiving a first image including a document area, generating a backward map for correcting distortion of the first image using a document image distortion correction model, and generating a second image in which distortion of the first image is corrected based on the backward map.
[0008] A computer program stored in a computer-readable recording medium may be provided to execute a method according to one embodiment of the present disclosure on a computer.
[0009] An information processing system according to one embodiment of the present disclosure comprises a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program may include instructions for receiving a first grid dividing a document area into a plurality of cells on a first image including a document area, generating a second grid representing a state in which distortion of the first image is corrected based on the first grid, generating a backward map calculating pixels in the first grid corresponding to each pixel in the second grid, and training a document image distortion correction model to generate the backward map from the first image.
[0010] According to some embodiments of the present disclosure, a document image distortion correction model can be trained to output a backward map that can be used to correct distortion in an original image. In this case, by using a backward map that defines a transformation relationship inversely to a forward map that defines a transformation relationship from an input image to a distortion-corrected output image, high-quality image distortion correction can be performed by preventing a phenomenon in which some pixels are not filled in a distortion-corrected image.
[0011] According to some embodiments of the present disclosure, distortion of a document can be corrected without relying on the shape of the document and / or text within the original image. Furthermore, even if the entire document area is not captured from the input image, distortion existing in the document area can be corrected. Since background information surrounding the document area is not lost even after distortion correction, a high-quality, distortion-corrected image can be generated.
[0012] According to some embodiments of the present disclosure, higher-quality images can be generated by post-processing distortion-corrected images. Furthermore, by extracting and translating text from the distortion-corrected images using a backward map, a higher-quality translation service can be provided.
[0013] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs (referred to as “one skilled in the art”) from the description of the claims.
[0014] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, wherein like reference numerals represent similar elements, but are not limited thereto.
[0015] FIG. 1 illustrates an example of a learning method for a document image distortion correction model according to one embodiment of the present disclosure.
[0016] FIG. 2 is a block diagram showing the internal configuration of an information processing system according to one embodiment of the present disclosure.
[0017] FIG. 3 is a diagram showing the internal configuration of a processor of an information processing system according to one embodiment of the present disclosure.
[0018] FIG. 4 is a diagram illustrating an example of the internal configuration of a document image distortion correction model according to one embodiment of the present disclosure.
[0019] FIG. 5 is a diagram illustrating an example of creating a grid in an original image according to one embodiment of the present disclosure.
[0020] FIG. 6 is a drawing showing an example in which multiple patches are displayed according to one embodiment of the present disclosure.
[0021] FIG. 7 is a diagram illustrating an example in which a second grid is generated based on a first grid according to one embodiment of the present disclosure.
[0022] FIG. 8 is a drawing showing an example in which multiple patches are displayed according to one embodiment of the present disclosure.
[0023] FIG. 9 is a diagram illustrating an example of generating a backward map according to one embodiment of the present disclosure.
[0024] FIG. 10 is a drawing showing an example of an image in which distortion has been corrected according to one embodiment of the present disclosure.
[0025] FIG. 11 is a diagram illustrating an example of post-processing an image generated according to one embodiment of the present disclosure.
[0026] FIG. 12 is a flowchart illustrating an example of a learning method for a document image distortion correction model according to one embodiment of the present disclosure.
[0027] FIG. 13 is a flowchart illustrating an example of a document image distortion correction method according to one embodiment of the present disclosure.
[0028] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions of widely known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.
[0029] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the description of the embodiments below, duplicate descriptions of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.
[0030] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the embodiments described below, along with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure the completeness of the disclosure and to fully inform those skilled in the art of the scope of the invention.
[0031] The terms used in this specification will be briefly explained, followed by a detailed description of the disclosed embodiments. The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.
[0032] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.
[0033] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.
[0034] According to one embodiment of the present disclosure, a 'module' or 'unit' may be implemented as a processor and a memory. 'Processor' should be broadly construed to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some circumstances, a 'processor' may also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), and the like. A 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination of configurations. In addition, 'memory' should be broadly construed to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.
[0035] In the present disclosure, the "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be comprised of one or more server devices. As another example, the system may be comprised of one or more cloud devices. As yet another example, the system may be configured and operated by a combination of a server device and a cloud device.
[0036] In the present disclosure, 'display' may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided from the computing device.
[0037] In the present disclosure, 'each of the plurality of As' or 'each of the plurality of As' may refer to each of all components included in the plurality of As, or may refer to each of some components included in the plurality of As.
[0038] In this disclosure, a "machine learning model" may include any model used to infer an answer to a given input. In one embodiment, the machine learning model may include an artificial neural network model including an input layer, multiple hidden layers, and an output layer. Here, each layer may include multiple nodes. In this disclosure, each of the multiple machine learning models is described as a separate machine learning model, but this is not limited thereto, and some or all of the multiple machine learning models may be implemented as a single machine learning model. Furthermore, a single machine learning model may include multiple machine learning models. In this disclosure, the terms "machine learning model" and "artificial neural network model" may be used interchangeably to refer to the same or similar models.
[0039] In the present disclosure, a 'document image' may include an image created by photographing or scanning a document in printed form, an image related to an electronic document (e.g., an electronic receipt, an electronic business card, an electronic ID card, etc.) created by a program, etc.
[0040] In the present disclosure, the term "document area" may refer to an area within a document image that contains text and is the subject of character recognition or translation. Furthermore, the term "background area" in the present disclosure may refer to an area within a document image excluding the document area.
[0041] FIG. 1 illustrates an example of a learning method for a document image distortion correction model according to one embodiment of the present disclosure. In one embodiment, a grid may be generated (120) on an original image (110) including a document area. Specifically, a first grid (122) may be generated that divides the document area of the original image (110) into a plurality of cells. Here, the inclination of the horizontal lines and the inclination of the vertical lines of the first grid (122) may be determined based on the direction of the text included in the document area. The first grid (122) generated in this way may roughly express distortions such as rotation, warping, and crumpling of the document area included in the original image (110) at a patch level. For example, the first grid (122) may include cells divided into a 10 x 10 grid, but is not limited thereto. The generation (120) of the first grid (122) may be performed using an image editing application or an appropriate grid generation algorithm.
[0042] In addition, the generated first grid (122) may be transformed to generate a second grid (132) (130). Here, the second grid (132) may express a state in which the distortion of the document area expressed by the first grid (122) is corrected at a patch level. The second grid (132) may be generated to include a plurality of cells in a horizontal right-angled shape based on an average of the coordinate values of the edge points of the first grid (122) and an average of a plurality of cell spacings.
[0043] In one embodiment, a plurality of patches may be generated (140) from each of the first grid (122) and the second grid (132). Specifically, a plurality of first patches representing a plurality of divided cells of the first grid (122) may be generated. Here, the plurality of first patches (142) may include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the first grid (122) and four vertices of the original image (110). Similarly, a plurality of second patches (144) representing a plurality of divided cells of the second grid (132) may be generated. Here, the plurality of second patches (144) may include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the second grid (132) and four vertices of the distortion-corrected image of the original image (110).
[0044] In one embodiment, a backward map may be generated (150) based on a plurality of first patches (142) and a plurality of second patches (144). Specifically, a transformation matrix defining a transformation relationship from each of the plurality of second patches to each of the plurality of corresponding first patches may be generated. A backward map for the entire image may be generated by combining the plurality of transformation matrices thus generated. Accordingly, the backward map may generate a pixel within the first grid (122) of the image before distortion correction corresponding to each pixel within the second grid (132) of the image after distortion correction.
[0045] In one embodiment, a document image distortion correction model may be trained to generate a backward map (160) from an original image (110). That is, the document image distortion correction model may be trained to use a pair of an original image (110) and an associated backward map (160) as training data to output a backward map capable of correcting distortion in an input image. Here, the document image distortion correction model may be, but is not limited to, an encoder-decoder-based deep learning model capable of being trained to output a backward map for correcting distortion of an image based on an input image.
[0046] The document image distortion correction model learned through this configuration can generate a backward map used to correct distortion in document regions from an input image. Furthermore, by applying the backward map to the input image, an output image with the distortion corrected can be generated. The distorted output image can be used in subsequent processes such as optical character recognition (OCR) or translation.
[0047] FIG. 2 is a block diagram illustrating the internal configuration of an information processing system (200) according to one embodiment of the present disclosure. The information processing system (200) may include a memory (210), a processor (220), a communication module (230), and an input / output interface (240). The information processing system (200) may be configured to communicate information and / or data with an external system via a network using the communication module (230).
[0048] The memory (210) may include any non-transitory computer-readable recording medium. According to one embodiment, the memory (210) may include a permanent mass storage device such as a read-only memory (ROM), a disk drive, a solid state drive (SSD), a flash memory, etc. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the information processing system (200) as a separate permanent storage device distinct from the memory. In addition, the memory (210) may store an operating system and at least one program code (e.g., code for executing document distortion correction, etc. installed and operated in the information processing system (200).
[0049] These software components may be loaded from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a recording medium directly connectable to the information processing system (200), for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory (210) through a communication module (230) other than a computer-readable recording medium. For example, at least one program may be loaded into the memory (210) based on a computer program (e.g., a program for learning a document distortion correction model, a program for executing document distortion correction, etc.) that is installed by files provided by developers or a file distribution system that distributes installation files of applications through the communication module (230).
[0050] The processor (220) may be configured to process commands of a computer program by performing basic arithmetic, logic, and input / output operations. The commands may be provided to a user terminal (not shown) or another external system via the memory (210) or the communication module (230). For example, the processor (220) may generate a second grid representing a state in which distortion of a first image is corrected based on a first grid, generate a backward map that calculates pixels in the first grid corresponding to each pixel in the second grid, and train a document image distortion correction model to generate the backward map from the first image. In addition, the processor (220) may generate a second image in which distortion of the first image is corrected using the backward map generated by the document image distortion correction model based on the first image. Additionally, the processor (220) of the information processing system (200) may be configured to manage, process, and / or store information and / or data received from a plurality of user terminals and / or a plurality of external systems.
[0051] The communication module (230) may provide a configuration or function for a user terminal (not shown) and the information processing system (200) to communicate with each other via a network, and may provide a configuration or function for the information processing system (200) to communicate with an external system (e.g., a separate cloud system, etc.). For example, control signals, commands, data, etc. provided under the control of the processor (220) of the information processing system (200) may be transmitted to the user terminal and / or the external system via the communication module (230) and the network through the communication module of the user terminal and / or the external system. For example, the information processing system (200) may transmit an image with document distortion corrected to the user terminal via the communication module (230).
[0052] In addition, the input / output interface (240) of the information processing system (200) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (200) or that the information processing system (200) may include. In FIG. 2, the input / output interface (240) is illustrated as an element configured separately from the processor (220), but is not limited thereto, and the input / output interface (240) may be configured to be included in the processor (220). The information processing system (200) may include more components than those illustrated in FIG. 2. However, there is no need to explicitly illustrate most of the conventional technology components.
[0053] FIG. 3 is a diagram illustrating the internal configuration of a processor (220) of an information processing system according to one embodiment of the present disclosure. According to one embodiment, the processor (220) may include a learning unit (310), a backward map generation unit (320), and an image generation unit (330). The internal configuration of the processor (220) of the information processing system illustrated in FIG. 3 is merely an example and may be implemented differently. For example, at least a portion of the configuration of the processor (220) may be omitted, other configurations may be added, and at least a portion of the operations or processes performed by the processor (220) may be performed by a processor of a user terminal that is communicatively connected to the information processing system.
[0054] The learning unit (310) can train a document image distortion correction model using learning data including one or more document images and one or more backward maps corresponding thereto. Specifically, the processor (220) can receive a first grid that divides a document area into a plurality of cells on a first image including a document area. In this case, the learning unit (310) can generate a second grid representing a state in which the distortion of the first image is corrected based on the first grid. Specifically, the learning unit (310) can generate a second grid including a plurality of cells in a horizontally orthogonal shape based on an average of the coordinate values of the edge points of the first grid and an average of the intervals between the plurality of cells.
[0055] In one embodiment, the learning unit (310) may generate a plurality of first patches representing a plurality of segmented cells of a first grid. Here, the plurality of first patches may include a plurality of triangle patches generated by applying Delaunay triangulation to a plurality of points included in the first grid and four vertices of the first image. Similarly, the learning unit (310) may generate a plurality of second patches representing a plurality of segmented cells of a second grid. The plurality of second patches may include a plurality of triangle patches generated by applying Delaunay triangulation to a plurality of points included in the second grid and four vertices of the second image.
[0056] In one embodiment, the learning unit (310) may generate a backward map that calculates a pixel in the first grid corresponding to each pixel in the second grid. Specifically, the learning unit (310) may generate a transition matrix from each of the plurality of second patches to each of the plurality of first patches. In addition, the learning unit (310) may generate a backward map based on the transition matrix. Here, the transition matrix may include, but is not limited to, an affine transformation matrix.
[0057] In one embodiment, the learning unit (310) may train a document image distortion correction model based on a first image and a backward map generated therefrom. Specifically, the learning unit (310) may train the document image distortion correction model to generate a backward map from the first image. Here, the document image distortion correction model may be an encoder-decoder-based deep learning model that can be trained to output a backward map based on the first image.
[0058] The backward map generation unit (320) can generate a backward map based on the input image. Here, the backward map generation unit (320) can include a document image distortion correction model learned by the learning unit (310). Alternatively, the backward map generation unit (320) can transmit the input image to a document image distortion correction model existing outside the processor (220) and receive a backward map generated from the document image distortion correction model. The backward map generation unit (320) can transmit the generated backward map to the image generation unit (330).
[0059] The image generation unit (330) can generate an image that corrects the distortion of the input image using the backward map received from the backward map generation unit (320). In addition, the image generation unit (330) can remove blank areas caused by distortion correction from the distortion-corrected image. Additionally, the image generation unit (330) can extract text from the distortion-corrected image using an appropriate OCR algorithm. In addition, the image generation unit (330) can translate the extracted text into text in a language different from the language of the text using an appropriate translation algorithm.
[0060] FIG. 4 is a diagram illustrating an example of an internal configuration of a document image distortion correction model (420) according to one embodiment of the present disclosure. In one embodiment, the document image distortion correction model (420) may receive a first image (410) including a background area and a document area. Here, the document image distortion correction model (420) may include an encoder (422), a classifier (424), and a decoder (426). For example, the document image distortion correction model (420) may be a deep learning model based on an encoder-decoder based on a U-Net, but is not limited thereto.
[0061] In one embodiment, the document image distortion correction model (420) can be trained to output a backward map by taking an original image as input. In this case, the document image distortion correction model (420) can encode the first image (410) using an encoder (422). Here, the encoder (422) can be configured as a transformer-based model, but is not limited thereto. In addition, the document image distortion correction model (420) can output a backward map (430) for the first image (410) using a decoder (426). The generated backward map (430) can be used to generate a second image (440) that corrects the distortion of the first image (410).
[0062] In one embodiment, the document image distortion correction model (420) can determine whether there is distortion in the first image (410) using the classifier (424). Specifically, the document image distortion correction model (420) can calculate a distortion score of the first image (410) using the classifier (424). Here, the distortion score can be normalized to have a value between 0 and 1, and a distortion score closer to 1 can indicate that there is distortion in the image.
[0063] In one embodiment, the document image distortion correction model (420) may be pre-trained to output either a backward map or an input image based on the calculated distortion score. Specifically, if the classifier (424) calculates a distortion score lower than a preset threshold (e.g., 0.5), the document image distortion correction model (420) may be trained to output the first image (410). Conversely, if the classifier (424) calculates a distortion score higher than the preset threshold (e.g., 0.5), the document image distortion correction model (420) may be trained to output a backward map (430) for the first image (410) using the decoder (426).
[0064] In one embodiment, the document image distortion correction model (420) can learn an original image and a backward map pair as a ground truth data pair or a training data pair. Here, the original images of the training data pair may include not only images with document distortion but also images without document distortion. Accordingly, the document image distortion correction model (420) can infer a backward map even for images without document distortion, and the classifier (424) can be trained using images without document distortion. In this case, for images without document distortion, an identity backward map using the locations of pixel values of the original image can be used. Additionally, for training data augmentation, the original image and the image with the grid displayed can be rotated and flipped to increase the amount of training data.
[0065] In one embodiment, an L1 loss (MAE loss) function may be used to train a decoder (426) that outputs a backward map. Additionally, a BCE (Binary Cross Entropy) loss function may be used to train a classifier (424).
[0066] FIG. 5 is a diagram illustrating an example of creating a grid (522) in an original image according to one embodiment of the present disclosure. A first example (510) illustrates an example of an original image. Here, the original image may include a document area (512) and a background area (514). If the original image includes multiple document areas, only one of the multiple document areas may be determined as the document area (512) to be subject to distortion correction, and the remaining document areas may be determined as the background area (514). In addition, a second example (520) illustrates an example of displaying a grid (522) in an original image.
[0067] In one embodiment, a grid (522) may be generated to cover the document area (512). Specifically, a main subject may be determined within the original image. For example, in the first example (510), a white poster may be determined as the main subject. In this case, the largest area of uncut text within the main subject may be determined as the area of the grid (522).
[0068] In one embodiment, the grid (522) may include a plurality of cells. For example, the grid (522) may include 10 x 10 cells by connecting 11 x 11 dots. In this case, the inclination of the horizontal lines and the inclination of the vertical lines of the grid (522) may be determined to correspond to the direction of the text, the direction of the lines, etc. included in the document area (512) to indicate the form of the document distortion. Accordingly, distortions such as rotation, bending, and crumpling of the document may be expressed at a patch level through the grid (522).
[0069] Although the grid (522) is illustrated as having a size of 10 x 10 in FIG. 5, this is not a limitation. For example, the number of grid cells may be determined based on the time and cost of generating the grid. By simplifying the document area (512) through this grid (522), a document distortion correction learning dataset can be constructed at a reasonable cost.
[0070] FIG. 6 is a diagram illustrating an example in which a plurality of patches (630) are displayed according to one embodiment of the present disclosure. In one embodiment, a plurality of patches (630) may be generated based on a plurality of points included in a grid (620) displayed in an original image (610). Specifically, the plurality of patches (630) may include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the grid (620) and four vertices of the original image (610). For example, when the grid (620) includes 11 x 11 points, Delaunay triangulation may be performed using a total of 125 points including the four vertices of the original image (610). Here, by using the four vertices of the original image (610), distortion correction may be performed on the entire original image (610), including not only the inside of the grid (620) but also the outside area. Additionally, distortion can be more smoothly corrected through triangle patches.
[0071] FIG. 7 is a diagram illustrating an example of generating a second grid (730) based on a first grid (712) according to one embodiment of the present disclosure. The first example (710) illustrates an example in which the first grid (712) is displayed on a first image (or original image) as described above with reference to FIG. 5. In addition, the second example (720) illustrates an example in which the second grid (730) is displayed on a second image (or target image).
[0072] In one embodiment, a second grid (730) representing a state in which distortion of the first image is corrected may be generated based on the first grid (712). Specifically, the second grid (730) may be generated in a horizontal right-angled shape based on an average of the coordinate values of edge points of the first grid (712) and an average of the spacing between a plurality of cells.
[0073] For example, the width (w) of the second grid (730) may be an average of the distance between the x-coordinate of the upper left vertex and the x-coordinate of the upper right vertex of the first grid (712) and the distance between the x-coordinate of the lower left vertex and the x-coordinate of the lower right vertex of the first grid (712). Similarly, the height (h) of the second grid (730) may be an average of the distance between the y-coordinate of the upper left vertex and the y-coordinate of the lower left vertex of the first grid (712) and the distance between the y-coordinate of the upper right vertex and the y-coordinate of the lower right vertex of the first grid (712).
[0074] In addition, based on the first grid (712), the coordinates of each vertex (732, 734, 736, 738) of the second grid (730) can be determined. Specifically, the x-coordinate of the upper left vertex (732) of the second grid (730) can be the average of the x-coordinates of the left vertices of the first grid (712). In addition, the y-coordinate of the upper left vertex (732) of the second grid (730) can be the average of the y-coordinates of the upper vertices of the first grid (712). After the coordinates of the upper left vertex (732) of the second grid (730) are calculated, the coordinates of the remaining vertices (734, 736, 738) can be calculated using the width (w) and height (h) of the second grid (730) calculated previously.
[0075] Additionally, the plurality of cell spacings of the second grid (730) may be determined based on the plurality of cell spacings of the first grid (712). Specifically, the spacing of horizontal lines of the second grid (730) may be calculated based on the spacing ratio of horizontal lines of the first grid (712). Similarly, the spacing of vertical lines of the second grid (730) may be calculated based on the spacing ratio of vertical lines of the first grid (712).
[0076] FIG. 8 is a diagram illustrating an example in which a plurality of patches (830) are displayed according to one embodiment of the present disclosure. In one embodiment, a plurality of patches (830) may be generated based on a plurality of points included in a grid (820) displayed in a target image (810). Specifically, the plurality of patches (830) may include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the grid (820) and four vertices of the target image (810). For example, when the grid (820) includes 11 x 11 points, Delaunay triangulation may be performed using a total of 125 points including the four vertices of the target image (810). Here, by using the four vertices of the target image (810), distortion correction may be performed on the entire target image (810), including not only the inside of the grid (820) but also the outside area. Additionally, distortion can be more smoothly corrected through triangle patches.
[0077] FIG. 9 is a diagram illustrating an example of generating a backward map according to one embodiment of the present disclosure, and FIG. 10 is a diagram illustrating an example of an image whose distortion has been corrected according to one embodiment of the present disclosure. The first example (910) illustrates an example in which a plurality of first patches (912) are generated for a first grid, as described above in FIG. 6. In addition, the second example (920) illustrates an example in which a plurality of second patches (922) are generated for a second grid, as described above in FIG. 8.
[0078] In one embodiment, a backward map may be generated that calculates a pixel in the first grid corresponding to each pixel in the second grid. Specifically, a transition matrix may be calculated from each of the plurality of second patches (922) to each of the plurality of first patches (912). For example, a first transition matrix may be calculated in which a first pixel (924) of the plurality of second patches (922) is converted to a first pixel (914) of the plurality of first patches (912). Similarly, a second transition matrix and a third transition matrix may be calculated in which a second pixel (926) and a third pixel (928) of the plurality of second patches (922) are converted to a second pixel (916) and a third pixel (918) of the plurality of first patches (912), respectively.
[0079] In one embodiment, a backward map can be generated by combining multiple transition matrices generated by the method described above. That is, by synthesizing backward maps that transition from each of the plurality of second patches (922) to each of the plurality of first patches (912), a backward map for the entire target image can be generated. In this way, by using the backward map to fill each pixel of the target image, no holes can be generated for any pixel within the target image, resulting in a high-quality image.
[0080] Referring to FIG. 10, a third example (1010) illustrates an example of an image in which values of all pixels within a plurality of second patches (1012) are generated using a backward map for the entire target image. In addition, a fourth example (1020) illustrates an example of an output image. As illustrated in FIG. 10, an image in which distortion of an original image is corrected can be output. In this case, the distortion-corrected image can maintain not only the document area of the original image but also the background area. In addition, distortion in the background area as well as the document area can be corrected.
[0081] This configuration allows for correction of document distortion without relying on the shape of the document and / or text within the original image. Furthermore, distortion can be corrected even if the entire document area is not captured, and background information is not lost, resulting in high-quality, distortion-corrected images.
[0082] FIG. 11 is a diagram illustrating an example of post-processing an image generated according to one embodiment of the present disclosure. In one embodiment, a second image (1120) with distortion corrected may be generated from a first image (1110) input by a user. Specifically, a processor (e.g., 220 of FIG. 2) may receive a first image including a document area. Here, the document area may refer to a text area included in the main subject. In addition, the processor may generate a backward map for correcting distortion of the first image (1110) using a pre-learned document image distortion correction model. Thereafter, the processor may generate a second image (1120) with distortion corrected of the first image (1110) based on the generated backward map.
[0083] In one embodiment, the second image (1120) may include a blank area (1122). This blank area (1122) may be generated at the edge of the image while correcting distortion of the first image (1120). In this case, the processor may generate the third image (1130) by cropping the blank area (1122). The process of cropping the blank area (1122) may be performed as long as it does not damage background information, etc., of the original image (1110).
[0084] In one embodiment, the processor may extract text contained in the second image (1120) or the third image (1130) using an appropriate OCR method. Furthermore, the processor may translate the extracted text into a language different from the language of the text using an appropriate translation method. For example, the processor may generate an image that translates Japanese text contained in the second image (1120) or the third image (1130) into Korean text.
[0085] This configuration allows for the creation of higher-quality images by post-processing the distortion-corrected images. Furthermore, by extracting and translating text from the distortion-corrected images, a higher-quality translation service can be provided.
[0086] FIG. 12 is a flowchart illustrating an example of a method (1200) for training a document image distortion correction model according to one embodiment of the present disclosure. In one embodiment, the method (1200) for training a document image distortion correction model may be performed by at least one processor. The method (1200) may begin with the processor receiving a first grid dividing a document area into a plurality of cells on a first image including a document area (S1210). Here, the inclination of horizontal lines and the inclination of vertical lines of the first grid may be determined based on the direction of text included in the document area.
[0087] Thereafter, the processor may generate a second grid representing a state in which the distortion of the first image is corrected based on the first grid (S1220). Specifically, the processor may generate a plurality of first patches representing a plurality of divided cells of the first grid. In addition, the processor may generate a second grid including a plurality of cells in a horizontally orthogonal shape based on an average of coordinate values of edge points of the first grid and an average of spacings of the plurality of cells. Here, the plurality of first patches may include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the first grid and four vertices of the first image.
[0088] Thereafter, the processor may generate a backward map that calculates a pixel in the first grid corresponding to each pixel in the second grid (S1230). Specifically, a transition matrix may be calculated from each of the plurality of second patches to each of the plurality of first patches. Furthermore, the processor may generate a backward map based on the transition matrix.
[0089] Thereafter, the processor may train a document image distortion correction model to generate a backward map from the first image (S1240). Here, the document image distortion correction model may be a deep learning model based on an encoder-decoder that can be trained to output a backward map based on the first image.
[0090] In one embodiment, the processor may generate a plurality of second patches representing a plurality of divided cells of the second grid, wherein the plurality of second patches may include a plurality of triangular patches generated by applying a Delaunay triangulation to a plurality of points included in the second grid and four vertices of the second image.
[0091] FIG. 13 is a flowchart illustrating an example of a document image distortion correction method (1300) according to one embodiment of the present disclosure. In one embodiment, the document image distortion correction method (1300) may be performed by at least one processor. The method (1300) may begin with the processor receiving a first image including a document area (S1310). In this case, if the first image includes multiple document areas, the processor may determine one of the multiple document areas as a document area to be subject to distortion correction. Additionally, the processor may determine the remaining document areas of the first image as background areas.
[0092] Thereafter, the processor may generate a backward map for correcting the distortion of the first image using a document image distortion correction model (S1320). Here, the document image distortion correction model may be a machine learning model learned based on a first grid that divides a document area in an input image into a plurality of cells, a second grid that represents a state in which the distortion of the first image generated based on the first grid is corrected, and a backward map that calculates a pixel in the first grid corresponding to each pixel in the second grid.
[0093] Thereafter, the processor can generate a second image that corrects the distortion of the first image based on the backward map (S1330). Furthermore, the processor can remove blank areas of the second image. Additionally, the processor can extract text from the document area of the generated second image. Thereafter, the processor can translate the extracted text into a language different from the language of the text.
[0094] In one embodiment, the backward map may be a transition matrix from each of a plurality of second patches generated from the second grid to each of a plurality of first patches generated from the first grid. In this case, the plurality of first patches may include a plurality of triangle patches generated by applying a Delaunay triangulation to a plurality of points included in the first grid and four vertices of the first image. Furthermore, the plurality of second patches may include a plurality of triangle patches generated by applying a Delaunay triangulation to a plurality of points included in the second grid and four vertices of the second image.
[0095] In one embodiment, the document image distortion correction model may include a classifier that determines whether an input image has distortion. Here, the classifier may calculate a distortion score of the first image, and the document image distortion correction model may be pre-trained to output either the backward map or the first image based on the calculated distortion score.
[0096] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0097] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software will depend on the particular application and the design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be construed as departing from the scope of the present disclosure.
[0098] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, GPUs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, a computer, or a combination thereof.
[0099] Accordingly, the various exemplary logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0100] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, a compact disc (CD), a magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described herein.
[0101] When implemented in software, the techniques may be stored on or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is suitably made to a computer-readable medium.
[0102] For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. Disk and disc, as used herein, includes compact discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks usually reproduce data magnetically, whereas discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0103] A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in the user terminal.
[0104] While the embodiments described above have been described as utilizing aspects of the presently disclosed subject matter in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the present disclosure may be implemented in multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include personal computers, network servers, and portable devices.
[0105] While the present disclosure has been described in connection with certain embodiments herein, various modifications and variations may be made without departing from the scope of the present disclosure, as would be understood by those skilled in the art. Furthermore, such modifications and variations are intended to fall within the scope of the claims appended to this specification.
Claims
1. A method for learning a document image distortion correction model, performed by at least one processor, A step of receiving a first grid dividing the document area into a plurality of cells on a first image including a document area; A step of generating a second grid representing a state in which the distortion of the first image is corrected based on the first grid; A step of generating a backward map that calculates a pixel in the first grid corresponding to each pixel in the second grid; and A step of training a document image distortion correction model to generate the backward map from the first image. A learning method for a document image distortion correction model, including:
2. In paragraph 1, A learning method for a document image distortion correction model, wherein the inclination of the horizontal line and the inclination of the vertical line of the first grid are determined based on the direction of the text included in the document area.
3. In paragraph 1, The step of generating the second grid is: A step of generating a plurality of first patches representing a plurality of divided cells of the first grid. A learning method for a document image distortion correction model, including:
4. In paragraph 3, The step of generating the second grid is: A step of generating the second grid including a plurality of cells in a horizontal right-angled shape based on the average of the coordinate values of the edge points of the first grid and the average of the spacings of the plurality of cells; and A step of generating a plurality of second patches representing a plurality of divided cells of the second grid. A learning method of a document image distortion correction model, which further includes:
5. In paragraph 3, A learning method for a document image distortion correction model, wherein the plurality of first patches include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the first grid and four vertices of the first image.
6. In paragraph 4, A learning method for a document image distortion correction model, wherein the plurality of second patches include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the second grid and four vertices of the second image.
7. In paragraph 4, The steps for generating the above backward map are: A step of calculating a transition matrix from each of the plurality of second patches to each of the plurality of first patches; and A step of generating the backward map based on the above transition matrix. A learning method for a document image distortion correction model, including:
8. In paragraph 1, A learning method for a document image distortion correction model, wherein the document image distortion correction model is a deep learning model based on an encoder-decoder that can be learned to output the backward map based on the first image.
9. In paragraph 1, When the first image includes a plurality of document areas, a step of determining one of the plurality of document areas as a document area to be subject to distortion correction; and Step of determining the remaining document area of the above first image as a background area A learning method of a document image distortion correction model, which further includes:
10. A method for correcting distortion of a document image, performed by at least one processor, A step of receiving a first image including a document area; A step of generating a backward map for correcting distortion of the first image using a document image distortion correction model; and A step of generating a second image in which the distortion of the first image is corrected based on the above backward map. A method for correcting distortion of a document image, comprising:
11. In paragraph 10, The above document image distortion correction model is, A method for correcting distortion of a document image, comprising: a first grid dividing a document area within an input image into a plurality of cells; a second grid representing a state in which distortion of the first image generated based on the first grid is corrected; and a machine learning model learned based on a backward map that calculates a pixel within the first grid corresponding to each pixel within the second grid.
12. In paragraph 11, A method for correcting distortion of a document image, wherein the backward map is a transition matrix from each of a plurality of second patches generated from the second grid to each of a plurality of first patches generated from the first grid.
13. In paragraph 12, A method for correcting distortion of a document image, wherein the plurality of first patches include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the first grid and four vertices of the first image.
14. In paragraph 12, A method for correcting distortion of a document image, wherein the plurality of second patches include a plurality of triangular patches generated by applying Delaunay triangulation to a plurality of points included in the second grid and four vertices of the second image.
15. In paragraph 10, A method for correcting distortion of a document image, wherein the above document image distortion correction model includes a classifier that determines whether or not there is distortion in an input image.
16. In paragraph 15, The above classifier calculates a distortion score of the first image, A method for correcting distortion of a document image, wherein the document image distortion correction model is pre-learned to output either the backward map or the first image based on the calculated distortion score.
17. In paragraph 10, Step of removing blank areas of the above second image A method for correcting distortion of a document image, the method further comprising:
18. In paragraph 10, A step of extracting text from the document area of the second image generated above; and A step of translating the extracted text into a text in a language different from the language of the text. A method for correcting distortion of a document image, the method further comprising:
19. A computer program stored on a computer-readable recording medium for executing the method according to any one of claims 1 to 18 on a computer.
20. As an information processing system, Communication module; memory; and At least one processor coupled to said memory and configured to execute at least one computer-readable program contained in said memory, At least one of the above programs, Receive a first grid dividing the document area into a plurality of cells on a first image including a document area, Based on the first grid, a second grid is generated to represent a state in which the distortion of the first image is corrected, Generate a backward map that produces a pixel in the first grid corresponding to each pixel in the second grid, An information processing system comprising commands for training a document image distortion correction model to generate the backward map from the first image.
Citation Information
Patent Citations
Method for correcting crease document image based on multiple views
CN115311160A
Sample warping document image generation method and device, equipment and medium
CN116523736A
Method and system for correcting distorted document image
JP2010171976A
Apparatus and Method for Geometric DistortionCorrection of Document Image using Affine Transform
KR1020060033973A
Mesh for rendering an image frame
US20070291233A1
Cited By
Camera image correction method and system based on artificial intelligence
CN121074348A