Image correction, identification and information retrieval method and related device
By detecting the single-page document area and adjusting the mapping information using the physical deformation correction model, the image distortion problem caused by the physical deformation of paper documents is solved, and the effects of image recognition and information retrieval are improved.
Patent Information
- Application Number
- CN202510843847.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies are unable to effectively correct image distortion caused by the physical deformation of paper documents, which affects the accuracy of image recognition and information retrieval.
By detecting the single-page area in the document image, the physical deformation correction model is used to learn the mapping information of the document, the mapping relationship of the pixel points is adjusted, and the image is restored to a flat state by combining perspective transformation and interpolation algorithm.
It improves the accuracy of image recognition and the efficiency of information retrieval, effectively corrects complex nonlinear deformations such as wrinkles and bulges, and improves image quality.
Smart Images

Figure CN120673425A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and more specifically, to an image correction, recognition and information retrieval method and related devices. Background Art
[0002] In some image recognition and image-based retrieval solutions, it is usually necessary to accurately identify the information contained in document images. For example, in a learning scenario, students take photos of book contents and retrieve the answers to test questions based on the photos.
[0003] In real-world photography, paper documents often exhibit physical deformations (such as wrinkles, bends, and bulges), making the captured image difficult to identify valid information and reducing subsequent usability. Furthermore, users may capture additional pages or portions of other pages in addition to the intended single document, negatively impacting image correction.
[0004] Traditional solutions typically use perspective transformation algorithms to correct images. This method works well for trapezoidal image deformation caused by the shooting angle, but it cannot correct the physical deformation of the document in the image. Summary of the Invention
[0005] In view of the above problems, this application proposes to provide an image correction, recognition and information retrieval method and related devices to alleviate the image geometric distortion problem caused by physical deformation of documents, thereby improving the image recognition accuracy and the efficiency of subsequent information retrieval. The specific solution is as follows:
[0006] In a first aspect, an image correction method is provided, comprising:
[0007] Acquire a target document image to be corrected;
[0008] detecting a single-page document area in the target document image, and obtaining a single-page document image including the single-page document area;
[0009] Processing the single-page document image using a configured physical deformation correction model to obtain mapping information output by the model, wherein the mapping information is used to describe a mapping relationship between pixels in the single-page document image and pixels in the corrected flattened image, wherein the physical deformation correction model is trained using image samples taken of single-page documents with physical deformation as training samples and the mapping information corresponding to the single-page document image samples as sample labels;
[0010] The single-page document image is corrected into a flat image according to the mapping information.
[0011] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of detecting a single-page document area in the target document image includes:
[0012] Using a corner detection model to identify contour corner information in the target document image, wherein the corner detection model is trained using document image training data marked with contour corners;
[0013] The single-page document area is determined based on the outline corner point information.
[0014] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of determining the single-page document area in combination with the outline corner point information includes:
[0015] The area enclosed by each outline corner point in sequence is used as the single-page document area.
[0016] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the number of the outline corner points is four. The process of determining the single-page document area based on the outline corner point information includes:
[0017] Determine the quadrilateral formed by the four contour corner points;
[0018] Along the two longitudinal sides of the quadrilateral, four contour corner points are offset outward by a set distance to obtain the four offset contour corner points, and the area surrounded by the four offset contour corner points is used as the single-page document area.
[0019] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of obtaining the single-page document image including the single-page document area includes:
[0020] predicting target size information of a document area without perspective distortion based on edge information of the single-page document area;
[0021] The single-page document region is subjected to perspective transformation processing according to the target size information and the contour corner point information to obtain a processed single-page document image containing the single-page document region.
[0022] In one possible design, in another implementation of the first aspect of the embodiments of the present application, before processing the single-page document image using the configured physical deformation correction model, the method further includes:
[0023] Processing the single-page document image, wherein the processing includes at least one of the following: normalization and proportional expansion;
[0024] The proportional expansion means that the single-page document image is proportionally filled according to the ratio of the input image required by the physical deformation correction model, and the pixel values of the filled pixels are set to a uniform pixel value.
[0025] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of correcting the single-page document image into a flattened image according to the mapping information includes:
[0026] According to the mapping information, the single-page document image is corrected into a flat image by setting an interpolation algorithm.
[0027] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the training process of the physical deformation correction model includes:
[0028] Obtaining an image sample taken of the single-page document with physical deformation, a mapping information tag corresponding to the image sample, and a three-dimensional point cloud information tag of the single-page document with physical deformation;
[0029] The image sample is fed into a backbone network for feature extraction, and the features extracted by the backbone network are fed into a first output module and a second output module respectively, wherein the first output module predicts mapping information and the second output module predicts three-dimensional point cloud information;
[0030] Calculating a first loss value based on the mapping information predicted by the first output module and the mapping information label, and calculating a second loss value based on the three-dimensional point cloud information predicted by the second output module and the three-dimensional point cloud information label;
[0031] A total loss value is calculated based on the first loss value and the second loss value, and the parameters of the backbone network, the first output module, and the second output module are updated according to the total loss value until the training end condition is reached. The physical deformation correction model is composed of the backbone network and the first output module.
[0032] In a second aspect, an image recognition method is provided, comprising:
[0033] Obtaining a target document image to be recognized;
[0034] Correcting the target document image using any one of the image correction methods described in the first aspect of the present application to obtain a corrected smooth image;
[0035] Content in the flattened image is identified.
[0036] A third aspect provides an information retrieval method, comprising:
[0037] Obtain target document image;
[0038] Recognize the target document image using the image recognition method described in the second aspect of the present application to obtain recognized content;
[0039] Relevant data is retrieved from the information source using the identification content as a retrieval condition.
[0040] In a fourth aspect, an image correction device is provided, comprising:
[0041] An image acquisition unit, configured to acquire an image of a target document to be corrected;
[0042] An image detection unit, configured to detect a single-page document region in the target document image and obtain a single-page document image containing the single-page document region;
[0043] a computing unit configured to process the single-page document image using a configured physical deformation correction model to obtain mapping information output by the model, wherein the mapping information is used to describe a mapping relationship between pixels in the single-page document image and pixels in the corrected flattened image, wherein the physical deformation correction model is trained using image samples taken of single-page documents with physical deformation as training samples, and using the mapping information corresponding to the single-page document image samples as sample labels;
[0044] An image processing unit is configured to correct the single-page document image into a flat image according to the mapping information.
[0045] In a fifth aspect, an electronic device is provided, comprising: a memory and a processor;
[0046] The memory is used to store programs;
[0047] The processor is used to execute the program to implement the various steps of the image correction method described in any one of the first aspects of the present application, or to implement the various steps of the image recognition method described in the second aspect of the present application, or to implement the various steps of the information retrieval method described in the third aspect of the present application.
[0048] In the sixth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the various steps of the image correction method described in any one of the first aspects of the present application, or implements the various steps of the image recognition method described in the second aspect of the present application, or implements the various steps of the information retrieval method described in the third aspect of the present application.
[0049] In the seventh aspect, a computer program product is provided, comprising a computer program. When the computer program is executed by a processor, the computer program implements the various steps of the image correction method described in any one of the first aspects of the present application, or implements the various steps of the image recognition method described in the second aspect of the present application, or implements the various steps of the information retrieval method described in the third aspect of the present application.
[0050] By means of the above technical solution, the present application detects a single-page document area from the target document image, and then obtains a single-page document image containing the single-page document area. Subsequently, the single-page document image can be corrected for physical deformation to avoid interference from the appendix or other page content in the original target document image, which helps to improve the image correction effect. In addition, the present application is equipped with a physical deformation correction model, which can output mapping information for describing the mapping relationship between the pixel points in the single-page document image and the pixel points in the corrected flat image based on the input single-page document image, that is, directly learning the mapping information required to restore the deformed target document image to the flat state before deformation through a neural network model. The physical deformation correction model is better at handling complex nonlinear deformations, such as various forms of physical deformation (wrinkles, bulges, etc.), and can implicitly learn prior knowledge of the document content to assist in correction, thereby improving the correction effect of document images with physical deformation. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0052] Figure 1 The following example shows a schematic diagram of an image taken of a textbook;
[0053] Figure 2 A schematic diagram of an implementation system architecture of the image correction method provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of a flow chart of an image correction method provided in an embodiment of the present application;
[0055] Figure 4 Example for Figure 1 Schematic diagram of contour corner points detected in an image;
[0056] Figure 5 An example of a process for offsetting contour corners is shown;
[0057] Figure 6 A schematic diagram of a single-page document image is illustrated;
[0058] Figure 7 A schematic diagram illustrating a process of correcting physical deformation of a single-page document image is provided;
[0059] Figure 8 An example of a flattened image after physical deformation correction is shown;
[0060] Figure 9 This example shows a schematic diagram of the training process of a physical deformation correction model;
[0061] Figure 10 This example illustrates another diagram of the training process for a physical deformation correction model.
[0062] Figure 11 A schematic diagram of the structure of an image correction device provided in an embodiment of the present application;
[0063] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0065] Image correction refers to the process of correcting distorted images to improve the quality of the corrected image. There are generally two types of distorted images. One is that the image has perspective distortion due to the shooting angle, that is, the document page in the image looks like a trapezoid or an irregular quadrilateral; the other is that the paper document has physical deformation, such as wrinkles, bends, bulges, etc., which causes physical deformation of the document in the image after shooting. This physical deformation is generally nonlinear distortion, which is prone to problems such as curved text lines and uneven text sizes on the page, resulting in the loss of effective information, such as text being blocked by wrinkles or severely distorted. Reference Figure 1 , which shows an example of an image taken of a textbook. It has both perspective distortion and physical deformation.
[0066] In addition, the image of the paper document taken by the user may contain not only the target page document but also the appendix or other page contents. This information may interfere with the image correction effect on the target page.
[0067] Some image correction schemes use perspective transformation algorithms to correct the target document image. This approach is effective for perspective distortion caused by the shooting angle, but it cannot correct the physical deformation of the document in the image.
[0068] Image correction solutions can be applied to a variety of scenarios. The following are some examples of optional scenarios:
[0069] Scenario 1: Image Recognition Scenario
[0070] Before performing OCR (Optical Character Recognition) on an image, you can first correct the distorted image to improve image quality and thus enhance OCR recognition accuracy.
[0071] Scenario 2: Image-based information retrieval scenario
[0072] In some scenarios, image-based retrieval of relevant information is possible. For example, in a shopping scenario, users can take a photo of an item of interest and then search for identical or similar products based on the image. Another example is a learning scenario where students can take a photo of printed documents like textbooks, exercises, and exams and search for identical or similar questions, knowledge points, answers, and other learning resources based on the image.
[0073] The embodiment of the present application provides an image correction solution that can alleviate the problem of image geometric distortion caused by the physical deformation of paper documents, and can further help improve the effect of the image in subsequent applications, such as improving the accuracy of image recognition and the efficiency of subsequent information retrieval.
[0074] This application provides an image correction and recognition method and an information retrieval method, which can be applied to Figure 2 The system architecture shown in FIG. 1 may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 2 (This section includes a server as an example).
[0075] Either the terminal 100 or the server 200 can be used independently to execute the image correction and recognition method and the information retrieval method provided in the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in conjunction to execute the image correction and recognition method and the information retrieval method provided in the embodiments of the present application.
[0076] Next describe Figure 2 The product form of the mid-terminal 100;
[0077] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a learning machine, a teaching screen, a wearable device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0078] The embodiment of the present application provides an image correction method, which is illustrated by applying the method to a computer device. The computer device may be Figure 2 The terminal 100 or the system consisting of the terminal 100 and the server 200. Figure 3 , the image correction method specifically includes the following steps:
[0079] Step S100: Acquire a target document image to be corrected.
[0080] The target document image is the image to be corrected. The target document image can be input by the user or obtained through other means. For example, the target document image is automatically captured by a camera, or is an image to be corrected imported by a third-party device.
[0081] Taking the learning scenario as an example, users can use mobile phones, learning machines and other electronic devices with shooting functions to take pictures of paper documents, and use the photographed image (or the image after pre-processing of the photographed image) as the target document image to be corrected.
[0082] The target document image contains the captured document information. Documents can be in various formats, such as textbooks, test papers, and advertising leaflets. Documents can include information in any one or more formats, such as text, images, and tables.
[0083] Step S110 : Detecting a single-page document region in the target document image to obtain a single-page document image containing the single-page document region.
[0084] In addition to the single-page document area that the user is interested in, the target document image may also contain other interference information, such as other page information that the user is not interested in, such as attached page information. Figure 1 As shown, the right side of the figure shows a single-page document that the user is interested in, and the left side of the figure also contains some information from the previous page, which is interference information.
[0085] In order to prevent the interference information from negatively impacting subsequent correction processing of the single-page document image, in this step, the single-page document region can be detected in the target document image to obtain a single-page document image containing the single-page document region.
[0086] Step S120: Process the single-page document image using the configured physical deformation correction model to obtain mapping information output by the model, where the mapping information is used to describe the mapping relationship between the pixel points in the single-page document image and the pixel points in the corrected flattened image.
[0087] Among them, the physical deformation correction model uses image samples taken of single-page documents with physical deformation as training samples, and uses the mapping information corresponding to the single-page document image samples as sample labels for training.
[0088] This application pre-trains a deep learning model (a physical deformation correction model) that directly predicts the mapping information needed to restore a deformed document to its flat state. The model inputs a single-page document image with physical deformations (such as wrinkles, bends, bulges, etc.) and directly outputs the mapping information.
[0089] The physical deformation correction model uses end-to-end learning to directly learn complex deformation correction mapping from raw pixels, improving data processing efficiency and avoiding error accumulation in multiple steps.
[0090] Physical deformation correction models can have deep learning model structures, such as Unet-based structures and Transformer structures. They can learn and process highly nonlinear and local deformations (such as bulges and dense wrinkles), greatly improving the correction effect of complex physical deformations.
[0091] Furthermore, during training, the physical deformation correction model implicitly learns prior knowledge about the input document content to aid correction. For example, it learns the statistical patterns and spatial structure of text, lines, and whitespace within the document, thereby improving the correction effect. During correction, even when severe deformation results in missing or blurred local texture information, the model can leverage this prior information to perform reasonable inference and reconstruction, improving correction robustness in challenging areas (such as deep wrinkles).
[0092] In an optional manner, before using the physical deformation correction model to process the single-page document image, a process of processing the single-page document image may be added. The processing operations include but are not limited to normalization and proportional expansion.
[0093] Normalization involves normalizing the pixel values in a single-page document image from [0, 255] to [0, 1]. This process improves numerical stability by consolidating the data range to a smaller, fixed interval ([0, 1]), thus avoiding instabilities (such as exploding or vanishing gradients) caused by excessively large or small values in subsequent model calculations.
[0094] Normalization can also accelerate model convergence. Many optimization algorithms (such as gradient descent) converge faster when features have similar scales. Normalizing pixel values to [0, 1] ensures that all input features (pixels) are on a similar scale.
[0095] Normalization can also simplify calculations. In some mathematical operations or loss functions, it may be more convenient to use values in the range [0,1].
[0096] Proportional expansion means that a single-page document image is proportionally filled according to the ratio of the input image required by the physical deformation correction model, and the pixel value of the filling pixel point is a set uniform pixel value, which can be a pixel value represented by black.
[0097] Step S130: Correct the single-page document image into a flat image according to the mapping information.
[0098] Since the mapping information describes the mapping relationship between the pixels in the single-page document image and the pixels in the corrected flat image, the single-page document image can be geometrically transformed based on the mapping relationship to obtain the corrected flat image.
[0099] The method provided in the embodiment of the present application detects a single-page document area from a target document image, and then obtains a single-page document image containing the single-page document area. Subsequently, the single-page document image can be corrected for physical deformation to avoid interference from appendices or other page content in the original target document image, which helps to improve the image correction effect. In addition, the present application is equipped with a physical deformation correction model, which can output mapping information for describing the mapping relationship between pixels in the single-page document image and pixels in the corrected flat image based on the input single-page document image. That is, the mapping information required to restore the deformed target document image to its pre-deformation flat state is directly learned through a neural network model. The physical deformation correction model is better at handling complex nonlinear deformations, such as various forms of physical deformation (wrinkles, bulges, etc.), and can implicitly learn prior knowledge of the document content to assist in correction, thereby improving the correction effect on document images with physical deformation.
[0100] In some embodiments of the present application, the process of detecting a single-page document area in the target document image in the aforementioned step S110 is described.
[0101] In some possible examples, a single-page document may be used as a target to be detected, and a target detection algorithm may be used to detect a single-page document region in a target document image.
[0102] In other possible examples, the present application may use a corner detection model to identify contour corner information (also referred to as key points) in the target document image, and then determine the single-page document area in combination with the contour corner information.
[0103] The corner detection model can be trained using document image training data marked with contour corners.
[0104] by Figure 1 As an example of the target document image shown in the figure, the contour corners identified by the corner detection model are as follows: Figure 4 shown.
[0105] The number of the contour corner points of the target document image can be a specified number, such as 4 or more. In addition, the number and positions of the contour corner points can also be determined by the corner point detection model. Figure 4 The example includes four contour corner points, which are located at the upper left corner, lower left corner, upper right corner, and lower right corner of the single-page document area.
[0106] After obtaining the outline corner point information, the single-page document area can be determined in combination with the outline corner point information.
[0107] In an optional implementation, the area enclosed by the contour corner points in sequence may be used as the single-page document area.
[0108] In another optional implementation, taking the number of contour corner points as four as an example, combined with Figure 5 As shown below, we introduce a way to determine the area of a single-page document:
[0109] Determine the quadrilateral formed by the four contour corner points;
[0110] Along the two longitudinal sides of the quadrilateral, the four contour corner points are offset outward by a set distance to obtain the four offset contour corner points. The area surrounded by the four offset contour corner points is used as the single-page document area.
[0111] Figure 5 The four corner points of the contour are defined as P0-P3. The four edges are defined as L0-L3. The angle formed by L0 and L3 is defined as θ.
[0112] like Figure 5As shown, the four corner points of the contour are offset outward along the two longitudinal edges by a set distance p. The coordinates of the four offset corner points can be calculated based on the geometric relationship. The offset distance p can be set by the user. For example, a value of 90px indicates an offset of 90 pixels. Of course, other values can also be set.
[0113] The offset distances of the four contour corner points can be the same, or the offset distances of P0 and P1 at the top remain the same, and the offset distances of P2 and P3 at the bottom remain the same, and the offset distance of P0 (P1) can be different from the offset distance of P2 (P3).
[0114] Next, taking the offset distance of the four contour corner points as p, and the contour corner points P0 and P1 on the left as examples, we will introduce an optional method for calculating the offset coordinates:
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123] direct=0 indicates the upper left corner point P0; direct=1 indicates the upper right corner point P1; direct=2 indicates the lower right corner point P2; direct=3 indicates the lower left corner point P3.
[0124] and The coordinates of the corner points of the offset contour. Figure 5 As shown, P0 is shifted to P0', P1 is shifted to P1', P2 is shifted to P2', and P3 is shifted to P3'.
[0125] The area enclosed by the four offset contour corner points P0'-P3' is used as the single-page document area.
[0126] The method provided in this embodiment can ensure that the upper and lower edge information of the document can be contained in the single-page document area by offsetting the detected contour corner points outward along the longitudinal edge. This edge information has strong guiding significance for the physical deformation correction processing in subsequent steps and can improve the image correction effect.
[0127] It should be noted that, considering that the left and right sides of the single-page document area in the target document image taken by the user are likely to contain other page information, this part of information is likely to have a negative impact on the physical deformation correction. Therefore, under normal circumstances, when the detected contour corner points are offset, they are only offset to the upper and lower sides along the longitudinal edge, and not to the left and right sides along the horizontal edge. This avoids the situation where when the contour corner points are offset along the horizontal edge, information from other pages is easily introduced into the single-page document area, affecting the subsequent image correction effect.
[0128] Of course, in some possible scenarios, if horizontal interference from other pages is not a concern, or if it is known that the single-page document area in the captured target document image does not contain any other page information in the horizontal direction, you can further offset the outline corner points outward along the horizontal edges. The area enclosed by the horizontally and vertically offset outline corner points is considered the single-page document area.
[0129] After obtaining the single-page document region in the target document image, the single-page document image containing the single-page document region can be further obtained.
[0130] In a possible implementation, a single-page document region may be captured from the target document image, thereby obtaining a single-page document image containing the single-page document region.
[0131] Compared with the target document image, the cropped single-page document image can remove some of the original background interference information, which is beneficial to the subsequent physical deformation correction processing.
[0132] In another possible implementation, considering that there may be perspective distortion due to the shooting angle during the shooting process, in this embodiment, the single-page document area in the target document image can be perspective transformed to obtain a processed single-page document image containing the single-page document area.
[0133] By performing perspective transformation on a single-page document area, the perspective distortion caused by the shooting angle can be improved, providing high-quality images for subsequent physical deformation correction, and improving the quality of the final corrected image.
[0134] for Figure 4 In the example shown, after obtaining the outline corner information, the outline corner offset process of the above steps is performed to obtain the single-page document area, and further the perspective transformation process is performed to obtain the processed single-page document image containing the single-page document area. Figure 6 shown.
[0135] This embodiment further explains the above perspective transformation process.
[0136] Perspective transformation is a linear transformation that projects an image from one perspective to another. By adjusting the perspective relationship of the image (such as rotation, tilt, scaling, etc.), the target plane presents a more regular geometric shape from the new perspective.
[0137] In this embodiment, target size information, such as aspect ratio, of the document region that has not been deformed may be predicted based on edge information of the single-page document region in the target document image.
[0138] There are many ways to implement the process of predicting the target size information of the document area without perspective distortion based on the edge information of the single-page document area. This embodiment introduces several optional implementation methods:
[0139] The first implementation method:
[0140] Combine Figure 5 As shown, the bottom edge L3 of the single-page document area can be considered to be free of perspective distortion; that is, the length of bottom edge L3 is the same as the bottom edge of the document area without perspective distortion. However, the top edge L1 of the single-page document area has undergone perspective distortion. The ratio of the lengths of L3 and L1 can be calculated as L3 / L1. The vertical distance h between L1 and L3 can be calculated as h' = h × L3 / L1, with h' being the true height of the document area without perspective distortion. This yields the aspect ratio of the document area without perspective distortion: L3:h'.
[0141] The second implementation method:
[0142] Combine Figure 5 As shown, the ratio of the lengths of the lines connecting the midpoints of opposite sides of the quadrilateral comprising the single-page document area is calculated as the aspect ratio (or height-to-width ratio) of the document area without perspective distortion. For example, the length of the line connecting the midpoints of L0 and L2 is calculated as the predicted width. The length of the line connecting the midpoints of L1 and L3 is calculated as the predicted height. The ratio of the predicted width to the predicted height is calculated as the aspect ratio of the document area without perspective distortion.
[0143] This method works well when the distortion is not severe.
[0144] The third implementation method:
[0145] Calculate the minimum bounding rectangle that encloses the single-page document area. Use the aspect ratio of the minimum bounding rectangle as the aspect ratio of the document area without perspective distortion.
[0146] After predicting the target size information of the document area without perspective distortion, the single-page document area can be perspective transformed according to the target size information and the contour corner point information of the single-page document area to obtain a processed single-page document image containing the single-page document area.
[0147] The target size information is taken as an example to illustrate the aspect ratio W / H.
[0148] A target rectangle can be defined based on the aspect ratio (W / H) of a single-page document area without perspective distortion. Typically, the height (H_target) of the target rectangle is set to a fixed value (e.g., 1000 pixels), and the width (W_target) is set to (H_target × (W / H). The four corner points of the target rectangle are fixed at (0, 0), (W_target, 0), (W_target, H_target), and (0, H_target).
[0149] The four corner points of the single-page document area detected in the previous step are matched one-to-one with the four corner point coordinates of the target rectangle defined in the previous step, and then the homography matrix is calculated.
[0150] A homography describes the perspective projection relationship between two planes (here, the plane corresponding to the target rectangle and the plane of the single-page document region undergoing perspective distortion). Given at least four sets of corresponding points, the homography can be solved using an algorithm, such as OpenCV's findHomography function.
[0151] The calculated homography matrix is used to perform perspective transformation processing on the single-page document region, so as to obtain a processed single-page document image containing the single-page document region.
[0152] Specifically, the homography matrix is used to map each pixel in the source image (single-page document area) to the corresponding position in the target image (single-page document image) according to the perspective relationship.
[0153] During the mapping process, some pixel positions of the target image may not have exactly corresponding source image pixels. In this case, an interpolation algorithm (such as bilinear interpolation, bicubic interpolation, etc.) can be used to calculate the color value of the pixel position to make the target image smoother.
[0154] In some embodiments of the present application, the process of correcting the single-page document image into a flat image according to the mapping information in step S130 is described in detail.
[0155] Combine Figure 7As shown, a single-page document image with physical deformation can be corrected into a flat image by using a set interpolation algorithm according to the mapping information after the physical deformation correction model is applied.
[0156] The interpolation algorithm can be bilinear interpolation, bicubic interpolation, or other interpolation algorithms. For example, bilinear interpolation is an image scaling or resampling method based on linear interpolation. It calculates the new pixel value by taking the weighted average of the four nearest neighboring pixels around the target point. Its core concept is to perform linear interpolation in two directions (horizontally and vertically) to ultimately produce a smooth output.
[0157] for Figure 6 The single-page document image shown in FIG. 1 is mapped using a physical deformation correction model to obtain mapping information. The single-page document image is corrected into a flat image according to the mapping information. The obtained corrected flat image is shown in FIG. Figure 8 As shown in the figure, the corrected flat image greatly improves physical deformation issues such as wrinkles, bulges, and text line distortion.
[0158] In some embodiments of the present application, the training process of the physical deformation correction model is described.
[0159] An optional training method, combined with Figure 9 As shown, the following steps may be included:
[0160] A1. Obtain an image sample taken of a single-page document with physical deformation, and a mapping information tag corresponding to the image sample.
[0161] The image sample can be obtained by photographing a single-page document with physical deformation using a camera. Alternatively, a virtual document can be created using 3D rendering software (such as Blender), and then controllable physical deformation can be applied to render the deformed image as the image sample.
[0162] For image samples rendered by 3D rendering software, the mapping relationship between each pixel in the deformed image and the pixel in the flattened image before physical deformation can be accurately calculated as the mapping information label corresponding to the image sample.
[0163] For image samples captured by a camera, a 3D point cloud can be obtained by scanning the image samples using devices such as structured light scanners and depth cameras. Mapping information labels can then be estimated using 3D reconstruction and unfolding algorithms. Alternatively, for a small amount of real data, some key points can be manually annotated, and then dense points can be estimated through interpolation and other methods to calculate mapping information labels.
[0164] To ensure the diversity of training data, the acquired image samples should include various types and degrees of physical deformation (slight bending, severe wrinkling, local bulging, multi-directional bending, etc.), as well as different lighting, backgrounds, document content (text, graphics, pictures), and paper materials, so that the trained physical deformation correction model is robust.
[0165] A2. Send the image sample to the physical deformation correction model to extract features through the model's backbone network, and use the output module to predict mapping information based on the features extracted by the backbone network.
[0166] A3. Calculate the loss value L based on the mapping information predicted by the model and the mapping information label.
[0167] A4. Update the parameters of the physical deformation correction model according to the loss value L until the training end condition is met, thereby obtaining the trained physical deformation correction model.
[0168] Through the above training process, the model can learn end-to-end from physically deformed single-page document images to the complex nonlinear relationship between mapping information for use in the inference stage.
[0169] Another optional training method is to combine Figure 10 As shown, the following steps may be included:
[0170] B1. Obtain image samples taken of the single-page document with physical deformation, mapping information labels corresponding to the image samples, and three-dimensional point cloud information labels of the single-page document with physical deformation.
[0171] Compared to step A1 in the previous embodiment, this step additionally obtains a 3D point cloud information tag of the single-page document with physical deformation. The 3D point cloud information tag can be obtained by scanning the single-page document with physical deformation using a device such as a structured scanner or a depth camera.
[0172] B2. Send the image sample to the backbone network for feature extraction, and send the features extracted by the backbone network to the first output module and the second output module respectively. The first output module predicts the mapping information, and the second output module predicts the 3D point cloud information.
[0173] B3. Calculate a first loss value L1 based on the mapping information and mapping information labels predicted by the first output module, and calculate a second loss value L2 based on the three-dimensional point cloud information and three-dimensional point cloud information labels predicted by the second output module.
[0174] B3. Calculate the total loss value based on the first loss value L1 and the second loss value L2, and update the parameters of the backbone network, the first output module, and the second output module according to the total loss value until the training end condition is reached. The physical deformation correction model is composed of the backbone network and the first output module.
[0175] Compared to the training method provided in the previous embodiment, this embodiment adopts a multi-task learning training strategy. During the training phase, the backbone network, the first output module, and the second output module are jointly trained. The first output module and the second output module share the same backbone network. The first output module is used to predict mapping information, and the second output module is used to predict the 3D point cloud information of the image sample.
[0176] The existence of the second output module (predicting 3D point cloud information) plays a crucial role in improving the accuracy of the first output module (predicting mapping information), mainly in the following aspects:
[0177] Learning more robust and physically meaningful feature representations:
[0178] Physical deformation is inherently three-dimensional: Wrinkling, bending, folding, and other physical deformations of a document occur in three-dimensional space. A two-dimensional image is simply a projection (or the result of light reflection) of this three-dimensional surface.
[0179] Backbone network guidance: Requiring the backbone network to extract features for two tasks (predicting mapping information and predicting 3D point clouds) at the same time is equivalent to giving the backbone network a strong constraint: it must learn features that can simultaneously explain the appearance of the 2D image and the 3D geometric structure behind it.
[0180] Intrinsic physicality of features: To accurately predict 3D point clouds, the backbone network's features must encode physical geometric information about surface normal directions, depth variations, curvature, and other aspects. This information is crucial for understanding physical deformation itself. When these geometric features with clear physical meaning are also used to predict mapping information, the physical deformation correction model can make inferences based on a deeper understanding of the nature of deformation, rather than relying solely on image texture or apparent pattern matching. This significantly improves the model's robustness (making it less sensitive to lighting, texture variations, camera angle, etc.).
[0181] Provide stronger supervision signals and regularization:
[0182] 3D point cloud labels provide more direct geometric supervision: Mapping information labels are two-dimensional representations that describe the deformation results. 3D point cloud labels, on the other hand, directly provide the absolute 3D spatial coordinates of the document surface, providing a more primitive, direct, and informative supervisory signal about the geometric nature of the deformation.
[0183] Mitigating Overfitting: Training a single network to simultaneously fit two related tasks (mapping information and 3D point cloud information) can be considered an effective regularization technique. The backbone network cannot simply "remember" image features useful for predicting mapping information (which may contain noise or information unrelated to deformation) but must also learn features that are also effective for predicting 3D shape. This forces the network to focus on learning more general underlying geometric features that are strongly related to physical deformation, reducing the risk of overfitting to specific noise or patterns in the training data and improving generalization.
[0184] Take advantage of the inherent dependencies between tasks:
[0185] Mapping information is the 2D projection of a 3D deformation: the change in the position of a point on a document in 3D space (the information described by the 3D point cloud) directly determines how its projection position on the 2D image plane changes (the information described by the mapping information). These two tasks are physically highly coupled and interdependent.
[0186] Knowledge Sharing: During joint training, the features learned by the backbone network and the knowledge learned by the second output module about how to infer 3D geometry from image features are implicitly passed to the first output module through the shared backbone network. For example, a feature learned by the backbone network may be very effective in detecting depth discontinuities at the edge of a paper, which is crucial for predicting the boundary of a 3D point cloud. This feature is also extremely useful for predicting the precise displacement (mapping information) of that edge in the image plane.
[0187] Resolving ambiguity in mapping predictions:
[0188] Monocular Figure 2 Ambiguity in 2D-to-2D mappings: Inferring 2D-to-2D mappings (e.g., from a deformed image to a flattened image) from a single 2D image is inherently ambiguous. The same 2D image pattern may correspond to multiple distinct 3D deformation states (especially in regions of occlusion or self-similar textures). For example, a localized image curvature may be a shallow and broad bend or a deep and narrow fold.
[0189] 3D information as a constraint: Requiring the model to also predict a 3D point cloud effectively introduces a 3D geometric consistency constraint. The model's predicted mapping information must be physically consistent with the predicted 3D shape (i.e., the mapping result should be the projection of the predicted 3D shape onto the image plane). This constraint significantly reduces the number of ambiguous solutions that can arise when predicting mappings based solely on 2D images, guiding the model to find a more physically plausible mapping relationship.
[0190] Improve spatial understanding of features:
[0191] Predicting an accurate 3D point cloud requires a deep understanding of the spatial relationships of every pixel in the image, including how local regions bend, stretch, or compress in 3D space. This understanding of 3D spatial relationships is naturally learned by the backbone network and shared with the mapping prediction task. This allows the first output module to better account for physical laws such as spatial continuity and neighborhood constraints when predicting pixel displacements, resulting in a smoother mapping field that better conforms to physical deformation laws.
[0192] After joint training is complete, deployment requires only the backbone network and the first output module to form a physical deformation correction model. This model has learned more powerful, essential, and robust feature representations than those trained on the mapping task alone, enabling it to more accurately predict the corrected mapping information for physically deformed documents. The 3D point cloud prediction branch (the second output module) acts like a "geometry teacher," guiding the backbone network during training to gain a deeper understanding of the physical nature of deformation, ultimately enhancing the performance of the "student" (the first output module) used for mapping prediction.
[0193] Based on the image correction methods described in the above embodiments, this embodiment of the present application further provides an image recognition method, including:
[0194] Obtain a target document image to be recognized, perform correction processing on the target document image using the image correction method described in any of the above embodiments to obtain a corrected flat image, and recognize content in the flat image.
[0195] The aforementioned image correction processing can improve image quality, thereby helping to improve the recognition accuracy of image content.
[0196] Furthermore, this embodiment also provides an image-based information retrieval method, which specifically includes:
[0197] Obtain a target document image. Use the above-mentioned image recognition method to recognize the target document image and obtain the recognized content. Use the recognized content as a retrieval condition to retrieve relevant data from the information source.
[0198] In some scenarios, image-based retrieval of relevant information is possible. For example, in a shopping scenario, users can take a photo of an object of interest and then search for identical or similar product information based on the captured image of the target document. Another example is a learning scenario where students can take a photo of printed documents such as textbooks, exercise books, and exams and search for identical or similar test questions, knowledge points, answer key explanations, and other learning resources based on the captured image of the target document.
[0199] Based on the aforementioned image correction processing, the quality of the corrected image is improved, thereby improving the recognition accuracy of the image content. On this basis, relevant data can be retrieved based on more accurate image recognition content, thereby improving the accuracy of information retrieval.
[0200] Taking textbook retrieval as an example, a test was conducted on a dataset of 1,537 textbook items. Without the image correction method proposed in this application, the textbook retrieval accuracy was 74.312%. After adopting the image correction method proposed in this application, the textbook retrieval accuracy increased to 98.279%.
[0201] The image correction method, image recognition method, and information retrieval method of the present application can be applied to mobile devices or fixed terminal devices, and have a wider range of applicable scenarios.
[0202] The image correction device provided in an embodiment of the present application is described below. The image correction device described below and the image correction method described above can be referenced to each other.
[0203] See also Figure 11 , Figure 11 This is a schematic diagram of the structure of an image correction device disclosed in an embodiment of the present application.
[0204] like Figure 11 As shown, the device may include:
[0205] An image acquisition unit 11 is used to acquire a target document image to be corrected;
[0206] An image detection unit 12 is configured to detect a single-page document region in the target document image and obtain a single-page document image containing the single-page document region;
[0207] A computing unit 13 is configured to process the single-page document image using a configured physical deformation correction model to obtain mapping information output by the model, wherein the mapping information is used to describe a mapping relationship between pixels in the single-page document image and pixels in the corrected flattened image. The physical deformation correction model is trained using image samples taken of single-page documents with physical deformation as training samples, and the mapping information corresponding to the single-page document image samples is used as sample labels.
[0208] The image processing unit 14 is configured to correct the single-page document image into a flat image according to the mapping information.
[0209] In one possible implementation, the process of the image detection unit detecting the single-page document area in the target document image includes:
[0210] Using a corner detection model to identify contour corner information in the target document image, wherein the corner detection model is trained using document image training data marked with contour corners;
[0211] The single-page document area is determined based on the outline corner point information.
[0212] In one possible implementation, the process of determining the single-page document area by the image detection unit in combination with the contour corner point information includes:
[0213] The area enclosed by each outline corner point in sequence is used as the single-page document area.
[0214] In a possible implementation, the number of the contour corner points is four, and the process of the image detection unit determining the single-page document area based on the contour corner point information includes:
[0215] Determine the quadrilateral formed by the four contour corner points;
[0216] Along the two longitudinal sides of the quadrilateral, four contour corner points are offset outward by a set distance to obtain the four offset contour corner points, and the area surrounded by the four offset contour corner points is used as the single-page document area.
[0217] In one possible implementation, the process of obtaining the single-page document image including the single-page document area by the image detection unit includes:
[0218] predicting target size information of a document area without perspective distortion based on edge information of the single-page document area;
[0219] The single-page document region is subjected to perspective transformation processing according to the target size information and the contour corner point information to obtain a processed single-page document image containing the single-page document region.
[0220] In one possible implementation, before the computing unit processes the single-page document image using the configured physical deformation correction model, it is further configured to:
[0221] Processing the single-page document image, wherein the processing includes at least one of the following: normalization and proportional expansion;
[0222] The proportional expansion means that the single-page document image is proportionally filled according to the ratio of the input image required by the physical deformation correction model, and the pixel values of the filled pixels are set to a uniform pixel value.
[0223] In one possible implementation, the image processing unit corrects the single-page document image into a flat image according to the mapping information, including:
[0224] According to the mapping information, the single-page document image is corrected into a flat image by setting an interpolation algorithm.
[0225] In one possible implementation, the apparatus of the present application may further include: a model training unit for training a physical deformation correction model, wherein the training process includes:
[0226] Obtaining an image sample taken of the single-page document with physical deformation, a mapping information tag corresponding to the image sample, and a three-dimensional point cloud information tag of the single-page document with physical deformation;
[0227] The image sample is fed into a backbone network for feature extraction, and the features extracted by the backbone network are fed into a first output module and a second output module respectively, wherein the first output module predicts mapping information and the second output module predicts three-dimensional point cloud information;
[0228] Calculating a first loss value based on the mapping information predicted by the first output module and the mapping information label, and calculating a second loss value based on the three-dimensional point cloud information predicted by the second output module and the three-dimensional point cloud information label;
[0229] A total loss value is calculated based on the first loss value and the second loss value, and the parameters of the backbone network, the first output module, and the second output module are updated according to the total loss value until the training end condition is reached. The physical deformation correction model is composed of the backbone network and the first output module.
[0230] Each unit in the above-mentioned image correction device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned units can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above-mentioned units.
[0231] An electronic device is also provided in an embodiment of the present application. Figure 12 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, learning machines, tablet computers, teaching large screens, wearable devices, etc. Figure 12 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0232] like Figure 12As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 2 or a program loaded from a storage device 8 into a random access memory (RAM) 3 to implement the image correction method, image recognition method, and / or information retrieval method of the aforementioned embodiments of the present application. When the electronic device is powered on, the RAM 3 also stores various programs and data required for the operation of the electronic device. The processing device 1, ROM 2, and RAM 3 are connected to each other via a bus 4. An input / output (I / O) interface 5 is also connected to the bus 4.
[0233] Typically, the following devices may be connected to the I / O interface 5: an input device 6 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 7 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 8 including, for example, a memory card, a hard disk, etc.; and a communication device 9. The communication device 9 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 12 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0234] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the image correction methods, image recognition methods and / or information retrieval methods provided in the embodiments of the present application.
[0235] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When one or more computer programs are executed by an electronic device, the electronic device can implement any image correction method, image recognition method and / or information retrieval method provided in the embodiment of the present application.
[0236] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0238] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0239] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0240] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referenced to each other.
Claims
1. An image correction method, characterized in that: include: Acquire a target document image to be corrected; detecting a single-page document area in the target document image, and obtaining a single-page document image including the single-page document area; Processing the single-page document image using a configured physical deformation correction model to obtain mapping information output by the model, wherein the mapping information is used to describe a mapping relationship between pixels in the single-page document image and pixels in the corrected flattened image, wherein the physical deformation correction model is trained using image samples taken of single-page documents with physical deformation as training samples and the mapping information corresponding to the single-page document image samples as sample labels; The single-page document image is corrected into a flat image according to the mapping information.
2. The method according to claim 1, characterized in that The process of detecting a single-page document area in the target document image includes: Using a corner detection model to identify contour corner information in the target document image, wherein the corner detection model is trained using document image training data marked with contour corners; The single-page document area is determined based on the outline corner point information.
3. The method according to claim 2, characterized in that The process of determining the single-page document area based on the outline corner point information includes: The area enclosed by each outline corner point in sequence is used as the single-page document area.
4. The method according to claim 2, characterized in that If the number of the outline corner points is four, the process of determining the single-page document area in combination with the outline corner point information includes: Determine the quadrilateral formed by the four contour corner points; Along the two longitudinal sides of the quadrilateral, four contour corner points are offset outward by a set distance to obtain the four offset contour corner points, and the area surrounded by the four offset contour corner points is used as the single-page document area.
5. The method according to claim 2, characterized in that The process of obtaining a single-page document image including the single-page document area includes: predicting target size information of a document area without perspective distortion based on edge information of the single-page document area; The single-page document region is subjected to perspective transformation processing according to the target size information and the contour corner point information to obtain a processed single-page document image containing the single-page document region.
6. The method according to claim 1, characterized in that Before using the configured physical deformation correction model to process the single-page document image, the method further includes: Processing the single-page document image, wherein the processing includes at least one of the following: normalization and proportional expansion; The proportional expansion means that the single-page document image is proportionally filled according to the ratio of the input image required by the physical deformation correction model, and the pixel values of the filled pixels are set to a uniform pixel value.
7. The method according to claim 1, characterized in that The process of correcting the single-page document image into a flat image according to the mapping information includes: According to the mapping information, the single-page document image is corrected into a flat image by setting an interpolation algorithm.
8. The method according to any one of claims 1 to 7, characterized in that The training process of the physical deformation correction model includes: Obtaining an image sample taken of the single-page document with physical deformation, a mapping information tag corresponding to the image sample, and a three-dimensional point cloud information tag of the single-page document with physical deformation; The image sample is fed into a backbone network for feature extraction, and the features extracted by the backbone network are fed into a first output module and a second output module respectively, wherein the first output module predicts mapping information and the second output module predicts three-dimensional point cloud information; Calculating a first loss value based on the mapping information predicted by the first output module and the mapping information label, and calculating a second loss value based on the three-dimensional point cloud information predicted by the second output module and the three-dimensional point cloud information label; A total loss value is calculated based on the first loss value and the second loss value, and the parameters of the backbone network, the first output module, and the second output module are updated according to the total loss value until the training end condition is reached. The physical deformation correction model is composed of the backbone network and the first output module.
9. An image recognition method, characterized in that: include: Acquire a target document image to be identified; Correcting the target document image using the image correction method according to any one of claims 1 to 8 to obtain a corrected smooth image; Content in the flattened image is identified.
10. An information retrieval method, characterized in that: include: Obtain target document image; Recognize the target document image using the image recognition method according to claim 9 to obtain recognized content; Relevant data is retrieved from the information source using the identification content as a retrieval condition.
11. An image correction device, characterized in that: include: An image acquisition unit, configured to acquire an image of a target document to be corrected; An image detection unit, configured to detect a single-page document region in the target document image and obtain a single-page document image containing the single-page document region; a computing unit configured to process the single-page document image using a configured physical deformation correction model to obtain mapping information output by the model, wherein the mapping information is used to describe a mapping relationship between pixels in the single-page document image and pixels in the corrected flattened image, wherein the physical deformation correction model is trained using image samples taken of single-page documents with physical deformation as training samples, and using the mapping information corresponding to the single-page document image samples as sample labels; An image processing unit is used to correct the single-page document image into a flat image according to the mapping information.
12. An electronic device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the image correction method according to any one of claims 1 to 8, or to implement the various steps of the image recognition method according to claim 9, or to implement the various steps of the information retrieval method according to claim 10.
13. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the various steps of the image correction method according to any one of claims 1 to 8, or implements the various steps of the image recognition method according to claim 9, or implements the various steps of the information retrieval method according to claim 10.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the steps of the image correction method according to any one of claims 1 to 8, or implements the steps of the image recognition method according to claim 9, or implements the steps of the information retrieval method according to claim 10.
Citation Information
Patent Citations
Photographing problem searching method and electronic device
CN108196782A
Method and device for correcting document image and storage medium
CN115082935A
Automatic copybook image correction method based on deep learning and related product
CN117765539A
Image correction method and device and dictionary pen
CN118135574A
Cited By
File image distortion correction method and system based on deep learning
CN122134598A