Document fold repairing method based on deep learning
Through a deep learning-based method, trace lines are used to divide sub-surfaces and perform local affine transformations, combined with image and geometric modal feature fusion, the problem of poor document wrinkle repair in existing technologies is solved, and efficient document wrinkle repair is achieved.
Patent Information
- Application Number
- CN202511236670.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-09-01
AI Technical Summary
Existing document wrinkle repair methods usually treat the deformed document as a single continuous surface and flatten it through a global affine transformation algorithm, resulting in loss of key details and poor repair results.
A deep learning-based method is used to obtain two-dimensional images and point cloud data, calculate trace lines and generate sub-surfaces for local affine transformation, combine image and geometric modal feature fusion, and cross-modal collaborative restoration of two-dimensional images.
It achieves efficient repair of document wrinkles, preserves key details, and improves the repair effect.
Smart Images

Figure CN120725929A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of document image restoration, and in particular to a document wrinkle repair method based on deep learning. Background Art
[0002] In image processing scenarios such as document restoration and cultural relic digitization, there is an urgent need to process wrinkled and deformed documents.
[0003] Current document wrinkle repair methods usually treat deformed documents as a single continuous surface and force flatten them through a global affine transformation algorithm, which results in the loss of key details and the fragmentation of the repair effect, resulting in poor repair results. Summary of the Invention
[0004] The main purpose of this application is to provide a document wrinkle repair method based on deep learning, aiming to solve the technical problem of poor repair effect of wrinkled and deformed documents.
[0005] To achieve the above objectives, this application proposes a document wrinkle repair method based on deep learning, comprising: Obtaining the two-dimensional image and point cloud data of the document to be processed, and aligning each pixel of the two-dimensional image with the point cloud data; Calculate the curvature change rate of the neighborhood of each point in the point cloud data, and calculate the trace line based on the curvature change rate of the neighborhood of each point; Divide the point cloud data into several sub-surfaces according to the trace lines, perform local affine transformation on each sub-surface to obtain the corresponding flattened sub-surface, and splice the flattened sub-surfaces to obtain a three-dimensional surface model; Extracting image modal features from a two-dimensional image and extracting geometric modal features from a three-dimensional surface model, and fusing the image modal features and the geometric modal features to obtain fused features; Based on the fused features, the 2D image is collaboratively repaired across modalities to obtain the target image.
[0006] In one embodiment, obtaining point cloud data of a document to be processed includes: Obtaining a first-perspective two-dimensional image, a second-perspective two-dimensional image, and TOF depth information of the document to be processed; Calculating binocular depth information based on the parallax between the first-view two-dimensional image and the second-view two-dimensional image; Calculate the noise variance of the binocular depth information and the noise variance of the TOF depth information, and calculate the fusion weight of the binocular depth information and the fusion weight of the TOF depth information based on the noise variance of the binocular depth information and the noise variance of the TOF depth information; According to the fusion weight of the binocular depth information and the fusion weight of the TOF depth information, the binocular depth information and the TOF depth information are fused by Kalman filtering to obtain fused depth information; Convert the fused depth information into point cloud data.
[0007] In one embodiment, calculating and obtaining binocular depth information based on the parallax between the first-view two-dimensional image and the second-view two-dimensional image includes: Acquire a first-perspective two-dimensional image and a second-perspective two-dimensional image showing structured light fringes; According to the first-perspective two-dimensional image and the second-perspective two-dimensional image showing the structured light stripes, a first-perspective structured light phase image and a second-perspective structured light phase image are calculated by Fourier transform method; Calculate the disparity of each matching point pair between the first-view two-dimensional image and the second-view two-dimensional image according to the first-view grayscale image, the second-view grayscale image, the first-view structured light phase map, and the second-view structured light phase map; The binocular depth information is obtained by calculating the disparity between the first-view two-dimensional image and the second-view two-dimensional image.
[0008] In one embodiment, calculating the curvature change rate of the neighborhood where each point in the point cloud data is located, and calculating and obtaining the trace line according to the curvature change rate of the neighborhood where each point is located, includes: Establish a neighborhood point set of each point in the point cloud data; Construct the covariance matrix of the neighborhood point set, and perform eigenvalue decomposition on the covariance matrix to obtain the discrete values of each point in the neighborhood point set in three orthogonal directions; According to the discrete values in three orthogonal directions, the principal curvature of the neighborhood of each point in the point cloud data is calculated, and according to the principal curvature of the neighborhood of each point in the point cloud data, the curvature change rate of the neighborhood of each point in the point cloud data is calculated; According to the curvature change rate of the neighborhood where each point in the point cloud data is located and a preset curvature change rate threshold, a trace point is calculated, wherein the curvature change rate of the neighborhood where the trace point is located is greater than the preset curvature change rate threshold; Perform straight line fitting on each trace point to obtain the trace line.
[0009] In one embodiment, generating a plurality of sub-surfaces includes: The sub-surfaces are reconstructed using the greedy triangulation reconstruction method.
[0010] In one embodiment, performing a local affine transformation on each sub-surface to obtain a corresponding flattened sub-surface includes: After performing a local affine transformation on the sub-surface, a flattened sub-surface is obtained; The flattened sub-surface is optimized by solving the minimization energy function method to obtain the optimized flattened sub-surface, wherein the energy function includes a first constraint term and a second constraint term. The energy function is used to measure the degree of distortion of the flattened sub-surface, the first constraint term is used to measure the deviation between the normal vector of the flattened sub-surface and the normal vectors of the corresponding sub-surface points, and the second constraint term is used to measure the deviation of the vector length before and after flattening.
[0011] In one embodiment, fusing image modality features and geometric modality features to obtain fused features includes: Based on the gating mechanism, image modality features and geometric modality features are fused to obtain fusion features; Obtaining an initial neighborhood point set for each point in the three-dimensional surface model and calculating the principal curvature of the initial neighborhood point set; Calculate the dynamic neighborhood radius based on the principal curvature of the initial neighborhood point set; Determine the neighborhood point set based on the dynamic neighborhood radius; Calculate the attention weight based on the point distance and normal vector angle of each point in the neighborhood point set; Based on the attention weights, the feature aggregation output is used to obtain optimized geometric features; Extract the main folding axis vectors of the 3D surface model; Perform curvature Fourier transform on the principal curvature of the initial neighborhood point set to obtain high-dimensional curvature features; After superimposing the main fold axis vector and curvature high-dimensional features of the 3D surface model, the geometric modal features in the fused features are enhanced based on the attention mechanism and optimized geometric features. In one embodiment, based on the fused features, cross-modal collaborative restoration of the 2D image to obtain the target image includes: Correcting the geometric modal features using the image modal features, and correcting the image modal features using the geometric modal features to obtain corrected image modal features and corrected geometric modal features; According to the corrected image modal features and the corrected geometric modal features, the two-dimensional image is geometrically repaired, texture-repaired, and structurally repaired to obtain a target image. In one embodiment, obtaining point cloud data of the document to be processed includes: Acquire first-view point cloud data and second-view point cloud data; Based on the point distance constraint and the normal vector gradient consistency constraint, the first-view point cloud data and the second-view point cloud data are registered and stitched to obtain the fused point cloud data; Locate the seam area of the fused point cloud data based on the density change in the fused point cloud data; RBF interpolation is used to process the seam area to obtain optimized fused point cloud data. In one embodiment, the method further includes: Perform curvature analysis on the point cloud data to screen out the curled sub-surfaces, where the curvature of the curled sub-surfaces is distributed in a ring shape; Split the curly sub-surfaces according to the growth algorithm; Establish the parametric equation of the rotation axis of the curled sub-surface; According to the parametric equation of the rotation axis of the curled sub-surface, the curled sub-surface is unfolded equidistantly to obtain the flattened sub-surface.
[0012] One or more technical solutions proposed in this application have at least the following technical effects: A deep learning-based document wrinkle repair method is proposed. Trace lines are identified, and then the point cloud data is divided into several sub-surfaces according to the trace lines. Each sub-surface is locally affine transformed to obtain the corresponding flattened sub-surface. The flattened sub-surfaces are spliced to obtain a three-dimensional surface model. The image modal features in the two-dimensional image are extracted, and the geometric modal features of the three-dimensional surface model are extracted. The image modal features and the geometric modal features are fused to obtain the fused features. Based on the fused features, the two-dimensional image is repaired collaboratively across modalities to obtain the target image, thereby improving the repair effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 This is a flowchart of the first embodiment of the document wrinkle repair method based on deep learning provided by this application; The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0016] It should be understood that the specific embodiments described herein are merely for explaining the technical solutions of the present application and are not intended to limit the present application. In order to better understand the technical solutions of the present application, the following detailed description will be given in conjunction with the accompanying drawings and specific implementation methods.
[0017] In traditional 3D reconstruction methods, parametric models are usually used to fit the wrinkled surface of a document. However, parametric models with fixed parameters cannot adapt to complex or multi-directional wrinkles, such as documents with multiple creases and random curls.
[0018] At present, there are several technical concepts for document repair methods based on deep learning methods. First, DocUNet uses stacked U-Net to directly regress the pixel displacement field. It is only trained through two-dimensional deformation, and it is easy to produce repair results that deviate from physical geometric principles in real scenes; second, DewarpNet introduces three-dimensional coordinate map regression, but the mid-level feature extraction uses a fixed radius query to aggregate neighborhood points, which cannot adapt to the scale difference between crease edges and flat areas, resulting in inaccurate aggregation of features such as local curvature and normal vectors; third, Transformer is a neural network model based on the attention mechanism. The position encoding of the encoder of this model only relies on the sine and cosine functions, which makes it difficult to model the overall deformation law of the document. The repair effect on local wrinkled areas without complete boundaries (such as local tears and edge curling) is poor.
[0019] In the first embodiment of the document wrinkle repair method based on deep learning of the present application, refer to Figure 1 , Figure 1 This is a flowchart of a first embodiment of a document wrinkle repair method based on deep learning in this application. The document wrinkle repair method based on deep learning may include steps S10 to S50: Step S10: Acquire the two-dimensional image and point cloud data of the document to be processed, and align each pixel of the two-dimensional image with the point cloud data.
[0020] It should be noted that the two-dimensional image of the document to be processed is an RGB image, and the point cloud data can be the complete point cloud data collected directly, that is, the complete point cloud data obtained by a single scan, or it can be the complete point cloud data spliced together from multi-view point cloud data.
[0021] In a feasible implementation, the step S10 of "obtaining point cloud data of the document to be processed" may include steps S101 to S105: Step S101 : obtaining a first-perspective two-dimensional image, a second-perspective two-dimensional image, and TOF depth information of a document to be processed.
[0022] It should be noted that the first perspective is different from the second perspective, that is, the first perspective two-dimensional image is different from the second perspective two-dimensional image.
[0023] In a specific embodiment, the first-perspective two-dimensional image is a left-eye two-dimensional image captured by a binocular camera, and the second-perspective two-dimensional image is a right-eye two-dimensional image captured by the binocular camera.
[0024] Step S102 : obtaining binocular depth information by calculation according to the parallax between the first-view two-dimensional image and the second-view two-dimensional image.
[0025] It should be noted that the disparity is the disparity of each matching point pair between the first-perspective two-dimensional image and the second-perspective two-dimensional image. The first-perspective two-dimensional image and the second-perspective two-dimensional image may only include the original image of the document to be processed, or may be superimposed with structured light stripes.
[0026] In a specific embodiment, structured light stripes are superimposed on the two-dimensional image, and step S102 may include steps S1021 to S1024: Step S1021 : collecting a first-perspective two-dimensional image and a second-perspective two-dimensional image showing structured light fringes.
[0027] Step S1022 , calculating and obtaining a first-perspective structured light phase image and a second-perspective structured light phase image by Fourier transform method based on the first-perspective two-dimensional image and the second-perspective two-dimensional image showing structured light fringes.
[0028] Step S1023 , calculating the disparity of each matching point pair of the first-view two-dimensional image and the second-view two-dimensional image according to the first-view grayscale image, the second-view grayscale image, the first-view structured light phase map, and the second-view structured light phase map.
[0029] Specifically, a preset stereo matching optimization model is used to calculate the disparity by taking the first-view grayscale image, the second-view grayscale image, the first-view structured light phase map, and the second-view structured light phase map as input.
[0030] The stereo matching optimization model includes a method for measuring pixel p In Parallax d The disparity cost calculation model of matching credibility is as follows:
[0031]
[0032] Where, To measure the pixel matching reliability p In Parallax d The matching confidence under , pixels p In pixel coordinates express, d is parallax, is the first-view grayscale image, is the second-view grayscale image, is the first-perspective structured light phase image, is the second-view structured light phase image, is the weight coefficient, is the phase difference attenuation coefficient; For the grayscale term, by calculating the sum of the L1 norms of the grayscale differences between the left and right images within a 3×3 window, and taking advantage of the L1 norm's insensitivity to illumination noise, the disparity calculation accuracy in texture-rich areas (such as text and seal areas) is ensured, avoiding the impact of single pixel errors on the overall matching. is the structured light term, which is used to extract the first-view structured light phase image and the second-view structured light phase image through Fourier transform, and measures the phase difference between the first-view structured light phase image and the second-view structured light phase image. The range is [0, 1]. The closer the value is to 1, the higher the similarity. When the phase difference between the first-view structured light phase image and the second-view structured light phase image is less than 5°, It approaches 1, providing additional phase constraints for low-texture areas (such as blank documents), solving the parallax ambiguity problem caused by "no features to match" in low-texture areas caused by traditional algorithms.
[0033] It should be noted that It is used to balance the contributions of grayscale terms and structured light terms. Texture-rich areas are dominated by grayscale terms, while low-texture areas are dominated by structured light terms. The disparity corresponding to the pixel p with the highest matching confidence is used as the optimal disparity, which improves the robustness of the overall disparity calculation and lays an accurate foundation for subsequent depth data generation.
[0034] Step S1024 : obtaining binocular depth information by calculation based on the disparity of each matching point pair between the first-view two-dimensional image and the second-view two-dimensional image.
[0035] Step S103 , calculating the noise variance of the binocular depth information and the noise variance of the TOF depth information, and calculating the fusion weight of the binocular depth information and the fusion weight of the TOF depth information based on the noise variance of the binocular depth information and the noise variance of the TOF depth information.
[0036] Specifically, the mathematical expression of S103 is:
[0037]
[0038] Where, is the fusion weight of binocular depth information, is the fusion weight of TOF depth information, is the noise variance of binocular depth information, is the noise variance of TOF depth information, is the confidence of binocular depth information, is the confidence of TOF depth information.
[0039] Step S104 , performing Kalman filtering fusion on the binocular depth information and the TOF depth information according to the fusion weight of the binocular depth information and the fusion weight of the TOF depth information, so as to obtain fused depth information.
[0040] Specifically, the mathematical expression of step S104 is:
[0041] Where, To fuse depth information, is the binocular depth information, It is TOF depth information.
[0042] It should be noted that the calculation accuracy of binocular depth information in reflective areas is low, and TOF depth information has large noise under strong light. By using Kalman filtering to achieve adaptive confidence fusion of binocular depth information and TOF depth information, more stable and higher-precision depth data can be output, thereby improving the accuracy of subsequent three-dimensional reconstruction.
[0043] Step S105: converting the fused depth information into point cloud data.
[0044] In step S20 , the curvature change rate of the neighborhood of each point in the point cloud data is calculated, and the trace line is obtained according to the curvature change rate of the neighborhood of each point.
[0045] In a feasible implementation, step S20 may include steps S201 to S205: Step S201: establishing a neighborhood point set of the neighborhood where each point in the point cloud data is located.
[0046] It should be noted that the neighborhood radius is set according to actual needs. In this embodiment, the neighborhood radius is set to 5 mm, that is, for each point in the point cloud data, the neighborhood points within a radius of 5 mm are selected to construct a neighborhood point set.
[0047] Step S202 : constructing a covariance matrix of the neighborhood point set, and performing eigenvalue decomposition on the covariance matrix to obtain discrete values of each point in the neighborhood point set in three orthogonal directions.
[0048] It should be noted that, by performing eigenvalue decomposition on the covariance matrix Cov, we can obtain ,in, 、 and They respectively represent the degree of discreteness of each point in the neighborhood point set in three orthogonal directions.
[0049] In step S203, the principal curvature of the neighborhood of each point in the point cloud data is calculated based on the discrete values in the three orthogonal directions, and the curvature change rate of the neighborhood of each point in the point cloud data is calculated based on the principal curvature of the neighborhood of each point in the point cloud data.
[0050] Specifically, the mathematical expression of step S203 is:
[0051] Where, is the first principal curvature, is the second principal curvature, is the curvature change rate, where the first principal curvature and the second principal curvature Reflects the degree of local surface curvature. The larger the value, the more severe the curvature. The curvature change rate is used to distinguish between areas with high curvature change and areas with low curvature change.
[0052] In step S204, trace points are calculated based on the curvature change rate of the neighborhood where each point in the point cloud data is located and a preset curvature change rate threshold, wherein the curvature change rate of the neighborhood where the trace point is located is greater than the preset curvature change rate threshold.
[0053] It should be noted that the preset curvature change rate threshold is a preset value. In this embodiment, the preset curvature change rate threshold is 0.8. If the curvature change rate of the neighborhood where a point is located is greater than 0.8, the point is set as a trace point.
[0054] Step S205: performing straight line fitting on each trace point to obtain a trace line.
[0055] In a specific embodiment, two trace pixels can be randomly selected as the initial straight line, and the distances from other trace points to the straight line are calculated. Trace points with a distance less than 0.3 mm are selected as inliers, and the process is iterated 50 times to retain straight line segments with a proportion of inliers greater than 60%. The angle between adjacent straight line segments is calculated. If the angle is less than 5° and the distance between the endpoints is less than 2 mm, the two straight line segments are segments of the same trace line. The trace lines are refitted using the least squares method to form a continuous trace line network.
[0056] Step S30 , dividing the point cloud data according to the trace lines and generating a number of sub-surfaces, performing a local affine transformation on each sub-surface to obtain a corresponding flattened sub-surface, and splicing the flattened sub-surfaces to obtain a three-dimensional surface model.
[0057] It should be noted that sub-surfaces are localized 3D surface models. After dividing the point cloud data into several local point clouds based on trace lines, 3D reconstruction can be performed based on each local point cloud to reconstruct the sub-surface. By first dividing the point cloud data by trace lines and then generating several sub-surfaces, the boundaries of the sub-surfaces can be aligned with the trace lines, preserving the key details of the government document. This solves the problem of traditional full-surface flattening ignoring trace line details.
[0058] In one feasible implementation, a greedy triangulation method is used to reconstruct the sub-surfaces. Specifically, with the trace line as the boundary, seed points are selected from the region with the minimum curvature variance (the mean squared error of the principal curvatures of all points within the sub-surface, which reflects the geometric flatness of the sub-surface) and gradually grow to the trace line boundary. The growth constraints are set as follows: the difference in the curvature change rate of adjacent points is less than 0.2, and the angle between the normal vector of the seed point and the seed point is less than 15°. The polygonal areas formed by the intersection of the trace lines are subdivided using Delaunay triangulation to ensure that the sub-surface boundaries completely overlap with the structural lines to avoid distortion caused by subsequent flattening.
[0059] In a feasible implementation, the step S30 of "performing a local affine transformation on each sub-surface to obtain a corresponding flattened sub-surface" includes steps S301 to S302: Step S301: Perform a local affine transformation on the sub-surface to obtain a flattened sub-surface. Step S302, optimize the flattened sub-surface by solving the minimization energy function method to obtain an optimized flattened sub-surface, wherein the energy function includes a normal constraint term and a length constraint term, wherein the energy function is used to measure the degree of distortion of the flattened sub-surface, the first constraint term is used to measure the deviation between the normal vector of the flattened sub-surface and the normal vectors of the corresponding sub-surface points, and the second constraint term is used to measure the deviation of the vector length before and after flattening.
[0060] Specifically, the mathematical expression of the energy function is:
[0061] Where, is the first constraint, is the second constraint, Characterizes the degree of distortion when mapping the sub-surface to the flattened sub-surface. The smaller the value, the lower the distortion. For sub-face ( face ) at any point within for point The unit normal vector in three-dimensional space, is the unit normal vector of the flattened sub-surface, Represents one of the sub-faces, is the weight coefficient, which is used to balance the influence of the first constraint and the second constraint. is the boundary line of the sub-surface,q For The other endpoint of the boundary line is The optimized points are p and point q , The midpoints of the sub-surfaces p and point q The three-dimensional coordinates of Represents the Euclidean distance, which is used to measure the deviation of the vector length before and after flattening.
[0062] It should be noted that the first constraint is used to force the unit normal vector of all points in the sub-surface in three-dimensional space to be Unit normal vector to the flattened subsurface The angle is smaller than the preset angle. In this embodiment, the preset angle is set to 10°, and cos10°≈0.98 is used to ensure that , to prevent the sub-surface from folding or flipping during the flattening process, in this embodiment, The second constraint is determined by the principal component analysis (PCA) of the sub-surface point cloud, that is, the plane vector that best fits the sub-surface. and the midpoint of the sub-surface The Euclidean distance deviation is used to ensure the consistent size ratio of semantic elements such as text and seals.
[0063] Step S40 , extracting image modal features from the two-dimensional image, extracting geometric modal features from the three-dimensional surface model, and fusing the image modal features and the geometric modal features to obtain fused features.
[0064] It should be noted that image modal features include text strokes, seal patterns, texture grayscale and text lines, while geometric modal features include wrinkle depth, coordinates of sub-surfaces after flattening, normal vector direction, etc.
[0065] In a feasible implementation, the “fusing image modality features and geometric modality features to obtain fusion features” in step S40 may include steps S401 to S409: Step S401: fusing image modality features and geometric modality features based on a gating mechanism to obtain fused features.
[0066] It should be noted that, in this embodiment, the gating mechanism is used to dynamically calculate the fusion weight of the image modality feature and the fusion weight of the geometric modality feature for each feature dimension to fuse the image modality feature and the geometric modality feature to obtain the fusion feature.
[0067] Step S402: obtaining an initial neighborhood point set of each point in the three-dimensional surface model, and calculating the principal curvature of the initial neighborhood point set.
[0068] Specifically, the mathematical expression of step S402 is:
[0069] Where, are the first two eigenvalues of the covariance matrix of the initial neighborhood point set, is the principal curvature of the initial neighborhood point set, which reflects the degree of discreteness of each point in the initial neighborhood point set in the orthogonal direction, where The larger it is, the more severe the local bending is. Step S403: Calculate the dynamic neighborhood radius according to the principal curvature of the initial neighborhood point set.
[0070] Specifically, the mathematical expression of step S403 is:
[0071] Where, for point i The dynamic neighborhood radius, Basic neighborhood radius, is a learnable weight matrix used to control the influence of curvature on radius.
[0072] Step S404: Determine a neighborhood point set according to the dynamic neighborhood radius.
[0073] Step S405: Calculate the attention weight based on the point distance and normal vector angle of each point in the neighborhood point set.
[0074] Specifically, the mathematical expression of step S405 is:
[0075] Where, is the attention weight of the point pair, is the first attenuation coefficient, which is used to control the attenuation speed of the distance to weight. is the second attenuation coefficient, which is used to control the attenuation speed of the angle to the weight. is the distance between the point pairs, is the angle between the normal vectors of the point pair.
[0076] Step S406: Based on the attention weights, feature aggregation output is performed to obtain optimized geometric features.
[0077] Specifically, the mathematical expression of step S406 is:
[0078] Where, is the feature vector of the aggregated point, It represents nonlinear mapping of the spatial difference between two points to extract local geometric features.
[0079] It should be noted that by dynamically configuring the neighborhood range, the differences between crease edges and flat areas can be adapted, thereby improving the accuracy of feature extraction.
[0080] Step S407: extracting the main folding axis vector of the three-dimensional surface model.
[0081] Specifically, the main folding axis vector is used to characterize the overall folding trend of the document. The main folding axis vector of the three-dimensional surface model is extracted through PCA as part of the position embedding to enhance the model's perception of the folding structure.
[0082] Step S408: Perform curvature Fourier transform on the principal curvature of the initial neighborhood point set to obtain a high-dimensional curvature feature.
[0083] In step S409, after superimposing the main folding axis vector and curvature high-dimensional features of the three-dimensional surface model, the geometric modal features in the fused features are enhanced according to the attention mechanism and the optimized geometric features.
[0084] Specifically, the mathematical expression of “superposition of the main folding axis vector and curvature high-dimensional features of the three-dimensional surface model” in step S409 is:
[0085] Where, for point p The position encoding vector, in this embodiment, has a dimension of 512. Sine-cosine position encoding is used to capture the periodicity of spatial coordinates. To perform Fourier transform on the local curvature and extract the frequency characteristics of the curvature distribution (such as the periodic bending of wrinkles), is the curvature feature weight matrix. In this embodiment, The dimensions are 512×64.
[0086] The mathematical expression of “enhancing the geometric modal features in the fusion features according to the attention mechanism and the optimized geometric features” in step S409 is:
[0087] Where, Q 、 K 、 V They are the query, key, and value matrices in the attention mechanism, which are used to capture the correlation between features; is the transposed matrix of the key matrix, which is multiplied by the query matrix to calculate the similarity with ; For characterization Q and KThe feature dimension is used as a scaling factor to prevent the calculation result from being too large, which would cause the softmax function gradient to disappear and ensure the stability of the attention calculation. A learnable weight coefficient is used to balance the contribution of geometric distance constraints in attention calculation and regulate the influence of geometric information on attention weight; The geometric distance matrix incorporates the spatial geometric distance information between points, which is used to strengthen the attention mechanism's focus on spatial consistency and help the model better capture the global structure of document folds (such as folding direction and curling trend); is an activation function used to normalize the attention score so that the output weight value is in the range of [0, 1] and the sum is 1. Finally, the attention output is obtained by weighted summation of the weights.
[0088] The deep learning-based document wrinkle repair method of this embodiment breaks through the limitations of single geometric features and innovatively reconstructs the PointNet++ and Transformer architecture: through the curvature-driven dynamic neighborhood aggregation mechanism and the adaptive adjustment of the neighborhood range of the local curvature of the point cloud, and the introduction of point pair attention, it improves the accuracy of capturing wrinkle details and the feature capture accuracy of global structures (such as curling trends), and optimizes the point cloud processing efficiency and feature expression capabilities.
[0089] Step S50: Based on the fused features, cross-modal collaborative restoration of the two-dimensional image is performed to obtain a target image. Specifically, step S50 may include steps S501 to S502: Step S501 : Correcting the geometric modal features using the image modal features, and correcting the image modal features using the geometric modal features to obtain corrected image modal features and corrected geometric modal features.
[0090] Specifically, the geometric modal features are used to correct the image modal features to obtain the mathematical expression of the corrected image modal features:
[0091] Where, is the modal feature of the restored image, The extracted image modal features, MLP (C) is the feature after nonlinear mapping of the geometric modal features. The attention mechanism is used to allow the image modal feature restoration process to refer to the geometric modal features to avoid the disconnection between the image modal features and the geometric modal features.
[0092] The geometric modal features are corrected using the image modal features to obtain the corrected geometric modal features. The mathematical expression is:
[0093] Where, is the restored geometric modal feature, C is the geometric modal feature, and MLP(C) is the feature after nonlinear mapping of the geometric modal feature. The attention mechanism is used to allow the geometric modal feature to repair the modal feature of the reference image, making the geometric modal feature more consistent with the real texture feature and compensating for the restoration deviation caused by the lack of geometric modal feature constraints in traditional U-Net++.
[0094] Step S502 : performing geometric restoration, texture restoration, and structural restoration on the two-dimensional image according to the corrected image modal features and the corrected geometric modal features to obtain a target image.
[0095] Specifically, the corrected image modal features and the corrected geometric modal features are used as inputs of a multi-task image processing model to perform geometric restoration, texture restoration, and structural restoration on the two-dimensional image to obtain a target image.
[0096] The total loss function of the multi-task image processing model is:
[0097]
[0098]
[0099]
[0100] Where, is the total loss function; is the geometric loss, used to constrain the three-dimensional coordinate offset; The 3D coordinate change predicted by the model, such as the difference between the 3D point coordinates before and after restoration, reflects the magnitude of geometric correction; is the actual 3D coordinate change, which is used to determine the geometric deviation after ideal restoration as a supervision target. The actual 3D coordinate change is manually labeled.
[0101] Geometric loss is used to pass (L1 norm), constraining the accuracy of 3D coordinate restoration and avoiding 3D structure offset after restoration.
[0102] To avoid physical loss and ensure visual quality; The 2D texture image after repair predicted by the model, such as the document surface texture after crease repair, including details such as text and seals; For the real original 2D texture image, the ideal texture without creases is used as the supervision target; It is a structural similarity index used to measure the visual similarity between two images. The closer the value is to 1, the more consistent the texture is. Physical losses are used to pass , while constraining the pixel-level difference (L1) and visual perception difference (SSIM) to ensure the natural texture after restoration (no stretching or blur).
[0103] It is a structural loss, used to ensure the continuity of table lines and text rows; The structure mask predicted by the model (such as a binary image of a table line or text row, where 1 indicates the presence of structure and 0 indicates background); To predict the gradient (edge change) of the structure mask (reflecting the edge continuity of table lines and text rows); is the gradient (edge change) of the true structure mask (ideal structure edge, as a supervision target); Structural loss through (L1 norm), constraining the continuity of the structure edge (avoiding table line breakage and text line dislocation after repair); combined (binary cross entropy), while constraining the correctness of the structure "existence or not".
[0104] The deep learning-based document wrinkle repair method of this embodiment recognizes trace lines, then divides the point cloud data according to the trace lines and generates several sub-surfaces, performs local affine transformation on each sub-surface to obtain a corresponding flattened sub-surface, splices the flattened sub-surfaces to obtain a three-dimensional surface model, extracts image modal features from the two-dimensional image, and extracts geometric modal features of the three-dimensional surface model, fuses the image modal features and the geometric modal features to obtain fused features, and based on the fused features, cross-modal collaborative repair of the two-dimensional image to obtain the target image, thereby improving the repair effect.
[0105] In a feasible implementation, the step S10 of "obtaining point cloud data of the document to be processed" includes steps A11 to A14: Step A11: Acquire first-viewpoint point cloud data and second-viewpoint point cloud data.
[0106] It should be noted that the first-view point cloud data is different from the second-view point cloud data. For example, the first-view point cloud data may be front-side scanning point cloud data, and the second point cloud data may be back-side scanning data.
[0107] In step A12, based on the point distance constraint and the normal vector gradient consistency constraint, the first-view point cloud data and the second-view point cloud data are registered and spliced to obtain fused point cloud data.
[0108] Specifically, the first-view point cloud data and the second-view point cloud data are registered and spliced according to the error function. The mathematical expression of the error function is:
[0109] Where, is the total energy error of ICP registration, the smaller the value, the better the alignment effect; is the point distance constraint, which is used to minimize the Euclidean distance of corresponding points to achieve global alignment. p is a point in the source point cloud, q For the target point cloud Matching points, for p and q The square of the Euclidean distance; is the square of the normal vector gradient difference, which serves as the gradient constraint term. is the weight coefficient, which regulates the weight of the gradient constraint term and constrains the normal vector gradient Consistency, avoid geometric mutations in the area of the seams, for point p and point q Step A13: locate the seam area of the fused point cloud data according to the density change in the fused point cloud data.
[0110] It should be noted that in normal continuous areas, adjacent points are evenly distributed and the density changes gently. However, at the seam, since the point clouds on both sides are not completely aligned, the spatial distance between adjacent points will suddenly increase, forming a density fault. The seam can be located by calculating the area where the point cloud density changes suddenly.
[0111] In step A14, the seam area is processed using RBF interpolation to obtain optimized fused point cloud data.
[0112] In one specific implementation, the coordinates of points within the transition zone are determined by the point clouds on both sides of the seam. Points closer to the left are more influenced by the left point cloud, while points closer to the right are more influenced by the right point cloud. The middle points balance the weights on both sides, achieving a smooth transition from left to right. Therefore, a Gaussian kernel RBF function is used to reconstruct the point cloud coordinates. The coordinates of points within the transition zone are generated by weighting the point clouds on both sides, with the weights dynamically adjusted based on the distance from the seam.
[0113] The deep learning-based document wrinkle repair method of this embodiment achieves seamless fusion of multi-view point clouds through FPFH+ICP registration and RBF interpolation to smooth the seams.
[0114] In a feasible implementation manner, after step S10, the following steps are further included: Curvature analysis is performed on the point cloud data to screen out curled sub-surfaces, where the curvature of the curled sub-surfaces is distributed in a ring shape.
[0115] Specifically, the 3D point cloud is projected into polar coordinate space, and voting is performed using candidate rotation axes as parameters. Specifically, the point pairs in the point cloud are traversed to generate potential rotation axes. The variance of the distance from each point to the axis is calculated. Points with a variance less than 0.3mm are considered inliers, and the cumulative number of inliers serves as the voting score. When the voting threshold exceeds 50% (i.e., more than half of the points conform to a cylindrical distribution), the axis is determined to be a valid rotation axis.
[0116] Split the curly sub-surfaces according to the growth algorithm; Establish the parametric equation of the rotation axis of the curled sub-surface; Specifically, the parametric equation of the rotation axis is as follows:
[0117] Where r is the curling radius, which is calculated by the average distance from the inner point to the axis. is the intercept of the rotation axis in the xy plane, is the tilt parameter in the z-axis direction, which is obtained by optimizing the interior point coordinates through the least squares fitting method. : The three-dimensional coordinates of a point on the warped sub-surface in the 3D point cloud.
[0118] According to the parametric equation of the rotation axis of the curled sub-surface, the curled sub-surface is unfolded equidistantly to obtain the flattened sub-surface.
[0119] Specifically, the 3D cylindrical surface is mapped to the 2D plane, and the mapping formula is: , ; The specific mapping formula is:
[0120] Where, is the polar angle, which is calculated by the relative position of the point (x, y) and the rotation axis intercept (a, b), reflecting the position of the point on the circumference of the cylindrical surface. x, y are the xy coordinates of a point on the curled sub-surface in the 3D point cloud, u, v are the 2D plane coordinates (u is the circumferential coordinate, v is the axial coordinate), realizing the isometric mapping from the cylindrical surface to the plane, and z is the 3D point cloud of a point on the curled sub-surface. The coordinates are directly used as the axial coordinates of the 2D plane to ensure that the axial length is conserved.
[0121] It should be noted that for the cylindrical curled area, isometric unfolding is first achieved through rotation axis determination and parametric modeling, which solves the problems of texture stretching and structural distortion caused by traditional overall flattening.
[0122] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A document wrinkle repair method based on deep learning, characterized in that: The method comprises: Acquire a two-dimensional image and point cloud data of a document to be processed, and align each pixel of the two-dimensional image with the point cloud data; Calculating the curvature change rate of the neighborhood where each point in the point cloud data is located, and calculating the trace line according to the curvature change rate of the neighborhood where each point is located; Dividing the point cloud data according to the trace lines to generate a plurality of sub-surfaces, performing a local affine transformation on each of the sub-surfaces to obtain a corresponding flattened sub-surface, and splicing the flattened sub-surfaces to obtain a three-dimensional surface model; Extracting image modal features from the two-dimensional image and extracting geometric modal features from the three-dimensional surface model, and fusing the image modal features and the geometric modal features to obtain fused features; Based on the fused features, the two-dimensional image is repaired collaboratively across modalities to obtain the target image.
2. The document wrinkle repair method based on deep learning according to claim 1, characterized in that: The step of obtaining the point cloud data of the document to be processed includes: Obtaining a first-perspective two-dimensional image, a second-perspective two-dimensional image, and TOF depth information of the document to be processed; Calculating binocular depth information based on the parallax between the first-perspective two-dimensional image and the second-perspective two-dimensional image; Calculating the noise variance of the binocular depth information and the noise variance of the TOF depth information, and calculating the fusion weight of the binocular depth information and the fusion weight of the TOF depth information according to the noise variance of the binocular depth information and the noise variance of the TOF depth information; Performing Kalman filtering on the binocular depth information and the TOF depth information according to the fusion weight of the binocular depth information and the fusion weight of the TOF depth information to obtain fused depth information; The fused depth information is converted into point cloud data.
3. The document wrinkle repair method based on deep learning according to claim 2, characterized in that: The calculating and obtaining binocular depth information according to the parallax between the first-view two-dimensional image and the second-view two-dimensional image includes: Acquire a first-perspective two-dimensional image and a second-perspective two-dimensional image showing structured light fringes; According to the first-perspective two-dimensional image and the second-perspective two-dimensional image showing the structured light stripes, a first-perspective structured light phase image and a second-perspective structured light phase image are calculated by Fourier transform method; Calculate the disparity of each matching point pair between the first-view two-dimensional image and the second-view two-dimensional image according to the first-view grayscale image, the second-view grayscale image, the first-view structured light phase map, and the second-view structured light phase map; Binocular depth information is obtained by calculation according to the disparity of each matching point pair of the first-perspective two-dimensional image and the second-perspective two-dimensional image.
4. The document wrinkle repair method based on deep learning as described in claim 1, characterized in that: The method of calculating the curvature change rate of the neighborhood of each point in the point cloud data and obtaining the trace line according to the curvature change rate of the neighborhood of each point includes: Establishing a neighborhood point set of the neighborhood of each point in the point cloud data; Constructing a covariance matrix of the neighborhood point set, and performing eigenvalue decomposition on the covariance matrix to obtain discrete values of each point in the neighborhood point set in three orthogonal directions; The principal curvature of the neighborhood of each point in the point cloud data is calculated based on the discrete values in the three orthogonal directions, and the curvature change rate of the neighborhood of each point in the point cloud data is calculated based on the principal curvature of the neighborhood of each point in the point cloud data; Calculating a trace point based on a curvature change rate of a neighborhood where each point in the point cloud data is located and a preset curvature change rate threshold, wherein the curvature change rate of the neighborhood where the trace point is located is greater than the preset curvature change rate threshold; Perform straight line fitting on each trace point to obtain the trace line.
5. The document wrinkle repair method based on deep learning as described in claim 1, characterized in that: Generating a plurality of sub-surfaces includes: The sub-surfaces are reconstructed using the greedy triangulation reconstruction method.
6. The document wrinkle repair method based on deep learning as described in claim 1, characterized in that: The performing a local affine transformation on each of the sub-surfaces to obtain a corresponding flattened sub-surface includes: After performing a local affine transformation on the sub-surface, a flattened sub-surface is obtained; The flattened sub-surface is optimized by solving a minimization energy function method to obtain an optimized flattened sub-surface, wherein the energy function includes a first constraint term and a second constraint term. The energy function is used to measure the degree of distortion of the flattened sub-surface, the first constraint term is used to measure the deviation between the normal vector of the flattened sub-surface and the normal vectors of the corresponding sub-surface points, and the second constraint term is used to measure the deviation of the vector length before and after flattening.
7. The document wrinkle repair method based on deep learning as described in claim 1, characterized in that: The fusing the image modality feature and the geometric modality feature to obtain a fusion feature includes: Based on the gating mechanism, image modality features and geometric modality features are fused to obtain fusion features; Obtaining an initial neighborhood point set for each point in the three-dimensional surface model and calculating the principal curvature of the initial neighborhood point set; Calculating a dynamic neighborhood radius according to the principal curvature of the initial neighborhood point set; Determining a neighborhood point set according to the dynamic neighborhood radius; Calculate the attention weight based on the point distance and normal vector angle of each point in the neighborhood point set; Based on the attention weights, feature aggregation outputs are used to obtain optimized geometric features; extracting the main folding axis vector of the three-dimensional surface model; Perform curvature Fourier transform on the principal curvature of the initial neighborhood point set to obtain high-dimensional curvature features; After superimposing the main folding axis vector and curvature high-dimensional features of the three-dimensional surface model, the geometric modal features in the fused features are enhanced according to the attention mechanism and optimized geometric features.
8. The document wrinkle repair method based on deep learning as described in claim 1, characterized in that: The cross-modal collaborative restoration of a two-dimensional image based on fusion features to obtain a target image includes: Correcting the geometric modal features using the image modal features, and correcting the image modal features using the geometric modal features to obtain corrected image modal features and corrected geometric modal features; According to the corrected image modal features and the corrected geometric modal features, the two-dimensional image is geometrically repaired, texture-repaired and structurally repaired to obtain a target image.
9. The document wrinkle repair method based on deep learning as described in claim 1, characterized in that: The step of obtaining the point cloud data of the document to be processed includes: Acquire first-view point cloud data and second-view point cloud data; Based on the point distance constraint and the normal vector gradient consistency constraint, the first-view point cloud data and the second-view point cloud data are registered and stitched to obtain the fused point cloud data; Locating a seam area of the fused point cloud data according to density changes in the fused point cloud data; RBF interpolation is used to process the seam area to obtain optimized fused point cloud data.
10. The document wrinkle repair method based on deep learning as claimed in claim 1, characterized in that: Also includes: Performing curvature analysis on the point cloud data to screen out curled sub-surfaces, wherein the curvature of the curled sub-surfaces is distributed in a ring shape; Segmenting the curled sub-surface according to a growth algorithm; Establishing a parametric equation of the rotation axis of the curled sub-surface; According to the parametric equation of the rotation axis of the curled sub-surface, the curled sub-surface is unfolded equidistantly to obtain a flattened sub-surface.
Citation Information
Patent Citations
Photographing document bending correction method and device based on key point guidance
CN116740720A
Robust document image geometric distortion correction method based on selective state space sequence modeling and related device
CN119477768A
Image crease removing method and device and storage medium
CN119991512A
Target paper identification method based on point cloud data
CN120564058A
Image processing device, image processing method, and image processing program
US20240046494A1
Cited By
Three-dimensional point cloud registration method and system based on two-dimensional visual large model
CN121482122A