Intelligent Document Detection Method and System Based on OCR and ES

By employing an intelligent document detection method based on OCR and ES, and utilizing multi-stage gradient analysis and reconstruction technology, the problem of incomplete text edge extraction in complex backgrounds was solved, achieving high-quality text output and improving the recognition capability of the OCR system.

CN120954012BActive Publication Date: 2026-03-31STATE GRID GANSU ELECTRIC POWER CO LANZHOU POWER SUPPLY CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately capture the gradient changes of text edges in complex backgrounds, resulting in incomplete or broken character outline extraction, which limits the ability of images to generate high-quality text output.

Method used

We employ an intelligent document detection method based on OCR and ES, and optimize the visual saliency of character outlines through multi-stage gradient analysis and reconstruction, including gradient magnitude and direction calculation, edge continuity judgment, breakpoint connection, outline refinement and gradient reconstruction, combined with dynamic gradient tracking and multi-source fusion strategies.

Benefits of technology

It significantly improves the accuracy and completeness of text edge extraction, is suitable for document image processing in complex backgrounds, enhances the recognition rate of OCR systems in blurry and damaged documents, and improves the processing quality and usability of document images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954012B_ABST
    Figure CN120954012B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of document image processing and character recognition, and discloses a document intelligent detection method and system based on OCR and ES. The method comprises the following steps: pre-processing an input fuzzy image, extracting and enhancing edge gradient information, performing a series of processing such as gradient amplitude and direction calculation, edge continuity judgment, breakpoint connection, contour thinning, gradient reconstruction, dynamic tracking and multi-source fusion, gradually optimizing the character edge information in the image, and finally generating a clear character contour gradient graph. The method can effectively process low-quality, fuzzy or low-contrast document images, significantly improves the extraction accuracy and integrity of the character edge, and is suitable for document digitization and OCR recognition tasks in a complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of document image processing and text recognition technology, and in particular to a document intelligent detection method and system based on OCR and ES. Background Technology

[0002] With the acceleration of the digital transformation of society, document image processing and text recognition occupy a key position in the digital transformation and are widely used in scenarios such as archive management, legal document analysis and historical document restoration. Its importance lies in its ability to transform massive amounts of unstructured image data into searchable and editable digital information.

[0003] In existing technologies, document image processing methods mainly rely on traditional image preprocessing and feature extraction procedures. For example, global or adaptive binarization is used to separate foreground text from the background, median filtering or Gaussian filtering is combined for noise suppression, and edge detection operators are used to extract text contours. These methods can achieve good results in scenes with high image quality and clean backgrounds.

[0004] Existing technologies largely rely on simple image preprocessing techniques, such as binarization or noise filtering, without fully considering the gradient changes and structural continuity of text edges in complex backgrounds. Especially when dealing with blurry, damaged, or low-contrast documents, these methods struggle to accurately capture edge gradient information, leading to incomplete or broken character outlines and severely limiting the ability to generate high-quality text output. In summary, existing technologies struggle to accurately capture the gradient changes of text edges in complex environments and reconstruct clear character outlines based on this. Summary of the Invention

[0005] This invention provides a document intelligent detection method and system based on OCR and ES, which can accurately capture the gradient change pattern of text edges in complex environments and reconstruct clear character outlines based on this.

[0006] Firstly, to address the aforementioned technical problems, this invention provides a document intelligent detection method based on OCR and ES, comprising:

[0007] The initial blurred image is acquired and preprocessed, edge gradient information is preserved and gradient magnitude and direction are calculated, gradient changes are processed and edge continuity is determined, and a gradient map containing potential text edges is obtained.

[0008] If the gradient at the edge of the gradient map is discontinuous, the broken gradient change parts are connected and the character contour extraction is improved to obtain an enhanced second image.

[0009] The second image is subjected to edge thinning processing, local gradient direction is calculated and character outline is highlighted to obtain a refined edge map;

[0010] If the refined edge map character outlines are blurred, the missing gradient information is reconstructed to obtain the reconstructed third image.

[0011] The gradient change dynamic tracking is performed on the third image to determine the stable contour of the text edges, thus obtaining the optimized note image;

[0012] The dynamic tracking results in the note image and the gradient map containing potential text edges are fused to obtain a fused intermediate gradient map.

[0013] If the dynamic changes in the intermediate gradient map do not match the initial edges, adjust the tracking gradient dynamics to obtain the adjusted enhanced image;

[0014] Based on the enhanced image, the gradient magnitude and direction are recalculated according to the optimized note image, and the edge continuity is further determined to obtain the final gradient map after fusion.

[0015] Secondly, the present invention provides a document intelligent detection system based on OCR and ES, comprising:

[0016] The image preprocessing and edge detection module acquires the initial blurred image and performs preprocessing, retains edge gradient information and calculates gradient magnitude and direction, processes gradient changes and determines edge continuity, and obtains a gradient map containing potential text edges.

[0017] The edge enhancement and contour restoration module connects the broken gradient change parts and improves character contour extraction if the gradient map edge gradient is discontinuous, thus obtaining an enhanced second image.

[0018] The edge refinement module performs edge refinement processing on the second image, calculates the local gradient direction and highlights the character outline to obtain a refined edge map.

[0019] The gradient reconstruction module reconstructs the missing gradient information to obtain the reconstructed third image if the refined edge map character outlines are blurred.

[0020] The dynamic gradient tracking module dynamically tracks the gradient changes in the third image to determine the stable contours of the text edges, thereby obtaining an optimized note image.

[0021] The multi-source gradient fusion module fuses the dynamic tracking results in the note image with the gradient map containing potential text edges to obtain a fused intermediate gradient map.

[0022] The gradient dynamic optimization module adjusts the tracking gradient dynamics if the dynamic changes in the intermediate gradient map do not match the initial edges, thus obtaining the adjusted enhanced image.

[0023] The final gradient generation module recalculates the gradient magnitude and direction based on the enhanced image and the optimized note image, further determines the edge continuity, and obtains the fused final gradient map.

[0024] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the document intelligent detection method based on OCR and ES as described above.

[0025] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the document intelligent detection method based on OCR and ES as described above.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] (1) This invention achieves accurate capture of text edges in low-quality, low-contrast documents by performing multi-stage gradient analysis and reconstruction on blurred images. This method significantly improves the accuracy and completeness of text edge extraction through a series of processes such as gradient magnitude and direction calculation, edge continuity judgment, breakpoint connection, contour thinning and gradient reconstruction, and is especially suitable for document image processing in complex backgrounds.

[0028] (2) This invention achieves stable capture and optimization of text edge change trajectories by introducing dynamic gradient tracking and multi-source fusion strategies. Through motion vector analysis, interference region suppression and continuity verification, it effectively distinguishes between real edges and dynamic interference, enhances the robustness of the system under degradation conditions such as jitter and blur, and further improves the processing quality and usability of document images.

[0029] (3) This invention optimizes the visual saliency of character outlines by intelligently adjusting the local contrast and gradient distribution of the image, thereby improving the recognition rate of the OCR system for blurry and damaged documents. While maintaining the naturalness of the edge structure, this method effectively suppresses noise and enhances the separation effect between text and background, making it suitable for the digital processing of complex scenarios such as historical documents and handwritten notes. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the document intelligent detection method based on OCR and ES provided in the first embodiment of the present invention;

[0031] Figure 2This is a schematic diagram of the document intelligent detection system based on OCR and ES provided in the second embodiment of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Reference Figure 1 The first embodiment of the present invention provides a document intelligent detection method based on OCR and ES, including the following steps:

[0034] S11: Obtain the initial blurred image and perform preprocessing, retain edge gradient information and calculate gradient magnitude and direction, process gradient changes and determine edge continuity to obtain a gradient map containing potential text edges.

[0035] S12, if the gradient at the edge of the gradient map is discontinuous, then connect the broken gradient change parts and improve the character contour extraction to obtain the enhanced second image;

[0036] S13, perform edge thinning processing on the second image, calculate the local gradient direction and highlight the character outline to obtain a refined edge map;

[0037] S14, if the refined edge map character outlines are blurred, then reconstruct the missing gradient information to obtain the reconstructed third image.

[0038] S15, perform gradient change dynamic tracking on the third image to determine the stable contour of the text edges and obtain the optimized note image;

[0039] S16, perform attribute fusion between the dynamic tracking results in the note image and the gradient map containing potential text edges to obtain the fused intermediate gradient map;

[0040] S17, if the dynamic changes in the intermediate gradient map do not match the initial edge, adjust the tracking gradient dynamics to obtain the adjusted enhanced image;

[0041] S18, based on the enhanced image, the gradient magnitude and direction are recalculated in combination with the optimized note image to further determine the edge continuity and obtain the final gradient map after fusion.

[0042] In step S11, an initial blurred image is acquired and preprocessed, edge gradient information is preserved and gradient magnitude and direction are calculated, gradient changes are processed and edge continuity is determined, resulting in a gradient map containing potential text edges, including:

[0043] S1101, Noise filtering is performed on the input blurred image, and a weighted average is performed on each pixel to retain edge gradient information, resulting in a smoothed first image;

[0044] S1102, Calculate the gradient values ​​of each pixel in the horizontal and vertical directions in the first image, and generate a gradient magnitude map containing the gradient magnitude and direction.

[0045] S1103, for each pixel in the gradient magnitude map, if the gradient magnitude value is greater than the preset gradient saliency threshold, the gradient direction angle is calculated to obtain the direction marker image.

[0046] S1104, Based on the directional marker image, compare the gradient magnitude of each pixel with the magnitude of the adjacent pixels in the gradient direction and take the maximum value to obtain the suppressed gradient image;

[0047] S1105, if the gradient magnitude of a pixel in the gradient image is between a low threshold and a high threshold and is adjacent to a strong edge, it is marked as a weak edge and connected to obtain a final gradient map containing potential text edges.

[0048] In step S1101, noise filtering is performed on the input blurred image, and a weighted average is performed on each pixel to retain edge gradient information, resulting in a smoothed first image.

[0049] It should be noted that Gaussian filtering is used to perform convolution operations on the blurred input image. A weighted average of the image pixels is then applied using a Gaussian kernel, with higher weights at the kernel center and gradually decreasing weights around the edges. This preserves the gradient information of the image edges while smoothing noise. The weighting coefficients and Gaussian kernel parameters involved in this invention, such as size or standard deviation σ, and gain factors, can all be pre-calibrated experimentally based on image resolution, noise level, and application scenario. For example, the standard deviation σ of the Gaussian kernel is typically set to 0.5-1.5, and the gain factor is typically set to 1.2-2.0 to ensure a balance between smoothing noise and preserving edges. These parameters can also be embedded in a pre-trained model and automatically adapted based on the input image features.

[0050] In one embodiment, a grayscale image containing salt-and-pepper noise is convolved using a 3×3 Gaussian filter kernel. The standard deviation σ of this Gaussian kernel is set to 0.8, the center weight is 0.36, and the weight values ​​are 0.06, 0.10, 0.06, 0.10, 0.36, 0.10, 0.06, 0.10, and 0.06, respectively. For a given pixel in the image, a weighted calculation is performed based on the grayscale values ​​of its neighboring pixels to obtain a smoothed grayscale value. This method effectively reduces noise interference, making subsequent edge detection more accurate. Compared to the original image, the smoothed first image has slightly blurred details, but the overall structure is clearer, laying the foundation for subsequent processing.

[0051] In step S1102, the gradient values ​​of each pixel in the first image in the horizontal and vertical directions are calculated to generate a gradient magnitude map containing the gradient magnitude and direction.

[0052] It should be noted that when using the Sobel operator to calculate the horizontal and vertical gradient values ​​for each pixel, edge features are captured by detecting grayscale changes, and the horizontal convolution kernel of the Sobel operator is used. and vertical convolution kernel Two-dimensional convolution operations are performed with the first image respectively. The horizontal convolution kernel... Used to detect horizontal grayscale changes in an image, responding to vertical edges; vertical convolution kernel. Used to detect grayscale changes in the vertical direction of an image, responding to horizontal edges. Based on the calculated horizontal gradient value. and vertical gradient value The gradient magnitude of each pixel is calculated using the Euclidean norm formula. The calculation formula is as follows:

[0053]

[0054] Wherein, the gradient magnitude value This reflects the intensity of grayscale change at that pixel; the larger the amplitude, the higher the probability that the location is an edge. Based on the calculated horizontal gradient value... and vertical gradient value The gradient direction of each pixel is calculated using the arctangent function. The calculation formula is as follows:

[0055]

[0056] in, Accurately represents the direction of the gradient vector. This direction is perpendicular to the edge direction. It represents the gradient magnitude at each pixel. and gradient direction The information is combined to form the final gradient magnitude graph.

[0057] In one embodiment, for a specific pixel in the image, its horizontal gradient value is obtained after convolution using the Sobel operator. =40, vertical gradient value =30, and the calculated gradient magnitude is 50. The gradient magnitude is in units of image grayscale intensity, which has no actual physical dimensions. The direction is determined to be approximately 37 degrees using the arctangent function. This method can clearly describe the intensity and direction of the edge, providing a foundation for subsequent processing.

[0058] In step S1103, for each pixel in the gradient magnitude map, if the gradient magnitude value is greater than a preset gradient saliency threshold, the gradient direction angle is calculated to obtain a direction marker image.

[0059] It should be noted that the preset gradient saliency threshold is a pre-defined critical value for the gradient magnitude of edge pixels selected from the gradient magnitude map of an image. Its setting is primarily based on adaptive adjustments to image resolution and noise level; for example, it can be set as a certain proportion of the maximum gradient magnitude. Directional marker images present the directional information of edge pixels in a discrete form, which helps to highlight the edge features of text regions.

[0060] In one embodiment, when generating a direction-marked image based on the gradient magnitude map, for pixels with gradient magnitudes greater than a preset gradient saliency threshold, such as an 8-bit grayscale image with a gradient magnitude range of 0-255, a typical preferred value for the gradient saliency threshold can be set to 12% to 18% of the maximum gradient magnitude, for example, 46, which is approximately 18% of the maximum possible gradient value. This value can be adaptively adjusted by statistically analyzing the image gradient histogram or according to the image noise level, for example, set to 1.5 to 2 times the median or mean of all pixel gradient magnitude values. If a pixel has a magnitude of 50 and a direction angle of 37 degrees, its direction is classified into a predefined direction range, such as 0 degrees, 45 degrees, 90 degrees, or 135 degrees, resulting in a direction-marked image. This classification simplifies subsequent processing and facilitates unified analysis of edge directions.

[0061] In step S1104, based on the directional marker image, the gradient magnitude of each pixel is compared with the magnitude of the adjacent pixels in the gradient direction, and the maximum value is taken to obtain the suppressed gradient image.

[0062] In one embodiment, for each pixel, the gradient magnitudes of neighboring pixels are compared along its gradient direction, and only local maxima are retained. If a pixel has a magnitude of 50 and its gradient direction is 45 degrees, and its neighboring pixels along this direction have magnitudes of 40 and 30 respectively, then the pixel with a magnitude of 50 is retained as the maximum value. This method effectively reduces interference in non-edge areas, highlighting the clarity of text edges. Non-maximum suppression ensures edge refinement and avoids redundant pixels in edge areas.

[0063] In step S1105, if the gradient magnitude of a pixel in the gradient image is between a low threshold and a high threshold and is adjacent to a strong edge, it is marked as a weak edge and connected to obtain a final gradient map containing potential text edges.

[0064] It should be noted that, for the suppressed gradient image, a dual-threshold connection is used to further optimize the edges, classifying and filtering edge pixels to distinguish between confirmed true edges and possible noise or weak edges. The low threshold is a low gradient magnitude threshold; any pixel with a gradient magnitude below the low threshold is directly suppressed, i.e., considered background or noise. The high threshold is a high gradient magnitude threshold; any pixel with a gradient magnitude above the high threshold is immediately marked as a strong edge. Pixels with gradient magnitudes between the low and high thresholds are marked as weak edges. The high and low thresholds can be determined based on the statistical distribution of the image gradient magnitude. For example, the high threshold can be taken as the 90th percentile value on the gradient magnitude histogram, and the low threshold as the 30th percentile value. In a typical embodiment, for a document image with moderate contrast, the high threshold can be 60, and the low threshold can be 20, approximately one-third of the high threshold. These values ​​can be linearly scaled and adjusted according to the global contrast of the specific image.

[0065] In one embodiment, a high threshold of 60 and a low threshold of 20 are set, which are suitable for typical 8-bit grayscale images, with a gradient magnitude range of (0, 255). A pixel with a magnitude of 70, greater than the high threshold, is directly marked as a strong edge; another pixel with a magnitude of 30, between the high and low thresholds, and adjacent to strong edge pixels, is marked as a weak edge and connected to form an edge portion. This method generates a final gradient map containing potential text edges by distinguishing and connecting strong and weak edges. The dual-threshold strategy balances edge integrity and noise suppression, ensuring continuous and clear text edges.

[0066] In step S12, if the gradient map edge gradients are discontinuous, the broken gradient change portions are connected and character contour extraction is improved to obtain an enhanced second image, including:

[0067] S1201, Analyze the neighborhood gradient value of each pixel in the gradient map, and connect the broken parts if there are discontinuous edges to obtain the repaired gradient map.

[0068] S1202, detect the grayscale value of the repaired gradient map. If it is lower than the preset low contrast judgment threshold, mark it as a low contrast pixel to obtain a marked image.

[0069] S1203, the grayscale values ​​of the neighborhood of low-contrast pixels in the marked image are enhanced and adjusted to obtain an edge-enhanced image;

[0070] S1204, extract the character outline features of the edge-enhanced image. If a pixel belongs to a strong edge, the feature is retained to obtain a second image containing complete character outlines.

[0071] In step S1201, the neighborhood gradient value of each pixel in the gradient map is analyzed. If there are discontinuous edges, the broken parts are connected to obtain the repaired gradient map.

[0072] It should be noted that analyzing the neighborhood gradient value of each pixel in the gradient map requires first defining a matrix template with a specific shape and size to detect the structuring elements of local regions in the input image. The center point of each structuring element is sequentially aligned with each pixel in the gradient map, and the gradient values ​​of all pixels covered by the structuring element in its neighborhood are analyzed. If pixels with significantly different gradient values ​​exist, they are identified as discontinuous edges. If discontinuous edges exist, a morphological dilation operation is performed to connect the broken parts. For each position in the gradient map, the maximum gradient value of all neighboring pixels covered by the structuring element is calculated, and then this maximum value replaces the gradient value of the current center pixel, resulting in the repaired gradient map.

[0073] In one embodiment, a structuring element such as a 3x3 circular kernel is defined to scan the neighborhood of each pixel. If a pixel has a broken edge with a gradient value difference exceeding 20 in its neighborhood, the gap is filled by expanding neighboring pixels. For example, in the original gradient map, a text edge breaks at coordinates (100, 150), with neighboring pixel gradient values ​​of 0 and 45. The 0-value region is expanded by 2 pixels towards the 45-value region, connecting them into a continuous edge. This operation repairs minor breaks in the text outline, ensuring a smooth transition of edge lines.

[0074] In step S1202, the grayscale value of the repaired gradient map is detected. If it is lower than the preset low contrast determination threshold, it is marked as a low contrast pixel, and a marked image is obtained.

[0075] It should be noted that the low contrast threshold is a preset grayscale value used to distinguish areas with insufficient contrast in the gradient map, and its setting is mainly based on image statistics.

[0076] In one embodiment, the light gray edge pixels in the text area have a gray level of 90, which is below the low contrast threshold and is marked as 1, while other pixels are marked as 0, resulting in a marked image. This marking highlights low-contrast areas, facilitating subsequent targeted enhancement.

[0077] In step S1203, the grayscale values ​​of the neighborhood of low-contrast pixels in the marked image are enhanced and adjusted to obtain an edge-enhanced image.

[0078] It should be noted that enhancing the grayscale values ​​of the neighborhood of low-contrast pixels in the labeled image requires sharpening the edges by increasing the difference between the pixels and their surrounding pixels, i.e., the gradient. The labeled image obtained in the previous step is iterated. For each pixel marked as 1, the corresponding position is found in the original restored gradient map and used as the center pixel for neighborhood enhancement. A convolution kernel is calculated, representing the difference between a pixel and its neighboring pixels, to detect the rate of change in that region. The center of this convolution kernel is aligned with the current low-contrast center pixel to be processed. Each value of the kernel is multiplied by the corresponding pixel grayscale value in the image it covers. The products are then summed to obtain a single value, which is multiplied by a gain factor and then superimposed onto the original center pixel value to obtain the edge-enhanced image.

[0079] In one embodiment, when adjusting the grayscale value of the neighborhood of a low-contrast pixel, a convolution kernel-weighted gain processing is performed on the 8-neighborhood of the marked pixel. If the grayscale value of the center pixel is 90 and the average value of the neighborhood is 110, then the center pixel is boosted to approximately 135 by a gain factor of 1.5, and the enhanced edge pixels are obtained to form an edge-enhanced image.

[0080] In step S1204, the character contour features of the edge-enhanced image are extracted. If a pixel belongs to a strong edge, the feature is retained to obtain a second image containing complete character contours.

[0081] It should be noted that when extracting character contour features from the edge-enhanced image, the gradient magnitude of the image is first calculated. Then, each pixel in the gradient magnitude map is traversed. When its gradient value exceeds the strong edge threshold, the pixel coordinates and gradient direction information are added to the feature set. Finally, the corresponding positions of all coordinate points in the feature set are used for reconstruction in the new image. If the strong edge threshold is 200, pixels with gradient values ​​exceeding 200 retain their contour coordinates and direction information. A second image containing complete character contours is then reconstructed using this information. The strong edge threshold is a pre-set gradient magnitude threshold used to select the most significant and reliable edge pixels from the edge-enhanced image; its setting is mainly based on image statistics.

[0082] In one embodiment, in the enhanced edge of the character "O", the pixels of the enhanced edge are connected to form a closed curve, and the coordinate information of approximately 150 pixels is retained as contour features using the method described above. This extraction method based on threshold judgment and coordinate retention focuses on high-intensity edges, generating an accurate representation of text boundaries.

[0083] In step S13, the second image undergoes edge thinning processing, local gradient directions are calculated, and character outlines are highlighted to obtain a refined edge map, including:

[0084] S1301, Scan the gradient magnitude of each pixel in the second image to obtain a gradient magnitude sequence;

[0085] S1302, if the gradient magnitude of a pixel in the gradient magnitude sequence exceeds a preset edge retention threshold, then the corresponding pixel position is retained.

[0086] S1303, for the pixel position, identify continuous edge ends to obtain an edge end set;

[0087] S1304, verify the gradient direction of adjacent edge segments in the edge end set; if they are consistent, merge the adjacent edge segments.

[0088] S1305, strip the boundary pixels of adjacent edge segments after merging to obtain the refined edge map.

[0089] In step S1301, the gradient magnitude of each pixel in the second image is scanned to obtain a gradient magnitude sequence.

[0090] It should be noted that after scanning the gradient magnitude of each pixel in the second image, the image is traversed row by row and column by column from the top left corner. The Sobel operator is used to calculate the local gradient direction and the gradient magnitude of each pixel to obtain the gradient magnitude sequence. This operator convolves the image with horizontal convolution kernels such as [-101;-202;-101] and vertical convolution kernels such as [-1-2-1;000;121] to highlight the gradient changes of the character contours.

[0091] In step S1302, if the gradient magnitude of a pixel in the gradient magnitude sequence exceeds a preset edge retention threshold, the corresponding pixel position is retained.

[0092] It should be noted that the edge preservation threshold is a threshold value used to filter gradient magnitude sequences. It is used to determine whether each pixel in the sequence belongs to an edge feature worth preserving. Its setting is mainly based on the global contrast of the image and can be dynamically adjusted according to the image contrast. For example, it can be set to 60 in low-light text to ensure that weak edges are not filtered out.

[0093] In one embodiment, in a sequence of character outlines, the edge pixel amplitude reaches 150, while the non-edge amplitude is only 20. This serialization facilitates unified analysis and avoids missing subtle changes. Based on the gradient amplitude sequence, a threshold comparison algorithm is used to determine whether the gradient amplitude exceeds a preset edge retention threshold, such as 80. If it does, the pixel position is retained.

[0094] In step S1303, for the pixel position, continuous edge ends are identified to obtain an edge end set.

[0095] It should be noted that, for the preserved pixel locations, adjacent pixels are checked using 8-connectivity, and independent blocks are marked to analyze the algorithm and identify continuous edge segments, resulting in a set of edge segments. For example, the two leg edges of the character "A" each form two segments, each containing approximately 50 pixels. This identification separates isolated noise, which is beneficial for focusing on the true contour.

[0096] In step S1304, the gradient directions of adjacent edge segments in the edge end set are verified. If they are consistent, the adjacent edge segments are merged.

[0097] It should be noted that the average gradient direction angle of each edge segment is calculated. If the difference in the average gradient direction angle of adjacent edge segments is less than a preset angle tolerance threshold, they are determined to have the same direction and the two segments are merged into a single continuous edge segment. The angle tolerance threshold is a preset upper limit of angle difference used to determine whether two adjacent edge segments have the same direction, so that they can be merged. Its setting is mainly based on image quality and noise level, and can be set from 5° to 15°.

[0098] In one embodiment, the average orientation angle of each edge segment is calculated. If the orientation angle of a segment is 0 degrees and the orientation angle of its adjacent segment is 10 degrees, the difference between the two is 10 degrees. If the preset angle tolerance threshold is 15 degrees, then since 10 degrees is less than 15 degrees, the merging condition is met, and the two edge segments are merged. In this way, two originally broken horizontal line segments are merged into a single continuous line segment, increasing its total length from 30 pixels to 80 pixels. This merging strategy based on orientation consistency determination effectively enhances the continuity of the edge structure, which is beneficial for capturing the shape features of text completely and accurately.

[0099] In step S1305, the boundary pixels of the merged adjacent edge segments are stripped to obtain the refined edge map.

[0100] It should be noted that when stripping the boundary pixels of adjacent edge segments after merging, all pixels within the edge segment are scanned. While maintaining topological connectivity, boundary pixels that meet the removal conditions are removed. This process is repeated iteratively until no further stripping is possible, and finally, a center line with a single pixel width is obtained, resulting in a refined edge map.

[0101] In one embodiment, a coarse edge segment is refined from a width of 5 pixels to a single-pixel line through three iterations, preserving the core path. This refinement makes the contour more uniform, which is beneficial to the accuracy of subsequent matching algorithms.

[0102] In step S14, if the refined edge map character outlines are blurred, the missing gradient information is reconstructed to obtain a reconstructed third image, including:

[0103] S1401, Scan the gradient magnitude of each pixel in the edge map. If it is lower than the preset contour clarity threshold, it is determined to be a blurred contour position, and a set of blurred contours is obtained.

[0104] S1402, extract the values ​​of the four neighboring pixels around each blurred pixel in the blurred contour set, calculate the preliminary estimated gradient, and obtain the preliminary estimated gradient sequence;

[0105] S1403, perform linear interpolation calculations on the preliminary estimated gradient sequence in the horizontal and vertical directions respectively, and fuse the neighboring pixel value weights to obtain the reconstructed gradient information;

[0106] S1404, verify the consistency of the reconstructed gradient information direction and texture. If they are consistent, merge the continuous regions to obtain the reconstructed third image.

[0107] In step S1401, the gradient magnitude of each pixel in the edge map is scanned. If it is lower than the preset contour clarity threshold, it is determined to be a blurred contour position, and a set of blurred contours is obtained.

[0108] It should be noted that the preset contour sharpness threshold is a preset gradient magnitude threshold value used to distinguish which contour pixels in the edge map are sharp and which are blurry. It can be dynamically adjusted according to image characteristics, such as reducing it to 60 in low-contrast scenes to ensure that more potential blurry points are captured. The threshold is scanned row by row and column by column when traversing the edge map.

[0109] In one embodiment, in a text edge detection scenario, the outline sharpness threshold can be set to 80. When scanning an image, if a pixel is found to have a gradient amplitude of only 50, it is marked as a blurry outline point. This method accurately distinguishes between sharp and blurry areas through pixel-by-pixel analysis, forming a set of blurry outlines.

[0110] In step S1402, the values ​​of the four neighboring pixels around each blurred pixel in the blurred contour set are extracted, and the preliminary estimated gradient is calculated to obtain the preliminary estimated gradient sequence.

[0111] It should be noted that when calculating the preliminary estimated gradient, the weighted average of the four neighboring pixel values ​​of each blurred pixel in the blurred contour set is calculated, and the four directional neighbors are given the same weight. The preliminary estimated gradient value calculated for each blurred point is stored, and finally a preliminary estimated gradient sequence is obtained.

[0112] In one embodiment, for a blurred point, its neighboring pixel values ​​are 100, 90, 110, and 80, respectively, and the weighting coefficient is set to 0.25, resulting in a preliminary gradient value of 95. This method smooths out the influence of noise through neighborhood information, forming a preliminary estimated gradient sequence, providing a reliable basis for subsequent processing.

[0113] In step S1403, linear interpolation is performed on the preliminary estimated gradient sequence in the horizontal and vertical directions respectively, and the neighboring pixel value weights are fused to obtain the reconstructed gradient information.

[0114] In one embodiment, in a text edge region, the horizontal neighborhood gradient values ​​of a blurred point are 90 and 100, and the vertical gradient values ​​are 85 and 95. A reconstructed gradient value of 92.5 is calculated using bilinear interpolation. This method enhances the continuity of gradient directions by weighted fusion of neighborhood information, which helps to recover the detailed features of the blurred region. The weights can be dynamically adjusted according to the texture complexity, such as increasing the weight of the center pixel in high-frequency texture regions.

[0115] In step S1404, the consistency of the reconstructed gradient information direction and texture is verified. If they are consistent, the continuous regions are fused to obtain the reconstructed third image.

[0116] It should be noted that directional consistency is verified by calculating the absolute difference between the average gradient direction angles of the two regions. If this difference is less than the directional tolerance threshold, directional consistency is satisfied. Texture consistency is verified by calculating the cosine similarity between the texture feature vectors of the two regions. The closer the similarity is to 1, the more similar the two vectors are. If the similarity is higher than the texture similarity threshold, texture consistency is satisfied. The directional tolerance threshold is an angle value that defines the maximum allowable angle difference for determining whether two edges have consistent directions. It is mainly based on the stroke characteristics of characters and is set within the range of 3 to 10 degrees, preferably 5 degrees. The texture similarity threshold is a value between 0 and 1, defining the minimum similarity score required to determine whether two regions have consistent textures. Its setting is mainly based on the statistical results of a large number of sample experiments. The optimal value is determined by adjusting this threshold on the training set and observing the naturalness and integrity of the fused region, preferably 0.85.

[0117] In one embodiment, when processing the edges of the character "A", the network analyzes that if the gradient direction angle of a certain edge is 30 degrees and that of the neighboring region is 32 degrees, and the texture feature similarity reaches 90%, it is determined to be consistent and merged into a continuous region, generating a reconstructed third image. This method utilizes a deep learning model to capture complex patterns, extracted by a pre-trained small convolutional neural network. The network accepts image patches as input and outputs a high-dimensional feature vector through a series of convolution, non-linear activation, pooling, and fully connected layer operations. For example, two convolutional layers with a kernel size of 3x3 and 32 and 64 channels respectively, each followed by a ReLU activation function and a 2x2 max pooling layer, finally outputting a high-dimensional feature vector with a dimension of 128. The network is trained using contrastive learning or triplet loss. The training data contains a large number of labeled pairs of similar and dissimilar image patches. The goal is to minimize the distance between features of similar patch pairs and maximize the distance between features of dissimilar patch pairs. The optimizer can be Adam, with an initial learning rate of 0.001, which decays based on the performance on the validation set. Hyperparameters such as feature dimensions are tuned on a reserved validation set through cross-validation to maximize the accuracy of texture discrimination; the pre-trained neural network model is embedded in the system program and is directly called during processing to calculate the feature vectors of image patches.

[0118] In step S15, gradient change dynamic tracking is performed on the third image to determine the stable contours of the text edges, resulting in an optimized note image, including:

[0119] S1501, For the third image, analyze the gradient change vector between adjacent pixels in space to obtain the dynamic feature set and the gradient change trajectory of the text edge;

[0120] S1502, based on the dynamic feature set, scan the fuzzy interference region in the gradient change trajectory. If the gradient change vector variance of the pixels in the fuzzy interference region is greater than the preset gradient change vector variance threshold, it is determined to be a dynamic interference location, and the fuzzy interference set is obtained.

[0121] S1503, calculate the weighted average value of the pixels in the blurred interference set and fuse the features of the neighboring pixels to obtain a stable contour sequence;

[0122] S1504, perform edge connection on the continuous regions in the stable contour sequence. If the edge continuity meets the preset continuity threshold, then fuse it into the reconstructed third image to obtain the optimized note image.

[0123] In step S1501, for the third image, the gradient change vector between adjacent pixels in space is analyzed to obtain the dynamic feature set and the gradient change trajectory of the text edge.

[0124] It should be noted that an image-based spatial gradient analysis method is used when calculating the spatial gradient change vector. This method estimates the gradient change trend of local regions by analyzing the gradient magnitude and direction changes of adjacent pixels in the spatial domain, thereby capturing the spatial distribution characteristics of text edges and providing a basis for the stability analysis of text edges. The dynamic feature set includes features such as the magnitude, direction, and continuity of the gradient change vector in multi-scale space; the gradient distribution trajectory of text edges is composed of the gradient change sequence of edge pixels in their spatial neighborhood, reflecting the spatial consistency and stability of the edge structure.

[0125] In one embodiment, in a document image containing handwritten notes, for a specific text edge pixel, the gradient change between it and its eight surrounding pixels is analyzed. If the horizontal gradient change is 2 pixels in intensity and the vertical gradient change is 1 pixel in intensity, then the spatial gradient change vector of that point can be represented as (2,1). By performing spatial gradient change analysis on all edge pixels in the image, a dynamic feature set can be formed, thereby reflecting the spatial distribution characteristics and gradient change trajectory of the text edge structure.

[0126] In step S1502, based on the dynamic feature set, the blurred interference region in the gradient change trajectory is scanned. If the gradient change vector variance of the pixels in the blurred interference region is greater than the preset gradient change vector variance threshold, it is determined to be a dynamic interference location, and the blurred interference set is obtained.

[0127] It should be noted that the gradient change vector variance threshold is used to distinguish between flat regions and drastically changing interference regions. This threshold can be obtained based on training with a large number of samples. Statistical analysis of the gradient change vector variance of clear text regions and dynamically blurred regions reveals that the variance of the former is mostly concentrated between 0 and 0.2, while that of the latter is greater than 0.5. Therefore, a value between 0.3 and 0.5 can be selected as the threshold, preferably 0.4, to achieve the best classification effect; the variance is a normalized, unitless quantity, with a value range of [0,1].

[0128] In one embodiment, a preset gradient change vector variance threshold is set to 0.4. When scanning the gradient change trajectory, if the variance of the motion vector of pixels in a certain region reaches 0.7, it is determined to be a dynamic interference location. In a handwritten note-taking scenario, if the gradient change vector direction of pixels in a certain region is inconsistent due to hand tremors, and the variance exceeds the preset gradient change vector variance threshold, it is classified into the fuzzy interference set.

[0129] In step S1503, a weighted average value is calculated for the pixels in the blurred interference set and the features of neighboring pixels are fused to obtain a stable contour sequence.

[0130] It should be noted that, for calculating the weighted average value of pixels in the aforementioned blurred interference set, the target pixel is first located and its neighborhood is defined, and the feature values ​​of the neighboring pixels, such as gradient magnitude, are obtained. Then, a weighting coefficient is assigned to each pixel within the neighborhood. This coefficient can be determined based on the spatial distance between the pixel and the center point or feature similarity; common methods include using a Gaussian weighted kernel, where the standard deviation σ is a hyperparameter set according to the image noise level, typically between 0.5 and 1.5 pixels. During calculation, the feature values ​​of each neighboring pixel are multiplied by their corresponding weighting coefficients, summed, and then divided by the sum of the weighting coefficients to obtain the new feature value of the target pixel, thereby achieving feature fusion and smoothing. This process is performed pixel-by-pixel, ultimately outputting a more consistent and stable contour sequence after noise suppression.

[0131] In one embodiment, the grayscale values ​​of neighboring pixels of a blurred pixel within an interference region are 120, 110, 130, and 100, respectively, with a weighting coefficient of 0.25, resulting in a smoothed grayscale value of 115. This method utilizes neighborhood texture information to smooth noise caused by dynamic interference, forming a stable contour sequence.

[0132] In step S1504, continuous regions in the stable contour sequence are edge-connected. If the edge continuity meets a preset continuity threshold, it is fused to the reconstructed third image to obtain an optimized note image.

[0133] It should be noted that the continuity threshold is a numerical criterion between 0 and 1, used to quantify the reliability of the visual and structural connection between a repaired edge and the edge in the existing image. Its setting is mainly based on the objective requirements of error tolerance in different application scenarios, and can be set to 0.8 to 0.9.

[0134] In one embodiment, when processing the edge of the character "B", if its upper part is broken due to dynamic interference, a complete single-pixel wide edge line can be formed by fusing neighboring edge segments with high continuity. The continuity threshold for a certain text edge region is 0.8. If a candidate edge segment has a continuity score of 0.9, it is fused into the reconstructed third image. This method significantly improves the clarity and integrity of the note image through edge connectivity, making it particularly suitable for real-time optimization scenarios for handwritten notes.

[0135] In step S16, the dynamic tracking results in the note image and the gradient map containing potential text edges are fused to obtain a fused intermediate gradient map, including:

[0136] S1601, Based on the note image, calculate the gradient magnitude and gradient direction of each pixel in the gradient map and obtain the initial edge features from them to determine the preliminary text edge features;

[0137] S1602, if the gradient magnitude of the pixel at the initial text edge position is greater than the gradient magnitude of the neighboring pixel, then the pixel at the initial text edge position is retained as the edge point to obtain the refined edge set.

[0138] S1603, if the gradient magnitude of the edge points in the refined edge set is greater than the high threshold, they are marked as strong edge points; if the gradient magnitude of the edge points is between the low threshold and the high threshold, they are marked as weak edge points, and the classified edge set is obtained.

[0139] S1604, check the connectivity between weak edge points and strong edge points in the classification edge set. If there is a continuous path in the neighborhood, merge it into the strong edge point to obtain an intermediate gradient map.

[0140] In step S1601, based on the note image, the gradient magnitude and gradient direction of each pixel in the gradient map are calculated, and initial edge features are obtained from them to determine preliminary text edge features.

[0141] It should be noted that the gradient magnitude and gradient direction of each pixel in the gradient map depend on the brightness difference of the pixel's neighborhood. The gradient magnitude reflects the intensity of the change, while the direction indicates the edge orientation, thereby determining the initial text edge position.

[0142] In one embodiment, for a note image, if a pixel has a gradient magnitude of 15 and a direction of 45 degrees, it is included in the initial edge feature set. This method helps to accurately pinpoint the boundary contours of handwritten handwriting.

[0143] In step S1602, if the gradient magnitude of a pixel at the initial text edge position is greater than the gradient magnitude of a neighboring pixel, then the pixel at the initial text edge position is retained as an edge point, thus obtaining a refined edge set.

[0144] In one embodiment, for the initial text edge locations, the gradient magnitude of each pixel is scanned, and only points with a magnitude greater than that of their neighboring pixels are retained as edge points, forming a refined edge set. At the "stroke turning point" of the note, if the center pixel magnitude is 20, and the surrounding pixels (top, bottom, left, and right) are 18, 16, 19, and 17 respectively, then that point is retained. This suppression process eliminates redundant responses, ensuring that the edge points are thinner and sharper, which is beneficial for subsequent edge refinement.

[0145] In step S1603, if the gradient magnitude of the edge points in the refined edge set is greater than the high threshold, they are marked as strong edge points; if the gradient magnitude of the edge points is between the low threshold and the high threshold, they are marked as weak edge points, and a classified edge set is obtained.

[0146] In one embodiment, a dual-threshold approach is used to classify edge points based on a refined edge set. The high threshold is set to 25, and the low threshold to 10. Points with an amplitude greater than 25 are considered strong edges, while those between 10 and 25 are considered weak edges, thus obtaining a classified edge set. For example, the main edge of the handwritten character "A" with an amplitude of 30 is marked as strong, while the subtle branches with an amplitude of 12 are marked as weak. This classification strategy balances noise suppression with edge integrity preservation.

[0147] In step S1604, the connectivity between weak edge points and strong edge points in the classification edge set is checked. If a continuous path exists in the neighborhood, it is merged into the strong edge point to obtain an intermediate gradient map.

[0148] It should be noted that the connectivity between weak and strong edge points is checked by verifying whether a weak edge point is connected to a strong edge point through a continuous path consisting of edge pixels. If a continuous neighborhood path exists, the weak edge point is merged into the strong edge point to generate an intermediate gradient map.

[0149] In one embodiment, if a weak edge point with an amplitude of 15 in the note is connected to an adjacent strong edge point with an amplitude of 28 via a horizontal path, the continuity of the path is improved after fusion. This verification enhances edge coherence and avoids breaks.

[0150] In step S17, if the dynamic changes in the intermediate gradient map do not match the initial edges, the tracking gradient change dynamics are adjusted to obtain the adjusted enhanced image, including:

[0151] S1701, obtain the comparison results between the dynamically changing region and the initial edge from the intermediate gradient map. If the pixel value difference of the dynamically changing region exceeds the preset threshold, determine the mismatch position and obtain the low contrast sub-region to obtain the dynamic gradient set to be adjusted.

[0152] S1702, if the continuity of the boundary pixels of the low-contrast sub-region in the gradient dynamic set is less than the preset connection threshold, then the boundary is expanded to obtain the adjusted continuous gradient path.

[0153] S1703, determine the integrity of the continuous gradient path. If the number of connected components is greater than a preset component threshold, then fuse the connected components into a single path to obtain a preliminarily enhanced fused image.

[0154] S1704, extract the edge features of the fused image, determine whether the final continuity meets the preset integrity threshold, and obtain the adjusted enhanced image.

[0155] In step S1701, the comparison results between the dynamically changing region and the initial edge are obtained from the intermediate gradient map. If the pixel value difference of the dynamically changing region exceeds the preset gradient contrast tolerance threshold, the mismatch position is determined and the low contrast sub-region is obtained to obtain the dynamic gradient set to be adjusted.

[0156] It should be noted that the preset gradient contrast tolerance threshold is a critical value used to determine whether pixel value differences are significant, and it is dynamically calculated based on the characteristics of the image being processed. First, the average and standard deviation of the global gradient magnitude of the entire intermediate gradient map are calculated. Then, the gradient contrast tolerance threshold is set as the sum of the global average and a scaling factor. The scaling factor k needs to be set in a balance between the false detection rate and the false negative rate, and its preferred range is 0.8 to 1.5. For images with clear text and low noise, k can be set to a lower value, such as 0.8 to 1.2, to improve detection sensitivity; for images with strong noise or complex backgrounds, k can be set to a higher value, such as 1.2 to 1.5, to enhance anti-interference capabilities.

[0157] In one embodiment, the average global gradient magnitude of the intermediate gradient map is first calculated to be 10, with a standard deviation of 3, and k=1.0 is set. The resulting gradient contrast tolerance threshold is 13. When the pixel value difference in a certain region reaches 15, it is marked as a mismatch location. Low-contrast sub-regions are extracted from these, such as blurred stroke connections in a notebook, forming a dynamic set of gradients to be adjusted. This adaptive contrast method based on image statistical characteristics helps to accurately isolate interference sources caused by noise or inconsistent enhancement.

[0158] In step S1702, if the continuity of the boundary pixels of the low-contrast sub-region in the gradient dynamic set is less than a preset connection threshold, the boundary is expanded to obtain an adjusted continuous gradient path.

[0159] It should be noted that the preset connectivity threshold is used to quantitatively determine whether a sequence of boundary pixels is sufficiently continuous to decide whether boundary expansion through morphological operations such as dilation is necessary. The connectivity threshold should be set in relation to the image resolution and the expected minimum physical length of the stroke. The connectivity threshold should not be less than the minimum number of pixels required to represent a valid character stroke, such as a dot or the shortest stroke, at the current image resolution. For example, for a 300 DPI image, a stroke with a physical length of 1 mm corresponds to approximately 12 pixels. Therefore, the connectivity threshold can be preset to an empirical value of 4, which ensures that most false breaks of less than 3 pixels in length caused by noise are filtered out, while retaining meaningful short strokes.

[0160] In one embodiment, the original boundary pixels consist of 5 consecutive points, with a preset connectivity threshold of 4. The current boundary is already considered substantially continuous, but it is still expanded to 8 points through a slight dilation operation to enhance boundary strength. If the continuity is less than the preset connectivity threshold of 4, an erosion operation removes the outer pixels, refining the boundary to 6 points, thus obtaining an adjusted continuous gradient path. This sequence of operations ensures a smooth transition of the path.

[0161] In step S1703, the integrity of the continuous gradient path is determined. If the number of connected components is greater than a preset component threshold, the connected components are fused into a single path to obtain a preliminarily enhanced fused image.

[0162] It should be noted that the preset component threshold is a critical value used to determine the severity of fragmentation in a continuous gradient path. It was established by analyzing a large number of standard document image datasets, such as ICDAR, and finding that the number of connected components corresponding to clear character strokes is less than or equal to 3 in the vast majority (greater than 95%). Therefore, setting the component threshold to 3 can effectively distinguish between genuine, slightly broken strokes and noise fragments that need to be suppressed.

[0163] In one embodiment, based on the adjusted continuous gradient path, the connected component analysis scans the break points of the scan path. If the number of components is greater than the analysis technique theme, the fused intermediate gradient map is used to compare the dynamically changing region with the initial edge. If the pixel value difference exceeds the component threshold, the mismatch position is marked. The preset component threshold is 3, then neighboring components are fused, such as merging two separate stroke segments into a single path to obtain a preliminarily enhanced fused image.

[0164] In step S1704, the edge features of the fused image are extracted, and it is determined whether the final continuity meets the preset integrity threshold to obtain the adjusted enhanced image.

[0165] It should be noted that the preset integrity threshold is a scale used to quantify the continuity and completeness of edges or strokes. Its setting is primarily based on the statistical distribution of a large number of test images and the desired edge continuity requirements. For example, analysis of clear document samples shows that their edge continuity indices are mostly above 0.85. Therefore, the integrity threshold can be set between 0.8 and 0.9, with 0.85 being the preferred value. This threshold can also be automatically determined by optimizing OCR recognition accuracy on a validation set. For the initially enhanced fused image, edge features are extracted through multi-layer filtering. For example, if the input image size is 256x256, the network outputs an edge intensity map, from which the final continuity is determined.

[0166] In one embodiment, if the horizontal and vertical lines of the character "A" in the note are divided into four components, they are merged one by one to improve path uniformity. If the continuity score is lower than the integrity threshold of 0.85, iterative adjustments are made to obtain an enhanced image. When the network detects that the note edge continuity is 0.78, weak connections are strengthened to 0.92.

[0167] In step S18, based on the enhanced image and combined with the optimized note image, the gradient magnitude and direction are recalculated to further determine edge continuity, resulting in the fused final gradient map, including:

[0168] S1801, Based on the enhanced image, obtain and calculate the magnitude values ​​and direction angles of the horizontal and vertical gradients from the note image to obtain a preliminary gradient magnitude map and a preliminary gradient direction map;

[0169] S1802, Select the pixel with the largest amplitude along the direction angle in the preliminary gradient direction map from the preliminary gradient amplitude map. If the amplitude value of the pixel is less than the preset amplitude threshold, it is suppressed to obtain a refined gradient amplitude map.

[0170] S1803, obtain high-threshold strong edge pixels and low-threshold weak edge pixels from the gradient magnitude map, and connect the weak edge pixels with the strong edge pixels to obtain an edge continuous path map;

[0171] S1804, extract the edge feature vector of the edge continuous path map. If the continuity index of the edge feature vector is greater than the preset integrity threshold, then fuse the path into a single gradient structure to obtain the final gradient map.

[0172] S1805. To facilitate understanding of the present invention, some preferred embodiments of the present invention will be described in further detail below.

[0173] In step S1801, based on the enhanced image, the magnitude values ​​and direction angles of the horizontal and vertical gradients are obtained and calculated from the note image to obtain a preliminary gradient magnitude map and a preliminary gradient direction map.

[0174] It should be noted that in the scenario of handwritten note image processing, the optimized note image provides a clear foundation for subsequent edge detection, based on which the horizontal and vertical gradients are calculated using the Sobel operator.

[0175] In one embodiment, the Sobel operator uses a 3x3 convolution kernel to detect pixel intensity changes in the horizontal and vertical directions of the image, generating horizontal and vertical gradients. Taking a 256x256 handwritten note image as an example, a pixel with a horizontal gradient value of 10 and a vertical gradient value of 6 calculates an amplitude value of approximately 11.66 and a directional angle of approximately 31 degrees, forming a preliminary gradient amplitude map and direction map. This step captures the edge changes of the note strokes, laying the foundation for subsequent processing. For the preliminary gradient amplitude map, a non-maximum suppression operation is used to refine the edges, retaining the strongest edge pixels.

[0176] In step S1802, the pixel with the largest directional angle along the preliminary gradient direction map is selected from the preliminary gradient magnitude map. If the magnitude value of the pixel is less than a preset magnitude threshold, it is suppressed to obtain a refined gradient magnitude map.

[0177] It should be noted that the preset amplitude threshold is a critical value for gradient amplitude used to distinguish between real edges and noise. It is established based on the statistical characteristics of image noise level and edge intensity distribution. By analyzing the global or local statistical characteristics of the preliminary gradient amplitude map, a critical value that can effectively distinguish between real edges and noise is calculated.

[0178] In one embodiment, in the preliminary gradient magnitude map, a pixel is examined along the gradient direction. If its magnitude value is not the maximum in its neighborhood (e.g., its magnitude value of 11.66 is less than the maximum value of 15 along the gradient direction in its neighborhood), and this value is lower than the magnitude threshold set according to the overall image noise model, then the pixel is determined not to constitute a significant edge and is suppressed to generate a refined gradient magnitude map. This non-maximum suppression operation can effectively reduce noise interference and highlight the clear boundaries of strokes in the notes.

[0179] In step S1803, high-threshold strong edge pixels and low-threshold weak edge pixels are obtained from the gradient magnitude map, and the weak edge pixels are connected to the strong edge pixels to obtain an edge continuous path map.

[0180] In one embodiment, a high threshold of 20 and a low threshold of 8 are set. Pixels with an amplitude value greater than 20 are marked as strong edges, and pixels with values ​​between 8 and 20 are marked as weak edges. Taking the handwritten letter "A" as an example, pixels with an amplitude value of 25 in the horizontal bar portion are strong edges, while pixels with an amplitude value of 12 at the connection point are weak edges. Through connection rules, weak edge pixels that are adjacent to strong edges are retained, generating a continuous path map. This ensures the connectivity at the break in the stroke.

[0181] In step S1804, the edge feature vectors of the edge continuous path map are extracted. If the continuity index of the edge feature vectors is greater than a preset integrity threshold, the path is fused into a single gradient structure to obtain the final gradient map.

[0182] In one embodiment, a 256x256 path map is input into the network, and a feature vector is generated through multiple layers of convolution to calculate the continuity index. If the index value of 0.9 is greater than the preset integrity threshold of 0.85, the fusion path is a single gradient structure. Taking the character '日' in the note as an example, the intersections of the horizontal and vertical strokes may form multiple separate paths, which are fused into a unified structure after fusion. This method strengthens the integrity of the edges.

[0183] To facilitate the understanding of the present invention, some preferred embodiments of the present invention will be further described below.

[0184] In summary, the present invention discloses a document intelligent detection method based on OCR and ES, including multi-stage gradient processing and dynamic tracking of the input blurred image, gradually extracting, enhancing, refining, reconstructing and fusing the text edge information, and finally generating clear character contours. By combining gradient magnitude and direction analysis, edge continuity judgment, dynamic interference suppression and multi-source information fusion strategies, the present invention realizes the accurate reconstruction and image generation of the text structure in blurred and low-contrast document images under complex backgrounds, and significantly improves the recognition accuracy, robustness and output image recognizability of the OCR system under harsh conditions such as noise and blur.

[0185] Refer to Figure 2 , the second embodiment of the present invention provides a document intelligent detection system based on OCR and ES, including:

[0186] An image preprocessing and edge detection module, which acquires an initial blurred image and performs preprocessing, retains edge gradient information and calculates gradient magnitude and direction, processes gradient changes and judges edge continuity, and obtains a gradient map containing potential text edges;

[0187] An edge enhancement and contour repair module, if the edge gradient of the gradient map is discontinuous, connects the disconnected gradient change parts and improves character contour extraction, and obtains an enhanced second image;

[0188] An edge refinement module, which performs edge refinement processing on the second image, calculates the local gradient direction and highlights the character contour, and obtains a refined edge map;

[0189] A gradient reconstruction module, if the character contour of the refined edge map is blurred, reconstructs the missing gradient information, and obtains a reconstructed third image;

[0190] A dynamic gradient tracking module, which performs dynamic tracking of gradient changes on the third image, determines the stable contour of the text edge, and obtains an optimized note image;

[0191] The multi-source gradient fusion module fuses the dynamic tracking results in the note image with the gradient map containing potential text edges to obtain a fused intermediate gradient map.

[0192] The gradient dynamic optimization module adjusts the tracking gradient dynamics if the dynamic changes in the intermediate gradient map do not match the initial edges, thus obtaining the adjusted enhanced image.

[0193] The final gradient generation module recalculates the gradient magnitude and direction based on the enhanced image and the optimized note image, further determines the edge continuity, and obtains the fused final gradient map.

[0194] It should be noted that the document intelligent detection system based on OCR and ES provided in this embodiment of the invention is used to execute all the process steps of the document intelligent detection method based on OCR and ES in the above embodiment. The working principles and beneficial effects of the two correspond one-to-one, so they will not be described again.

[0195] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a document intelligent detection method program based on OCR and ES. When the processor executes the computer program, it implements the steps in the various embodiments of the document intelligent detection method based on OCR and ES described above, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, such as the dynamic gradient tracking module.

[0196] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.

[0197] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0198] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.

[0199] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0200] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0201] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0202] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for intelligent detection of documents based on OCR and ES, characterized in that, The method comprises the following steps: If the edge gradient of the gradient map is discontinuous, the disconnected gradient change part is connected and the character contour extraction is improved to obtain an enhanced second image; If the character contour of the refined edge map is fuzzy, the missing gradient information is reconstructed to obtain a reconstructed third image; If the dynamic change in the intermediate gradient map does not match the preliminary edge, the tracking gradient change dynamic is adjusted to obtain an adjusted enhanced image; Based on the enhanced image, the gradient amplitude and direction are recalculated in combination with the optimized note image, the edge continuity is further judged, and a final fused gradient map is obtained. The method comprises the following steps: If the edge gradient of the gradient map is discontinuous, the disconnected gradient change part is connected and the character contour extraction is improved to obtain an enhanced second image; If the character contour of the refined edge map is fuzzy, the missing gradient information is reconstructed to obtain a reconstructed third image; If the dynamic change in the intermediate gradient map does not match the preliminary edge, the tracking gradient change dynamic is adjusted to obtain an adjusted enhanced image; 2. The OCR and ES based document intelligence method of claim 1, wherein, Based on the enhanced image, the gradient amplitude and direction are recalculated in combination with the optimized note image, the edge continuity is further judged, and a final fused gradient map is obtained. The method comprises the following steps: If the edge gradient of the gradient map is discontinuous, the disconnected gradient change part is connected and the character contour extraction is improved to obtain an enhanced second image; If the character contour of the refined edge map is fuzzy, the missing gradient information is reconstructed to obtain a reconstructed third image; If the dynamic change in the intermediate gradient map does not match the preliminary edge, the tracking gradient change dynamic is adjusted to obtain an adjusted enhanced image; Based on the enhanced image, the gradient amplitude and direction are recalculated in combination with the optimized note image, the edge continuity is further judged, and a final fused gradient map is obtained.

3. The OCR and ES based document intelligence method of claim 1, wherein, The method comprises the following steps: If the edge gradient of the gradient map is discontinuous, the disconnected gradient change part is connected and the character contour extraction is improved to obtain an enhanced second image; If the character contour of the refined edge map is fuzzy, the missing gradient information is reconstructed to obtain a reconstructed third image; If the dynamic change in the intermediate gradient map does not match the preliminary edge, the tracking gradient change dynamic is adjusted to obtain an adjusted enhanced image; Based on the enhanced image, the gradient amplitude and direction are recalculated in combination with the optimized note image, the edge continuity is further judged, and a final fused gradient map is obtained.

4. The OCR and ES based document intelligence method of claim 1, wherein, ​ Scanning the gradient amplitudes of each pixel point of the second image, a gradient amplitude sequence is obtained; If the gradient amplitude of a pixel point in the gradient amplitude sequence exceeds a preset edge retention threshold, the corresponding pixel point position is retained; For the pixel point position, continuous edge ends are identified to obtain an edge end set; The gradient directions of adjacent edge segments in the edge end set are verified, and if consistent, the adjacent edge segments are merged; The boundary pixels of the adjacent edge segments after merging are stripped to obtain a refined edge map.

5. The OCR and ES based document intelligence method of claim 1, wherein, If the character outline of the refined edge map is blurred, missing gradient information is reconstructed to obtain a reconstructed third image, including: Scanning the gradient amplitudes of each pixel point of the edge map, if lower than a preset outline clarity threshold, the pixel point is determined as a blurred outline position to obtain a blurred outline set; Extracting the surrounding four neighborhood pixel values of each blurred pixel point in the blurred outline set, a preliminary estimated gradient is calculated to obtain a preliminary estimated gradient sequence; Linear interpolation calculation is performed on the preliminary estimated gradient sequence in horizontal and vertical directions respectively, and the neighborhood pixel value weight is fused to obtain reconstructed gradient information; The direction consistency and texture consistency of the reconstructed gradient information are verified, and if consistent, the continuous regions are fused to obtain the reconstructed third image.

6. The OCR and ES based document intelligence method of claim 1, wherein, Gradient change dynamic tracking is performed on the third image to determine the stable outline of the character edge to obtain an optimized note image, including: For the third image, the gradient change vectors between adjacent pixels in space are analyzed to obtain a dynamic feature set and a gradient change trajectory of the character edge; According to the dynamic feature set, the blurred interference region in the gradient change trajectory is scanned, and if the gradient change vector variance of the pixel points in the blurred interference region is greater than a preset gradient change vector variance threshold, the pixel points are determined as dynamic interference positions to obtain a blurred interference set; A weighted average value of the pixel points in the blurred interference set is calculated and the features of the neighborhood pixels are fused to obtain a stable outline sequence; The continuous regions in the stable outline sequence are connected, and if the edge continuity meets a preset continuity threshold, the continuous regions are fused to the reconstructed third image to obtain the optimized note image.

7. The OCR and ES based document intelligence method of claim 1, wherein, The dynamic tracking result in the note image and the gradient map containing potential character edges are attribute fused to obtain a fused intermediate gradient map, including: According to the note image, the gradient amplitude and gradient direction of each pixel point in the gradient map are calculated to obtain initial edge features and determine preliminary character edge features; If the gradient amplitude of a pixel point in the preliminary character edge position is greater than the gradient amplitude of a neighborhood pixel point, the pixel point in the preliminary character edge position is retained as an edge point to obtain a refined edge set; If the gradient amplitude of an edge point in the refined edge set is greater than a high threshold, the edge point is marked as a strong edge point, and if the gradient amplitude of the edge point is between a low threshold and a high threshold, the edge point is marked as a weak edge point to obtain a classified edge set; The connectivity of the weak edge points and the strong edge points in the classified edge set is checked, and if there is a continuous path in the neighborhood, the weak edge points are fused to the strong edge points to obtain an intermediate gradient map.

8. The OCR and ES based document intelligence method of claim 1, wherein, If the dynamic change in the intermediate gradient map does not match the preliminary edge, adjust the gradient change dynamic to obtain an adjusted enhanced image, including: Obtain a contrast result of the dynamic change region and the preliminary edge from the intermediate gradient map. If the pixel value difference of the dynamic change region exceeds a preset gradient contrast tolerance threshold, determine a mismatch position and obtain a low-contrast sub-region to obtain a gradient dynamic set to be adjusted; If the continuity of the boundary pixels of the low-contrast sub-region in the gradient dynamic set is less than a preset connection threshold, expand the boundary to obtain an adjusted continuous gradient path; Determine the integrity of the continuous gradient path. If the number of connected components is greater than a preset component threshold, fuse the connected components into a single path to obtain a preliminary enhanced fusion image; Extract the edge features of the fusion image and determine whether the final continuity meets a preset integrity threshold to obtain an adjusted enhanced image.

9. The OCR and ES based document intelligence method of claim 1, wherein, Based on the enhanced image, recompute the gradient amplitude and direction in combination with the optimized note image to further determine the edge continuity to obtain a fused final gradient map, including: Based on the enhanced image, obtain and compute the amplitude values and direction angles of the horizontal gradient and the vertical gradient from the note image to obtain a preliminary gradient amplitude map and a preliminary gradient direction map; Select a pixel point with the maximum amplitude along the direction angle of the preliminary gradient direction map from the preliminary gradient amplitude map. If the amplitude value of the pixel point is less than a preset amplitude threshold, suppress it to obtain a refined gradient amplitude map; Obtain high-threshold strong edge pixels and low-threshold weak edge pixels from the gradient amplitude map, connect the weak edge pixels with the strong edge pixels to obtain an edge continuous path map; Extract the edge feature vector of the edge continuous path map. If the continuity index of the edge feature vector is greater than a preset integrity threshold, fuse the path into a single gradient structure to obtain a final gradient map.

10. A system for intelligent detection of documents based on OCR and ES, characterized in that, Including: An image preprocessing and edge detection module obtains an initial blurred image and performs preprocessing, retains edge gradient information, computes gradient amplitude and direction, processes gradient change, and determines edge continuity to obtain a gradient map containing potential character edges; An edge enhancement and contour repair module. If the edge gradient of the gradient map is discontinuous, connect the disconnected gradient change parts and improve character contour extraction to obtain an enhanced second image; An edge refinement module performs edge thinning processing on the second image, computes local gradient direction, and highlights character contours to obtain a refined edge map; A gradient reconstruction module. If the character contour of the refined edge map is fuzzy, reconstruct the missing gradient information to obtain a reconstructed third image; A dynamic gradient tracking module determines the stable contour of the character edge by tracking the gradient change dynamic of the third image to obtain an optimized note image; A multi-source gradient fusion module attributes fuse the dynamic tracking result in the note image and the gradient map containing potential character edges to obtain a fused intermediate gradient map. A gradient dynamic optimization module adjusts the dynamic of tracking gradient change to obtain an adjusted enhanced image if the dynamic change in the intermediate gradient map does not match the preliminary edge. A final gradient generation module re-computes the gradient amplitude and direction based on the enhanced image and the optimized note image, further judges the edge continuity, and obtains a fused final gradient map.

Citation Information

Patent Citations

  • Methods and systems for text detection in mixed-context documents using local geometric signatures

    US7043080B1

  • Sharpness-based frame selection for OCR

    US9576210B1