A scanning file fuzzy retrieval self-adaptive splicing method

CN122265026BActive Publication Date: 2026-08-21BOWENDE (BEIJING) TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610290731.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-11
Publication Date
2026-08-21
Estimated Expiration
2046-03-11

AI Technical Summary

Technical Problem

[0003]当前主流文档图像拼接方法,普遍采用基于局部特征匹配的全局单应性刚性变换进行配准,大多面临以下技术性缺陷:仅依赖重叠区局部特征完成粗匹配,无法结合文档全局版面拓扑进行约束,难以对模糊图像块实现局部自适应形变修正

Benefits of technology

因为采用局部模糊检索与全局模糊检索相结合的技术手段,所以克服了模糊图像下局部特征鲁棒性差、初始位姿估计不准的问题,进而提升了图像块匹配与初步拼接的稳定性与精度;因为采用局部刚性纹理基元与全局拓扑约束流形特征点联合约束的技术手段,所以克服了仅依赖局部特征、缺乏全局版面约束的问题,进而保证文本行与线条连续完整,避免断裂、错位与结构扭曲;因为采用自适应形变网格与逐网格位姿补偿修正的技术手段,所以克服了全局刚性变换无法适配局部形变、拼接误差难以消除的问题,进而实现非均匀精准位姿修正,有效消除重影与接缝;因为采用移动终端非固定式分块扫描、自动捕获与模糊度筛选的技术手段,所以克服了手持拍摄不稳定、图像质量参差不齐的问题,进而适配移动办公、档案数字化、工程图纸电子化等实际场景,实现模糊条件下长幅文档的高质量自动拼接。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265026B_ABST
    Figure CN122265026B_ABST
Patent Text Reader

Abstract

The application provides a scanning file fuzzy retrieval self-adaptive splicing method, and relates to the technical field of document digitization and computer vision. The method comprises the following steps: obtaining a plurality of fuzzy image blocks obtained by non-fixed block scanning of a long document by a mobile terminal, wherein the long document comprises a layout structure composed of text lines and lines; performing local fuzzy retrieval on each fuzzy image block to extract local features of each image block, and determining an initial pose relationship and an overlapping area between the image blocks according to the local feature matching; determining a local rigid texture primitive in the overlapping area of each image block; performing preliminary splicing on the plurality of fuzzy image blocks according to the initial pose relationship and the overlapping area to obtain a preliminary splicing image, and determining a global topological constraint manifold feature point in the preliminary splicing image. The application can realize accurate splicing of fuzzy image blocks and layout structure restoration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of document digitization and computer vision technology, and in particular to a fuzzy retrieval and self-adaptive stitching method for scanned documents. Background Technology

[0002] In scenarios such as mobile office, document digitization, and engineering drawing digitization, users often use mobile terminals such as smartphones and tablets to perform non-fixed handheld block scanning of long contracts, engineering drawings, and scroll documents to obtain complete images of long documents. Due to the limitations of the field of view and shooting stability of mobile terminal cameras, the actual acquired multi-frame image blocks generally have motion blur and slight distortion, and adjacent image blocks are only stitched together by relying on local overlapping areas.

[0003] Current mainstream document image stitching methods generally employ global homography rigid transformation based on local feature matching for registration. However, these methods often suffer from the following technical limitations: they rely solely on local features in overlapping areas for coarse matching, failing to incorporate constraints based on the document's overall layout topology, and making it difficult to achieve local adaptive deformation correction for blurred image blocks. Due to these limitations, in practical applications, existing technologies not only exhibit poor robustness to local features under blurred conditions, but also experience a cumulative amplification of initial pose estimation errors, leading to broken text lines, misaligned lines, and distorted table structures after stitching, directly reducing the accuracy of subsequent OCR recognition and information extraction. Furthermore, global rigid transformation cannot adapt to local deformations, easily resulting in ghosting, misalignment, or seams in overlapping areas. Even pixel fusion struggles to eliminate structural errors, failing to meet the requirements of layout integrity and geometric consistency for long documents. These issues make it difficult for traditional methods to achieve high-quality automatic stitching of blurred, segmented scanned images under non-fixed and non-ideal shooting conditions on mobile terminals. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a fuzzy retrieval and self-adaptive stitching method for scanned documents, which can realize the high-precision automatic stitching of images of long documents under non-fixed, segmented, and motion-blurred conditions by mobile terminals.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: A fuzzy search and self-adaptive stitching method for scanned documents, the method comprising: Multiple blurred image blocks are obtained by non-fixed block scanning of a long document by a mobile terminal, wherein the long document contains a layout structure composed of text lines and lines; Local fuzzy retrieval is performed on each blurred image patch to extract local features of each image patch, and the initial pose relationship and overlapping region between each image patch are determined based on local feature matching; local rigid texture primitives are determined within the overlapping region of each image patch. Based on the initial pose relationship and overlapping area, multiple blurred image blocks are initially stitched together to obtain a preliminary stitched image. In the preliminary stitched image, global topological constraint manifold feature points are determined. A global fuzzy search is performed on the preliminary stitched image to extract global structural features, and the global structural features are compared with the local features of each image block to identify local areas with stitching errors. Within the local area, based on the spatial distribution relationship between local rigid texture primitives and global topological constraint manifold feature points, multiple adaptive deformation meshes are constructed with local rigid texture primitives and global topological constraint manifold feature points as vertices. Based on the topological consistency between adjacent adaptive deformation meshes, the local deformation gradient of each mesh is calculated. By comparing the vertex offsets of the theoretical mesh and the measured mesh, the pose compensation coefficients corresponding to each mesh are generated. Based on the pose compensation coefficient, the initial pose relationship of the corresponding image blocks is corrected grid by grid to obtain the corrected pose relationship, which is used to re-stitch multiple blurred image blocks to obtain the final stitched image.

[0006] The above-described solution of the present invention has at least the following beneficial effects: By employing a combination of local and global fuzzy retrieval techniques, the problems of poor robustness of local features and inaccurate initial pose estimation in blurred images are overcome, thereby improving the stability and accuracy of image block matching and preliminary stitching. By using a joint constraint technique combining local rigid texture primitives and global topological constraints on manifold feature points, the problem of relying solely on local features and lacking global layout constraints is overcome, ensuring the continuity and integrity of text lines and avoiding breaks, misalignments, and structural distortions. By employing adaptive deformation meshes and grid-by-grid pose compensation correction techniques, the problems of global rigid transformations being unable to adapt to local deformations and the difficulty in eliminating stitching errors are overcome, thereby achieving non-uniform and precise pose correction and effectively eliminating ghosting and seams. By employing non-fixed block scanning, automatic capture, and fuzziness filtering techniques on mobile terminals, the problems of unstable handheld shooting and inconsistent image quality are overcome, thus adapting to practical scenarios such as mobile office, archive digitization, and electronic engineering drawings, achieving high-quality automatic stitching of long documents under blurred conditions. Attached Figure Description

[0007] Figure 1 This is a flowchart illustrating an adaptive splicing method for fuzzy retrieval of scanned documents provided by an embodiment of the present invention. Detailed Implementation

[0008] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0009] like Figure 1 As shown, an embodiment of the present invention proposes a fuzzy retrieval and self-adaptive splicing method for scanned documents, the method comprising the following steps: Step 1: Obtain multiple blurred image blocks obtained by non-fixed block scanning of a long document by a mobile terminal. The long document contains a layout structure composed of text lines and lines. Step 2: Perform local fuzzy search on each blurred image block to extract local features of each image block, and determine the initial pose relationship and overlapping area between each image block based on local feature matching; determine local rigid texture primitives within the overlapping area of ​​each image block. Step 3: Based on the initial pose relationship and overlapping area, perform preliminary stitching of multiple blurred image blocks to obtain a preliminary stitched image. In the preliminary stitched image, determine the global topological constraint manifold feature points. Step 4: Perform global fuzzy search on the preliminary stitched image to extract global structural features, and compare the global structural features with the local features of each image block to identify local areas with stitching errors; within the local area, construct multiple adaptive deformation meshes with local rigid texture primitives and global topological constraint manifold feature points as vertices based on the spatial distribution relationship between local rigid texture primitives and global topological constraint manifold feature points. Step 5: Based on the topological consistency between adjacent adaptive deformation meshes, calculate the local deformation gradient of each mesh, and generate the pose compensation coefficients corresponding to each mesh by comparing the vertex offsets of the theoretical mesh and the measured mesh. Step 6: Based on the pose compensation coefficient, the initial pose relationship of the corresponding image blocks is corrected grid by grid to obtain the corrected pose relationship, which is used to re-stitch multiple blurred image blocks to obtain the final stitched image.

[0010] In this embodiment of the invention, it should be noted that the fuzzy retrieval refers to the operation of feature extraction and matching of the acquired image blocks with motion blur, rather than blurring the image itself; its purpose is to effectively detect layout structural features such as text lines and lines even under blurred conditions. By performing local and global dual-layer fuzzy retrieval on blurred image blocks obtained by non-fixed block scanning of mobile terminals, and constructing an adaptive deformation mesh by combining local rigid texture primitives and global topological constraint manifold feature points, and calculating deformation gradients and pose compensation coefficients based on mesh topological consistency, the mesh-by-mesh pose correction and re-stitching of blurred image blocks can be achieved. This can effectively improve the accuracy and stability of image stitching under motion blur conditions, suppress text line breaks, line misalignment, and stitching ghosting problems, ensure the integrity and geometric consistency of long document layout structures, and better adapt to the high-quality automatic stitching requirements of mobile terminals in non-ideal shooting scenarios.

[0011] In a preferred embodiment of the present invention, step 1 above may include: Step 1.1: In response to the scan start command, the mobile terminal's camera is invoked to capture video stream image frames of the document in continuous preview mode. Specifically, this includes: When responding to the scan start command, which is generated by the operation triggered for scanning long documents, it is used to start the complete processing flow of non-fixed block scanning. After the command is responded to, it immediately enters the continuous preview mode. This mode is characterized by continuous acquisition, real-time transmission, and real-time processing. It can continuously acquire image information of the target document area and simultaneously perform stable acquisition at a pre-set frame rate adapted to the mobile acquisition scenario. This frame rate takes into account both the smoothness of image acquisition and the real-time nature of subsequent processing, ensuring that the long document information, including text lines and line layout structure, is completely recorded during the acquisition process, forming a continuous and time-consistent video stream image frame. Each frame image is synchronously incorporated into the subsequent edge recognition, automatic cropping, and quality assessment processing flow.

[0012] Step 1.2 involves real-time detection of document edges in each video stream image frame. When a document boundary is detected to fall completely within a preset area of ​​the current frame, image capture of that frame is automatically triggered to obtain the original image frame containing local document content. Specifically, this includes: performing real-time, pixel-by-pixel document edge recognition and contour extraction on each video stream image frame; firstly, preprocessing the current frame image to reduce ambient light interference and suppress image noise; then capturing pixels with abrupt changes in grayscale values ​​in the image; connecting these pixels according to spatial continuity to form preliminary edge lines; performing denoising, smoothing, and connectivity restoration on the preliminary edge lines to remove discrete invalid edge points and retain complete document contour lines; and comprehensively judging whether the current frame contains complete document edge content based on the distribution of the extracted edge lines, the geometric regularity of the document region (such as the rectangularity and corner sharpness of the document contour), and the closure integrity of the contour. The system continuously tracks the positional changes of document boundaries across multiple consecutive frames to determine if they are stable and free from significant shifts or jitter. A preset effective acquisition range is a pre-defined image area capable of completely encompassing a single local area of ​​the document. When all four boundaries of the document fall within this preset effective acquisition range, and the positional shift of the document boundaries remains stable and without significant jitter across 3 to 5 consecutive frames, an image cropping operation is automatically triggered. The image area containing the complete local document content in the current frame is cropped as the original image frame. During cropping, all text lines, lines, and other layout structure information of the local document are preserved without losing any details. The cropped original image frames are then temporarily stored according to the acquisition sequence to establish a temporary storage sequence, providing a basis for subsequent overlap calculations, region comparisons, and stitching processes with other acquired original image frames.

[0013] Step 1.3: Based on the document area covered by the captured original image frames, generate a coverage heatmap of the currently scanned area, and calculate the directional prompts for the unscanned areas based on the coverage heatmap to guide the user's mobile terminal to cover the unscanned areas. Specifically, this includes: analyzing the physical coverage position and actual coverage range of each original image frame in the long document based on all acquired original image frames, accurately locating the coordinate range of each image frame in the overall layout of the long document, spatially overlaying and integrating the coverage areas of all original image frames, confirming the boundaries of the scanned document area and the range of unscanned blank areas; and generating a directional prompt for the user's mobile terminal in real time within the virtual stitching space. The coverage heatmap provides a clear visual indication of the acquisition progress. Based on the overall size of the long document, it uses a gradient color coding method to present the acquisition status. The areas that have been scanned are filled with dark colors (such as dark blue and dark green), and the color depth increases with the acquisition frequency. The unscanned blank areas are presented with light colors (such as light blue and light green) or transparent colors. The overlapping parts of adjacent scanned areas are marked with transition colors. At the same time, the outline of the scanned areas, the acquisition time sequence number of each original image frame, and the approximate area of ​​the unscanned areas are marked on the heatmap, which makes it easy to clearly distinguish between covered and uncovered areas and make the acquisition progress clear at a glance.

[0014] Based on the specific location and actual size of the uncovered areas presented by the coverage heatmap, as well as the orientation and distance of the uncovered areas relative to the current acquisition viewpoint, and combined with the distribution pattern of the acquired areas, a two-dimensional spatial coordinate system is first established with the center point of the current acquisition viewpoint as the origin. The geometric center of the uncovered area and the edge coordinates of the acquired areas are then mapped to this coordinate system. Next, the straight-line distance and azimuth angle between the geometric center of the uncovered area and the center point of the current acquisition viewpoint are calculated. Combined with the actual size of the uncovered area, the minimum required movement is determined. At the same time, the distribution density of the acquired areas is taken into account to avoid duplicate scanning areas and prioritize the selection of areas adjacent to the acquired areas that maximize the coverage of the unscanned areas. The optimal acquisition path for the domain is calculated using coordinate differences. The specific process is as follows: compare the geometric center coordinates of the uncovered area with the coordinates of the center point of the current acquisition viewpoint, calculate the coordinate differences between the two in the horizontal and vertical directions, determine the approximate direction of movement based on the sign of the coordinate difference, determine the basic range of movement based on the absolute value of the coordinate difference, and then adjust the movement range appropriately based on the actual size of the uncovered area to ensure that the core part of the uncovered area can be completely covered after movement, while taking into account the distribution of the already acquired area to avoid the movement path crossing the already scanned area and causing repeated acquisition. Finally, determine the optimal acquisition path from the current acquisition position to the target acquisition position that is non-repetitive and has the highest efficiency.

[0015] Based on the optimal acquisition path, the location prompts are transformed into concise and easy-to-understand information: the direction corresponding to the optimal acquisition path (determined by the sign of the coordinate difference) is translated into simple directional guidance, such as a positive horizontal coordinate difference indicating movement to the right and a negative one indicating movement to the left, and a positive vertical coordinate difference indicating movement downwards and a negative one indicating movement upwards; the movement range corresponding to the optimal acquisition path (determined by adjusting the absolute value of the coordinate difference combined with the size of the uncovered area) is translated into intuitive range prompts, such as a small range indicating a small movement, a moderate range indicating a constant movement, and a large range indicating a slow movement to the designated area; simultaneously, combined with the markings of uncovered areas on the coverage heatmap, the approximate location of the uncovered area is supplemented in the prompts, such as a small movement to the right if the unscanned area is on the right side of the current field of view. The location prompts are displayed synchronously with the coverage heatmap to guide the gradual adjustment of the acquisition direction, ensuring the orderly progress of the acquisition process and gradually covering the complete long document area, avoiding missed or repeated scanning.

[0016] Step 1.4: During the user's mobile terminal operation, continuously capture new raw image frames and calculate the overlap between each newly captured raw image frame and previously captured raw image frames. Only raw image frames with a preset overlap range with previously captured image frames are retained. Specifically, this includes: as the acquisition position gradually moves according to the directional prompts, continuously acquiring and temporarily storing raw image frames of newly entered long document areas within the acquisition field of view. For each newly acquired raw image frame, based on the document boundary features identified in the image frame, determine its corresponding two-dimensional coordinate range within the overall layout of the long document, i.e., the coordinate values ​​of the upper left and lower right corners, and delineate the outline of the document area corresponding to the image frame based on this coordinate range; then... The spatial position of the region contour of the new image frame is compared one by one with the region contours of all the original image frames that have been temporarily stored. The pixel-level region overlap calculation process is initiated: the coordinate system of the new image frame and the single temporarily stored image frame is unified into the global coordinate system of the long document. All pixels of the new image frame are traversed pixel by pixel. It is determined whether the coordinates of each pixel fall within the region contour range of the temporarily stored image frame. The total number of pixels that meet this condition is the number of overlapping pixels. Based on the number of overlapping pixels and the physical size of a single pixel, the actual area of ​​the overlapping region between the two image frames is calculated. Then, the area of ​​the overlapping region is divided by the total area of ​​the document region corresponding to the new image frame to obtain the proportion of the overlapping region to the total area of ​​the new image frame.

[0017] The preset reasonable overlap range is a numerical range pre-defined based on the requirements of long document stitching, typically covering an overlap ratio range of 20% to 40%. This range ensures that adjacent image frames have sufficient registration feature areas while avoiding data redundancy. During the overlap determination stage, if the overlap ratio of a new image frame with any temporarily stored image frame is within this preset range, the new image frame is determined to have registration value and is retained. If the overlap ratio is 0 (no overlap), less than 20% (insufficient overlap, unable to meet feature matching requirements), or greater than 40% (excessive overlap, causing data redundancy), the new image frame is determined to have no registration value and is automatically removed from the temporary queue. Only the original image frames that meet the overlap ratio requirements are retained, providing an effective data source for subsequent blur detection and stitching processing.

[0018] Step 1.5: Perform real-time blur detection on each retained original image frame to extract the gradient magnitude distribution features of the image frame. When the gradient magnitude distribution features are lower than a preset threshold, the image frame is determined to be a blurred image block and is removed; otherwise, the image frame is marked as a valid blurred image block. Specifically, this includes: performing a real-time sharpness evaluation process on each retained original image frame after overlap filtering, converting the image frame to a grayscale image to reduce color interference, and then performing block processing on the grayscale image to divide it into several local image blocks of uniform size, such as 8×8 pixels or 16×16 pixels. For each local image block, the gradient value is calculated pixel by pixel: taking the pixel as the center, the gray values ​​of its horizontally adjacent pixels are subtracted to obtain the horizontal gradient magnitude of the pixel; then the gray values ​​of its vertically adjacent pixels are subtracted to obtain the vertical gradient magnitude of the pixel; the sum of the squares and the square root of the horizontal and vertical gradient magnitudes of each pixel are processed to obtain the comprehensive gradient magnitude of the pixel. After calculating the gradient magnitudes of all pixels in a single local image block, the mean, maximum value and dispersion of the comprehensive gradient magnitudes of all pixels in the local image block are statistically analyzed.

[0019] After calculating the gradients of all local image blocks, a global statistical analysis is performed: the pixel gradient magnitudes of all local image blocks are summarized, and the arithmetic mean of the gradient magnitudes of all pixels in the entire image is calculated, i.e., the gradient magnitude mean. Then, the average of the sum of squares of the deviations of all pixel gradient magnitudes from this mean is calculated, i.e., the gradient magnitude variance, to reflect the overall distribution of gradient magnitudes in the entire image. Simultaneously, a local statistical analysis is performed: the gradient magnitude mean, maximum value, and dispersion of each local image block are recorded, and the differences in gradient magnitudes in different regions are analyzed, such as the difference in gradient magnitude distribution between text lines and blank areas, and the difference in gradient magnitude distribution between image edges and the center area, forming the distribution characteristics of gradient magnitudes in each region. Then, the gradient magnitude mean and variance obtained from the global statistical analysis are integrated with the gradient magnitude mean, maximum value, dispersion, and regional difference characteristics obtained from the local statistical analysis to form a set of gradient magnitude distribution characteristics that can comprehensively reflect the overall and local sharpness of the image. This feature reflects both the overall sharpness of the entire image and can accurately capture the blurry or sharp state of local areas.

[0020] The preset threshold is a judgment benchmark determined after training with a large number of long document image samples with different degrees of blur. It includes three core thresholds: gradient magnitude mean threshold, gradient magnitude variance threshold, and local gradient difference threshold. The mean threshold is used to judge the overall gradient strength of the image, the variance threshold is used to judge the uniformity of the gradient distribution, and the local gradient difference threshold is used to judge whether the gradient features of key areas such as text lines and lines are preserved. The three types of thresholds work together to effectively distinguish between slight blur that can be used for stitching and severe motion blur that cannot be registered. The extracted gradient magnitude distribution features are compared with the preset threshold item by item. If the feature values ​​are all below the threshold, it means that the pixel gradient changes in the image are weak and edge details are lost. The image frame is judged to have obvious motion blur and is directly excluded. If the feature values ​​reach or exceed the preset threshold, it means that although the image may have slight blur, it retains enough layout structure features such as text lines and lines, which meets the basic requirements for subsequent local blur retrieval and stitching. The image frame is then marked as a valid blurred image block that can be used for subsequent stitching processing and stored in the designated storage area according to the acquisition time sequence.

[0021] In a preferred embodiment of the present invention, step 2 above may include: Step 2.1: Perform multi-scale corner detection on the overlapping area of ​​each blurred image block to obtain corner response maps at each scale. Pixels with local maxima in their response values ​​are retained as corner candidate points through non-maximum suppression. Specifically, for each selected effective blurred image block, the overlapping area between the image block and adjacent effective blurred image blocks is accurately located by comparing the recorded coverage coordinate range of each effective blurred image block in the overall layout of the long document. The specific coordinates of the upper left and lower right corners of the overlapping area are confirmed, and a clear boundary of the overlapping area is defined. This strictly limits the scope of subsequent corner detection, ensuring that the operation is only performed on the overlapping area to avoid invalid detection of non-overlapping areas, effectively reducing computational redundancy and improving processing efficiency. Perform multi-scale corner detection on the overlapping area by pre-setting a set of reasonable detection window scale parameters, covering small scale (suitable for small features such as text stroke turning corners and fine line corners), medium scale (suitable for text line corners and medium-thickness line intersection corners), and large scale (suitable for document edge corners and thick line intersection corners), ensuring comprehensive capture of corner features of different sizes and types within the overlapping area.

[0022] During detection, detection windows of different sizes are used to traverse the overlapping region pixel by pixel in ascending order of scale. During this traversal, for each pixel, the corner response intensity is quantified by calculating the grayscale difference, gradient magnitude change, and grayscale contrast between that pixel and its neighboring pixels within the window, generating a corner response map for the corresponding scale. The corner response map is the same size as the overlapping region. The response value of each pixel directly represents the probability that the pixel is a corner; a higher response value indicates a more dramatic grayscale change, a more obvious corner feature, and stronger resistance to blurring interference. After detection at all scales is completed, non-maximum suppression is applied to the corner response maps generated at all scales. First, a local neighborhood window of the corresponding size is defined based on the current detection scale. The smaller the scale, the smaller the neighborhood window, ensuring accurate capture of local maxima; the larger the scale, the larger the neighborhood window, avoiding the omission of corner features over a wide range. Then, each pixel is traversed pixel by pixel, comparing the response value of each pixel with all pixels in its neighborhood. At the same time, combined with the preset corner response threshold, only pixels whose response value is the maximum value in the local range and whose response value is higher than the threshold are retained, while pixels with low response values, inconspicuous corner features, non-local maxima, and severe blur interference are removed. These pixels, after undergoing multi-scale detection and non-maximum suppression, are uniformly marked as corner candidate points. These corner candidate points can comprehensively cover corner features of different scales and types in the overlapping area, and have strong stability and recognizability, effectively avoiding the impact of motion blur on feature extraction.

[0023] Step 2.2 involves performing text line direction projection analysis on the same overlapping area. The starting and ending lines of the text lines are located using the troughs of the projection curve. Pixels at the starting and ending columns of each text line are used as candidate endpoints. Specifically, this includes: for the overlapping areas of the same valid blurred image block, preprocessing is performed. Noise smoothing and edge enhancement are used to suppress grayscale interference caused by image blurring, enhance the grayscale contrast between the text lines and the background, ensure the accuracy of subsequent projection analysis, and avoid projection deviations caused by text stroke adhesion and blurring. After preprocessing, the arrangement direction of the text lines is determined. By traversing the text pixel distribution within the overlapping area, it is determined whether the text lines are arranged horizontally or vertically. If horizontally arranged, projection is performed along the vertical direction; if vertically arranged, projection is performed along the horizontal direction, ensuring that the projection direction is perpendicular to the text line arrangement direction, maximizing the difference between the text lines and the blank areas.

[0024] Along the determined projection direction, the grayscale values ​​of each column (horizontal text line) or each row (vertical text line) within the overlapping area are accumulated. The accumulated grayscale value of each column / row is used as the vertical axis, and the column / row number is used as the horizontal axis to generate a projection curve that can intuitively reflect the distribution pattern of the text lines. In the projection curve, the peaks correspond to the text line areas with high accumulated grayscale values ​​because the grayscale of the text pixels differs greatly from the background. The valleys correspond to the blank areas between text lines with low accumulated grayscale values ​​and no obvious text pixels. To accurately locate the valleys, the projection curve is first smoothed to remove the interference of spikes in the curve. Then, a valley determination threshold is set (this threshold is an empirical determination value obtained based on the projection feature statistics of a large number of blurred document images, used to distinguish between effective blank intervals and noise fluctuations). Areas with accumulated values ​​lower than this threshold and located between two peaks are determined as effective valleys. By using the start and end positions of the effective valleys, the upper and lower boundaries (horizontal text lines) or left and right boundaries (vertical text lines) of each text line are accurately located, thereby determining the specific coordinates of the start and end lines of each text line.

[0025] After determining the start and end lines of a text line, traverse all columns or all rows of the text line. By comparing grayscale values, find the starting column where the text line pixels begin to appear and the ending column where the text line pixels disappear. Locate pixels at the intersection coordinates of the starting column and the corresponding starting row, and the ending column and the corresponding ending row. At the same time, filter these located pixels to remove false endpoints caused by image blurring or stroke adhesion, such as pixels with too low grayscale values ​​or no continuous text pixels around them. Use the filtered pixels as candidate endpoints of the text line. These candidate points can accurately represent the boundary features of the text line, supplement the shortcomings of the corner candidate points in step 2.1, and enrich the text structure features in the overlapping area.

[0026] Step 2.3 involves extracting line skeletons from the same overlapping region. Pixels at the intersections and endpoints of the line skeletons are located as candidate points for line structures. Specifically, this includes: extracting candidate points for line structures from overlapping regions of the same effective blurred image block; performing targeted preprocessing on the overlapping region; first, using adaptive Gaussian filtering to smooth noise and filter out random noise points caused by motion blur and ambient light interference to prevent noise from being mistakenly identified as lines; then, performing edge enhancement to strengthen the grayscale contrast between lines and the background, highlighting the outline boundaries of the lines and laying a clear image foundation for subsequent skeleton extraction; after preprocessing, performing skeletonization on the lines within the overlapping region. First, the image is converted into a black and white binary image through binarization, with the line area as the foreground and the background area as the background; then, the binarized lines are peeled off layer by layer, with each iteration removing only redundant pixels at the line edges and retaining the center pixels of the lines, until all lines are converted into continuous central skeletons of single-pixel width. This ensures that the skeleton can accurately restore the original line direction, intersection relationships, and topological structure, while avoiding the impact of uneven line thickness caused by blur on feature extraction.

[0027] After skeleton extraction, the line skeletons with a single pixel width are analyzed pixel by pixel and segment by segment. First, the skeleton feature point determination rules are set: the intersection position is a pixel where three or more skeleton branches meet, the branch position is a pixel where a secondary skeleton extends from the main skeleton, and the endpoint position is an isolated endpoint with only one adjacent skeleton pixel. The feature positions on the skeleton are identified one by one according to the rules, and the corresponding pixel is accurately located at each intersection position, branch position, and endpoint position that meets the rules. Then, the effectiveness of these located pixels is screened to remove false feature points caused by skeleton breakage, noise residue, or isolated points without continuous skeleton pixel support. Only the effective pixels that can truly reflect the line topology are retained. These effective pixels are uniformly marked as line structure candidate points. These candidate points can accurately represent the key structural features of lines such as table lines and separators in the document. They complement the corner candidate points and text line endpoint candidate points, further enriching the structural feature types in the overlapping area and improving the comprehensiveness and stability of the subsequent candidate point set.

[0028] Step 2.4 merges the corner candidate points, text line endpoint candidate points, and line structure candidate points to form an initial candidate point set. Specifically, this includes: summarizing the obtained corner candidate points, text line endpoint candidate points, and line structure candidate points; performing spatial deduplication on all the summarized points to remove duplicate points with completely overlapping coordinates, thus avoiding feature redundancy and repeated calculations; and then integrating the deduplicated feature points together to form an initial candidate point set that includes multiple structure types and has comprehensive coverage, providing a complete basic point set for subsequent feature selection and matching.

[0029] Step 2.5: For each point in the initial candidate point set, a local neighborhood window is defined centered on its coordinates. The gradient magnitude of all pixels within the window is calculated, and points with an average gradient magnitude greater than a first preset threshold are selected to form a strong gradient candidate point set. Specifically, this includes: for each feature point in the constructed initial candidate point set, first confirming the precise coordinates of the feature point within the overlapping region, and defining a fixed-size local neighborhood window centered on these coordinates. The window size is pre-determined and adjusted to fit the document feature scale, typically set to 5×5 pixels or 7×7 pixels. This ensures that the window can completely cover the local texture and gradient change information around the feature point, while avoiding excessively large windows that cause interference from irrelevant pixels or excessively small windows that fail to capture sufficient gradient features. After the window is defined, the gradient magnitude of each pixel within the window is calculated: the pixel and its horizontal... The gradient component of a pixel is calculated by subtracting the gray value of its left-hand neighbor from the gray value of its right-hand neighbor in the horizontal direction. A larger absolute value of this component indicates a more drastic change in gray level in the horizontal direction. Similarly, a larger absolute value of the gradient component in the vertical direction is obtained by subtracting the gray value of its upper-hand neighbor from the gray value of its lower-hand neighbor. The gradient components in both directions are then combined. First, each component is squared, then the squared results are added together, and finally, the square root of the sum is taken to obtain the overall gradient amplitude of the pixel. This amplitude comprehensively reflects the overall intensity of gray level changes in both the horizontal and vertical directions; a higher value indicates more prominent edge and texture features of the pixel.

[0030] After calculating the gradient magnitude of all pixels within the window, the gradient magnitudes of all pixels are statistically analyzed, and their arithmetic mean is calculated. This mean directly reflects the overall gradient intensity of the local area surrounding the feature point. The higher the mean, the more dramatic the gray-level changes around the feature point, the more obvious the feature, and the stronger the resistance to blurring interference. The calculated mean gradient magnitude is then compared one by one with a pre-set first threshold. This first threshold is an empirical value obtained based on the gradient feature statistics of a large number of effective blurred image blocks, which can effectively distinguish between strong and weak gradient feature points and is suitable for feature extraction of blurred images. To avoid gradient decay caused by ambiguity affecting the screening accuracy, after comparison, only candidate points with a mean gradient magnitude greater than the first preset threshold and high feature strength are retained, while candidate points with a mean gradient magnitude lower than or equal to the first preset threshold, weak gradient, susceptible to motion blur, and low feature recognition are removed. At the same time, the retained candidate points are verified a second time to remove points with abnormal gradient distribution within the window (such as excessively large gradient magnitude dispersion or no obvious gradient peak), ensuring that the retained candidate points all have stable and clear gradient features, and finally forming a set of strong gradient candidate points with stronger stability, higher recognition, and stronger anti-interference ability.

[0031] Step 2.6: Between adjacent image blocks, feature descriptors are constructed and matched for points in the strong gradient candidate point set. The Euclidean distance between matched point pairs is calculated, and point pairs with a distance less than a second preset threshold are retained. The coordinates of the matched point pairs in the two image blocks are recorded. Specifically, this includes: extracting the corresponding strong gradient candidate point sets between two adjacent effective blurred image blocks after screening; constructing a feature descriptor with rotation invariance, grayscale invariance, and anti-blurring interference capability for each strong gradient candidate point; the feature descriptor is a fixed-dimensional feature vector, and the construction process conforms to the features of the blurred image; selecting a local neighborhood of a fixed size around the strong gradient candidate point as the center, and connecting it with the features of the strong gradient candidate point. Step 2.5 adapts the local neighborhood window size to ensure feature consistency. First, the gradient direction and gradient magnitude of all pixels in the neighborhood are statistically analyzed. The gradient direction is divided into several uniform intervals, and the cumulative value of the gradient magnitude is calculated according to the interval. At the same time, the texture gray-level distribution features in the neighborhood are incorporated. After standardizing and encoding these statistical information, they are integrated to form a feature vector. This vector can completely and uniquely represent the gradient changes, texture distribution and structural features of the local area around the candidate point. It also has rotation invariance and gray-level invariance, which can effectively resist feature interference caused by motion blur, illumination changes and image rotation, and ensure that the descriptors of feature points corresponding to the same physical location in different image blocks have high similarity.

[0032] After constructing the feature descriptors, each strong gradient candidate point in the first blurred image block is paired with all strong gradient candidate points in the second blurred image block. The Euclidean distance is calculated for each pair of descriptors to be matched. The magnitude of the Euclidean distance is used to measure the similarity between the two sets of features. The smaller the distance, the more similar the features are and the higher the matching confidence. At the same time, a second preset threshold is set. This threshold is obtained through a large number of blurred image matching experiments and is used to distinguish between correct and incorrect matches. Only matching point pairs with an Euclidean distance less than the second preset threshold are retained, while mismatched point pairs with excessively large distances, obvious feature differences, or susceptibility to blurring are removed. The high-confidence matching point pairs are then subjected to consistency verification to further exclude abnormal matches with strong dispersion. For each valid matching point pair, its precise coordinate position in the first and second blurred image blocks is recorded, forming a stable set of matching point pairs with one-to-one coordinate correspondence.

[0033] Step 2.7: Mark the points located on the current image patch from the retained matching point pairs as local rigid texture primitives. Specifically, this includes: determining the affiliation of each pair of points in two adjacent valid blurred image patches based on the obtained high-confidence matching point pairs, determining the image patch affiliation of each point, and filtering out all matching points belonging to the currently processed blurred image patch. These filtered points are the local rigid texture primitives. The local rigid texture primitives are feature points with stable spatial structure characteristics and strong anti-blurring interference capabilities. Their core feature is that they can maintain their relatively stable spatial position relationship and local structural features during the motion and slight deformation of the image acquisition process. It will not exhibit significant feature shifts or distortions due to motion blur, slight displacement, or image deformation. These primitives are all derived from high-confidence matching point pairs selected in step 2.6. After multiple rounds of selection based on gradient strength and feature similarity, they retain the core features of corner points, text line endpoints, and line structure points, while possessing strong recognizability and stability. They can effectively resist interference such as feature weakening and edge blurring caused by motion blur, accurately reflecting the true spatial correspondence between two adjacent blurred image blocks. At the same time, their stable spatial structure characteristics can serve as the core constraint for subsequent initial pose estimation and deformation correction of adjacent image blocks, ensuring the accuracy of pose calculation and deformation correction.

[0034] In a preferred embodiment of the present invention, step 3 above may include: Step 3.1 involves performing text line connected component analysis on the preliminary stitched image, extracting the bounding rectangles of each connected component, and identifying continuous text lines across image blocks based on the aspect ratio of the bounding rectangles and the arrangement of adjacent rectangles. Specifically, this includes: performing targeted preprocessing operations on the preliminary stitched image formed through local rigid texture primitive matching and preliminary stitching; converting the preliminary stitched image to grayscale to eliminate interference from color channels and ensure consistency in subsequent text region analysis; and performing noise reduction to filter out noise points and pseudo-pixels caused by motion blur and image stitching, preventing noise from interfering with the aggregation of text pixels. The grayscale contrast between the text and the background is slightly enhanced, laying a clear image foundation for text line connected component analysis. After preprocessing, text line connected component analysis is performed on the text region in the image. A reasonable grayscale threshold is set, and the grayscale image is binarized so that the text pixels appear as foreground and the background appears as white, accurately distinguishing the text and background regions. Relying on the connected component labeling algorithm, all foreground pixels in the image are traversed, and text pixels with the same grayscale value and adjacent spatial position are aggregated into independent connected components. Each connected component corresponds to a continuous text segment, and a unique label is assigned to each connected component to facilitate subsequent differentiation and analysis.

[0035] After connecting component labeling, the smallest bounding rectangle that can completely enclose all pixels of each connected component is calculated. The coordinates of the top-left and bottom-right corners of each bounding rectangle are accurately obtained, and the length, width, and aspect ratio of the rectangle are calculated accordingly. The center coordinates and height of each rectangle are recorded simultaneously. All bounding rectangles are then filtered according to the set text line feature criteria, selecting rectangles with a reasonable aspect ratio that fits the document text line format. The heights of adjacent rectangles are compared, retaining those with strong height consistency and excluding those with excessive height differences or that do not belong to the same line of text. At the same time, the horizontal spacing of adjacent rectangles is analyzed, retaining those with uniform spacing and horizontal alignment. This process forms a set of rectangular boxes that conform to the characteristics of normal text lines. By combining the recorded boundary coordinates of each image block, the continuity and alignment of adjacent rectangular boxes are further judged. If multiple rectangular boxes are separated by the boundaries of image blocks, but still maintain horizontal alignment, consistent height, uniform spacing, and continuous text content, then the connected components corresponding to these rectangular boxes can be identified as belonging to the same line of text. This completes the identification of continuous text lines across image blocks. Simultaneously, the starting and ending image blocks, coverage coordinate range, and text distribution characteristics of the continuous text lines across image blocks are recorded. This provides a precise and accurate text structure target for text line refinement and global candidate point extraction, ensuring that global feature extraction can fit the actual text distribution pattern of the document.

[0036] Step 3.2 involves refining the continuous text lines across each image block, extracting the central skeleton line of the text line, and locating pixels at the turning points, endpoints, and intersections of two or more central skeleton lines as first-class global candidate points. Specifically, this includes: performing targeted preprocessing and refining operations on each identified continuous text line across the image block; firstly, performing edge enhancement processing on the region where the text line is located, strengthening the grayscale difference between the text stroke edges and the background through gradient operators; then, setting a specific binarization threshold based on the grayscale features of the text line to complete the binarization processing of the text line region, making the text pixels the foreground and the background white, further highlighting the contour features of the text line; based on the binarized text line region, performing pixel-by-pixel stripping, removing only redundant pixels at the edges of the text strokes in each iteration, retaining the core pixels at the center of the strokes, until all text line strokes are converted into a single-pixel-width central skeleton line. This skeleton line can accurately restore the original direction, curvature, and connection relationship of the text line, making the overall structure of the text line clearer and more discernible.

[0037] After extracting the central skeleton line, each skeleton line is traversed and analyzed segment by segment. A criterion for determining directional changes is set. The corresponding pixels are accurately located at turning points where the angle of the skeleton line changes beyond a preset angle threshold (where the preset angle threshold is an empirical value obtained from the statistics of the conventional turning angles of the text lines, which can effectively distinguish between normal changes in the direction of the text lines and minor directional shifts caused by noise), the starting and ending points of the skeleton line where there are no adjacent pixels at the beginning and end, and the intersection points where two or more skeleton lines meet. At the same time, false feature points caused by text blurring and breakage are removed. These effective pixels that can accurately represent the global topological structure of the text lines are marked as the first type of global candidate points. These candidate points, starting from the overall structural level of the text lines, provide stable text structure constraints for the subsequent global splicing accuracy.

[0038] Step 3.3 involves performing line segment detection on the preliminary stitched image, extracting continuous lines with a length greater than a preset length threshold, and selecting lines that cross image block boundaries as cross-block continuous lines. Specifically, this includes: performing line segment detection on the preliminary stitched image; first, preprocessing the image by converting it to grayscale and Gaussian smoothing to filter out high-frequency noise and prevent noise from being misidentified as short line segments; then, traversing all pixel segments in the image, identifying all line segments that conform to the characteristics of a straight line by calculating the gradient direction and grayscale consistency of the pixel segments, and recording the starting point, ending point coordinates, and direction of each line segment; calculating the actual physical length of each line segment based on its coordinate information, and then... Line segments exceeding a pre-set length threshold are retained as valid straight line segments. This pre-set length threshold, combined with the standard size settings of table lines and separators in the document, can effectively exclude invalid linear structures such as short, messy lines and noisy line segments. Based on the recorded boundary coordinates of each image block, it is determined whether the coordinate range of each valid straight line segment crosses the splicing boundary of adjacent image blocks. If the straight line segment covers the image areas on both sides of the splicing boundary, and the line segments on both sides of the boundary maintain a consistent direction without obvious offset or breakage, it can be filtered as a continuous line across blocks. These lines are mostly rigid structural lines such as table borders and separators in the document, which can provide strong positional constraints for global splicing.

[0039] Step 3.4: Skeletonize each continuous line spanning a block. Locate pixels at the branch points, inflection points, and intersections of the line skeleton with the image block boundary as second-class global candidate points. Specifically, this includes: binarizing each selected continuous line spanning a block; setting a specific threshold based on the line's grayscale features to make line pixels the foreground and background white; then using median filtering for smoothing to fill in minor breaks caused by blurring and eliminate noise interference with the line structure; and performing skeletonization on the preprocessed continuous lines spanning a block by peeling away the edge pixels layer by layer to uniformly convert lines of different thicknesses into a single-pixel width center. The skeleton fully preserves the original line direction, length, and intersection with the image block boundary. After skeleton extraction, the line skeleton is analyzed point by point. At the branch points where the skeleton branches into secondary branches, the inflection points where the direction changes angle beyond a set threshold, and the intersection points where the skeleton intersects with the image block splicing boundary, the corresponding pixels are accurately located. At the same time, the continuity of the surrounding skeleton of each point is checked, and false feature points formed by line blurring and breakage are eliminated. These effective pixels that can reflect the global straight line structure and the spatial relationship of the splicing boundary are marked as the second type of global candidate points. These candidate points supplement the global features at the straight line structure level, further improving the stability of the splicing result.

[0040] Step 3.5 merges the first and second types of global candidate points to form a global candidate point set. Specifically, this includes: aggregating the first and second types of global candidate points into the same dataset, using the global two-dimensional coordinate system of the initially stitched image as a reference, unifying the coordinate labeling rules of all candidate points to ensure that the coordinate references of the two types of candidate points are without deviation; performing spatial deduplication on the aggregated candidate point set, comparing the horizontal and vertical coordinates of candidate points point by point, and if two or more candidate points have completely overlapping coordinates, then points with higher feature recognition are retained first, such as line skeleton intersections having higher priority than text line skeleton turning points, and other duplicate points are removed to avoid feature redundancy and subsequent repeated calculations; classifying and labeling all deduplicated candidate points according to feature type, and orderly integrating them into a global candidate point set. This point set simultaneously covers the topological structure and linear structure features of text lines, is evenly distributed and without redundancy, and provides a complete global feature foundation for subsequent association of local rigid texture primitives and extraction of global constraint features.

[0041] Step 3.6: Calculate the Euclidean distance between each point in the global candidate point set and each local rigid texture primitive, and filter out points whose shortest distance is less than a preset radius threshold to form an associated candidate point set. Specifically, this includes: for each point in the constructed global candidate point set, using the global coordinates of the point on the initial stitched image as a reference, calculating the Euclidean distance between the point and all local rigid texture primitives in spatial position, recording each set of distance values ​​completely and forming a distance sequence; finding the smallest value in the distance sequence to determine the shortest distance between the current global candidate point and the nearest local rigid texture primitive, and comparing this shortest distance with a preset radius threshold. A certain radius threshold is used for comparison. The radius threshold is an empirical value obtained by comprehensively statistically analyzing the effective range of local features and the degree of image blurring. It is used to screen out global points that are closely related to the local stable feature space. Only global candidate points with a shortest distance less than the radius threshold are retained. These points can form a stable spatial constraint relationship with local rigid texture primitives. Isolated points that are not effectively related to local features and are prone to offset are excluded. All points that meet the conditions are integrated in an orderly manner according to their original coordinate information to form a set of associated candidate points that have both local stability and global structure. This provides a reliable set of intermediate points for subsequent structural tensor analysis and topological feature screening.

[0042] Step 3.7: For each point in the candidate set of associated points, a neighborhood window is defined to calculate the structure tensor. The anisotropy of the region is determined based on the eigenvalues ​​of the structure tensor. Points with anisotropy greater than a third preset threshold are retained as global topological constraint manifold feature points. Specifically, this includes: For each point in the constructed candidate set of associated points, the precise pixel coordinates of the point in the global coordinate system of the initial stitched image are first confirmed. A square neighborhood window of a fixed size is defined with these coordinates as the center. The window size is pre-calibrated based on extensive document feature analysis, usually 7×7 or 9×9 pixels. This ensures that the gradient changes and structural information around the point are fully covered, while avoiding interference caused by irrelevant background pixels due to an excessively large window, ensuring that the analysis scope focuses on the core structural region of the feature point. Within the defined neighborhood window, the gradient magnitude and gradient direction of each pixel are calculated pixel by pixel. Based on this gradient information, the structure tensor matrix corresponding to the point is constructed. The matrix dimension is adapted to the window size, which can accurately quantify the gradient distribution pattern, texture direction, and structural concentration of pixels within the window, fully reflecting the linear structural features and distribution pattern of the local region where the point is located.

[0043] After constructing the structure tensor matrix, solve for the two eigenvalues ​​of the matrix, denoted as . and ,and First, calculate the difference between the two eigenvalues. This reflects the absolute difference between the eigenvalues, and then the ratio of the two eigenvalues ​​is calculated. This reflects the relative difference in eigenvalues, and the degree of anisotropy in the region is calculated using a weighted fusion formula: Degree of Anisotropy ,in and These are weighting coefficients, and The weighting coefficients are set based on the statistical results of the structural features of the blurred document image, and are used to balance the absolute contribution of the feature value difference and the relative contribution of the ratio. The larger the feature value difference and the higher the ratio, the more concentrated the pixel gradient direction in the region, the more prominent the linear structure (such as text lines and line segments) features, and the stronger the constraint on the splicing position.

[0044] The calculated anisotropy level is compared one by one with a third preset threshold. The third preset threshold is set in combination with the statistical results of the structural features of the blurred document image. It can effectively distinguish between structural feature points with strong constraints and isotropic points without obvious linear structure. Only the associated candidate points with anisotropy level greater than the third preset threshold are retained. Invalid points with loose structure, messy gradient direction, and easy to be affected by motion blur are eliminated at the same time. The points finally selected are the global topological constraint manifold feature points. These feature points have both local gradient stability and global structural constraints. They can be used as the core constraint nodes of the global optimization algorithm to perform block-by-block pose correction and overall deformation correction on the initial stitched image. This effectively eliminates the stitching misalignment problem caused by local matching deviation and motion blur, and effectively improves the consistency, accuracy and visual continuity of long document cross-image block stitching.

[0045] In a preferred embodiment of the present invention, step 4 above may include: Step 4.1 involves performing a global fuzzy search on the preliminary stitched image to extract the global text line skeleton network and the global line topology, which together constitute the global structural features. Specifically, this includes performing a global structural feature retrieval and extraction operation on the preliminary stitched image formed by local rigid texture primitive matching and preliminary stitching. Based on the image grayscale distribution and gradient field information, connected component tracking and skeleton extraction are performed on the text regions in the entire image. The scattered text line skeletons are connected and integrated according to horizontal direction, spacing distribution, and paragraph arrangement to form a global text line skeleton network that can represent the layout of the entire text. At the same time, global detection is performed on structural linear features such as straight lines, table lines, and separator lines in the image, preserving the intersection, branching, collinearity, and extension relationships between lines to construct a complete global line topology. The global text line skeleton network and the global line topology are then integrated and unified to form a global structural feature that can completely represent the overall layout, direction, and relative positional relationships of the document, providing a unified and stable global reference benchmark for subsequent global registration, offset calculation, and stitching error detection.

[0046] Step 4.2 involves registering the global structural features with the local features of each image patch block by block. The offset of the local rigid texture primitives within each image patch relative to the corresponding global structural position is calculated to form an offset vector field. Specifically, this includes: registering the constructed global structural features (global text line skeleton network and global line topology) with each independent image patch involved in the stitching process block by block. A layered registration method combining coarse and fine registration is used to ensure accurate alignment between local features and the global structure. First, coarse registration is performed, using key anchor points in the global structural features (intersections of cross-block text lines and endpoints of long lines) as a reference. Affine transformations are then used to map each image patch to the global coordinate system, initially correcting the overall translation, rotation, and scaling deviations of the image patches. The affine transformation relationship can be expressed as: ; Where (x, y) are the original coordinates of the pixels within the image patch. The coordinates are the mapped global coordinates. a, b, c, and d are used to control rotation and scaling. a is the scaling and rotation blending coefficient in the horizontal direction, b is the shearing and rotation blending coefficient in the horizontal direction, c is the shearing and rotation blending coefficient in the vertical direction, and d is the scaling and rotation blending coefficient in the vertical direction. , Used to control translation. The horizontal translation component is the x-axis component. The vertical axis represents the translation component. Fine registration is performed, based on the coarse registration result, locating local feature regions within each image patch that correspond to the global structural orientation, position, and layout. Registration parameters are iteratively optimized to achieve the optimal overlap between local feature regions and the global structure. Point-to-point precise matching is then performed on all local rigid texture primitives obtained within these local feature regions, finding the theoretical topological position of each local rigid texture primitive on the global structural features. This is compared with the actual coordinates of the local rigid texture primitives in the image patch. The theoretical coordinates on the global structure are as follows The positional deviation between the two elements in a two-dimensional plane is calculated and an offset vector is formed; the offset vector is represented in two-dimensional Cartesian coordinates, denoted as . ;in This represents the offset value of the local rigid texture primitive in the horizontal axis direction; This represents the offset value of the local rigid texture primitive in the vertical axis direction.

[0047] The offset vectors within all image blocks are arranged in an ordered manner according to their spatial positions, and the sparse offset points are densified by bilinear interpolation to form a dense offset vector field that covers the entire preliminary stitched image, is evenly distributed and continuous. This offset vector field can intuitively, completely and meticulously reflect the magnitude and direction of the deviation of each local feature relative to the global structure, providing a stable and reliable quantitative basis for subsequent stitching error judgment, error area marking and deformation correction.

[0048] Step 4.3: Based on the offset magnitude of each point in the offset vector field, set an error judgment threshold. Mark continuous regions where the offset magnitude exceeds the error judgment threshold as local regions with splicing errors. Specifically, this includes: traversing each offset vector in the obtained offset vector field and calculating the offset magnitude corresponding to that vector. The offset magnitude is used to quantify the spatial deviation between the local rigid texture primitive and the global structure. Its calculation formula is as follows: In the formula, L is the offset modulus of the current local rigid texture primitive; Δx is the offset value in the horizontal direction; Δy is the offset value in the vertical direction; combining the maximum reasonable error range allowed for document image stitching, the degree of image blurring and noise level, a corresponding error judgment threshold T is set, and the region connectivity traversal is performed on the preliminary stitched image. Connected regions whose offset modulus L continuously exceeds the error judgment threshold T are marked. These regions have obvious stitching misalignment, stretching, compression or distortion, and are uniformly marked as local regions with stitching errors, providing a precise target range for subsequent refined feature extraction and deformation correction.

[0049] Step 4.4: Within each marked local region, extract the coordinates of local rigid texture primitives and the coordinates of global topological constraint manifold feature points. Specifically, for each marked local region with splicing errors, first, based on the minimum and maximum x and y coordinates of the region boundary, confirm the spatial range and global coordinate system of the region. Perform dual feature point extraction and verification operations within the boundary-defined region. On the one hand, traverse all pixels within the local region to extract the precise global coordinates of all obtained local rigid texture primitives and verify whether the coordinates of each primitive fall within the region boundary, eliminating invalid primitives that exceed the region range. On the other hand, simultaneously extract the precise global coordinates of all obtained global topological constraint manifold feature points within the local region, similarly verifying their region affiliation and eliminating invalid points outside the region. The coordinates of the two types of feature points are then formatted in a unified way, converting all coordinates to floating-point data of the same precision and performing coordinate calibration to ensure that the two types of point sets are standardized, accurate, and complete under the same global coordinate system. Simultaneously, record the type label of each valid point: local rigid texture primitive or global topological constraint manifold feature point, providing high-quality and highly reliable point set data for subsequent geometric topological relationship modeling.

[0050] Step 4.5 involves modeling the geometric topological relationship between the extracted local rigid texture primitives and the global topological constraint manifold feature points, constructing a manifold constraint graph structure based on the connectivity of point sets, so that the vertices of each graph unit are composed of the two types of points. Specifically, this includes: performing geometric topological relationship modeling on the extracted local rigid texture primitives and the global topological constraint manifold feature points, using a manifold constraint graph construction method based on an improved graph neural network architecture. This method integrates local geometric constraints and global topological consistency constraints into the traditional graph structure, enhancing the model's understanding of document structure and adapting to the local deformation correction requirements of long document splicing; dividing the spatial neighborhood of the two types of feature points, setting a neighborhood window with a fixed radius centered on each point. The window radius is set according to the document feature scale and splicing accuracy requirements. All points within the window are considered potential connection objects for that point. The spatial distance and directional consistency of each potential connection point are calculated. The closer the spatial distance and the more consistent the direction with the overall orientation of the global text line skeleton network and the global line topology, the stronger the constraint between the two points, providing a stable quantitative basis for the subsequent edge construction.

[0051] An initial graph structure is constructed by using all validated local rigid texture primitives and global topological constraint manifold feature points as vertices. Pairs of points with constraint strength greater than a preset threshold are selected to form initial edges. The edge weights are determined by a weighted fusion of spatial distance and directional consistency; the closer the distance and the more consistent the direction, the higher the edge weight and the stronger the constraint. This forms an initial manifold constraint graph containing vertices and edges. The initial graph structure is then trained and optimized using a modified network structure derived from the standard graph attention network architecture. This modified network structure optimizes the calculation logic of attention weights to meet the topological constraint requirements of document stitching, strengthens global structural consistency constraints, and better achieves stable binding between local features and the global structure.

[0052] During training, the weights of edges are dynamically adjusted through an attention mechanism to strengthen effective connections consistent with the global topological structure and weaken redundant edges caused by noise or mismatches. Simultaneously, a local rigidity constraint loss function is introduced to ensure the relative stability of local rigid texture primitives during subsequent deformation correction, preventing distortion and misalignment of local features. The specific expression of the local rigidity constraint loss function is as follows: ; In the formula, represents the local rigid constraint loss value. The smaller the loss value, the more stable the relative positions of the local rigid texture primitives in the graph structure. N is the total number of valid point pairs participating in the local rigid constraint calculation. This represents the actual spatial distance vector between the i-th pair of points during the training process. Let be the theoretical spatial distance vector determined for the i-th pair of points based on global structural features. The L2 norm is used to calculate the Euclidean distance between the actual distance vector and the theoretical distance vector, thus quantifying the degree of deviation between the two. The training process focuses on minimizing the local rigid constraint loss value as the core optimization objective, combined with global topological consistency loss for auxiliary optimization. The weight parameters of the graph attention network are iteratively updated through the backpropagation algorithm, gradually improving the graph structure's ability to represent splicing errors and its constraint stability. After each iteration, the edge weights of the graph structure are re-evaluated, redundant edges with weights below a preset threshold are removed, core connections with strong constraints are retained, and new effective connections generated by parameter updates are added to maintain the integrity and reliability of the graph structure. After multiple iterations of training, the training is completed when the local rigid constraint loss value tends to stabilize and no longer shows a significant decrease, resulting in a manifold constraint graph structure with stable constraint capabilities.

[0053] The stable manifold constraint graph structure is expressed as follows: In the formula, V is the set of vertices of the graph structure, denoted by the extracted and verified local rigid texture primitives. The feature points of the global topologically constrained manifold are denoted as Together constitute, that is Each vertex contains its own global coordinates, feature type label, and other attributes; E is the set of edges in the graph structure, representing the effective topological constraints between vertices after training and optimization. It only includes connections with constraint strength higher than a preset threshold, and the edge types are divided into two categories: (Connections between local rigid texture primitives) (The connection between local rigid texture primitives and global topologically constrained manifold feature points), i.e. W is the set of edge weights that corresponds one-to-one with the edge set E, where , As vertex and The weights of edges between vertices are determined by both the attention mechanism and the constraint strength; the larger the weight, the stronger the constraint. The weight directly reflects the constraint strength between vertices. It is a set of topological constraint rules for graph structures, including local rigid constraints and global topological consistency constraints, used to regulate the deformation rules of graph structures and ensure that the subsequent deformation correction process can conform to the actual layout characteristics of the document.

[0054] Through this enhanced manifold constraint graph structure, local stability features and global structural features are tightly bound together, forming a constraint system that combines local robustness and global consistency. This system can accurately characterize the topological relationships and constraint strength between feature points, providing a reliable and accurate topological foundation for subsequent adaptive deformation mesh construction and local splicing error correction.

[0055] Step 4.6: Based on the direction and magnitude of the offset vector field, adaptively refine or sparse the edges of the constructed manifold constraint graph structure to form multiple adaptive deformation meshes that match the degree of local deformation. Specifically, this includes: based on the direction and offset magnitude of each vector in the obtained offset vector field, performing adaptive refinement or sparse adjustment on the edges between vertices in the manifold constraint graph structure to achieve a precise match between constraint density and the degree of local deformation; defining adaptive adjustment coefficients. Its calculation formula is Where L is the offset magnitude corresponding to the current vertex, and T is the error judgment threshold. Reflects the degree of deformation in a local area; the offset modulus is large, and the adaptive adjustment coefficient is large. Higher elevations indicate severe local deformation. Edge densification is applied by increasing the number of connections between vertices and their neighborhoods, thus improving constraint density and correction accuracy. Smaller offset modulus and adaptive adjustment coefficients are also used. Lower deformation indicates gentle local deformation. The edges are sparsed, and redundant edges with low weights are removed to reduce computational complexity while maintaining the constraint effect. The edge density is dynamically allocated by adaptively adjusting the coefficients, and finally multiple adaptive deformation meshes that are precisely matched with the degree of local deformation are formed. The resolution of the meshes changes dynamically with the degree of deformation, providing a hierarchical, graded and high-precision constraint basis for subsequent local deformation optimization, misalignment correction and global smoothing of image blocks.

[0056] In a preferred embodiment of the present invention, step 5 above may include: Step 5.1: Extract the vertex coordinate set for each adaptive deformable mesh. The vertices are composed of local rigid texture primitives and global topological constraint manifold feature points. Specifically, this includes: extracting vertex information for each constructed adaptive deformable mesh; reading the global coordinate data of all vertices in the mesh one by one according to the spatial distribution and structural composition of the adaptive deformable mesh. These vertices are all composed of the local rigid texture primitives and global topological constraint manifold feature points. During the extraction process, the vertex coordinates are validated, and invalid vertices that are outside the image range, have abnormal coordinates, or are affected by noise are removed. The validated vertex coordinates are classified and organized according to mesh affiliation to form a standardized vertex coordinate set corresponding to each adaptive deformable mesh, denoted as […]. ,in The coordinates are the global coordinates of the nth vertex, where n is the total number of valid vertices within a single grid. This set provides accurate basic data for subsequent similarity calculations and deformation analysis between grids.

[0057] Step 5.2: For two adjacent adaptive deformable meshes, calculate the geometric similarity of their shared boundary, including the length ratio and angle difference of the shared edge, and use this as a measure of topological consistency between adjacent meshes. Specifically, for any two adaptive deformable meshes that are spatially adjacent and have a common boundary region, perform boundary matching based on the global coordinates of the mesh vertices and the boundary distribution information to determine the actual shared boundary line segments and common vertex set between the two meshes. Within the shared boundary range, extract the vertex sequence, line segment direction and boundary length information of the corresponding boundary of the two meshes to complete the accurate alignment of boundary features.

[0058] The geometric similarity of shared boundaries is quantified in multiple dimensions. The actual lengths of two adjacent meshes at the shared edge are counted and the length ratio is calculated. ,in , These represent the corresponding edge lengths of two adjacent grids at the shared boundary. Simultaneously, the direction angles between corresponding line segments of adjacent grids at the shared boundary are calculated, and the angle difference is obtained. ,in , These are the orientation angles of the shared edges between two adjacent grids; the length ratio and the angle difference are weighted and fused according to preset weights, and the fusion formula is: Topology Consistency Measurement ,in and For the preset weighting coefficients, and , The contribution weights used to balance the length ratio Contribution weights used to balance the differences in included angles The angle difference is normalized to the range of 0 to 1 to ensure that the length ratio and angle difference are integrated under the same dimension. The topological consistency measurement result can intuitively reflect the smoothness, matching accuracy and topological rationality of adjacent grids in structural connection, effectively distinguish between normal connection areas and abnormal connection areas with misalignment and distortion, and provide a reliable quantitative basis for the stable construction of deformation transfer relationship diagram.

[0059] Step 5.3: Based on the topology consistency metric, construct a deformation transfer graph between meshes, and calculate the local deformation gradient of each mesh by propagating along the deformation transfer graph. The local deformation gradient represents the rate of change of the relative displacement of vertices within the mesh. Specifically, this includes: constructing a deformation transfer graph to describe the constraint transfer relationship between meshes based on the topology consistency metric results between adjacent adaptive deformation meshes. This deformation transfer graph is a weighted undirected graph used to accurately represent the adjacency relationship and deformation constraint transfer strength between each adaptive deformation mesh. Each node in the graph corresponds to an adaptive deformation mesh, and the connection between nodes corresponds to the constraint transfer relationship between adjacent meshes. The weight of the connection corresponds to the topology consistency metric value of the adjacent meshes. The higher the weight, the tighter the topology connection between adjacent meshes, and the stronger the stability and continuity of deformation transfer.

[0060] The process of constructing the deformation transfer graph is as follows: First, the spatial distribution and adjacency relationships of all adaptive deformation meshes are analyzed, treating each adaptive deformation mesh as an independent node to ensure that each mesh corresponds to a unique node in the graph, clearly distinguishing the identities of different meshes. Second, all adaptive deformation meshes are traversed, and it is determined whether any two meshes share a boundary. For adjacent meshes with shared boundaries, a connection is drawn between the corresponding two nodes; this connection is the constraint transfer edge between meshes. Connections are only established between mesh nodes with actual adjacency relationships to avoid invalid connections affecting the accuracy of transfer. Third, the topological consistency metric of each pair of adjacent meshes is assigned to the corresponding constraint transfer edge as the edge weight, thereby quantifying the constraint strength of deformation transfer between meshes. The higher the topological consistency metric, the greater the edge weight, and the stronger the constraint of deformation transfer between meshes. Finally, the initial graph structure is simply verified, and weak connection edges with too low weights are removed. These weak connection edges correspond to poor topological connections between adjacent meshes, resulting in low reliability of deformation transfer. Removing these edges ensures that the deformation transfer graph accurately reflects the effective constraint transfer relationships between meshes.

[0061] After constructing the deformation transfer graph, the calculation is performed layer by layer along the connections between nodes and edges in the graph. During the propagation process, the distribution patterns of global structural features and the spatial distribution information of the obtained offset vector field are combined to solve for the local deformation gradient within each adaptive deformation mesh. The correct expression for the vertex displacement field function U is as follows: In the formula, U represents the overall displacement state of the vertices of the adaptive deformation mesh; P is the actual displacement of the vertices of the adaptive deformation mesh in the horizontal direction; Q is the actual displacement of the vertices of the adaptive deformation mesh in the vertical direction; and the local deformation gradient is also represented by U. middle, This represents the rate of change of the vertex displacement field function in the horizontal direction. The ratio represents the rate of change of the vertex displacement field function in the vertical direction. The combination of the two can accurately quantify the deformation change law inside the mesh. This local deformation gradient can accurately characterize the rate and trend of change of relative displacement of each vertex inside the mesh under the influence of splicing error, and intuitively reflect the intensity and distribution law of deformation within the mesh area.

[0062] Step 5.4: Based on the obtained initial pose relationship, the theoretical coordinate position of each vertex is inverted and compared point-by-point with the measured vertex coordinates in the current adaptive deformation mesh to obtain the offset vector of each vertex. Specifically, this includes: using the global structural features of the initially stitched image as a unified benchmark, and combining the initial pose relationship between image blocks with global topological constraints, the theoretical position inversion calculation is performed on each vertex in the adaptive deformation mesh. The specific process is as follows: using the global text line skeleton network and the global line topology as the core reference, the global structure to which the current vertex belongs, such as the text line, table line, or separator line, is confirmed, and the position is calculated based on the global structure. By considering the direction, spacing, and layout patterns, the approximate position of the vertex is initially determined under conditions without stitching errors. Then, combined with the initial pose relationship between image blocks and referring to the positional connection logic of adjacent image blocks, the initially determined position is coarsely adjusted to ensure that the vertex position remains consistent with the corresponding structure of adjacent image blocks. Combined with global topological constraints, and based on the reasonable relative positional relationship between the vertex and other effective vertices (local rigid texture primitives and global topological constraint manifold feature points), the vertex position is finely adjusted to correct any possible deviations, so that the inverted vertex position not only conforms to the global structural layout but also maintains normal topological connection with surrounding vertices.

[0063] After multiple rounds of fine-tuning and calibration, the precise theoretical coordinate position of each vertex was determined under a state without splicing errors, denoted as . The theoretical coordinates obtained from the inversion are compared with the measured coordinates actually acquired from the vertices in the current adaptive deformation mesh. Perform precise point-by-point comparison, calculate the deviation between coordinates according to a unified global coordinate system, and form an offset vector corresponding to each vertex. This offset vector fully records the positional deviation of the vertex in the horizontal and vertical directions, providing direct and accurate data support for the subsequent fitting and solution of the pose compensation coefficient.

[0064] Step 5.5: Taking each adaptive deformation mesh as a unit, jointly optimize and solve the offset vectors and local deformation gradients of all vertices within the mesh to fit the pose compensation coefficients corresponding to this mesh. The pose compensation coefficients include rotation, scaling, and translation components. Specifically, this involves: using a single adaptive deformation mesh as an independent processing unit, jointly optimizing and solving the offset vectors and corresponding local deformation gradients of all valid vertices within the mesh. The specific implementation process is as follows: first, summarize the measured coordinates, theoretical coordinates, offset vectors, and local deformation gradients of all valid vertices within the mesh, confirming that the core objective of the joint constraint optimization is to minimize the deviation between the theoretical and measured coordinates of the vertices, while also considering the smoothness of deformation within the mesh and local rigidity constraints to avoid local texture distortion or structural deformation during the correction process; the solution is obtained by stepwise optimization. The iterative approach involves first setting initial iteration parameters and a precision threshold. In each iteration, the offset deviation of each vertex is calculated one by one. Combined with the corresponding local deformation gradient, the direction and magnitude of parameter adjustment are determined. Priority is given to adjusting the vertex parameters in areas with larger offset deviations and more severe deformation gradients, gradually narrowing the gap between theoretical and measured coordinates. During the iteration process, two constraints are monitored in real time: first, the smoothness of deformation within the mesh, ensuring continuous displacement changes between adjacent vertices without abrupt distortions; second, local rigidity constraints, ensuring the relative positions between local rigid texture primitives are stable and without stretching or compression deformation. If, after a certain iteration, the coordinate deviation of all vertices is less than the preset precision threshold and both constraints are met, the iteration stops, the joint constraint optimization solution is completed, and the optimal displacement, angle, and size adjustment parameters are obtained.

[0065] After completing the joint constraint optimization solution, the fitting of the complete pose compensation coefficients corresponding to the current adaptive deformation mesh begins. The specific process is as follows: Based on the optimal parameters obtained through iterative convergence, adjustment information related to angle correction is extracted. Combining the angle deviation distribution of vertices within the mesh, rotation components are determined through methods such as angle mean calibration and deviation smoothing correction to accurately correct the angle deviation of the mesh. Next, adjustment information related to size correction is extracted. By comparing the size ratio of the theoretical coordinates and the measured coordinates of vertices within the mesh, and combining the distribution law of local deformation gradients, a unified size adjustment ratio is calculated to determine the scaling component, which is used to correct the size deviation of the mesh and ensure that the corrected mesh size matches the global structure. Finally, adjustment information related to position translation is extracted. The average displacement deviation of all vertices within the mesh is statistically analyzed. Combined with the position reference of the global structure, the translation amounts in the horizontal and vertical coordinate directions are determined to form translation components, which are used to correct the position deviation of the mesh.

[0066] After fitting, the rotation, scaling, and translation components are normalized to unify the parameter scale. Then, stability is checked to remove fluctuations caused by local abnormal deviations. The three types of components are integrated to obtain the complete pose compensation coefficients corresponding to the current adaptive deformation mesh. The pose compensation coefficients include rotation components for correcting angular deviations, scaling components for correcting dimensional deviations, and translation components for correcting positional deviations. The three types of components work together to form a complete pose correction system, which can achieve comprehensive compensation for multiple types of errors such as mesh translation, rotation, and scaling, ensuring that the corrected mesh can be accurately aligned with global structural features.

[0067] Step 5.6: Aggregate the pose compensation coefficients of each grid by image patch, establishing a mapping relationship between each blurred image patch and the compensation coefficients of each grid within its coverage area. Specifically, this includes: classifying and aggregating the pose compensation coefficients corresponding to all adaptive deformation grids according to the spatial range of their respective blurred image patches, establishing a stable mapping relationship between each blurred image patch and the pose compensation coefficients of all grids within it. This mapping relationship is a one-to-one or one-to-many association, specifically: each blurred image patch serves as the input end of the mapping, and the pose compensation coefficients (including rotation, scaling, and translation components) corresponding to all adaptive deformation grids it covers serve as the output end of the mapping, and each image patch is associated with its respective grid. The compensation coefficients of the grids are precisely bound together through spatial coordinate ranges, which can clearly define the pose compensation type and intensity required for different regions (corresponding to different adaptive deformation grids) within a certain image block. When establishing this mapping relationship, based on the boundary coordinates of the image block and the coverage of the adaptive deformation grid, all grids contained in each image block are identified. Then, the pose compensation coefficients of these grids are arranged in order of their position within the image block to form structured associated data. The final established mapping relationship can determine the compensation intensity, compensation direction, and compensation type required for different image regions, providing a complete and standardized parameter basis for performing globally unified pose optimization, misalignment correction, and deformation smoothing on the entire initially stitched image.

[0068] In a preferred embodiment of the present invention, step 6 above may include: Step 6.1: Based on the established mapping relationship, determine the adaptive deformation meshes covered by each blurred image block and their corresponding pose compensation coefficients. Specifically, this includes: based on the established mapping relationship between blurred image blocks, adaptive deformation meshes, and pose compensation coefficients, traversing all blurred image blocks one by one, using the spatial boundary coordinates of each blurred image block as the retrieval basis, accurately matching all adaptive deformation meshes covered by the image block in the mapping relationship, and simultaneously extracting the pose compensation coefficients corresponding to each adaptive deformation mesh completely. The extracted compensation coefficients are then validated to exclude abnormal values ​​and missing parameters, ensuring that the rotation, scaling, and translation components corresponding to each mesh are real, effective, and meet the preset accuracy requirements, thus providing stable and reliable parameter support for sub-region division and region-by-region pose correction.

[0069] Step 6.2: For each blurred image block, its coverage area is divided into multiple sub-regions corresponding to the adaptive deformation mesh, and the pose compensation coefficients of the mesh to which each sub-region belongs are associated. Specifically, this includes: for each blurred image block whose parameters have been confirmed, first reading the complete spatial boundary coordinates, pixel size, and detailed distribution information of the internal adaptive deformation mesh of the image block, confirming the spatial position, boundary range, and distribution density of all adaptive deformation meshes within the image block, ensuring a precise understanding of the spatial layout of the mesh; when dividing the sub-regions, strictly according to the spatial boundary coordinates of each adaptive deformation mesh, using the vertex connection of the mesh as the dividing basis, the overall imaging area of ​​the image block is precisely cut into multiple sub-regions. During the division process, the principles of mutual adjacency, no overlap, and no omission are strictly followed, ensuring that the geometric contour and spatial range of each sub-region completely coincide with the boundary of the corresponding adaptive deformation mesh, without any overlap or omission. Even beyond the grid's boundaries, no pixel area covered by the grid is overlooked, while ensuring seamless connection of adjacent sub-region boundaries to avoid boundary breaks or overlapping redundancy. After sub-region division, a unique region identifier is assigned to each sub-region to distinguish its affiliation. Then, through precise matching of region identifiers and grid identifiers, a unique association is established between each sub-region and the pose compensation coefficients of the corresponding adaptive deformation grid. During the association process, the parameter integrity of the compensation coefficients is checked one by one to ensure that each sub-region can be accurately bound to its own rotation, scaling, and translation components. At the same time, the association information is recorded in a structured manner, accurately marking the grid identifier, compensation parameters, and spatial range corresponding to each sub-region. This provides a precise parameter matching basis for the independent pose correction of each sub-region, ensuring that subsequent pose correction operations can accurately act on the corresponding local area and avoid parameter confusion or deviation in the area of ​​application.

[0070] Step 6.3: For each sub-region, construct the local homography transformation matrix based on the rotation, scaling, and translation components in its associated pose compensation coefficients. Specifically, this includes: for each sub-region that has been successfully divided and associated with the pose compensation coefficients, stably read all pose compensation parameters bound to the current sub-region from the mapping relationship, and then completely separate the rotation component for angle deviation correction, the scaling component for size deviation correction, and the translation component for position deviation correction from the compensation parameters. Perform numerical rationality verification on the three types of components respectively, and eliminate abnormal fluctuations and erroneous data that exceed the effective range. When constructing the local homography transformation matrix, use the measured coordinate distribution of all vertices within the current sub-region as the underlying basis, and combine the global structural feature constraints provided by the global text line skeleton network and the global line topology to determine the overall direction and boundary conditions of the matrix construction. At the same time, refer to the transformation trends and boundary connection requirements of adjacent sub-regions to predict the connection status between the current sub-region and the surrounding regions after transformation in advance, so as to avoid splicing discontinuities caused by conflicts between the transformation rules of a single sub-region and the surrounding regions.

[0071] Following the priority and logical order of geometric transformations, three types of components are gradually integrated into the matrix construction: The core parameters of the angle transformation matrix are determined based on the rotation component. Using the geometric center of the sub-region as the origin of rotation, the angle values ​​corresponding to the rotation components are converted into rotation parameters of the matrix, ensuring that the sub-region can accurately correct angular deviations around the center. The scaling component is integrated into the matrix as a scaling factor. Using the origin of rotation as the scaling center, the scaling parameters of the matrix are adjusted according to the values ​​of the scaling components to adapt to the size correction requirements of the sub-region in both horizontal and vertical dimensions, ensuring that the texture ratio of the scaled sub-region is consistent with the global structure. Finally, the position offset parameters corresponding to the translation component are superimposed, converting the overall translation requirements of the sub-region into translation terms of the matrix, completing the construction of the basic homography transformation matrix. Based on this, further... By incorporating deformation smoothing constraints within sub-regions and local rigid texture preservation constraints, the base matrix is ​​iteratively optimized: the deformation smoothing constraint verifies the displacement change rate of adjacent vertices within the sub-region after the matrix is ​​applied. If the displacement abruptly exceeds a preset threshold, it indicates a risk of texture distortion, so the scaling or rotation parameters of the matrix are finely adjusted to reduce the magnitude of the displacement abruptness. Then, the local rigid texture preservation constraint verifies the relative position of local rigid texture primitives after the matrix is ​​applied. If the spacing or angle between primitives deviates from the theoretical value, the compensation magnitude of the translation component is corrected. Simultaneously, the boundary transformation parameters of the current sub-region and adjacent sub-regions are compared. If there is a connection deviation in the coordinates of the transformed boundary pixels, the edge compensation coefficient of the matrix is ​​adjusted synchronously to ensure seamless connection of the transformed pixels at the boundary.

[0072] After multiple rounds of parameter calibration and constraint verification, the matrix effect is recalculated and the constraints are checked after each calibration until all constraints are met and the transformation effect meets expectations. Finally, a local homography transformation matrix that is specific to the current sub-region and adapted to its local error characteristics is generated. This matrix can take into account the pose compensation requirements of rotation, scaling and translation, while ensuring the integrity of the texture inside the sub-region and the continuity of the boundary of adjacent regions. It has complete and accurate coordinate transformation capabilities and can be directly used for unified pose correction of all pixels in the sub-region.

[0073] Step 6.4: Apply the local homography transformation matrix to the coordinates of all pixels within the sub-region to correct the pose of this sub-region, gradually converging the local rigid texture primitives and global topologically constrained manifold feature points within this sub-region to their theoretical coordinate positions. Specifically, this includes: applying the constructed local homography transformation matrix to every pixel within the current sub-region; first, reading the original image coordinates of all pixels within the sub-region; then, according to the set pixel traversal order, performing precise point-by-point transformation and position update on the coordinates of each pixel; during the transformation process, strictly adhering to the rotation, scaling, and translation rules integrated in the local homography transformation matrix to ensure that the coordinate adjustment of each pixel accurately corresponds to its pose deviation; and during the coordinate transformation... While performing the transformation, the positional changes of local rigid texture primitives and global topological constraint manifold feature points within the sub-region are monitored in real time to ensure that these key feature points can be adjusted synchronously with the pixel coordinates and gradually converge to the theoretical coordinate positions obtained in step 5.4. During this process, the positional deviation of the feature points is continuously checked. If the deviation does not reach the preset accuracy, the transformation parameters are finely adjusted until the key feature points are precisely aligned with the theoretical coordinates. Through this refined point-by-point transformation, various splicing errors such as positional offset, angle distortion, and size deviation within the sub-region are effectively eliminated. At the same time, the integrity and continuity of the texture structure within the sub-region are ensured, avoiding problems such as pixel stretching, distortion, or discontinuity, thereby achieving high-precision pose correction for the entire sub-region.

[0074] Step 6.5: Traverse all sub-regions covered by the same blurred image block to complete the grid-by-grid non-uniform pose correction of the image block, obtaining the corrected pose relationship of the image block. Specifically, this includes: processing all sub-regions covered by the same blurred image block sequentially according to a preset region traversal order, i.e., from the upper left corner to the lower right corner of the image block, or according to the grid distribution order; repeatedly performing local homography transformation and pose correction operations on each sub-region; using a grid-by-grid non-uniform correction method throughout; and implementing differentiated correction based on the error magnitude and error type of different sub-regions: for sub-regions with large deformation gradients and obvious stitching errors, appropriately increasing the correction accuracy and refining the transformation parameters; for sub-regions with smaller errors and relatively stable structures... To ensure the correction effect, the operation is simplified and the processing efficiency is improved. During the correction process, the focus is on the boundary connection between adjacent sub-regions. The coordinates and texture features of the boundary pixels of adjacent sub-regions are compared in real time, and the transformation parameters at the boundary are adjusted to ensure that the boundaries of adjacent sub-regions are seamlessly connected and the texture transition is natural. This avoids problems such as new boundary breaks, misalignments, or texture discontinuities caused by independent correction. After all sub-regions in the blurred image block have completed pose correction, the correction effect of the entire image block is checked as a whole. The alignment degree between the overall pose of the image block and the global structural features is verified, and local residual errors are corrected. Finally, the overall error of the entire blurred image block is eliminated, and a stable and accurate corrected pose relationship of the image block is obtained.

[0075] Step 6.6: Reproject all blurred image blocks onto the same stitching plane according to the corrected pose relationships, and perform pixel-level fusion on the overlapping areas of adjacent image blocks to generate the final stitched image. Specifically, this includes: projecting all blurred image blocks that have undergone local correction and overall pose optimization onto the same common stitching plane according to their corrected pose relationships. During the projection process, the global coordinate system is strictly followed to ensure that the projection position of each image block accurately corresponds to its preset position in the overall stitched image, so that the image blocks form an orderly, continuous, and misaligned overall layout in space. At the same time, the global topological consistency of all image blocks is checked to ensure that the global structures such as text lines and lines between image blocks can be connected coherently without breaks or misalignments. After projection, the overlapping areas between adjacent image blocks are accurately identified, and the range, number of pixels, and positional distribution of the overlapping areas are determined. Pixel-level fusion processing is performed on the overlapping areas: First, the brightness and color of corresponding pixels in the overlapping areas are compared point by point, the pixel difference is calculated, and a smooth transition method is used to adjust the gradient of pixels with large differences to eliminate brightness differences and color banding; the texture features of the overlapping areas are matched to correct texture misalignment problems so that the textures of adjacent image blocks can be connected naturally; the stitching effect is previewed in real time during the fusion process, and if there are obvious stitching marks or texture discontinuities, the fusion parameters are fine-tuned in time until the overlapping areas transition naturally and there are no obvious stitching marks; after the overlapping areas of all adjacent image blocks have been fused, the entire stitched image is globally optimized to correct the overall brightness and contrast of the image, remove redundant pixels at the edges, and finally generate a final stitched image with complete structure, clear texture, no obvious stitching marks, and consistent global topology, which meets the accuracy and visual effect requirements of long and blurred document stitching.

[0076] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A self-adaptive splicing method for fuzzy retrieval of scanned documents, characterized in that, The method includes: Multiple blurred image blocks are obtained by non-fixed block scanning of a long document by a mobile terminal, wherein the long document contains a layout structure composed of text lines and lines; Local fuzzy retrieval is performed on each blurred image patch to extract local features of each patch, and the initial pose relationship and overlapping region between each image patch are determined based on local feature matching; local rigid texture primitives are determined within the overlapping region of each image patch. Based on the initial pose relationship and overlapping area, multiple blurred image blocks are initially stitched together to obtain a preliminary stitched image. In the preliminary stitched image, global topological constraint manifold feature points are determined. A global fuzzy search is performed on the preliminary stitched image to extract global structural features, and the global structural features are compared with the local features of each image block to identify local areas with stitching errors. Within the local area, based on the spatial distribution relationship between local rigid texture primitives and global topological constraint manifold feature points, multiple adaptive deformation meshes are constructed with local rigid texture primitives and global topological constraint manifold feature points as vertices. Based on the topological consistency between adjacent adaptive deformation meshes, the local deformation gradient of each mesh is calculated. By comparing the vertex offsets of the theoretical mesh and the measured mesh, the pose compensation coefficients corresponding to each mesh are generated. Based on the pose compensation coefficient, the initial pose relationship of the corresponding image blocks is corrected grid by grid to obtain the corrected pose relationship, which is used to re-stitch multiple blurred image blocks to obtain the final stitched image.

2. The self-adaptive splicing method for fuzzy retrieval of scanned documents according to claim 1, characterized in that, Acquire multiple blurred image blocks obtained by non-fixed block scanning of a long document by a mobile terminal, wherein the long document contains a layout structure composed of text lines and lines, including: In response to the scan start command, the mobile terminal's camera is invoked to capture video stream image frames of the document in continuous preview mode; In each video stream image frame, document edges are detected in real time. When a document boundary is detected to fall completely into the preset area of ​​the current frame, image capture of this frame is automatically triggered to obtain the original image frame containing local document content. Based on the document area covered by the captured original image frame, a coverage heatmap of the currently scanned area is generated, and the location prompt information of the unscanned area is calculated based on the coverage heatmap to guide the user's mobile terminal to cover the unscanned area. During the process of using the user's mobile terminal, new original image frames are continuously captured, and the overlap between each newly captured original image frame and the already captured original image frames is calculated. Only original image frames that have a preset overlap range with the already captured image frames are retained. Real-time blur detection is performed on each retained original image frame to extract the gradient magnitude distribution features of the image frame. When the gradient magnitude distribution features are lower than a preset threshold, the image frame is determined to be a blurred image block and is removed; otherwise, the image frame is marked as a valid blurred image block.

3. The self-adaptive splicing method for fuzzy retrieval of scanned documents according to claim 2, characterized in that, Within the overlapping regions of each image patch, the local rigid texture primitives are determined, including: Multi-scale corner detection is performed on the overlapping area of ​​each blurred image block to obtain corner response maps at each scale, and non-maximum suppression is used to retain pixels with local maxima of response values ​​as corner candidate points. Text line direction projection analysis is performed on the same overlapping area. The start and end lines of the text lines are located by the trough of the projection curve. Pixels at the start and end columns of each text line are located as candidate endpoints of the text lines. Line skeletons are extracted from the same overlapping area, and pixels are located at the intersections and endpoints of the line skeletons as candidate points for line structures. The candidate points for corners, candidate points for text line endpoints, and candidate points for line structures are merged to form an initial set of candidate points; For each point in the initial candidate point set, a local neighborhood window is defined with its coordinates as the center. The gradient magnitude of all pixels in the window is calculated, and points with a gradient magnitude greater than the first preset threshold are selected to form a strong gradient candidate point set. Between adjacent image blocks, feature descriptors are constructed and matched for points in the strong gradient candidate point set. The Euclidean distance between matched point pairs is calculated, point pairs with a distance less than a second preset threshold are retained, and the coordinates of the matched point pairs in the two image blocks are recorded respectively. Points in the retained matching point pairs that lie on the current image patch are marked as local rigid texture primitives.

4. The self-adaptive splicing method for fuzzy retrieval of scanned documents according to claim 3, characterized in that, In the initial stitched image, global topological constraint manifold feature points are identified, including: Text line connectivity analysis is performed on the preliminary stitched image to extract the bounding rectangles of each connectivity region. Based on the aspect ratio of the bounding rectangles and the arrangement of adjacent rectangles, continuous text lines across image blocks are identified. For each consecutive text line across an image block, refine the text line, extract the central skeleton line, and locate the pixel points at the turning points, endpoints, and intersections of two or more central skeleton lines as the first type of global candidate points. Line segment detection is performed on the preliminary stitched image to extract continuous lines with a length greater than a preset length threshold, and lines that cross the boundary of image blocks are selected as cross-block continuous lines. For each continuous line spanning a block, skeletonization is performed. At the branch points, inflection points, and intersection points of the line skeleton with the image block boundary, pixels are located as second-class global candidate points. The first type of global candidate points and the second type of global candidate points are merged to form a global candidate point set; Calculate the Euclidean distance between each point in the global candidate point set and each local rigid texture primitive, and select the points whose shortest distance is less than a preset radius threshold to form an associated candidate point set; For each point in the set of associated candidate points, a neighborhood window is defined to calculate the structure tensor. The degree of anisotropy of the region is determined based on the eigenvalues ​​of the structure tensor. Points with anisotropy greater than a third preset threshold are retained as global topological constraint manifold feature points.

5. The self-adaptive splicing method for fuzzy retrieval of scanned documents according to claim 4, characterized in that, Identify local regions with splicing errors; within these local regions, based on the spatial distribution relationship between local rigid texture primitives and global topological constraint manifold feature points, construct multiple adaptive deformation meshes with local rigid texture primitives and global topological constraint manifold feature points as vertices, including: A global fuzzy search is performed on the preliminary stitched image to extract the global text line skeleton network and the global line topology, which together constitute the global structural features. The global structural features are registered with the local features of each image block one by one, and the offset of the local rigid texture primitives in each image block relative to the corresponding global structural position is calculated to form an offset vector field. Based on the offset magnitude of each point in the offset vector field, an error judgment threshold is set, and continuous regions whose offset magnitude exceeds the error judgment threshold are marked as local regions with splicing errors. Within each marked local region, the coordinates of the local rigid texture primitives and the coordinates of the global topologically constrained manifold feature points within that region are extracted respectively. Geometric topological relationship modeling is performed on the extracted local rigid texture primitives and global topological constraint manifold feature points. A manifold constraint graph structure based on the connectivity of point sets is constructed so that the vertices of each graph unit are composed of the two types of points. Based on the direction and magnitude of the offset vector field, the edges of the constructed manifold constraint graph structure are adaptively densified or sparsed to form multiple adaptive deformation meshes that match the degree of local deformation.

6. The self-adaptive splicing method for fuzzy retrieval of scanned documents according to claim 5, characterized in that, Based on the topological consistency between adjacent adaptive deformation meshes, the local deformation gradient of each mesh is calculated. By comparing the vertex offsets of the theoretical mesh and the measured mesh, pose compensation coefficients corresponding to each mesh are generated, including: Extract the set of vertex coordinates for each adaptive deformation mesh, where each vertex is composed of local rigid texture primitives and global topologically constrained manifold feature points; For two adjacent adaptive deformation meshes, calculate the geometric similarity of their shared boundary, including the length ratio and angle difference of the shared edge, and use this as a measure of topological consistency between adjacent meshes; Based on the topological consistency metric, a deformation transfer graph between meshes is constructed, and the local deformation gradient of each mesh is calculated by propagating along the deformation transfer graph. The local deformation gradient represents the rate of change of the relative displacement of vertices within the mesh. Based on the obtained initial pose relationship, the theoretical coordinate position of each vertex is inverted and compared point by point with the measured vertex coordinates in the current adaptive deformation mesh to obtain the offset vector of each vertex. For each adaptive deformation mesh, the offset vectors of all vertices in the mesh and the local deformation gradient are jointly optimized and solved to fit the pose compensation coefficients corresponding to this mesh. The pose compensation coefficients include rotation components, scaling components and translation components. The pose compensation coefficients of each grid are aggregated by image block to establish a mapping relationship between each blurred image block and the compensation coefficients of each grid within its coverage area.

7. The self-adaptive splicing method for fuzzy retrieval of scanned documents according to claim 6, characterized in that, Based on the pose compensation coefficients, the initial pose relationship of the corresponding image blocks is corrected grid by grid to obtain the corrected pose relationship, which is used to re-stitch multiple blurred image blocks to obtain the final stitched image, including: Based on the established mapping relationship, determine the adaptive deformation meshes covered by each blurred image block and their corresponding pose compensation coefficients; For each blurred image patch, its coverage area is divided into multiple sub-regions corresponding to the adaptive deformation mesh, and the pose compensation coefficient of the mesh to which each sub-region belongs is associated. For each sub-region, a local homography transformation matrix is ​​constructed based on the rotation, scaling, and translation components in its associated pose compensation coefficients. The local homography transformation matrix is ​​applied to the coordinates of all pixels in the sub-region to achieve pose correction of the sub-region, so that the local rigid texture primitives and global topological constraint manifold feature points in the sub-region gradually converge to the theoretical coordinate positions. Traverse all sub-regions covered by the same blurred image block to complete the grid-by-grid non-uniform pose correction of the image block, and obtain the corrected pose relationship of the image block. All blurred image blocks are reprojected onto the same stitching plane according to the corrected pose relationship, and the overlapping areas of adjacent image blocks are fused at the pixel level to generate the final stitched image.

Citation Information

Patent Citations

  • Panoramic image real-time splicing algorithm and system based on multi-sensor fusion

    CN120823091A

  • Scanning image processing method and device based on single camera, equipment and medium

    CN121056571A