A method for automatic splicing of document fragments based on a LoFTR-SIFT hybrid strategy

CN122780977APending Publication Date: 2026-09-18LANZHOU INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611000002.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

无检测器方法LoFTR借助Transformer全局注意力机制实现像素级稠密匹配,在弱纹理及重复纹理场景中表现极佳,但当碎片缺损超过30%或大视角变化超过30°时,匹配点数急剧下降,成功率降至50%至60%

Benefits of technology

(1)本发明提出了一种基于LoFTR-SIFT混合策略的文献碎片自动缀合方法,基于深度学习与传统特征匹配的协同互补,通过设计自适应混合匹配策略和全局拓扑优化技术构建图像拼接模型,从而为文献修复提供智能化辅助工具,实现对破损、低纹理、边缘残缺的文献碎片进行高鲁棒性自动缀合。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780977A_ABST
    Figure CN122780977A_ABST
Patent Text Reader

Abstract

This invention discloses a method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy, belonging to the field of image processing technology. First, the original image is augmented using enhancement techniques such as Gaussian blur, rotation, noise reduction, and color shifting to create a dedicated dataset. Then, ResNet-FPN is used as the backbone network to extract multi-scale features, enabling the model to have hierarchical representation capabilities from edges to semantics. Subsequently, detector-free dense matching is achieved through the self-attention and cross-attention mechanisms of Transformer, and an adaptive backoff strategy based on dynamic evaluation of inlier rate is introduced, automatically switching to SIFT when the matching quality is insufficient, achieving a synergistic complementarity of deep learning-first and traditional methods as a fallback. Finally, the optimal splicing path is planned using the maximum spanning tree to suppress error accumulation, multi-band fusion is used to eliminate visual seams, and a full-process web interactive system is built based on the Flask framework. This invention can significantly improve the effect of document fragment joining in scenarios such as damage, low texture, and large viewpoint changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy. Background Technology

[0002] Historical document fragments, due to their age and poor preservation conditions, often exhibit low texture, incomplete edges, and severe damage. Image stitching technology can integrate multiple overlapping images of the same scene into a wide-field panoramic view, widely used in remote sensing mapping, urban security, virtual reality, and medical image analysis. However, in the digital preservation of cultural heritage, especially in the task of piecing together fragments of Dunhuang documents, stitching faces unique challenges. For example, the documents unearthed from the Dunhuang Mogao Caves' Library Cave are ancient and poorly preserved; the paper manuscripts are mostly broken into irregularly shaped fragments, generally exhibiting low texture features such as blank paper surfaces, faded ink, incomplete edges, and severe damage. Traditional image stitching methods like SIFT plus RANSAC rely on manually designed local gradient extrema points to extract features, making it difficult to obtain stable and reliable matching points in low-texture areas, with a stitching success rate of only 50% in weak-texture scenes. While deep learning-based detector methods like SuperPoint plus SuperGlue improve keypoint density, their underlying principle remains a detection-then-match paradigm, facing the fundamental problem of having no detectable points in textureless areas. The detector-free LoFTR method achieves pixel-level dense matching by leveraging the global attention mechanism of the Transformer, which performs exceptionally well in scenes with weak or repetitive textures. However, when fragment loss exceeds 30% or the field of view changes exceed 30°, the number of matching points drops sharply, and the success rate falls to 50% to 60%. Furthermore, sequential stitching of multiple images can lead to error accumulation, resulting in distortion and misalignment of the panoramic image. Summary of the Invention

[0003] To overcome the shortcomings of the prior art, this invention provides a method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy, which enables highly robust automatic joining of damaged, low-texture, and edge-damaged document fragments.

[0004] To achieve the above objectives, the present invention adopts the following technical solution, including: A method for automatic document fragment concatenation based on a LoFTR-SIFT hybrid strategy includes the following steps: S1. Preprocess the original images of the Dunhuang document fragments obtained, screen valid images, remove damaged images, expand the sample diversity through image augmentation, crop image blocks of uniform size, and construct a dataset. S2. Multi-scale feature extraction and dense feature matching are completed by combining the ResNet-FPN backbone network with the Transformer attention mechanism, realizing global and local feature enhancement and pixel-level accurate matching of fragmented images; S3. Construct an adaptive hybrid matching mechanism that combines LoFTR deep learning dense matching with SIFT traditional feature matching. Based on the inlier rate and the total number of matching points, dynamically switch matching branches to improve the robustness of matching damaged document fragments. S4. The maximum spanning tree algorithm is used to complete the global topology optimization of multiple fragments, eliminate the accumulation of sequential stitching errors, and combine the multi-mode image fusion strategy to remove stitching seams and generate a seamless panoramic restoration image. S5. Construct quantitative evaluation indicators for adapting to the splicing of literature fragments, and build a Web interactive system to realize the full-process engineering application of intelligent fragment grouping, automatic splicing, process visualization and result export.

[0005] Preferably, step S1 is as follows: filter complete and valid fragment images with a long side not lower than a set value, and remove damaged and invalid images; perform image augmentation using Gaussian noise, Gaussian blur, HSV color shift, and random rotation; crop the images to a uniform size and construct a dataset.

[0006] Preferably, step S2 includes: a ResNet-FPN backbone network, a LoFTR module, a coarse matching module, and a fine matching module; The ResNet-FPN backbone network is used to output global coarse feature maps and local fine feature maps; The LoFTR module flattens the coarse feature map and overlays it with a two-dimensional sinusoidal positional code. The coarse feature maps of the two positionally encoded images are simultaneously input into the Transformer module, which consists of multiple layers of self-attention and cross-attention that are stacked alternately. The self-attention layer performs global aggregation on feature points within the same image to enhance the feature representation of key regions; the cross-attention layer establishes dense correspondences between feature points in the two images; through the alternating stack of multiple attention layers, global feature interaction and enhancement across images are completed. The coarse matching module calculates the similarity matrix between coarse feature points of two images based on the enhanced coarse feature maps. After obtaining the similarity matrix, a bidirectional Softmax operation is used to obtain the matching confidence matrix between the coarse feature points of the two images. A confidence threshold is set and a nearest neighbor condition is applied to select matching pairs that simultaneously satisfy bidirectional optimality and have a confidence score greater than the confidence threshold, thus forming the coarse matching point set.

[0007] The fine matching module uses fine feature maps to perform local window sub-pixel fine-tuning on coarse matching points. Each pixel position within the window corresponds to a fine feature vector. By calculating the dot product of the feature vectors within two windows, a correlation response map is obtained. The response map is normalized to obtain a probability distribution. Based on this probability distribution, the expected value is calculated to obtain the sub-pixel precision offset, thereby correcting the coarse matching points to fine matching points, and finally obtaining the fine matching point set.

[0008] Preferably, the similarity matrix in the coarse matching module Matching confidence matrix The calculation method is as follows:

[0009] in, and Let i and j be the coarse feature vectors corresponding to coarse feature points i and j in images A and B, respectively, and τ be the temperature coefficient. In the fine matching module, based on the offset Coarse matching points Corrected to fine matching points , .

[0010] Preferably, step S3 is as follows: First, perform LoFTR dense matching on the input image, i.e., proceed to step S2, to obtain the initial set of matching points; The RANSAC algorithm is used to calculate the total number of matching points N and the inlier rate r in the initial matching point set. If r ≥ the inlier rate threshold and N ≥ the total number of matching points threshold, the matching quality is deemed acceptable, and the LoFTR dense matching result is used to output the corresponding homography matrix. Otherwise, the matching quality is deemed unacceptable, and the process is switched to SIFT matching. After key point detection, FLANN matching, Lowe ratio test and RANSAC geometric verification, the matching result and homography matrix are output.

[0011] Preferably, in step S4, the global topology optimization specifically involves: constructing an undirected weighted graph with each fragment as a node and the number of matching points between fragments as edge weights; constructing a maximum spanning tree using the Kruskal algorithm; selecting the node with the highest degree as the global reference root node; and recursively obtaining the global unified homography matrix of all fragments layer by layer to suppress the accumulation of sequential splicing errors and geometric drift.

[0012] Preferably, in step S4, the weighted average fusion method completes image fusion by generating a binary mask, performing Gaussian blur processing, and weighted summation of pixels, which is used to achieve fast preview; the multi-band fusion method constructs a Gaussian pyramid and a Laplacian pyramid, and reconstructs a complete image by weighted fusion of the images at each level of the pyramid, which is used to completely eliminate visual seams.

[0013] Preferably, in step S5, the quantitative evaluation index is the normalized weighted projection error. The calculation formula is:

[0014] Where D is the length of the image diagonal, in pixels; For the weight of the matching points, , These are the original coordinates and projected coordinates of the matching point, respectively.

[0015] Preferably, in step S5, the Web interaction system is a B / S architecture platform built on Flask, which supports SSE real-time progress push, batch image upload and verification, fragment intelligent clustering and grouping, matching and alignment visualization, detailed magnification observation, multi-format result export and task log recording.

[0016] The present invention also provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for automatic document fragment joining based on the LoFTR-SIFT hybrid strategy.

[0017] The advantages of this invention are: (1) This invention proposes an automatic document fragment stitching method based on the LoFTR-SIFT hybrid strategy. Based on the synergistic complementarity of deep learning and traditional feature matching, an image stitching model is constructed by designing an adaptive hybrid matching strategy and global topology optimization technology, thereby providing an intelligent auxiliary tool for document restoration and realizing highly robust automatic stitching of damaged, low-texture, and edge-damaged document fragments.

[0018] (2) This invention addresses the characteristics of low texture, scarce samples, and complex degradation types of document fragments by using Gaussian blur, rotation, noise, and color shift enhancement to augment the original images, creating a dedicated dataset and performing size and pixel normalization, effectively solving the problems of insufficient data and uneven distribution. Multi-scale features are extracted using the ResNet-FPN backbone network, enabling the model to have hierarchical representation capabilities from edges to semantics. Detector-free dense matching is achieved through the self-attention and cross-attention mechanisms of the Transformer, thus providing reliable geometric invariance priors for extremely damaged samples. An adaptive backoff strategy based on LoFTR-SIFT with dynamic evaluation of inlier rate is introduced, automatically switching to SIFT when matching quality is insufficient, achieving a synergistic complementarity of deep learning priority and traditional methods as a fallback. The optimal splicing path is planned using the maximum spanning tree to suppress error accumulation, and multi-band fusion is used to eliminate visual seams, achieving automatic splicing of document fragments. A full-process web interactive system is built based on the Flask framework, providing batch upload, intelligent grouping, real-time progress push, and panoramic image download functions.

[0019] (3) This invention innovatively designs a LoFTR-SIFT adaptive hybrid matching mechanism, which breaks through the limitations of a single matching algorithm. Relying on the dense matching advantage of LoFTR, it can accurately adapt to low-texture, weak-edge conventional fragments of documents. For extreme scenarios such as large-area missing fragments, very few effective overlapping areas, and scarce text, it automatically switches to SIFT feature matching as a fallback. It uses the geometric invariance of traditional features to ensure the effectiveness of matching, completely solves the problems of single algorithm matching failure and homography matrix estimation drift, and greatly improves the applicable scenarios and fault tolerance of the algorithm.

[0020] (4) This invention relies on ResNet-FPN multi-scale feature extraction combined with Transformer global attention mechanism to achieve pixel-level dense matching without the need for explicit key point detection in traditional algorithms; through bidirectional Softmax confidence constraint and sub-pixel fine-tuning dual optimization, it effectively suppresses mismatches and greatly improves the feature alignment accuracy of weak texture document fragments, adapting to the inherent characteristics of sparse texture and incomplete handwriting in document fragments.

[0021] (5) This invention abandons the traditional sequential splicing method and introduces the global topology optimization strategy of maximum spanning tree. It constructs a topology graph with the number of points in fragment matching as the weight, unifies the global reference coordinate system, and transmits transformation parameters layer by layer. It eliminates the unidirectional geometric drift and error accumulation problems of multi-fragment splicing from the root, and significantly improves the overall consistency of multi-fragment splicing.

[0022] (6) The present invention designs a dual-mode mechanism of weighted average fast fusion and multi-band high-precision fusion. Fast fusion meets the real-time preview requirements of the Web terminal, with high computation efficiency and low latency. Multi-band fusion completely eliminates the problems of brightness jump, splicing seam and color discontinuity in the overlapping area of ​​fragments through pyramid layer reconstruction, and outputs a high-definition and smooth panoramic image of document restoration, which is suitable for different business use scenarios.

[0023] (7) This invention is based on Flask to build a B / S architecture Web interactive system, which integrates intelligent fragment grouping, real-time progress push, visual verification, detailed magnification observation and one-click result export functions. It can be operated without professional algorithm foundation. At the same time, it has built-in log recording, responsive layout and standardized interface, which can be compatible with multiple terminals and facilitate subsequent platform integration, truly realize the implementation of algorithm technology and meet the actual work needs of document digital restoration.

[0024] (8) Targeted data augmentation is performed to address the degradation characteristics of real documents, simulating real defects such as scanning noise, blurring, and color shift, and the samples are closely matched to real scenarios; the data set division method is standardized to effectively improve the model's generalization ability and robustness on real damaged fragments.

[0025] (9) The algorithm deployment stage simplifies the reasoning structure. Under the premise of ensuring that the splicing accuracy meets the needs of cultural relic restoration, redundant fine matching modules are disabled to save video memory, improve reasoning speed, and balance accuracy and real-time performance. It is suitable for online deployment and batch processing of Web interactive systems. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating the overall framework of automatic document fragment joining in an embodiment of the present invention.

[0027] Figure 2 This is a flowchart of the data preprocessing process according to an embodiment of the present invention.

[0028] Figure 3 This is an embodiment of the LoFTR dense matching process based on Transformer feature enhancement.

[0029] Figure 4 This is a diagram showing the internal structure of the LoFTR dense matching module based on Transformer feature enhancement in an embodiment of the present invention.

[0030] Figure 5 This is a flowchart of the LoFTR-SIFT hybrid matching decision-making process according to an embodiment of the present invention.

[0031] Figure 6 This is a schematic diagram of global optimization of multi-fragmentation based on the maximum spanning tree according to an embodiment of the present invention.

[0032] Figure 7This is a schematic diagram of the multi-band fusion process according to an embodiment of the present invention.

[0033] Figure 8 This is a schematic diagram of the main interface of the Web interactive system according to an embodiment of the present invention.

[0034] Figure 9 This is a schematic diagram of intelligent grouping in a Web interaction system according to an embodiment of the present invention.

[0035] Figure 10 This is a schematic diagram of spatial correction and alignment in a Web interaction system according to an embodiment of the present invention.

[0036] Figure 11 This is a schematic diagram of the final splicing result of the Web interaction system according to an embodiment of the present invention. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] This embodiment provides an automated and robust fragment joining service based on a deep learning and traditional feature collaboration mechanism for the scenario of piecing together Dunhuang manuscript fragments. The overall framework and process of automatic fragment joining are as follows: Figure 1 As shown.

[0039] This invention discloses a method for automatic document fragment stitching based on a LoFTR-SIFT hybrid strategy. The method first normalizes the size and pixels of the input fragments, extracts multi-scale features using a ResNet-FPN backbone network, and then dynamically evaluates the matching quality based on the inlier rate using LoFTR-SIFT adaptive hybrid matching. A weighted graph is constructed using the number of matched inliers, and the optimal stitching path is determined using the maximum spanning tree, outputting a global homography matrix. Subsequently, multi-band fusion is used to eliminate visual seams and generate a seamless panoramic image. Finally, interaction and result output are achieved through a web system.

[0040] The following sections fully elaborate on the design concept and implementation process of the method of this invention, focusing on five core components: data preprocessing, feature extraction and Transformer dense matching, LoFTR-SIFT adaptive hybrid matching, global topology optimization and multi-band fusion, loss function and Web interaction system.

[0041] (1) Data preprocessing

[0042] The images used in this embodiment are from the Dunhuang manuscript fragment reassembly dataset publicly released by the China Scientific Data Platform. The original Dunhuang manuscript dataset only contains 95 groups of 366 color images, each with varying sizes and containing scanning noise and varying degrees of blur. Therefore, images with a side length of at least 640 pixels were retained, while damaged samples were removed, resulting in 1277 valid images. During training, Gaussian noise, Gaussian blur, HSV offset, and random rotation were added. The resulting images were then randomly cropped into uniform 640×640 training blocks, generating 9000 patches. All patches were divided into a training set of 7200 images, a validation set of 900 images, and a test set of 900 images in an 8:1:1 ratio, increasing the diversity of training samples and ensuring that the images input to the feature extractor had uniform sizes and degradation forms highly consistent with reality. The complete preprocessing workflow from the original images to the training patches is described in [link to documentation]. Figure 2 .

[0043] (2) Feature extraction and dense matching of Transformer

[0044] After data preprocessing, the images are input into the feature extraction network. This method uses ResNet-FPN as the backbone network, with a total of 11.6M parameters. ResNet-18 extracts hierarchical features from edges to semantics, and FPN outputs two branches. The coarse feature map has a resolution of 80×80 and 192 channels, which is one-eighth of the original image, and is used for global context modeling. The fine feature map has a resolution of 320×320 and 64 channels, which is half of the original image, and is used for sub-pixel fine-tuning. The fine matching module is disabled during actual inference to save GPU memory. The coarse feature map is flattened and then fed into the Transformer module after being encoded with 2D sinusoidal positional encoding. This module alternately stacks 4 layers of self-attention and 4 layers of cross-attention to establish pixel-level global dependencies between the two images. The overall process is as follows: Figure 3 As shown.

[0045] exist Figure 3 In this example, using Dy001v1_1.png and Dy001v1_2.png as examples, the diagram demonstrates inputting Dunhuang manuscript fragment images from the left and right branches respectively. A ResNet-FPN backbone network with shared weights extracts multi-scale features, obtaining coarse and fine feature maps. The coarse feature map is expanded and positionally encoded before being fed into a Transformer, where it undergoes alternating self-attention and cross-attention enhancements to achieve global context awareness and alignment. The enhanced features are then processed by bidirectional Softmax in the coarse matching stage to obtain a confidence matrix. Coarse matching point pairs are selected based on a confidence threshold of 0.2 and the nearest neighbor condition. Subsequently, sub-pixel registration is performed on a 5×5 local window around each coarse matching point on the fine feature map to obtain the final matching point pairs with sub-pixel precision. The entire process eliminates the need for explicit keypoint detection, effectively handling issues such as low texture and missing edges.

[0046] This method is in Figure 3 The coarse matching stage shown uses formulas (1) and (2) to calculate similarity and confidence.

[0047] Formula (1) is as follows, where This is a similarity matrix. and These are the coarse feature vectors of images A and B, respectively, with a temperature coefficient τ of 0.1.

[0048] Formula (1)

[0049] Formula (2) is as follows, where To match the confidence matrix, bidirectional Softmax normalization is used to ensure bidirectional consistency of the matching. This method retains matching pairs with a confidence score greater than 0.2 and satisfying the nearest neighbor condition as coarse matching results, obtaining the coarse matching point set Mc. Equations (1) and (2) are applied precisely at... Figure 3 The coarse matching stage, which is after Transformer feature enhancement and before filtering matching point pairs.

[0050] Formula (2)

[0051] The internal structure of the LoFTR dense matching module in this method is decomposed as follows: Figure 4 As shown, it is divided into four parts: ResNet-FPN backbone network, LoFTR module (Local Feature Matching Module based on Transformer), coarse matching module and fine matching module.

[0052] The ResNet-FPN backbone network outputs multi-scale features, including a coarse feature map at 1 / 8 scale and a fine feature map at 1 / 2 scale. The coarse feature map has a resolution of 80×80 and 192 channels, with each feature point corresponding to the original... Figure 8 An 8-pixel region is suitable for capturing global contextual information. The fine feature map has a resolution of 320×320 and 64 channels, with each feature point corresponding to the original... Figure 2 The ×2 pixel area retains finer spatial details, providing high-resolution feature support for subsequent sub-pixel retouching.

[0053] Secondly, in the feature enhancement stage of the LoFTR module, the coarse feature map is first flattened, transforming the 80×80 two-dimensional space into 6400 one-dimensional feature points. Then, a two-dimensional sinusoidal positional code is added to each feature point. Positional encoding enables the Transformer to perceive the relative spatial relationship of feature points in the original image, avoiding the loss of positional information due to the flattening operation. Subsequently, the coarse feature sequences of the two position-encoded images are simultaneously input into the Transformer module. This module consists of four layers of self-attention and four layers of cross-attention, stacked alternately. The self-attention layers perform global aggregation of feature points within the same image, strengthening the feature representation of key regions. The cross-attention layers establish dense correspondences between feature points in the two images, enabling feature points in one image to perceive information from all feature points in the other image. Through the alternating stacking of multiple layers of attention processing, the features of the two images achieve full interaction and global context alignment, greatly improving the feature matching ability in weakly textured regions.

[0054] Then, in the coarse matching module, the enhanced coarse features are represented as follows: and Each feature point has a dimension of 192. This method first calculates the similarity matrix between coarse feature points of two images. Similarity calculation uses temperature-scaled cosine similarity, with a temperature coefficient. To control the smoothness of the similarity distribution, a bidirectional Softmax operation is used after obtaining the similarity matrix to obtain the matching confidence matrix. Bidirectional Softmax ensures high confidence only when two feature points are each other's optimal match, effectively suppressing false matches. Subsequently, a confidence threshold of 0.2 is set, and a nearest neighbor condition is applied to select matching pairs that simultaneously satisfy bidirectional optimality and have a confidence greater than 0.2, forming a coarse matching point set. For low-texture fragments of Dunhuang documents, a considerable number of coarse matching points with high reliability can be obtained without explicit keypoint detection.

[0055] Finally, in the fine-matching module, sub-pixel accuracy optimization of the coarse-matching points is performed using a 1 / 2-scale fine feature map. For each coarse-matching point... A 5×5 local window is cropped from the fine feature maps of both images, centered at the corresponding locations. Each pixel position within the window corresponds to a fine feature vector. The correlation response map is obtained by calculating the dot product of the feature vectors within the two windows. The response map is then normalized using Softmax to obtain a probability distribution, representing the probability of sub-pixel offset at each candidate position. The expected value is then calculated based on this probability distribution to obtain the sub-pixel precision offset. This corrects the coarse matching points into more precise fine matching points. Finally, a fine matching point set is obtained. .

[0056] It should be noted that, in order to save video memory and improve running speed during the actual inference stage, the fine matching module in the lower right corner is disabled. Figure 4 The dashed boxes and text labels indicating that the inference phase is disabled are used to illustrate this; therefore, only a coarse matching point set is output in the final deployed system. The subsequent hybrid matching strategy was used, and experiments showed that its accuracy met the requirements for splicing Dunhuang fragments.

[0057] (3) LoFTR-SIFT adaptive hybrid matching

[0058] Although LoFTR already exhibits excellent performance in weakly textured scenes, experiments have shown that its matching point count drops sharply to less than 10 pairs when the fragmentation area exceeds 30% or the effective overlapping text region is extremely small, leading to homography estimation failure. Therefore, this method designs an adaptive hybrid strategy for dynamic evaluation of the intimacy rate. The specific process is as follows: Figure 5 As shown, LoFTR dense matching is first performed on the input image to obtain an initial set of matching points. The RANSAC algorithm is then used to calculate the inlier rate and the total number of matching points N. If the inlier rate r ≥ 0.3 and the total number of matching points N ≥ 10, the matching quality is considered acceptable, and the homography matrix is ​​directly output using the LoFTR result. And the interior point set. If the condition is not met, it automatically falls back to the SIFT matching process, and performs keypoint detection, FLANN feature matching, Lowe ratio test and RANSAC geometric verification in sequence, finally outputting the SIFT homography matrix. And matching point pairs.

[0059] Regardless of the branch, the final output will be a set of matching point pairs with no fewer than 10 interior points. The corresponding homography matrix H is also derived. The inlier rate r ≥ 0.3 and the total number of matching points N ≥ 10 were rigorously derived through extensive preliminary experiments. When the inlier rate is below 0.3, the homography matrix estimation exhibits significant drift, and fewer than 10 matching points cannot provide sufficient geometric constraints. This strategy achieves a complementary mechanism of prioritizing deep learning and using traditional methods as a fallback. It leverages the high accuracy advantage of LoFTR in conventional scenarios while significantly improving the robustness of extremely damaged samples through the geometric invariance of SIFT.

[0060] (4) Global topology optimization and multi-band fusion

[0061] After adaptive hybrid matching is completed, the reliable homography matrix and the number of interior points between each pair of fragments can be obtained. To address the error accumulation problem in multi-fragment splicing, this method further performs global topology optimization and multi-band fusion. Specifically, a maximum spanning tree is used for global topology planning. Each fragment is treated as a node, and the number of matching interior points between two fragments is used as the edge weight to construct an undirected graph. Then, the Kruskal algorithm is used to select N-1 edges in descending order of weight, actively avoiding the formation of cycles, to obtain the maximum spanning tree. The node with the highest degree is selected as the root node of the global system, and the homography matrix of each fragment relative to the global system is solved recursively along the tree edges, thereby eliminating the unidirectional drift problem commonly found in sequential splicing. Figure 6 Using four fragments as an example, the entire process of graph construction, maximum spanning tree construction, and transformation propagation is fully demonstrated.

[0062] After geometric alignment, brightness abrupt changes often occur in overlapping areas. This method offers two fusion modes. First, a weighted average fusion is used to generate a binary mask, followed by Gaussian blurring, and then weighted summation. This method is relatively fast and suitable for real-time web preview. Multi-band fusion first constructs a Gaussian pyramid and a Laplacian pyramid, then independently performs weighted fusion on each layer before reconstructing the image. This method can completely eliminate visual seams and is suitable for high-quality output. Figure 7 The specific process of multi-band fusion is given, including three steps: pyramid decomposition, hierarchical fusion, and image reconstruction.

[0063] (5) Loss function and Web interaction system

[0064] Since the model fine-tuning stage only optimizes the coarse matching loss, i.e., the negative log-likelihood, in order to objectively compare the accuracy of image stitching at different resolutions, this method proposes a dimensionless index normalized weighted projection error, as detailed in formula (3), where D is the length of the image diagonal in pixels.

[0065] Formula (3)

[0066] Finally, because the algorithm module needed to be transformed into a tool that cultural relic restorers could directly use, this method developed a web-based interactive system using a B / S architecture based on the Flask framework. The overall system interface layout is clear, the operation entry points are well-defined, and the main view is designed as follows: Figure 8 As shown.

[0067] The backend uses SSE technology to push the processing progress of each stage—feature extraction, matching, global optimization, and multi-band fusion—to the frontend in real time. Users can intuitively see the current splicing task's execution status without refreshing the page. The frontend interface is built on HTML5 and JavaScript, supporting batch uploading of Dunhuang document fragment images. The system automatically recognizes image formats and performs size verification. The intelligent grouping function automatically clusters multiple uploaded fragments based on the number of matching interior points and the consistency of the homography matrix between images, grouping fragments belonging to the same original document appearance into the same splicing group, and automatically selecting a reference image within each group. Users can view the grouping results on the interface and manually adjust the groups, improving interactive flexibility. The intelligent grouping interface is shown below. Figure 9 As shown.

[0068] The stitching execution module invokes the complete LoFTR-SIFT hybrid matching and global optimization process from the backend. The system features a visualization panel for spatial correction and alignment, displaying the feature point matching effect and the alignment result after homography transformation in real time, facilitating user evaluation of stitching quality. The real-time progress push interface during the spatial correction and alignment process is shown below. Figure 10 As shown.

[0069] After all the stitching groups are processed, the interface automatically redirects to the final stitching result comparison and download page. The main area of ​​this page uses a two-column comparison view. The left side displays the stitched effect of this design, while the right side displays the original fragment image for comparison. Users can hover the mouse over either image, which automatically activates the magnifying glass function for detailed observation of the stitching edges and text continuity. The stitched image on the left clearly shows Dunhuang text such as "Rinan has a tavern" and "fine wine," while the original image on the right retains the original state of the fragments. A download button is located at the bottom of the page, allowing users to download the complete panoramic image with one click, supporting JPEG and PNG formats. This page serves as a direct embodiment of the algorithm's application in practical cultural relic preservation and lays the foundation for its widespread application in the digital restoration of Dunhuang documents. The final stitching result comparison and download page is shown below. Figure 11 As shown.

[0070] Furthermore, this method integrates a logging module, automatically saving key parameters and runtime for each stitching task. The front-end adopts a responsive layout to adapt to different display devices. The back-end interface follows the RESTful style, facilitating future integration into larger digital protection platforms.

[0071] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for automatic document fragment reassembly based on a LoFTR-SIFT hybrid strategy, characterized in that, Includes the following steps: S1. Preprocess the original images of the Dunhuang document fragments obtained, screen valid images, remove damaged images, expand the sample diversity through image augmentation, crop image blocks of uniform size, and construct a dataset. S2. Multi-scale feature extraction and dense feature matching are completed by combining the ResNet-FPN backbone network with the Transformer attention mechanism, realizing global and local feature enhancement and pixel-level accurate matching of fragmented images; S3. Construct an adaptive hybrid matching mechanism that combines LoFTR deep learning dense matching with SIFT traditional feature matching. Based on the inlier rate and the total number of matching points, dynamically switch matching branches to improve the robustness of matching damaged document fragments. S4. The maximum spanning tree algorithm is used to complete the global topology optimization of multiple fragments, eliminate the accumulation of sequential stitching errors, and combine the multi-mode image fusion strategy to remove stitching seams and generate a seamless panoramic restoration image. S5. Construct quantitative evaluation indicators for adapting to the splicing of literature fragments, and build a Web interactive system to realize the full-process engineering application of intelligent fragment grouping, automatic splicing, process visualization and result export.

2. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 1, characterized in that, Step S1 is as follows: Select complete and valid fragment images with a long side not lower than the set value, and remove damaged and invalid images; use Gaussian noise, Gaussian blur, HSV color shift, and random rotation to augment the images; crop the images to a uniform size and construct a dataset.

3. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 1, characterized in that, Step S2 includes: ResNet-FPN backbone network, LoFTR module, coarse matching module, and fine matching module; The ResNet-FPN backbone network is used to output global coarse feature maps and local fine feature maps; The LoFTR module flattens the coarse feature map and overlays it with a two-dimensional sinusoidal positional code. The coarse feature maps of the two positionally encoded images are simultaneously input into the Transformer module, which consists of multiple layers of self-attention and cross-attention that are stacked alternately. The self-attention layer performs global aggregation on feature points within the same image to enhance the feature representation of key regions; the cross-attention layer establishes dense correspondences between feature points in the two images; through the alternating stack of multiple attention layers, global feature interaction and enhancement across images are completed. The coarse matching module calculates the similarity matrix between coarse feature points of two images based on the enhanced coarse feature maps. After obtaining the similarity matrix, a bidirectional Softmax operation is used to obtain the matching confidence matrix between the coarse feature points of the two images. A confidence threshold is set and a nearest neighbor condition is applied to select matching pairs that simultaneously satisfy bidirectional optimality and have a confidence score greater than the confidence threshold, thus forming the coarse matching point set. The fine matching module uses fine feature maps to perform local window sub-pixel fine-tuning on coarse matching points. Each pixel position within the window corresponds to a fine feature vector. By calculating the dot product of the feature vectors within two windows, a correlation response map is obtained. The response map is normalized to obtain a probability distribution. Based on this probability distribution, the expected value is calculated to obtain the sub-pixel precision offset, thereby correcting the coarse matching points to fine matching points, and finally obtaining the fine matching point set.

4. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 3, characterized in that, Similarity matrix in coarse matching module Matching confidence matrix The calculation method is as follows: in, and Let i and j be the coarse feature vectors corresponding to coarse feature points i and j in images A and B, respectively, and τ be the temperature coefficient. In the fine matching module, based on the offset Coarse matching points Corrected to fine matching points , .

5. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 1, characterized in that, Step S3 is as follows: First, perform LoFTR dense matching on the input image, i.e., proceed to step S2, to obtain the initial set of matching points; The RANSAC algorithm is used to calculate the total number of matching points N and the inlier rate r in the initial matching point set. If r ≥ the inlier rate threshold and N ≥ the total number of matching points threshold, the matching quality is deemed acceptable, and the LoFTR dense matching result is used to output the corresponding homography matrix. Otherwise, the matching quality is deemed unacceptable, and the process is switched to SIFT matching. After key point detection, FLANN matching, Lowe ratio test and RANSAC geometric verification, the matching result and homography matrix are output.

6. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 1, characterized in that, In step S4, the global topology optimization specifically involves: constructing an undirected weighted graph with each fragment as a node and the number of matching points between fragments as edge weights; constructing a maximum spanning tree using the Kruskal algorithm; selecting the node with the highest degree as the global reference root node; and recursively obtaining the global unified homography matrix of all fragments layer by layer to suppress the accumulation of sequential splicing errors and geometric drift.

7. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 1, characterized in that, In step S4, the weighted average fusion method completes image fusion by generating a binary mask, Gaussian blurring, and pixel weighted summation to achieve fast preview; the multi-band fusion method constructs a Gaussian pyramid and a Laplacian pyramid, and reconstructs a complete image by weighted fusion of images at each level of the pyramid to completely eliminate visual seams.

8. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 1, characterized in that, In step S5, the quantitative evaluation index is the normalized weighted projection error. The calculation formula is: Where D is the length of the image diagonal, in pixels; For the weight of the matching points, , These are the original coordinates and projected coordinates of the matching point, respectively.

9. The method for automatic document fragment joining based on a LoFTR-SIFT hybrid strategy according to claim 1, characterized in that, In step S5, the Web interaction system is a B / S architecture platform built on Flask, which supports SSE real-time progress push, batch image upload and verification, fragment intelligent clustering and grouping, matching and alignment visualization, detailed magnification observation, multi-format output export and task log recording.

10. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for automatic document fragment joining based on the LoFTR-SIFT hybrid strategy as described in any one of claims 1 to 9.